What Is a Behaviorally Anchored Rating Scale?
In short
- A behaviorally anchored rating scale ties every point on the scale to a specific, observable behavior instead of a vague label like "excellent" or "poor"
- Industrial psychologists Smith and Kendall built the format in 1963 to fix graphic rating scales, which were easy to skew with a halo effect
- A 2022 study of medical students rated on simulated emergencies found inter-rater agreement still ranged from 0.256 to 0.529 across eight domains, below what the field treats as reliable
- The anchors remove one source of disagreement, but a rater still needs training and familiarity with the skill being scored to use the scale well
Contents
A behaviorally anchored rating scale, BARS for short, is a scoring method that ties each point on a rating scale to a specific, observable behavior instead of a vague label like "excellent" or "poor." A rater does not decide how skilled someone seemed. They decide which described behavior most closely matches what they actually saw, which is a narrower and more checkable judgement.
The difference sounds small until you have sat on both sides of a review. A generic five point scale for "communication" asks a rater to translate an impression into a number, and two careful raters can land on different numbers for the same performance because each is averaging a different mental picture. A BARS replaces that translation step with a description: "asked a clarifying question before restating the deadline back to the team" scores differently, on purpose, from "waited to be told twice." The rater is matching evidence, not forming an opinion.
Where the method came from
Industrial psychologists Patricia Cain Smith and Lorne Kendall built the format in 1963 as a fix for graphic rating scales, which had been the industry standard for decades and were notoriously easy to skew with a halo effect: rate someone highly on one trait and every other trait on the same form tends to drift upward with it. Their method starts by collecting real critical incidents, actual descriptions of things people did on the job, from the practitioners who supervise or perform that work, then sorting those incidents onto scale points through a process called retranslation, where an independent group of raters sorts the same incidents back onto the dimensions they were meant to describe. An anchor only earns its place on the scale once independent raters agree, separately, on where it belongs.
That construction cost is the whole point. A BARS is expensive to build precisely because it forces the vague half of a rating, what "good communication" even looks like in this specific job, to get settled once, in public, before a single candidate is ever scored.
See how your own judgement scores against a behavioral anchor
Twenty four scenarios, seven minutes, one behaviorally anchored score per domain instead of a resume adjective.
Does it actually close the gap
The honest answer is that a well built BARS narrows disagreement between raters, but it does not remove the rater from the equation, and a 2022 study of second-year medical students rated on non-technical skills during simulated pediatric emergencies, published in Medical Education Online, shows exactly how much still rides on who is holding the scale. Three trained raters scored the same recorded scenarios using a BARS across eight domains, and inter-rater agreement in that study, measured by Krippendorff's alpha, ranged from 0.256 for team communication to 0.529 for individual decision making, below what the field treats as an acceptable reliability floor. The same paper states plainly that the BARS "demonstrated limited reliability when assessing medical students during their pediatric clerkship."
That is not a verdict against the format itself. The same study points to an earlier trial of a similar behaviorally anchored tool with experienced anesthesiologists, where good inter- and intra-rater reliability came out of only two hours of rater training, and reasons that its own novice medical students, still building the baseline clinical judgement the scenarios assumed, were a harder population to rate consistently than the training time alone could fix. The scale format was the same kind of tool in both cases. The gap was who was using it and what they already knew going in.
Why the raters matter more than the rubric
That is the part a BARS cannot do by itself. Anchoring a scale to real behavior removes the vaguest source of disagreement, whether two people mean the same thing by "excellent," but it still needs raters who are trained on the anchors and, ideally, familiar enough with the underlying skill to recognize it when it shows up in an unfamiliar form. A rubric without a calibrated rater behind it is still a guess, just a more organized one.
This is why serious use of the format pairs it with a scenario that gives the rater something concrete to score in the first place, the same logic behind a situational judgement test: present a realistic problem, then grade the response against anchors that practitioners already agreed on, rather than asking a candidate to self-report a trait or a rater to guess at one. It is also why structured interviews built on a fixed scoring key outperform unstructured ones, and why unscored self-assessment correlates only weakly with actual performance: the format of the question matters less than whether the answer gets scored against something checkable. The six-domain check behind this site follows the same rule, scoring recorded responses against behavioral anchors rather than asking anyone, rater or candidate, to just know.
FAQ
What makes a rating scale "behaviorally anchored"?
Is a behaviorally anchored rating scale more accurate than a normal rating scale?
Can anyone use a behaviorally anchored rating scale correctly with no training?
Where is this method used outside of medicine?
Find out where you actually stand.
Six domains, twenty four items, one score. It takes about seven minutes and tells you which parts of your work AI is closest to, and which parts it is not.
Take the checkFree · about 7 minutes · no account