How testing works

What Is a Behaviorally Anchored Rating Scale?

Published 23 August 2026 5 min read All articles
In short
  • A behaviorally anchored rating scale ties every point on the scale to a specific, observable behavior instead of a vague label like "excellent" or "poor"
  • Industrial psychologists Smith and Kendall built the format in 1963 to fix graphic rating scales, which were easy to skew with a halo effect
  • A 2022 study of medical students rated on simulated emergencies found inter-rater agreement still ranged from 0.256 to 0.529 across eight domains, below what the field treats as reliable
  • The anchors remove one source of disagreement, but a rater still needs training and familiarity with the skill being scored to use the scale well
Contents

A behaviorally anchored rating scale, BARS for short, is a scoring method that ties each point on a rating scale to a specific, observable behavior instead of a vague label like "excellent" or "poor." A rater does not decide how skilled someone seemed. They decide which described behavior most closely matches what they actually saw, which is a narrower and more checkable judgement.

The difference sounds small until you have sat on both sides of a review. A generic five point scale for "communication" asks a rater to translate an impression into a number, and two careful raters can land on different numbers for the same performance because each is averaging a different mental picture. A BARS replaces that translation step with a description: "asked a clarifying question before restating the deadline back to the team" scores differently, on purpose, from "waited to be told twice." The rater is matching evidence, not forming an opinion.

Where the method came from

Industrial psychologists Patricia Cain Smith and Lorne Kendall built the format in 1963 as a fix for graphic rating scales, which had been the industry standard for decades and were notoriously easy to skew with a halo effect: rate someone highly on one trait and every other trait on the same form tends to drift upward with it. Their method starts by collecting real critical incidents, actual descriptions of things people did on the job, from the practitioners who supervise or perform that work, then sorting those incidents onto scale points through a process called retranslation, where an independent group of raters sorts the same incidents back onto the dimensions they were meant to describe. An anchor only earns its place on the scale once independent raters agree, separately, on where it belongs.

That construction cost is the whole point. A BARS is expensive to build precisely because it forces the vague half of a rating, what "good communication" even looks like in this specific job, to get settled once, in public, before a single candidate is ever scored.

See how your own judgement scores against a behavioral anchor

Twenty four scenarios, seven minutes, one behaviorally anchored score per domain instead of a resume adjective.

Take the check

Does it actually close the gap

The honest answer is that a well built BARS narrows disagreement between raters, but it does not remove the rater from the equation, and a 2022 study of second-year medical students rated on non-technical skills during simulated pediatric emergencies, published in Medical Education Online, shows exactly how much still rides on who is holding the scale. Three trained raters scored the same recorded scenarios using a BARS across eight domains, and inter-rater agreement in that study, measured by Krippendorff's alpha, ranged from 0.256 for team communication to 0.529 for individual decision making, below what the field treats as an acceptable reliability floor. The same paper states plainly that the BARS "demonstrated limited reliability when assessing medical students during their pediatric clerkship."

That is not a verdict against the format itself. The same study points to an earlier trial of a similar behaviorally anchored tool with experienced anesthesiologists, where good inter- and intra-rater reliability came out of only two hours of rater training, and reasons that its own novice medical students, still building the baseline clinical judgement the scenarios assumed, were a harder population to rate consistently than the training time alone could fix. The scale format was the same kind of tool in both cases. The gap was who was using it and what they already knew going in.

Why the raters matter more than the rubric

That is the part a BARS cannot do by itself. Anchoring a scale to real behavior removes the vaguest source of disagreement, whether two people mean the same thing by "excellent," but it still needs raters who are trained on the anchors and, ideally, familiar enough with the underlying skill to recognize it when it shows up in an unfamiliar form. A rubric without a calibrated rater behind it is still a guess, just a more organized one.

This is why serious use of the format pairs it with a scenario that gives the rater something concrete to score in the first place, the same logic behind a situational judgement test: present a realistic problem, then grade the response against anchors that practitioners already agreed on, rather than asking a candidate to self-report a trait or a rater to guess at one. It is also why structured interviews built on a fixed scoring key outperform unstructured ones, and why unscored self-assessment correlates only weakly with actual performance: the format of the question matters less than whether the answer gets scored against something checkable. The six-domain check behind this site follows the same rule, scoring recorded responses against behavioral anchors rather than asking anyone, rater or candidate, to just know.

FAQ

What makes a rating scale "behaviorally anchored"?
Each point on the scale is tied to a description of an actual behavior, drawn from real critical incidents and sorted onto the scale through a validation process called retranslation, rather than a general label like "good" or "poor" that every rater interprets differently.
Is a behaviorally anchored rating scale more accurate than a normal rating scale?
It removes one major source of disagreement, differing interpretations of a vague label, but a 2022 study found that agreement between raters using a BARS still varied widely depending on how experienced and well calibrated those raters were.
Can anyone use a behaviorally anchored rating scale correctly with no training?
No. The anchors reduce ambiguity in the scale itself, but a rater still has to recognize the described behavior when it happens, which research on the format ties to both training time and prior familiarity with the skill being rated.
Where is this method used outside of medicine?
Structured hiring assessments, performance reviews, and simulation-based training across fields from aviation to customer service use behaviorally anchored scales for the same reason: it is the difference between grading an impression and grading evidence.
How much of you can AI replace?

Find out where you actually stand.

Six domains, twenty four items, one score. It takes about seven minutes and tells you which parts of your work AI is closest to, and which parts it is not.

Take the check

Free · about 7 minutes · no account