Calibrate the ratings: fair scores between inflation and a forced curve
1. Before you start
Rating inflation is what happens to performance scores when every manager rates independently, with nothing to check them against each other: the scores drift upward and spread unevenly, so a “4 out of 5” from one manager and a “4” from another do not mean the same thing. Calibration is the fix — a meeting where managers compare their proposed ratings against a single shared standard and adjust until a given score means the same thing no matter who gave it.
A tiny example. Two managers each rate ten people who do the same work at the same level. One manager calls seven of her ten “exceeds expectations”; the other calls two of his ten the same. Their teams are not three-and-a-half times apart in talent — most of that gap is the rater, not the rated. Calibration exists to squeeze that rater difference out, so the score reflects the person’s work and not their manager’s habits.
Three honest statements before you start:
- This is a Decide course. You read a situation, learn the ideas, and make a call. You do not write or run any code, and there is no calculator here — the whole task is judgment.
- The company in the taught case, Vantic, and the organisation in your final call, Rivermark Regional Hospital, are composite — invented from ordinary, realistic dynamics so the reasoning is clean. No number here is a claim about any real employer.
- This is not a certification. It proves, to you, that you can steer a set of ratings between the two ways they go wrong — everyone scored high, or everyone crammed into a forced curve — and defend the call.
The hard part is that both failures feel responsible from the inside: rating everyone highly feels kind, and enforcing a strict curve feels rigorous. Neither is fair on its own.
2. The Situation
You lead a 40-person product-engineering department at Vantic, a composite mid-size business-software company, and the annual ratings your five team leads submitted just landed on your desk: half the department sits in the top two score levels, and one lead’s team looks twice as strong as another’s. You have to sign off a set of ratings that splits a fixed merit budget — and if you rubber-stamp what came in, you reward the leads who scored generously and short-change everyone under a stricter one. The pressure from above is to “just apply a curve and be done,” which only trades one unfairness for the other.
3. What you’ll be able to do
After this course you will be able to:
- Spot rating inflation and rater drift — read a set of submitted scores and say whether a team really outperformed or was simply rated by a more generous manager, and name the evidence that would tell them apart.
- Run a calibration to a shared standard — decide how to move a rating up or down so it means the same thing across managers, comparing people to a fixed bar rather than to each other.
- Choose a fair distribution without a quota — decide when an unusual distribution is genuine and when it needs a harder look, and say why a forced curve and a rubber stamp are both the wrong answer.
4. Prerequisites & time box
Prerequisites: none beyond having seen a performance rating before — the idea that people get scored once a year on how they did. No spreadsheet or setup; the Decide hall is read-and-decide in the browser, so see the Decide hall’s how-to-read page if this is your first concept course. No prior Decide course is assumed.
Time box: about 25 minutes of reading (measured), plus your own thinking time on the call in section 7. That is at the 25-minute cap for a concept course.
Difficulty: 5 / 8 — a manager-level decision. There is no formula that hands you the answer: you weigh people against a standard, judge whether a lopsided distribution is real or a rater artefact, and hold a line between two opposite failures while a budget and other managers push on you. It sits a step above a course with one clean rule and one tidy call, and below a director-level call made under a genuine crisis. What would make it harder is a second interacting pressure — say, a layoff quota riding on the same scores.
Free-tier honesty: no signups, no paid tools, no special hardware. Nothing here costs money to learn.
5. The case & where the numbers come from
Vantic (the taught case, section 6) and Rivermark Regional Hospital (the transfer, section 7) are composite organisations: invented from typical performance-review dynamics. Every score, headcount, share, and budget figure below is an in-course illustrative assumption, chosen so the problem is easy to see and reproduce by hand — not drawn from, or claimed about, any real employer.
What is cited, in section 11, is the standard definition of each idea the course teaches — the leniency and halo biases that inflate ratings, behaviourally anchored rating scales, the forced-distribution / vitality-curve method, and how merit pay attaches money to the score — because those are established concepts with public references, not invented here.
Vantic rates its people on a five-level scale, from lowest to highest: Below expectations (1), Partially meets (2), Meets expectations (3), Exceeds (4), Outstanding (5). Here is what the five team leads submitted for the 40-person department this year (illustrative figures):
| Rating level | People | Share of department |
|---|---|---|
| Outstanding (5) | 6 | 15% |
| Exceeds (4) | 14 | 35% |
| Meets expectations (3) | 18 | 45% |
| Partially meets (2) | 2 | 5% |
| Below expectations (1) | 0 | 0% |
| Total | 40 | 100% |
Two of the five leads, each with eight reports, submitted very different pictures:
| Team lead | Reports | Rated Exceeds or Outstanding | Share |
|---|---|---|---|
| Lead A (generous) | 8 | 5 | 62.5% |
| Lead E (strict) | 8 | 1 | 12.5% |
One more figure drives the money: a fixed merit pool of $120,000 is set aside for the people rated above the standard (Exceeds or Outstanding) to share. Hold these facts — every point in section 6 reads off them.
6. The Concepts
Rating inflation
Rating inflation is the upward, uneven drift of scores that happens when managers rate independently, with no shared reference and no moderation. It is not that a manager lies; every quiet pressure simply points the same way. Rating someone highly avoids an uncomfortable conversation and keeps a valued person’s bonus and morale safe (leniency bias); one standout project colours the whole year (the halo effect); a strong recent month outweighs a weak earlier one. And scoring high costs the manager nothing — the bill for inflation lands on the department, not on them. Across five managers, scores creep up, at different rates for each.
Read Vantic’s numbers with that lens. Half the department — 20 of 40, the 6 Outstanding plus the 14 Exceeds — sits above the “meets the standard” line. When half the group is labelled above-standard, “above standard” has quietly come to mean “standard,” and the top two levels stop telling anyone apart. The sharper signal is the gap between leads: Lead A rated 5 of 8 (62.5%) above the standard, Lead E just 1 of 8 (12.5%) — a five-to-one difference in “top” ratings between two teams doing the same kind of engineering. Lead A’s team could be genuinely far stronger, but it is far more likely that Lead A rates generously and Lead E strictly: the gap is mostly the rater, not the rated.
Why it matters. Inflated ratings are unfair to the genuine top performers, whose “Outstanding” now looks like a generous “Exceeds.” They misdirect the merit money: the $120,000 pool split across 20 “above-standard” people is $6,000 each; if only about half of them were truly above the bar, the same pool would be roughly double that each — inflation roughly halves what real excellence is worth. (Calibration below lands Vantic’s honest number at 12; the point here is only the direction — fewer genuine top performers, more money each.) And they corrode trust: once people learn the score depends on which manager they drew, the exercise reads as a lottery. Inflation is the first failure mode; you will meet its opposite shortly.
What a calibration session is
A calibration session is a moderated meeting where managers put their proposed ratings side by side and adjust them, together, against one shared definition — so that a given score means the same thing regardless of who assigned it. Its single job is to remove the rater from the rating. It is not a meeting where each manager defends their own people, and it is not a negotiation where everyone trades a notch to keep the peace.
What good calibration looks like in the Vantic room:
- Evidence, not adjectives. “She is a rock star” is not a rating; “she led the billing migration two quarters early and mentored two juniors to independence” is. A score has to be earned by what the person did.
- Like compared with like. You line up people doing comparable work — the senior engineers across all five teams — so a soft “Exceeds” on one team sits next to a demanding “Meets” on another.
- Challenge in both directions. Calibration corrects Lead A’s over-scoring and Lead E’s under-scoring. A session that only ratchets scores down is a disguised cost-cut; a real one moves ratings toward the truth, up or down.
Calibration fixes inflation — but note what it does not do: it never sets a target for how many people may land in each level. That distinction is the whole game, and the next two ideas draw it.
The shared standard
Calibration only works if there is something to calibrate to. That something is a shared standard: a common, behaviour-based description of what each rating level actually looks like, so every manager is holding people to the same bar. The strongest form of this is a set of behaviourally anchored descriptions — each level tied to concrete, observable behaviour rather than a vague adjective. “Exceeds” is not “really good”; it is, say, “consistently delivers beyond the scope of the role and lifts the people around them,” with examples of what that has meant.
The standard forces the one move that makes ratings fair: compare each person to the bar, not to each other. Comparing people to each other is ranking — it can only ever produce a first, a second, and a last, even in a room full of strong performers. Comparing each person to a fixed standard lets two people clear the same bar and both earn the same score, and lets a whole team fall short if that is the truth. In the Vantic case, the shared standard is what lets you say Lead E’s “Meets” engineer and Lead A’s “Exceeds” engineer are, on the evidence, doing the same level of work — and should carry the same rating once the rater is taken out.
Without a shared standard, a calibration session is just five managers arguing from private definitions, and the loudest or most senior one wins. With it, the argument has an anchor outside any one manager’s head.
Forced ranking and its failure mode
Forced ranking — also called forced distribution, stack ranking, or the vitality curve made famous at General Electric — mandates that ratings fit a fixed shape no matter what: for example, the top 20% of people must be rated high, the middle 70% medium, and the bottom 10% low, in every team, every year. Its appeal is real and worth naming: it kills inflation instantly and forces managers to differentiate instead of rating everyone “Exceeds.”
But it fails in the mirror image of inflation, and the failure is built into the method: a forced curve is a quota, so it decides the distribution before looking at the people. That means:
- It marks down people who met the standard. If a strong team has six people clearing the bar but the curve allows three top ratings, three people who earned the score lose it — on the arithmetic of the quota, not the evidence.
- It punishes strong teams and rewards weak ones. Every team must produce a “bottom” person even if its weakest is solid, and even a struggling team hands out its quota of top ratings. The rating stops measuring the work.
- It corrodes trust and collaboration. When a fixed share must be labelled low regardless of results, people compete instead of cooperating, and strong performers leave rather than gamble on next year’s curve. This is why a string of large companies abandoned stack ranking.
Hold the two failures together. Inflation says everyone is a 4. Forced ranking says exactly 10% must be a 1, whether or not anyone earned it. Both detach the score from the person’s actual work — one by flattering, one by quota. The fair answer is neither, and finding it is the last idea.
Deciding a fair distribution
So where does a fair set of ratings land? Between the two failures — and the thing that puts it there is the shared standard, not a target. You calibrate every rating to the bar, honestly, in both directions, and then you accept the distribution that falls out, whatever shape it takes. That is the core judgment of this course: the distribution is an output of calibrating to the standard, never an input you enforce.
A reference distribution — “a typical on-standard team is maybe a quarter to a third above the bar” — still has a use, but only as a prompt to look harder, not a cap to enforce. If a team comes in at 60% above standard, you do not delete the excess; you ask the manager to walk you through the evidence person by person. Sometimes it holds and the team really is exceptional — that survives. Usually the exercise surfaces two or three “Exceeds” that are really strong “Meets,” and they move on the evidence, not to hit a number.
Run Vantic through it. Calibrated against the shared standard, the department’s above-the-bar share falls from the raw 50% (20 people) to about 30% (12 people): several of Lead A’s “Exceeds” ratings were strong “Meets” on the evidence, while two of Lead E’s people were under-rated and moved up. Now the money means something — the $120,000 pool across 12 genuine above-standard performers is $10,000 each instead of $6,000 spread thin across an inflated 20. Note what you did not do: you did not force the number to a curve. A hard 15% cap (6 people) would have denied six people a rating they could defend and over-concentrated the pool at $20,000 each — the forced-ranking failure. The fair distribution is the honest one: calibrated to the standard, defensible person by person, with no particular shape fixed in advance.
7. Your Call
You have seen how rating inflation, a calibration session, a shared standard, and the forced-ranking trap combine into one fair-distribution judgment at Vantic. Now a different call lands on your desk.
Rivermark Regional Hospital is a composite hospital, and you are its Director of Nursing — a completely different organisation and sector from Vantic’s software department. You are calibrating this year’s ratings across six charge nurses who together oversee 60 nurses, and the raw scores came in with 55% of the division in the top box (the single “Outstanding” rating); the ICU charge nurse rated 7 of her 10 nurses “Outstanding,” while the float-pool charge nurse rated 2 of 10. Midway through your calibration, Human Resources hands down a new mandate: no more than 15% of nursing staff may be rated in the top box this year, to hold the merit budget — a hard cap regardless of the evidence. A clinical-ladder advancement decision rides on a top-box rating this cycle.
How this differs from the taught case (the transfer): this is a different company and sector (a hospital nursing division, not Vantic’s corporate software team); the figures are different, so you must reason over new numbers; it is a different kind of decision (how you respond to an imposed quota handed down mid-process, not an open calibration you run from scratch); and it adds a new constraint (the forced-distribution mandate, plus a promotion riding on the score). The core concept is the same: steering ratings between inflation and a forced curve by calibrating to a shared standard.
8. Self-check
Before you write the memo, make sure you can say each of these in one sentence, without an answer key:
- What is your call on the Rivermark ratings, and what single piece of evidence — one nurse’s defensible record against the standard — would change where a given rating lands?
- Why is a five-to-one gap in top ratings between two units more likely to be the rater than the rated, and what would you look at to tell the two apart?
- Why is calibrating to a shared standard different from both rubber-stamping the raw scores and enforcing a forced curve — and what does each of those two failures cost?
- Where does your honest distribution come from — is it something you enforce or something that falls out of calibrating each person to the bar?
If any of these is fuzzy, reread the last two headings in section 6 — the shared standard and the fair-distribution judgment are the heart of the course.
9. Stretch
Push the thinking further on your own:
- The ICU unit really might be stronger — hospitals do concentrate senior nurses in intensive care. Sketch what evidence would let you keep ICU above the reference share honestly, and what would tell you the high scores are leniency instead.
- A promotion on the clinical ladder rides on a top-box rating this cycle. Describe how that stake quietly pushes charge nurses toward inflation, and one change to the process — not the people — that would lower the pressure without lowering the standard.
- The hard one: HR’s cap exists because the merit budget is fixed. Design a way to keep the honest, calibrated distribution and live within a fixed pool by separating the rating (an honest measure of the work) from the reward (how the money is split). What breaks if you let a budget constraint set the ratings instead?
10. Ship it — your decision memo
Write a one-page memo to Rivermark’s HR partner. State the call — for example: hold the calibrated ~30% top-box distribution, defensible nurse by nurse against the shared standard, and decline the 15% cap as a quota that would mark down people who met the bar; if the merit budget is the real constraint, solve it in how the money is split, not by bending the ratings. Show the reasoning in two or three lines: the raw 55% and the seven-of-ten-versus-two-of-ten gap are signs of rater drift, calibration corrects it in both directions, and the honest distribution is an output of that, not a target. Name what you rejected — rubber-stamping the inflated raw scores, and enforcing HR’s forced 15% curve. Name the one thing that would change a given rating: evidence, against the standard, that a specific nurse’s score is wrong. Keep it to a single page an HR partner grasps in two minutes. This memo is your own argued claim — not a credential.
11. Sources
Vantic and Rivermark Regional Hospital, and every score, share, headcount, and budget figure attached to them, are composite and illustrative — invented from ordinary performance-review dynamics for clean teaching, not drawn from or claimed about any real employer. What is cited below are the standard definitions of the concepts the course teaches.
| Concept / claim | Source (publisher) | URL | Accessed |
|---|---|---|---|
| Performance appraisal: rating employees against expectations | Wikipedia — Performance appraisal | https://en.wikipedia.org/wiki/Performance_appraisal | 2026-07-20 |
| Leniency error: raters scoring more generously than warranted | Wikipedia — Leniency error | https://en.wikipedia.org/wiki/Leniency_error | 2026-07-20 |
| Halo effect: one strong trait colouring an overall judgement | Wikipedia — Halo effect | https://en.wikipedia.org/wiki/Halo_effect | 2026-07-20 |
| Behaviourally anchored rating scales: levels tied to observable behaviour | Wikipedia — Behaviorally anchored rating scales | https://en.wikipedia.org/wiki/Behaviorally_anchored_rating_scales | 2026-07-20 |
| Forced ranking / vitality curve and its history | Wikipedia — Vitality curve | https://en.wikipedia.org/wiki/Vitality_curve | 2026-07-20 |
| Forced distribution as a rating method | Wikipedia — Forced distribution | https://en.wikipedia.org/wiki/Forced_distribution | 2026-07-20 |
| Grading on a curve: fitting scores to a fixed distribution | Wikipedia — Grading on a curve | https://en.wikipedia.org/wiki/Grading_on_a_curve | 2026-07-20 |
| Merit pay: attaching pay to a performance rating | Wikipedia — Merit pay | https://en.wikipedia.org/wiki/Merit_pay | 2026-07-20 |
| Composite case basis (Vantic, Rivermark): invented illustrative figures, no real employer | Provesmith concept-course contract (case_basis: composite) | https://en.wikipedia.org/wiki/Fictitious_entry | 2026-07-20 |
Next up
Finished this call? Continue the People & HR track:
Set the pay band: fix compression without breaking it for everyone → · Browse all courses