Value-added ranking instability
A Socratic walk-through of value-added ranking instability — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why can the same teacher rank near the top one year and near the bottom the next?
A value-added score is meant to be the fair measure. It does not ask how clever a teacher's pupils are; it asks how much they gained, given where they started. That is a real improvement on ranking schools by raw attainment, and it is why the method spread.
Then the same teacher, with the same training and the same practice, appears in the top fifth one year and the bottom fifth the next. Nobody thinks she taught brilliantly and then badly. So either the measure is capturing something that genuinely changes each year, or it is capturing something that was never about her at all. Which — and how would we tell the difference?
Reasoning it through
REASONING #Write the measured score as a sum of what could be in it. There is a persistent teacher effect, the thing we want. There is who she was given — a cohort with more prior knowledge than the model's controls know about. There is whatever happened to this class this year: a disruptive pupil, a building site outside the window, a fortnight of illness. And there is the fact that each child's test score on the day is partly the day.
Now ask which of those the arithmetic can wash out, because that is the crux and it is genuinely surprising.
Suppose each child's score has a personal, unrepeatable component with a spread of s. Average twenty-five children and the spread of that average is s divided by the square root of twenty-five — that is, s over five. Averaging a class buys a fivefold reduction in child-level noise, and with a hundred pupils it would be tenfold. That is why value-added is not hopeless.
But look at the class-level shock. The building site is not twenty-five independent draws; it is one draw landing on all twenty-five children at once. So it does not shrink at all when you average — the square root has nothing to work on. Only repeating the year reduces it, because next year brings a fresh, independent class shock.
That single observation explains the instability. The part of the score that fails to repeat is not mostly child-level noise, which averaging already handled. It is a shock attached to the class, and a teacher gets one class a year, so she gets one draw a year. However many pupils she teaches, her measured score rests on a sample of one classroom-year.
Now put that together with what a ranking does. regression-to-the-mean.md sets out the general result and asks at its close how reliability interacts with it; this is that question in its live form. If two years' measures correlate at r, a teacher measured two standard deviations above average this year is expected, next year, to sit at r times two. Set r at one half and she is expected at one standard deviation — a large apparent decline produced by nothing. And ranks exaggerate this, because the middle of a teacher distribution is crowded: a modest move in score can cross many percentiles. The top-to-bottom swings that make the newspapers are the tail behaviour of exactly this arithmetic.
I will not give you a value for r. Reported year-to-year correlations vary with subject, test, grade, and how many years are pooled, and quoting one number as the stability of value-added is how this discussion goes wrong.
The analogy
THE ANALOGY #Think of judging a farmer by one field's yield in one season. The soil and the skill are his and they persist; the twenty thousand plants in the field average out the luck of any individual seed beautifully. But the hailstorm fell on the whole field at once, and no amount of counting plants divides it away. To learn about the farmer you do not need a bigger field. You need more seasons.
fields are assigned by ownership and largely fixed, whereas classes are assigned by a timetabler who may deliberately give the difficult cohort to the strong teacher — so a teacher's "weather" is partly chosen for her, which is a problem farming does not have and which no amount of repetition solves.
Clarifying the model
THE MODEL #Three refinements, and one of them is uncomfortable.
First, instability is not the same defect as bias. Averaging several years fixes instability — it is exactly the "more seasons" move, and pooled multi-year scores are markedly steadier. It does nothing for non-random assignment, because if a teacher is given weaker cohorts every year, the error repeats and pooling makes it more confident, not less. A measure can be highly reliable and consistently wrong, which is test-validity.md's point arriving from a different direction.
Second, some of the year-to-year change is real. A teacher does have better and worse years, and teachers improve sharply early in their careers, so the measure is not obliged to be constant. That only sharpens the question of how to separate real change from an unlucky class.
Third, the honest state of the evidence. Whether value-added scores capture something durable about teachers is genuinely contested, and the central exchange is worth knowing about rather than resolving here: Chetty, Friedman and Rockoff argued from very large administrative datasets that teacher value-added predicts pupils' later outcomes, while Rothstein argued that the identifying assumptions about pupil assignment do not hold, on the evidence that value-added appears to "predict" pupils' prior attainment — which no teacher can have caused. That last test is a beautiful piece of reasoning and it is the kind of thing this whole area needs more of.
A picture of it
THE PICTURE #How to readRead the cardinality marks, because they carry the whole argument. The double bar between TEACHER and CLASS_YEAR means one class-year each: whatever attaches at that level is sampled once. The crow's foot from CLASS_YEAR to PUPIL_SCORE means many pupils per class, so PUPIL_NOISE is drawn about twenty-five times and averaging shrinks it fivefold. CLASS_SHOCK hangs off CLASS_YEAR with a one-to-one mark — a single draw shared by every pupil in the room — which is why more pupils cannot reduce it and only more years can. TRUE_SKILL is the only thing in the diagram that persists across years, and it is the only thing the score was supposed to be about.
What became clearer
WHAT CLEARED #Value-added is not unstable because tests are noisy at the level of the child; averaging a class already handles that, and handles it well. It is unstable because the teacher is the unit being judged while the class is the unit being sampled, and a teacher supplies one class per year. A sample of one is the whole problem, dressed up as a sample of twenty-five.
The load-bearing claim is that the unrepeated part is class-level rather than pupil-level, and it has a clean test. Split each teacher's own class at random into two halves and compute a score on each. If the instability were pupil-level noise, the two halves should disagree about as much as two consecutive years do. If they agree closely while the years do not, the shared class shock is confirmed — and the corollary is that no test refinement will help, only more years. Should split-half agreement come out no better than across-year agreement, this account is wrong and the trouble really is in the measurement of individual pupils.
Where to go next
ONWARD #- How many years must be pooled before a ranking is stable enough to act on.
- Why a falsification test on prior attainment is the sharpest available check on non-random pupil assignment.
Key terms
TERMS #| Term | What it means |
|---|---|
| Value-added model | an estimate of a teacher's or school's contribution, based on pupils' gains relative to what their prior attainment predicts. |
| Class-level shock | an influence common to every pupil in one class in one year, which averaging across pupils cannot reduce. |
| Reliability | the extent to which a measure would repeat itself on another occasion, distinct from whether it measures the right thing. |
| Non-random assignment | systematic allocation of particular pupils to particular teachers, which produces error that repeats rather than averaging out. |
Every term the collection defines is gathered in the glossary.