THIS EXPLANATION
THE ROOM
EDU·33 Education & Learning 6 MIN · 8 STATIONS

Test validity

A Socratic walk-through of test validity — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why can a scrupulously fair test still misjudge what a student knows?

Imagine an exam administered impeccably. Every student gets the same paper, the same hour, the same silence; scripts are anonymised, two markers agree closely, and a student who sat it twice would score within a point or two. By every procedural standard it is fair. And it can still tell you something false about what a student knows.

That should be uncomfortable, because most of what we call "making a test fair" is aimed at the procedure. So where is the remaining gap — and what exactly is the thing that can be wrong?

b

Reasoning it through

REASONING #

Start with what the test actually produces. It produces a number, and a number by itself makes no claim at all. The claim arrives when someone says this student has grasped fractions, or this candidate is ready for the next course. So the thing that can be right or wrong is not the paper. It is the step from the number to the claim.

That reverses a habit of speech. We say "a valid test", and the phrase hides where the error lives. The same score can support one inference perfectly well and another not at all — a reading-comprehension score might be sound evidence for placing a child in a reading group and poor evidence for judging her teacher. Nothing about the paper changed between those two sentences. Validity is a property of an inference from a score, for a stated purpose and population.

And what is the claim usually about? Something you cannot see: mathematical reasoning, clinical judgement. Call that the construct. It has no ruler; you can only build tasks you believe require it and read the construct off performance. Every test is an argument by proxy, and an argument by proxy can fail in exactly two ways.

The first: the tasks cover only part of what you meant. Test writing with multiple-choice grammar items and you have measured something real, but not composition — not planning, not revising, not sustaining a paragraph. The construct has been under-represented, and a student strong in what you left out is scored as weaker than she is. Notice that this failure is perfectly consistent: ask her again next week and you get the same wrong answer.

The second: something other than the construct is moving the scores. Set your mathematics problems in dense prose and weak readers lose marks for reasons unrelated to mathematics. Make the paper heavily speeded and you have partly measured working pace; set a word problem about cricket and you have partly measured familiarity with cricket. This is construct-irrelevant variance — real differences between students, reliably measured, tracing to the wrong cause.

So where does reliability sit? Reliability is consistency: would the same performance earn about the same score again. It is genuinely necessary — a number that bounces around can carry no inference, and a test's agreement with an outside criterion is bounded by how much it agrees with itself. But it is not sufficient, and both failures above show why. A narrow test is consistently narrow; a contaminated test is consistently contaminated. Consistency measures the steadiness of the instrument, not its aim.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a bathroom scale that is beautifully repeatable but reads six pounds heavy, used to certify aircraft cargo. Weigh the same crate ten times and you get the same figure to the ounce; the instrument is faultless in its consistency. The trouble is entirely in the sentence "therefore the aircraft is within limits."

WHERE IT BREAKS DOWN

a scale's error is a fixed offset you could measure and subtract, whereas construct-irrelevant variance differs from student to student — the reading load costs the weak reader marks and the strong reader nothing — so there is no single correction. And a scale has an agreed standard kilogram behind it; a construct like clinical judgement has no such reference object, which is why validity is argued rather than calibrated.

d

Clarifying the model

THE MODEL #

Three refinements hold the pieces together.

First, "the test is invalid" is almost always the wrong sentence. The right one names an inference: using this score to decide X, for these students, is not well supported. That makes the claim testable, because you can then ask what evidence would support it — does the content sample the domain, do the tasks provoke the reasoning you intended, do scores relate to other things as the construct predicts.

Second, fairness of procedure and validity of inference are separate. Identical conditions for everyone is exactly what makes construct-irrelevant variance invisible: everyone met the same reading load, so nothing looks unequal, yet the score means something different for a weak reader than a strong one.

Third, a caveat. This unified framing — validity as one argument about score interpretation, with the old "content, criterion, construct" types folded in — is the position of the field's Standards and of Messick's work, and is broadly but not unanimously accepted; some argue it has grown so encompassing that it no longer tells a developer what to do. Either way the load-bearing point survives: the number is not the claim.

e

A picture of it

THE PICTURE #
Test validity
Test validity Start at V0, the claim someone actually wants to make from a score. The four arrows out of it point at the conditions it rests on -- read "derives" as derived from, so V0 holds only where all four do. The two elements at the bottom are uses of a score, and what matters is which arrows are missing. The grammar-only writing test is reliable and uncontaminated, yet has no arrow to V2: it measures a corner of writing and calls it writing. The wordy mathematics paper covers the mathematics fully and scores consistently, but has no arrow to V3: part of what it measures is reading. Each fails once, and neither failure is visible in how fairly the exam was run. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/test-validity.md","sourceIndex":1,"sourceLine":4,"sourceHash":"4bd27bb0991dfd808bb6850feccba94b5ed5c8a8cde8c3e28c32024af203a62d","diagramType":"requirement","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":2044,"height":616},"qa":{"passed":true,"findings":[]}} derives derives derives derives satisfies satisfies satisfies satisfies <<Requirement>> construct_defined ID: V1 Text: the attribute being claimed is stated clearly Risk: Medium Verification: Analysis <<Requirement>> full_coverage ID: V2 Text: the tasks sample the whole construct, not one corner Risk: High Verification: Analysis <<Requirement>> no_irrelevant_variance ID: V3 Text: differences in score come from the construct alone Risk: High Verification: Analysis <<Requirement>> reliable_scores ID: V4 Text: the same performance would score the same on a retest Risk: Medium Verification: Test <<Requirement>> valid_inference ID: V0 Text: this score supports this claim for this purpose Risk: High Verification: Demonstration <<Element>> grammar_only_writing_test Type: score use <<Element>> wordy_maths_paper Type: score use

How to readStart at V0, the claim someone actually wants to make from a score. The four arrows out of it point at the conditions it rests on — read "derives" as derived from, so V0 holds only where all four do. The two elements at the bottom are uses of a score, and what matters is which arrows are missing. The grammar-only writing test is reliable and uncontaminated, yet has no arrow to V2: it measures a corner of writing and calls it writing. The wordy mathematics paper covers the mathematics fully and scores consistently, but has no arrow to V3: part of what it measures is reading. Each fails once, and neither failure is visible in how fairly the exam was run.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Validity is not a badge a test earns and keeps. It is a claim about one inference, from one score, for one purpose — and it fails in two specific ways: the tasks leave out part of what you meant, or something you did not mean is moving the marks. Reliability rules out a third failure, random noise, but cannot tell you whether the instrument is pointed at the right thing. A scrupulously fair exam is scrupulously fair about the procedure, and the procedure was never where the error was.

g

Where to go next

ONWARD #
  • How differential item functioning detects an item behaving differently for two groups matched on ability.
  • Why attaching stakes to a measure tends to corrupt it, and what that does to a validity argument.
h

Key terms

TERMS #
TermWhat it means
Constructthe unobservable attribute a test is meant to be evidence about, such as reading comprehension.
Construct under-representationthe test samples only part of the construct, so genuine strength outside that part goes unscored.
Construct-irrelevant variancescore differences caused by something other than the construct, such as reading load on a mathematics paper.
Reliabilitythe consistency of scores across occasions, forms, or markers; necessary for a defensible inference but not sufficient.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4