Screening false positives
A Socratic walk-through of screening false positives — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why can a test that is right ninety-nine times in a hundred still flag mostly healthy people?
You are told a screening test is 99% accurate, your result is positive, and you are handed a leaflet about next steps. It seems obvious what to conclude: you are almost certainly ill. Hold that thought, because it can be wrong by a factor of five — and it is wrong for a reason that has nothing to do with the test being poor. The test may be excellent. So what else is in the calculation that we have not been told?
Reasoning it through
REASONING #Ask what "99% accurate" actually names. Two different questions hide behind it. Sensitivity asks: of people who have the condition, what fraction does the test catch? Specificity asks: of people who do not, what fraction does it correctly clear? Both are properties of the test measured on people whose status is already known.
Now notice the direction. Both figures start from the disease and ask about the result. What you want is the reverse: given the result, what about the disease? Those are not the same question, and swapping them is the whole error.
Why does the direction matter? Because you also need to know how many people of each kind were fed into the test. Let us do it with counts rather than percentages, which is where this becomes obvious. Take a test with 99% sensitivity and 95% specificity, and screen 10,000 people for a condition that 1% of them have.
Of the 10,000, one hundred have the condition. The test catches 99 of them and misses one. So far, so good.
But 9,900 do not have it. The test clears 95% of them — and gets the other 5% wrong. Five per cent of 9,900 is 495. Four hundred and ninety-five healthy people are told they may be ill.
Now count the positives: 99 true, 495 false, 594 in total. What fraction of the people holding a positive result actually have the condition? Ninety-nine out of 594 — about 17%. Five in six are fine.
Sit with where that came from. The test made almost no errors as a proportion: it was right 99 times out of 100 on the sick and 95 out of 100 on the well. But the well outnumbered the sick ninety-nine to one, so a small error rate applied to a huge group swamped a large success rate applied to a tiny one. Nothing here is a flaw in the test. The base rate did it.
Does that mean the test was useless? No, and this is the part people overshoot. Before the test your chance was 1 in 100. After it, 1 in 6. The test moved you a long way — seventeen-fold — it just started from very far down. That is what a test does: it multiplies the odds you walked in with. It cannot manufacture a probability out of nothing.
The analogy
THE ANALOGY #Think of a metal detector at a beach that beeps for 99 of every 100 coins and also for 5 of every 100 bottle caps. It is a fine detector. But the beach is mostly bottle caps, so most of what you dig up is a bottle cap — and no improvement in how it responds to coins will change that, because the problem is what is in the sand.
Digging is cheap and harmless, whereas a false positive in medicine carries real cost — weeks of fear, a biopsy, sometimes treatment for something that would never have caused harm — so the arithmetic here has a weight the beach does not.
Clarifying the model
THE MODEL #The single most useful move is to stop thinking in percentages and think in people. "99% accurate" is almost impossible to reason with; "of 594 positives, 99 are real" is not. The counts carry the base rate with them automatically, which is exactly the thing percentages hide.
Three refinements. First, the number we computed — 17% — is the positive predictive value, and unlike sensitivity and specificity it is not a property of the test. Change who you screen and it changes. Run the same test in a population where 20% have the condition and the positive predictive value rises to about 83%. This is why symptomatic patients and screened well people get such different answers from an identical result, and why raising the base rate by testing higher-risk groups improves a programme more reliably than improving the test does.
Second, note which figure the false positives came from. Sensitivity had almost nothing to do with the answer; specificity did all the damage, because it is applied to the large group. In rare conditions, specificity is the number to interrogate.
Third, this is the same reasoning as Bayesian updating — prior odds multiplied by the strength of the evidence — and if you have met that idea, screening is its most consequential everyday instance. But you never need the theorem. Take a round population, split it by the base rate, apply each error rate to its own column, and read off the answer.
A picture of it
THE PICTURE #How to readThe widths are people, drawn to scale. The first split is the base rate: a thin ribbon of 100 with the condition, a vast one of 9,900 without. Follow each ribbon into the test and watch which stream feeds the positive result — almost the entire true-positive group arrives from the thin ribbon, yet it is dwarfed by the 5% sliver bleeding out of the wide one. The positive box holds 594 people, and only 99 of them are ill. The visual point is that a narrow slice of an enormous group beats nearly all of a tiny one.
What became clearer
WHAT CLEARED #A test result is not a verdict, it is an update. How much you should believe it depends on how many people like you were sick before anyone tested anything — and when a condition is rare, even a very good test produces mostly false alarms, because the healthy majority contributes more errors in absolute number than the sick minority contributes correct hits. The right question on receiving a positive result is not "how accurate is this test?" but "how many people who test positive on this actually have it?"
Where to go next
ONWARD #- Why screening programmes are judged on mortality rather than on cases found, and what overdiagnosis means.
- How a second, independent test changes the arithmetic — and why "independent" is doing heavy lifting there.
Key terms
TERMS #| Term | What it means |
|---|---|
| Sensitivity | of those who have the condition, the fraction the test correctly flags. |
| Specificity | of those who do not have it, the fraction the test correctly clears. |
| Prevalence (base rate) | the fraction of the tested population that actually has the condition. |
| Positive predictive value | of those testing positive, the fraction who truly have the condition; depends on prevalence, not on the test alone. |
| False positive | a positive result in someone who does not have the condition. |
Every term the collection defines is gathered in the glossary.