Random audits
A Socratic walk-through of random audits — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does a tax authority that inspects a few returns at random collect more honest filings than one that inspects only suspicious ones?
An audit is expensive and most returns are honest, so spending a check on a random filer looks like waste. The obvious policy is to audit where the money is: score every return for suspicion, work down the list.
Yet tax authorities keep a random component, examining people against whom nothing is suspected, and defend it as improving compliance overall. That should look like a mistake. Under what conditions is it not?
Reasoning it through
REASONING #One answer is that a known rule can be routed around: a filer who learns the score's shape sits just below it. That is the unpredictability argument, worked through for a different setting in randomized-screening.md. It applies here, but taking it as the whole answer misses the mechanism peculiar to auditing — so set it aside.
Ask instead what a purely targeted programme knows. It audits the returns it scored highly and learns whether those were wrong. What does it learn about the rest? Nothing. Under-reporting there is not merely unmeasured but unmeasurable from the programme's own records, because error is only ever observed inside the selected set.
Now follow that forward. Next year's scoring model is built from the audits it has — which are the returns the last model chose. So the model is trained on its own output. Whatever it already suspected, it confirms and refines; whatever it never suspected produces no cases, and the absence of cases reads exactly like an absence of fraud. The model does not merely stay ignorant of its blind spot. It acquires evidence that the blind spot is empty.
That is the structural reason for a random draw, and it is not about deterrence at all. A random sample is the only unbiased picture of who under-reports, because selection into it depends on nothing the authority believes. It is what lets you estimate how much revenue is lost, where, and by whom — and therefore what lets the targeted programme be any good. The United States has run such programmes for decades, the older Taxpayer Compliance Measurement Program and its successor the National Research Program, and the risk scores selecting ordinary audits are built from that data (names recalled; I quote no revenue or coverage figures, since published tax-gap numbers vary substantially with method and vintage).
Second mechanism, and now deterrence proper. What deters is a filer's perceived chance of examination. Under pure targeting, anyone who believes themselves unremarkable can rationally set that near zero. A random component makes it strictly positive for everyone, and it cannot be argued down by looking ordinary.
How much does that buy? A large Danish field experiment — Kleven and colleagues, around 2011 — assigned audits at random and separately sent audit-threat letters. Reported evasion was very low on income reported by a third party, such as wages, and substantial on self-reported income; threat letters moved the self-reported component and barely touched the rest (recalled). Audits mainly discipline the part of the return nobody else is watching. Where an employer or bank already files the same number, honesty comes from the cross-check, not the fear of an audit.
So what does the random component cost? First, yield: most of those audits find little, so revenue per hour falls far below a targeted audit, and a ministry judging the programme on collections sees a losing line. Second, burden and legitimacy: the people examined are disproportionately compliant and did nothing to attract it. The American measurement programme of the 1980s acquired a reputation for intrusiveness and was suspended, its successor designed to lean more on existing records. Measurement quality against burden on the innocent is the real constraint on sample size, and it is not a technical one.
The comparison is evidence. Systems built on comprehensive third-party reporting and pre-filled returns, as in Denmark or Chile, need much less audit effort for the same compliance, because the information architecture does the work an audit would otherwise do. Change what is observed independently and the value of examination changes with it.
The analogy
THE ANALOGY #Think of a quality inspector who only opens the boxes a supervisor has flagged. He gets very good at the faults the supervisor knows about, and his records show no others — which the supervisor reads as proof no others exist. Opening a few boxes at random tells him almost nothing about any particular box, and everything about whether the flagging is any good.
Boxes do not choose their contents, whereas filers respond to what they believe is inspected, so the sample measures a population partly reacting to the measurement. And the inspector could eventually open every box, whereas a tax authority never can.
Clarifying the model
THE MODEL #Three refinements, and one correction.
The correction: random audits do not, case for case, catch more. They catch less. The claim is about the system — honest filings overall — and runs through two channels, better targeting and broader perceived risk, neither of which shows up as a catch rate. A programme judged on yield per audit will always conclude the random component should be cut.
First refinement: the point is not that random beats targeted, but that random is the input targeting needs. The sensible design is a mixture, most effort steered by risk scores and a minority drawn at random; the argument here is only about why the minority cannot go to zero.
Second, the fixed point of difference from unpredictability. That argument concerns an adversary who observes your rule and moves around it, and would apply even if you could measure the population perfectly. This one applies even to filers who never think about audits, because the defect it fixes is in the authority's evidence, not the filer's strategy.
Third, honest limits. The deterrent effect appears concentrated where information reporting is absent, so a random programme is worth far more in a system with weak third-party data — and the size of the lasting compliance response, as against a one-off recovery, is contested.
That gives the test. If selection bias is the mechanism, a targeted-only programme should keep high detection on its familiar patterns while newly emerging schemes go unrepresented for years, and introducing a random sample should reveal under-reporting the model ranked low risk. The refuting observation: a random sample turning up under-reporting distributed just as the existing model predicted, which would mean there was no blind spot and no measurement value.
A picture of it
THE PICTURE #How to readStart at the store at the top and follow the two paths out of it. The right-hand path is ordinary targeting: the model scores, the gate sends returns above the threshold to audit and everything below into the red box, where error stays unobserved. The left-hand path is the random draw, drawn as a preparation shape because it uses no information at all. Both audits feed back into the model, and those two back-edges are the point: one carries information about the whole population, the other only about what the model already suspected. The dotted edge is the sole route by which an unexamined return can ever be seen.
What became clearer
WHAT CLEARED #The case for auditing at random is less about being unpredictable than about being able to see. A programme that only inspects what it suspects observes error nowhere else, trains next year's suspicions on this year's, and reads its own blind spot as clean ground. The random draw is the only sample whose selection does not depend on what the authority already believes, so it makes the targeting trustworthy — and separately, it stops an ordinary-looking filer concluding their chance of examination is zero. Both benefits are invisible in the statistic the programme is judged by, which is why the random share is perennially the first thing cut.
Key terms
TERMS #| Term | What it means |
|---|---|
| Selection bias | distortion arising when what gets observed is chosen by a rule related to what is being measured. |
| Third-party reporting | income reported to the authority by an employer, bank or platform as well as by the taxpayer. |
Every term the collection defines is gathered in the glossary.