THIS EXPLANATION
THE ROOM
SPT·09 Sports, Exercise & Recreation 7 MIN · 7 STATIONS

Doping arms race

A Socratic walk-through of the doping arms race — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why do athletes who would rather compete clean still feel pushed to dope even where testing is real and penalties are severe?

Ask athletes in a heavily tested sport whether they would prefer a clean field, and most say yes without hesitation. Testing exists, it is not a sham, and the penalties — years of ineligibility, forfeited results, a permanently attached reputation — are severe by any ordinary standard.

And yet the pressure does not go away. Athletes who genuinely would rather not still describe feeling pushed. That combination is the puzzle: a population that mostly wants the rule, a regime that mostly enforces it, and a persistent pull in the opposite direction. Either the athletes are confused about their own preferences, or the situation has a structure that makes wanting the rule and following it come apart.

b

Reasoning it through

REASONING #

Start with the individual decision, stated carefully.

An athlete considering a prohibited method faces a gamble. If they use it and are not caught, they gain something — call it G, the value of the performance advantage, which in a sport where places are separated by fractions is not small. If they use it and are caught, they lose L: the ban, the results, the sponsorship, the standing. If they do not use it, they get neither.

Deterrence works when the expected loss exceeds the expected gain. Write it out: the gamble is worth taking when (1 − p)·G is greater than p·L, where p is the probability of being caught. Rearranged, deterrence requires L > G·(1 − p)/p.

Now put a number in, and the shape of the problem appears. Suppose an athlete using a well-managed protocol faces a five per cent chance of detection. Then (1 − p)/p is 0.95/0.05 = 19, so the penalty must be worth more than nineteen times the advantage to deter them. Raise detection to twenty per cent and the required penalty falls to four times the gain. Raise it to fifty per cent and it falls to parity.

That arithmetic says something specific: penalties scale linearly, detection probability scales the whole problem. Doubling bans doubles L. Doubling detection roughly halves the penalty needed. Since there is a practical ceiling on how much a sport can punish someone, most of the available leverage lives in p — exactly the term that is hardest to move.

Now add the second athlete, and the private preference becomes a trap. If everyone else is clean, doping wins races. If everyone else is doping, staying clean means losing to people you could otherwise beat. Whatever the field does, the individually best response is the same — and the outcome where everyone dopes is worse for every athlete than the one where nobody does, since the advantages cancel and only the costs remain. Nobody needs to be a cynic for this to bite.

Then the element that makes it an arms race rather than a static problem. Every test is a test for something specific, and each new assay creates an incentive to find what it does not catch: a substance not yet listed, a dose below a threshold, a timing that clears before sampling, a method leaving no foreign molecule at all. Detection improves, evasion adapts, detection improves again. So p is not a dial the authorities can turn up and leave.

And this has a nasty statistical consequence. Those caught are, by construction, the ones using methods the tests can see, so the well-resourced programme is underrepresented among positives. The visible caught population is a biased sample of the doping population — and the catch rate, the most natural measure of how well the system works, is exactly the wrong thing to read as reassurance.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a queue at a ferry gate where anyone can gain by stepping half a pace out of line, and everyone can see everyone else. One person edges out and gains. The next matches to avoid losing their place. Within a minute the queue is a wedge, nobody is further forward relative to anyone else, and boarding is slower and worse-tempered for all of them.

Now station a marshal who can see part of the queue and turns back anyone caught. Doubling the penalty — sending offenders to the back rather than to their old place — changes something, but far less than doubling the number of marshals would.

WHERE IT BREAKS DOWN

A queue-jumper gains nothing if everyone jumps, so the wedge is merely wasteful, whereas doping's gains and costs fall differently on different athletes — those with better medical support, more money, or more tolerance for risk come out ahead even in a fully doping field, which gives the equilibrium a persistence a queue does not have.

d

Clarifying the model

THE MODEL #

This is not the same problem as the detection science. The neighbouring account of the athlete biological passport is about the inference: building a baseline from an athlete's own history so a deviation becomes visible without identifying a substance. That is a genuine improvement in p, and the right response to an arms race precisely because it targets the physiological effect rather than the molecule, so a new molecule does not defeat it. What this piece adds is the strategic layer above it: why p must be pushed so hard, and why gains in it get partly eroded by the adaptation they provoke.

The expected-value model is a simplification, and a flattering one. Real athletes discount future losses against present gains, overestimate their ability to avoid detection, and operate inside teams where the decision may not be fully theirs. All push toward more doping than the arithmetic predicts — so the model understates the problem.

Both sides of the empirical dispute are interested parties. Prevalence estimates come, on one side, from bodies whose legitimacy rests on the claim that testing works, and on the other from researchers and former athletes whose accounts are vivid but hard to verify. Anonymous surveys have generally produced far higher estimates than testing does, which is consistent with the sampling bias above but does not settle the size. I am not going to quote a figure; the gap between the two methods is itself the most informative observation available.

The falsification test. If the mechanism is a collective-action trap driven mainly by detection probability, then an intervention that raises p without touching penalties should reduce doping more than one that raises penalties without touching p. Long-term sample storage with retrospective re-testing is close to this experiment: it raises the effective p years after the event while leaving the sanction unchanged, and it has produced retrospective positives in cases that passed at the time. If lengthening bans had shifted behaviour while retrospective testing had not, the account here would be wrong.

e

A picture of it

THE PICTURE #
Doping arms race
Doping arms race Each point is a style of regime placed by argument, not by measurement -- read the positions relative to each other rather than as values. The horizontal axis is the term that does the work; the vertical is the one that is easy to increase and does less. The two upper-left regimes are the instructive ones: severe sanctions with weak detection catch mainly the careless, which is why that quadrant is labelled for the unlucky rather than the guilty. Moving right is expensive and provokes adaptation, which is why no point sits at the far edge. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/doping-arms-race.md","sourceIndex":1,"sourceLine":4,"sourceHash":"338d3e45b7a59971fa8e75fd4e3fc6ace56eb13dff07934ff72b950355422d14","diagramType":"quadrantChart","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":720,"height":621},"qa":{"passed":true,"findings":[]}} Deters Q1 Punishes the unlucky Q2 Nominal only Q3 Deters cheaply Q4 Passport plus storage Whereabouts plus OOC In-competition only Announced tests No programme Low detection High detection Light sanction Heavy sanction Where a testing regime sits

How to readEach point is a style of regime placed by argument, not by measurement — read the positions relative to each other rather than as values. The horizontal axis is the term that does the work; the vertical is the one that is easy to increase and does less. The two upper-left regimes are the instructive ones: severe sanctions with weak detection catch mainly the careless, which is why that quadrant is labelled for the unlucky rather than the guilty. Moving right is expensive and provokes adaptation, which is why no point sits at the far edge.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The pressure athletes describe is not weakness of will and not hypocrisy. It is the predictable output of a structure where the individually best response is the same whatever everyone else does, where the deterrent term that matters most is the one hardest to raise, and where every improvement in detection creates the incentive that erodes it. That also explains why the countermeasures that have changed behaviour are the ones that raise the chance of being caught — out-of-competition testing, whereabouts requirements, longitudinal profiling, retrospective analysis of stored samples — rather than the ones that lengthen the ban.

g

Where to go next

ONWARD #
  • How the athlete biological passport raises detection probability without needing to name a substance.
  • Why retrospective testing of stored samples changes the calculation years after a result.
  • How the same collective-action structure appears in professional cycling's team-level decisions rather than individual ones.

Nearby on the shelf

4