Near misses
A Socratic walk-through of near misses — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why do serious accidents usually follow a long trail of small ones that nobody acted on?
Read almost any accident investigation and the same sentence appears: this had happened before, in smaller form, and had been seen. Not once — often many times, documented, sometimes discussed in meetings. The organisation was not blind, so the puzzle is not why nobody noticed. It is why noticing did not produce action.
That is stranger than it looks, because a near miss is unusually good information: the whole causal chain ran, and only the last link failed to close. What could make an organisation systematically discard the best warning it will ever get for free?
Reasoning it through
REASONING #Start by asking what an organisation actually looks at when deciding whether something was a problem. Almost always: the outcome. Someone was hurt, so we investigate; nobody was hurt, so we do not. But the outcome is the one part of the event that luck had a hand in — the same sequence, run again, might have ended differently. Judging by outcome means judging by the noisiest available signal and ignoring the clean one.
Now push harder, because there is a subtler trap underneath. How does a near miss feel from inside? Not like a warning. It feels like a success. The valve stuck, and the operator caught it. The O-ring eroded, and the joint held anyway. Each is a story in which the system worked — and on that story it is entirely reasonable to conclude the defences are sound. Research on how people interpret near misses, notably by Tinsley, Dillon and colleagues, finds this directly: observing one tends to make people more willing to accept the risk afterwards, because they encode it as evidence of resilience rather than of a barrier nearly breached.
Follow that one step further and you get the mechanism that matters. The anomaly is observed, nothing bad results, so it is accepted — and once accepted it is no longer an anomaly. The expectation quietly moves to include it, and next time it is not a deviation at all; it is what happens. Diane Vaughan called this normalization of deviance in her study of the Challenger launch decision, and the uncomfortable part of her finding is that the people involved were not breaking rules. They were following rules, inside a standard that had shifted a step at a time, each step defensible on the evidence then available. Erosion of the booster joint seals had been seen on earlier flights, outside the original design expectation, and was reclassified over time as an acceptable, understood condition.
Notice what makes this so hard to catch from inside. There is no moment where anyone decides to lower the bar. Each decision compares the current event to the current standard, and passes. Only comparing today's standard to the original one makes the drift visible, and nothing in the routine of the work performs that comparison.
Add the ordinary organisational reasons and the discarding is overdetermined. Reporting costs the reporter time and sometimes exposure to blame, while the benefit is diffuse; nobody is credited for the accident that did not happen. And under production pressure, the anomaly that stops nothing gets an explanation rather than an investigation.
The analogy
THE ANALOGY #Think of a road crossing where a hedge blocks the view, so drivers routinely brake hard and stop in time. From the council's records the crossing is safe — no incidents. From the drivers' side, each near thing confirms that braking works. The hedge is never cut, because the only evidence that would fund cutting it is the collision that has not yet occurred, and every day without one is filed as proof it need not be.
The hedge is a single, fixed, visible cause that could be removed in an afternoon, whereas the conditions behind most organisational accidents are distributed across many people's reasonable decisions and have no equivalent object to point at — which is why "just fix the obvious hazard" is a poor model of the real problem.
Clarifying the model
THE MODEL #A caution about the most-quoted piece of this subject. Heinrich's triangle, from 1931, is usually stated as a fixed ratio — one major injury for every twenty-nine minor injuries and three hundred no-injury incidents. Those numbers are repeated everywhere and are poorly supported: the underlying data were never published in a form anyone could check, the ratios do not hold stably across industries, and the framing has done real harm by conflating personal injury with process safety. Organisations have posted excellent slip-and-trip statistics while their catastrophic risk went unmanaged, and investigations after major process accidents have made exactly that point. If the triangle is used at all, use it for the idea that small events outnumber large ones — not as a rate you can multiply.
What survives without the disputed arithmetic is a claim about information rather than frequency. A near miss and an accident can share the same causal chain, separated only by a last link that is often luck. Near misses are therefore samples from the same process that produces disasters, arriving far more often and at no cost, and discarding them means choosing to learn only from the rare, expensive instances.
One more refinement, because the obvious remedy is wrong. "Investigate every anomaly" is not achievable — complex operations generate anomalies constantly, and an organisation that halted for each would not operate. The real task is discrimination: telling the anomalies that indicate a weakened barrier from the noise of ordinary variation. That is genuinely difficult, it is where the expertise lives, and it is why the useful institutional moves are a cheap non-punitive reporting channel, plus a habit of comparing today's accepted conditions against the original design intent rather than against last month's.
A picture of it
THE PICTURE #How to readStart at the top and follow a single deviation. The branch out of the anomaly is the whole subject: taken as a signal it leads to repair and the original standard survives, while taken as a good outcome it returns to the expected state — but a wider version of it, since the deviation now counts as normal. That back-edge is a loop that runs many times, each pass costing nothing and each widening the target, until the branch to loss is finally taken from a position far from where the system started.
What became clearer
WHAT CLEARED #Serious accidents follow a trail of small ones because the small ones are read by their outcomes, and their outcomes are good. Each near miss arrives disguised as evidence that the defences work, gets accepted, and in being accepted redefines what normal means — so the drift is invisible to anyone comparing today's event to today's standard. The information was never missing. It was free, plentiful, and thrown away, and the trail exists precisely because nothing about a near miss forces anyone to look at it.
Where to go next
ONWARD #- Why confidential, blame-free reporting systems in aviation changed the flow of this information, and what they cost.
- How high-reliability organisations deliberately treat quiet periods as a reason for suspicion rather than reassurance.
Key terms
TERMS #| Term | What it means |
|---|---|
| Near miss | an event in which the causal sequence of an accident occurred but the harmful outcome did not. |
| Normalization of deviance | the gradual reclassification of an out-of-specification condition as acceptable, after it repeatedly produces no harm. |
| Outcome bias | judging the quality of a decision or condition by how it happened to turn out rather than by the risk it carried. |
| Process safety | the management of low-frequency, high-consequence hazards, distinct from the personal-injury statistics often used as a proxy for it. |
Every term the collection defines is gathered in the glossary.