Failure after repair
A Socratic walk-through of failure after repair — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why is a machine often most likely to break just after it has been serviced?
The boiler ran all winter without complaint. It was serviced on Tuesday and failed on Sunday. Everyone who has owned machinery has this story, and it is usually told one of two ways: the engineer broke it, or the mind is playing tricks and any week is as likely as any other.
Both hold something true. Neither holds the part that matters, which is that opening a working machine is itself an operation with a failure rate — not a complaint about workmanship but a structural feature of intervening in anything.
Reasoning it through
REASONING #Grant the sceptics their share first, because a good deal of the pattern really is bookkeeping. Attention concentrates after a service, and any fault now has an obvious thing to be blamed on. Worse, machines are frequently serviced because somebody already noticed something, so the serviced population is pre-selected for being suspect.
But the pattern survives where bookkeeping cannot explain it — on fleets under scheduled maintenance with no preceding complaint, where every failure is logged whether or not anyone was watching. So ask a sharper question: what changes at the moment the cover comes off?
Three clocks start, and they are genuinely different from one another.
The first is the one this collection already covers under the bathtub curve. A component that has run for years has sat an examination its replacement has not sat; swap it and you move from a proven item back into the unfiltered population, where a small share were born defective. That is real — but notice how narrow it is. It applies only to parts actually replaced, and only to their infant-mortality share.
The second clock usually dominates, and it is not about parts at all. A service is a manufacturing operation performed in the worst available conditions: in place rather than on a bench, under time pressure, often in poor light, by one person, with no test rig at the end. So it generates the errors manufacturing generates — a fastener at the wrong torque, a seal pinched on reassembly, a connector pushed home but not latched, swarf left in a passage, the wrong grade of oil. Their clock starts at the service, and most surface within the first few duty cycles, because that is when the assembly is first loaded, heated and cycled.
The third is the disturbance of everything that was not the target. A settled machine is full of parts that have bedded in, seized gently, or corroded into place. Do the job and you move all of them: a bolt never undone before shears, a brittle clip breaks, an adjacent hose cracks when flexed for the first time in a decade. You repaired one thing and perturbed twenty.
Is this real mechanism or a story that fits? Two predictions separate it. The effect should scale with how invasive the work was — a filter change and a full teardown should not carry the same post-service hazard — and post-service failures should be disproportionately assembly and human-error modes rather than wear modes. Both hold, and the second is what failure analysis finds when it opens the units up.
The strongest evidence that this matters is that it changed doctrine. The reliability-centred maintenance work done for the airline industry in the 1960s and 1970s — Nowlan and Heap's report for United Airlines is the standard reference — sorted components by how their failure rate behaved with age and found that most showed no wear-out age at all. For those items scheduled overhaul had nothing to prevent, so it could only add fresh opportunities for error, and practice shifted from fixed-interval stripping toward condition monitoring.
That conclusion is routinely over-read. The finding was about which items, not that maintenance is harmful. Where wear-out is genuine and the consequence severe — brake friction material, tyres, lubrication, fatigue-critical structure — interval-based replacement remains correct, because there the wear curve really does rise.
The analogy
THE ANALOGY #Think of an old sash window that has swollen shut over the years and stopped rattling. Take it apart to fix the catch and you get the catch fixed, along with a window that no longer sits in its frame, a pane whose putty you have just disturbed, and a cord that was fine until it was moved.
the window announces its new problems immediately, whereas a mis-torqued bolt or a pinched seal gives no sign until it is loaded and heated — which is why the machine version arrives days later and feels like a mystery rather than a consequence.
Clarifying the model
THE MODEL #The three clocks call for different remedies. Infant mortality in new parts is a population effect, answered by burn-in, supplier quality, and not replacing what is working. Human error in assembly is answered by procedure, torque records and post-maintenance testing. Disturbance damage is answered by intervening less often and less deeply. Lumping them together produces the crude conclusion that servicing is bad, which is not what any of them says.
That is also the honest reconciliation with the neighbouring account of the bathtub curve. That file is right that swapping a proven part for a fresh one resets a selection effect, and names this pattern as a consequence. What is added here is that the reset is usually the smaller half: the sharper post-service hazard tends to come from the intervention itself, in modes population statistics do not contain, which is why the fix is procedural rather than statistical.
One caution about the question's own wording. "Most likely to break" means most likely relative to the week before the service, not relative to every moment in the machine's life. A neglected machine deep into wear-out is riskier still. The claim is that the servicing interval carries a hazard of its own, not that servicing is worse than neglect.
A picture of it
THE PICTURE #How to readStart at Settled, the state a machine spends most of its life in and the safest one it occupies. Every path to Broken is labelled with its reason, and the three arrows leaving Refitted are the three clocks in the text — human error, a fresh part's own infant mortality, and collateral disturbance — all running only in the window just after the cover goes back on. The single arrow from Settled to Broken is ordinary wear-out, the only failure the service was there to prevent. Note the back-edge from Broken to Stripped: a repair returns the machine to the risky window, which is why one repaired twice in quick succession is not unlucky so much as re-exposed.
What became clearer
WHAT CLEARED #Maintenance is not free, and part of its price is paid in reliability rather than money. Opening a settled machine resets a selection effect on any part replaced, introduces the error rate of an assembly job done in bad conditions, and disturbs neighbours that had stopped moving years ago — three separate hazards, all with their clock starting at handover. That is why modern practice tries to intervene on evidence rather than on the calendar, and why "we serviced it and then it broke" deserves to be taken seriously rather than laughed off.
Where to go next
ONWARD #- How condition monitoring decides when there is enough evidence to intervene at all.
- Why post-maintenance testing exists, and what it can and cannot catch before handover.
Key terms
TERMS #| Term | What it means |
|---|---|
| Maintenance-induced failure | a failure whose cause was introduced by the maintenance work rather than by use. |
| Infant mortality | the elevated early failure rate of a fresh population, driven by the small share built with a latent defect. |
| Reliability-centred maintenance | choosing tasks from how each item actually fails and how bad the consequence is, rather than from a calendar. |
| Wear-out | a failure mode whose rate genuinely rises with age or use, and the only kind scheduled replacement can pre-empt. |
Every term the collection defines is gathered in the glossary.