Border queue surges
A Socratic walk-through of border queue surges — reasoned out one step at a time, not lectured.
The question we started with
THE QUESTION #Why does a passport desk that handles the day's arrivals easily on average still produce hour-long queues?
Take an arrivals hall with more capacity than it needs. Over a full day the desks could clear half again as many passengers as turn up, and the daily figures show it: hours of slack, desks standing idle in the afternoon. Yet at seven in the morning the hall holds a thousand people and the wait is an hour.
The obvious readings — too few desks, too slow an officer — are both refused by the daily arithmetic. So something is wrong with the arithmetic itself. What does dividing a day's passengers by a day's capacity fail to capture?
Reasoning it through
REASONING #Start with the smallest possible hall: one desk, one officer. Suppose he averages forty-five seconds a passenger, and passengers arrive at an average of one a minute. He is busy three quarters of the time. Does a queue ever form?
It does, and seeing why is the whole argument. Average is not steady. Some passengers take twenty seconds; some take four minutes, because of a query or a document that will not read. Arrivals are equally irregular. So there will be moments when two people arrive during one long case, and a queue of two exists.
Now the crucial question: how does that queue disappear? Only during a gap — an interval when nobody arrives and the officer can catch up. And here is the trap. As he gets busier, the events that build a queue become more frequent and the gaps that drain one become rarer and shorter, at the same time. The two effects multiply rather than add.
That multiplication produces the shape everyone underestimates. For the simplest model in queueing theory — one server, arrivals falling randomly and independently, randomly distributed service times, utilisation written as the fraction of time the server is busy — the average number waiting is that fraction squared, divided by one minus it. At half utilisation, half a person waits. At eighty percent, about three. At ninety, about eight. At ninety-five, eighteen. At ninety-eight, close to fifty. The curve does not bend gently; it runs to a vertical wall at full utilisation. The last few percent of a server's capacity cost more to use than all the rest put together.
So a hall sized against the daily average is sized against a number that never occurs. Utilisation is not something a day has; it is something each moment has. If the day's mean puts you at forty percent you are comfortably on the flat part of the curve — but during the morning bank, when arrivals run at four times the daily mean, the instantaneous figure is above one. Beyond one the model stops applying at all, because the queue does not settle at some large value. It grows for as long as the overload lasts. The hour of waiting is not a level; it is an accumulation.
Then the last piece: why do arrivals bunch like that? Passengers do not walk to a border, they land on it, three hundred at a time, released over about a quarter of an hour. And the aircraft are correlated by design — airlines schedule banks so connections work, slots cluster in the hours travellers accept, and long-haul services from a region land at similar clock times because of time zones and curfews. The burstiness is structural, not noise. Recovery is slow too: a twenty-minute overload takes far longer than twenty minutes to clear, because the backlog drains only at the rate that already failed to keep up.
The analogy
THE ANALOGY #Think of a basin with a slow drain. Run the tap gently and the basin looks empty. Turn it up towards the drain's rate and a level appears, rising steeply the closer the two rates get. Tip in a bucket all at once — one the drain could easily have handled spread over five minutes — and the basin fills, then takes far longer than the pouring took to empty again.
A basin's level responds only to the net rate, so a perfectly steady tap below the drain's capacity never backs up at all — whereas a desk with a steady average still forms queues, because it is the irregularity of arrivals and service times, not just their mean, that creates the waiting.
Clarifying the model
THE MODEL #Three refinements matter, and one is a correction worth making explicitly.
The correction first: this is not the same problem as how to organise a queue. Whether travellers form one long line feeding many desks or a line per desk is a genuine question, and pooling helps — it stops you being trapped behind one slow case while a neighbouring desk runs free. But that fixes variance in who waits. It cannot help here. No arrangement of a line changes how fast desks process people, and when arrivals exceed processing capacity, discipline is irrelevant; the hall simply fills.
Second, the numbers above come from a deliberately simple model and should be read as a shape, not a measurement. Desks sharing a queue do better at the same utilisation than the single-server formula suggests, because one officer's idle moment absorbs another's long case. Pulling the other way, real arrivals are batched rather than independent, which is worse than the model assumes, and service times have a long tail — the referral to secondary inspection. Which effect dominates in a given terminal is empirical, not something the formula settles.
Third, what survives the simplification is the structure, and Kingman's approximation states it plainly: waiting time is roughly a variability term, times utilisation over one-minus-utilisation, times the mean service time. Three levers, and only three. Cut variability — stagger the schedule, meter passengers off the aircraft. Cut utilisation — open desks before the bank arrives rather than once the queue has formed. Cut service time — automate routine cases so officers face only exceptions. Adding staff works on the second term, but it is the least efficient lever once you are near the wall, which is why terminals that respond only by opening more desks keep having the same morning.
A picture of it
THE PICTURE #How to readEach point is the average queue the simplest single-server model predicts at that utilisation. Reading left to right, the first thirty percentage points of extra busyness cost almost nothing — the line barely lifts off the axis — while the last three, from ninety-five to ninety-eight, add thirty people on their own. That steepening answers the question: a hall planned at the left of this chart is briefly operating past its right-hand edge every morning, and beyond that edge the queue does not level off, it keeps growing until the bank of arrivals ends.
What became clearer
WHAT CLEARED #An average utilisation figure describes a hall that never exists. The hour-long queue comes from two facts colliding: waiting time rises non-linearly as a server approaches capacity, running to infinity rather than to some tolerable ceiling, and arrivals come in aircraft-sized batches that push instantaneous demand far above the daily mean. A desk comfortable on the day's arithmetic can therefore be hopelessly overloaded for forty minutes — and then spend considerably longer recovering.
Where to go next
ONWARD #- Why hospitals deliberately run wards below full occupancy, and what the same curve implies about the cost of "efficiency".
- How metering — holding aircraft on stand, or releasing passengers in waves — trades a small certain delay for a large uncertain one.
Key terms
TERMS #| Term | What it means |
|---|---|
| Utilisation | the fraction of time a server is busy; the ratio of arrival rate to service rate. |
| Burstiness | the tendency of arrivals to cluster in time rather than spread evenly, here imposed by flight schedules. |
| Kingman's formula | an approximation giving mean waiting time as a variability term times a utilisation term times the mean service time. |
Every term the collection defines is gathered in the glossary.