THIS EXPLANATION
THE ROOM
LAN·26 Language, Media & Communication 6 MIN · 8 STATIONS

Simultaneous interpreting

A Socratic walk-through of simultaneous interpreting — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does a simultaneous interpreter start speaking before the sentence being translated has finished?

Watch an interpreter in a booth and you notice something that looks like recklessness. She begins rendering a sentence she has not heard the end of. In a language whose verb comes last, she has committed to what the speaker did before the speaker has said what he did. Why accept that risk, when waiting a few seconds would remove it? The obvious answer — that she is simply fast — cannot be right, because speed does not create information that has not arrived yet. Something else is forcing her hand.

b

Reasoning it through

REASONING #

Start by taking the cautious strategy seriously and asking what it costs. Suppose she waits for the full sentence, then renders it. What is she doing during the wait? Holding the sentence — in a memory that leaks. Unrehearsed verbal material fades in seconds, and rehearsing it is not free, because the speaker has not stopped and the incoming stream needs the same machinery. The wait buys certainty about the end of the sentence at the price of losing the beginning, and of falling behind a speaker already into the next clause.

Now the impatient strategy. She begins almost at once, the memory burden nearly vanishes — but she has to guess, committing to a syntactic frame, and sometimes a content word, that the rest of the sentence may contradict.

So this is not a question with a right answer; it is a trade-off with a control variable. The variable has a name in the field: the ear-voice span, or décalage — the lag between hearing something and saying it. Push it short and memory load falls while commitment risk rises; push it long and the reverse. An interpreter is not choosing between waiting and not waiting; she is continuously tuning a lag. Values commonly reported sit in the range of roughly two to four seconds, though the spread is wide and I am reporting that range rather than deriving it.

Does that make a prediction we could check? A strong one. If interpreting were reactive transcoding — hear a chunk, convert it, say it — the lag would be a processing latency, roughly constant, indifferent to what is being said. If it is a managed trade-off, the lag should move with the structure of the source, stretching where the target language cannot be built without information the source has not yet delivered.

That is the German case. In a subordinate clause German holds its verb to the end; English cannot build a clause without one. So an interpreter working from German into English must either wait for the verb — lengthening the lag, and the memory load with it — or supply one before it arrives. Studies of ear-voice span do find it varying with source structure in this way, and, more strikingly, find interpreters producing verbs their speaker has not yet produced. That is anticipation, and it is not a flourish; it is what keeps the lag survivable.

Anticipation of what kind, though? Not clairvoyance. By the time a German clause approaches its verb, most of the sentence's constraints are on the table: the subject, the objects, the case marking, the topic, the fact that the speaker has said something similar three times this morning. The set of verbs compatible with all of that is often small. This is not translating but inference under uncertainty, using the redundancy of the discourse to license an early commitment — which is why interpreters demand the text and the slides in advance. Those are not conveniences; they are the priors that make guessing cheap.

And when the guess is wrong? A repair — a hedge, a recast, a swallowed clause. Which gives a second prediction: anticipation errors should exist, be visible as repairs, and cluster where the source is least predictable. They do.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of catching a ball rather than receiving one. A fielder does not watch the ball land and then move to it — by then it is past. He starts running on the first fraction of the trajectory, on an estimate, and updates as the ball comes on. Standing still until the outcome is certain is the one strategy guaranteed to fail.

WHERE IT BREAKS DOWN

a ball's path is governed by physics, so early evidence constrains the endpoint lawfully, whereas a speaker can change his mind mid-sentence — and unlike the fielder, who simply misses, an interpreter who commits wrongly has already said something in someone else's name and must unsay it while still listening.

d

Clarifying the model

THE MODEL #

Three refinements.

The first corrects the natural picture of the booth. It is tempting to imagine a queue: words in, translate, words out. But the interpreter is listening, holding, producing and monitoring her own output at the same time, and those tasks contend for the same resources — which is why the lag is a load-management decision, not a stylistic one. Danica Seleskovitch's influential claim was that what is held in the gap is not the words but the sense stripped of its form.

The second is about expertise, where the popular version outruns the evidence. It is often said that interpreters have exceptional working memory. Studies comparing interpreters with matched bilinguals on general memory span have given mixed results, and the more defensible reading is that expertise shows up as task-specific skill — better anticipation, better chunking, more efficient monitoring — rather than a larger general-purpose buffer. Treat that as contested rather than settled.

The third is the limit of my own account. I have described the lag as tuned by a single trade-off between memory and risk, and it is plainly not only that: delivery rate, the need to sound fluent, the density of numbers and names, and simple fatigue all pull on it. What would refute the core claim is easy to state — measure ear-voice span across many language pairs and find it constant, indifferent to how much the target language needs information the source has withheld. Find also that interpreters never produce content in advance of the source, or do so no better than chance, and the anticipation mechanism goes with it.

e

A picture of it

THE PICTURE #
Simultaneous interpreting
000204060810Subject and objects arrive Listening, nothing said yet Renders subject and objects Verb guessed from context Verb finally arrives Next clause begins Repair if the guess was wrong Already listening ahead SpeakerInterpreterThe lag across one verb-final sentence

How to readRead left to right as about eleven seconds of one clause; the two sections are the speaker's utterance and the interpreter's, and the horizontal offset between them is the ear-voice span. The interpreter's rendering of the subject and objects begins while the speaker is still delivering them, and the bar marked as guessing the verb starts before the speaker's verb bar does — that overlap is the anticipation. The repair bar is conditional rather than routine, the price paid when a guess is contradicted; the last bar shows why waiting is not an option, because the next clause is already arriving.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The interpreter starts early not out of confidence but necessity: waiting does not preserve the sentence, it destroys the front of it while the speaker keeps going. The real object of skill is a lag — long enough to have heard something worth saying, short enough that memory holds and the speaker is not lost — and where a language withholds what the target needs, the lag is bridged by inference rather than patience. Simultaneous interpreting is less a feat of translation than of prediction under load, which is why the measurable thing is a span of seconds.

g

Where to go next

ONWARD #
  • Why interpreters work in pairs and rotate on a fixed interval, and what that says about where the cost falls.
h

Key terms

TERMS #
TermWhat it means
Ear-voice span (décalage)the lag between hearing a stretch of source speech and producing its rendering; the field's main measurable for this trade-off.
Anticipationproducing an element of the target sentence before the corresponding source element has been uttered, licensed by context rather than signal.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4