THIS EXPLANATION
THE ROOM
LAN·28 Language, Media & Communication 6 MIN · 8 STATIONS

Telephone bandwidth

A Socratic walk-through of telephone bandwidth — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why does a voice stay intelligible on a line that removes the very pitch you hear it at?

The classic telephone channel passes roughly 300 Hz to 3400 Hz and discards everything outside. A typical adult male voice has a fundamental frequency somewhere near 100 to 120 Hz. So the line throws away the very component you would name if asked what note the man is speaking on — and yet he does not sound as though he has been transposed, does not sound genderless, and is understood without effort.

Two things are worth separating here before we start. One is why the pitch survives its own deletion. The other is why the words survive, given that a good deal of speech energy also lies outside the band. They have different answers, and mixing them is where most accounts go wrong.

b

Reasoning it through

REASONING #

Take the pitch first. What the vocal folds produce is not a single frequency but a periodic train of puffs. A periodic waveform repeating 100 times a second contains energy at 100, 200, 300, 400 Hz and upward — harmonics spaced by the repetition rate. Remove the 100 Hz component and the spacing of everything that remains is untouched: 200, 300, 400 are still 100 apart, and the whole signal still repeats 100 times a second. The period never depended on the presence of the first harmonic.

So the question becomes what the ear is actually measuring. If it were reading off the lowest present component, a filtered voice would jump an octave — 200 Hz instead of 100. It does not. The auditory system recovers a pitch corresponding to the pattern's periodicity even when nothing is there at that frequency: the missing fundamental, also called residue pitch.

Can we test that against the obvious rival explanation? The rival is that the ear's own nonlinearity regenerates the missing component as a distortion product, so there really is 100 Hz there after all. Schouten's pitch-shift experiment settles it. Take harmonics at 1800, 2000 and 2200 Hz — spacing 200, so periodicity 200 Hz — and shift them all up by 40 Hz, to 1840, 2040, 2240. The spacing is unchanged at 200. A distortion-product account predicts the difference tone stays at 200 and the pitch does not move. What listeners report is a small but definite rise. That is the refuting observation for the distortion story, and evidence that the system is fitting a pattern to the whole set of components rather than extracting one.

Now the words. Vowels are distinguished not by the fundamental but by the resonances of the vocal tract — the formants. The first two formants, which do most of the work of telling one vowel from another, sit in a range from a few hundred hertz up to around 2.5 kHz for adult speakers, comfortably inside the passband. This is why the channel is drawn where it is: the band was chosen around the formants, not around the voice's pitch.

What does get lost? Chiefly the high-frequency noise of the voiceless fricatives. The energy distinguishing /s/ from /f/ lies substantially above 3.4 kHz, which is precisely why "S for Sugar, F for Freddie" spelling alphabets exist — an operational fact that would be pointless if the band were transparent.

The refuting observation for the whole account: if intelligibility lived in the low end, then low-pass filtering speech and high-pass filtering it should not be comparably damaging. The classical measurements — French and Steinberg's articulation-index work, recalled rather than recomputed here — found a crossover in the neighbourhood of 1.8 to 2 kHz, above and below which the two halves of the spectrum contribute about equally to intelligibility. A voice cut off at 2 kHz and a voice with everything below 2 kHz removed are each about half-intelligible. That result is fatal to any picture in which the low fundamental carries the message.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a picket fence seen through a narrow window. You cannot see the ends of the fence, or the gatepost it starts from, but you can see enough pickets to read off the spacing exactly — and from the spacing you know where every picket you cannot see must be, including the first one. Widening the window tells you nothing new about the spacing.

WHERE IT BREAKS DOWN

the fence's spacing is the only thing you were after, whereas in speech the spacing (pitch) and the pattern of loud and quiet regions (formants) are two independent pieces of information carried by the same components, and the window keeps one of them far better than the other.

d

Clarifying the model

THE MODEL #

The structural facts and the engineering choice need to be kept apart, because the band is often discussed as though it were a discovery about speech. It is not. It is an economic decision. Analogue carrier systems stacked telephone channels 4 kHz apart, and the digital successor sampled at 8 kHz — Nyquist then caps the representable content at 4 kHz, and the roll-off and guard band bring the usable top down to 3.4 kHz. Eight thousand samples a second at eight bits each gives exactly 64,000 bits per second, the familiar channel rate. Every one of those numbers is a choice about how many conversations fit down a cable, constrained by what would still be understood. Speech did not ask for 300 Hz to 3400 Hz; accountancy did.

Two corrections follow. First, "the telephone removes the fundamental so pitch is inferred" overstates it: the fundamental is attenuated by the filter's skirt rather than surgically removed, and for many speakers — most adult women, and children — it lies inside the band. The inference mechanism is real and demonstrable, but it is not doing rescue work on every call.

Second, the band's cost is real even though intelligibility survives. Narrowband speech is measurably worse for confusable letters and for unfamiliar names, and speaker recognition suffers. Wideband telephony extends the range, roughly 50 Hz to 7 kHz, and the audible improvement is mostly in exactly the places this account predicts: sibilants, and the naturalness of the low end.

e

A picture of it

THE PICTURE #
Telephone bandwidth
Telephone bandwidth Each point is one part of the speech signal. Read across for whether the passband keeps it, and up for how much the words depend on it. The top-right quadrant is why the call works at all: the two formants that identify vowels are both kept and both essential. The top-left quadrant is the channel's one genuine loss -- sibilant noise above the cutoff, which is needed and is not there, hence the spelling alphabets. The bottom-left holds the male fundamental: removed, and unmissed, because the harmonics still inside the band fix the periodicity for you. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/telephone-bandwidth.md","sourceIndex":1,"sourceLine":4,"sourceHash":"171e61820190f6f10f9b8b0242e5913fadab2686a3983fb322ee81fc43aa5a3f","diagramType":"quadrantChart","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":720,"height":621},"qa":{"passed":true,"findings":[]}} Kept and needed Q1 Lost and needed Q2 Lost, minor Q3 Kept, minor Q4 Voice warmth Sibilant noise Second formant First formant Low harmonics Male fundamental Below or above band Inside 300-3400 Hz Little intelligibility Carries the words Speech components against the telephone passband

How to readEach point is one part of the speech signal. Read across for whether the passband keeps it, and up for how much the words depend on it. The top-right quadrant is why the call works at all: the two formants that identify vowels are both kept and both essential. The top-left quadrant is the channel's one genuine loss — sibilant noise above the cutoff, which is needed and is not there, hence the spelling alphabets. The bottom-left holds the male fundamental: removed, and unmissed, because the harmonics still inside the band fix the periodicity for you.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

Two independent facts make the narrowband call work, and neither is the one people reach for. The pitch survives because pitch was never carried by a single component — it is the spacing of a harmonic series, and the surviving harmonics preserve the spacing exactly. The words survive because vowel identity is carried by tract resonances that happen to lie inside the band, and the band was drawn around them deliberately. What is genuinely lost is the high fricative noise, and the workarounds telephony invented for that loss are the visible evidence.

g

Where to go next

ONWARD #
  • How low-bitrate codecs go further still, transmitting the filter's shape as a handful of coefficients and resynthesising the source entirely.
  • Why the same missing-fundamental effect makes a small loudspeaker sound as though it has bass it cannot physically produce.
h

Key terms

TERMS #
TermWhat it means
Fundamental frequencythe repetition rate of the vocal folds, heard as the voice's pitch.
Formanta resonance of the vocal tract; the first two largely determine which vowel is heard.
Missing fundamentalthe pitch heard at a harmonic series' spacing when no energy is present at that frequency.
Nyquist limitthe highest frequency representable at a given sampling rate, equal to half of it.

Every term the collection defines is gathered in the glossary.

Nearby on the shelf

4