THIS EXPLANATION
THE ROOM
LAN·34 Language, Media & Communication 6 MIN · 7 STATIONS

Unwritten vowels in abjads

A Socratic walk-through of unwritten vowels in abjads — reasoned out one step at a time, not lectured.

abcdefgh
a

The question we started with

THE QUESTION #

Why can a script leave most of its vowels unwritten without leaving fluent readers lost?

Hebrew and Arabic newspapers are printed without most of their vowels, and fluent adults read them at ordinary speed. Try the same trick on English — th mn rd th bk — and you can decode it, slowly, with effort, and with a fair chance of getting it wrong.

The usual explanation is that the readers are simply used to it. That is a restatement rather than an explanation, and it makes a checkable prediction: if habituation were the whole story, any language could be written this way given enough practice. So the question worth asking is not "how do they cope?" but "what is it about these particular languages that makes the omission cheap?"

b

Reasoning it through

REASONING #

Begin with what a vowel is doing in an English word. In bit, bat, but, bet, bought, the consonants are constant and the vowel is the entire difference between five unrelated words. Delete it and the word's identity is destroyed, with no rule that recovers it — English has no systematic relationship between bit and bat; they simply happen to share consonants.

Now the Semitic case, where something structurally different is going on. Words there are built from a root — characteristically three consonants — carrying the lexical idea, into which a pattern of vowels is inserted to make a particular word of it. From the Arabic root k-t-b come the words for he wrote, book, writer, office; from the Hebrew root m-l-k come king, to reign, kingdom. The vowels are not idle, but notice which job they have: they mark which grammatical shape of the root you are looking at. What the word is about lives entirely in the consonants.

Sit with that, because it changes the accounting. Writing the consonants of a Semitic word is not writing part of the word and hoping. It is writing the lexeme completely and leaving the inflection unmarked. And inflection is what a reader is best placed to reconstruct, because the surrounding syntax constrains it: whether a verb is active or passive, past or present, whose it is, is largely determined by what else is in the sentence. The omitted material is precisely the recoverable material.

Does the writing system know this, or is it luck? History suggests fit rather than design. The Phoenician script that gave rise to these systems wrote consonants alone; when Greeks adopted it for a language that is not root-and-pattern — where, as in English, consonants underdetermine the word — they repurposed spare letters as vowels, and the alphabet begins there. These are not a primitive and an improved version, but two solutions matched to two morphologies.

Is the omission ever costly? Continuously, and the scripts record their own patches. Certain consonant letters came to double as vowel indicators — matres lectionis, "mothers of reading" — so long vowels are often written after all. Later, full pointing systems were invented: Hebrew niqqud, Arabic harakat, marks that fix every vowel. Those are the tell. Look at where pointing is used: children's books, learners, poetry, sacred text where the exact reading is doctrinally load-bearing, and dictionary headwords — a precise inventory of the cases where context is thin or a wrong reading is costly, exactly as the account predicts.

Here is the falsification test, and there is a historical experiment for it. If the omission works because the language supplies consonantal roots, then imposing an abjad on a non-Semitic language should be uncomfortable. Persian, Urdu, Ottoman Turkish and Malay all adopted the Arabic script; none has root-and-pattern morphology, and all lean much harder on vowel letters than Arabic does. Ottoman Turkish spelling was a standing complaint about ambiguity, and that complaint formed part of the case for the Latin alphabet in the Turkish script reform of 1928 — a recalled date, and one of the few worth stating here.

What would refute the account? Two observations. Find that fluent readers of unpointed Hebrew or Arabic stumble no more on unfamiliar foreign names and loanwords — items with no root to fall back on — than on native vocabulary, and the claim that roots do the work collapses. Or find a non-Semitic language written with a bare abjad for centuries with no expansion of vowel marking and no reported difficulty, and the morphological explanation gives way to habituation after all.

c

The analogy

THE ANALOGY #
THE FIGURE

Think of a filing system in which the folder label carries the client's name and the coloured tab carries only the year. An unlabelled-tab folder is not a mystery: the name identifies the case completely, and the year is obvious from where the folder sat and what is inside. Take away the name instead and no amount of context recovers which client it is.

WHERE IT BREAKS DOWN

filing labels are a designed convention someone chose, whereas nobody arranged Semitic morphology or the script to suit each other — the fit is historical, and it is loose enough that both languages had to add vowel marks anyway for the cases the folder-and-tab picture would say were already solved.

d

Clarifying the model

THE MODEL #

Two corrections, one of them widespread enough to be worth naming as simply false.

The false one first: abjads are routinely described as deficient alphabets, systems that "forgot" vowels or had not yet worked out how to write them. That is a prescriptive judgement wearing structural clothing — it takes the alphabet as the standard and scores everything else by distance from it. Structurally, a script is fitted to a morphology, and marking every vowel in a root-and-pattern language spends ink on the most predictable material in the word. The historical order is also the reverse of the usual story: the abjad came first, and the alphabet is what happened when the system was borrowed by speakers whose language could not use it as it stood.

The second correction is to my own simplification. I have written as though unpointed reading were effortless, and it is not quite. Studies report real costs, particularly for beginners, for words ambiguous out of context, and for anything the reader must sound out rather than recognise — and how much residual cost fluent adults pay is an open research question. The mechanism explains why the cost is survivable, not why it is zero.

It is also worth separating this from a similar-looking trick. A newspaper headline drops the, a and is because those are the most predictable items in that sentence; the deletion is contextual, and it fails when a predictable word was also holding the structure up. An abjad's omission is not contextual. It is fixed by the morphology before any sentence is written: the same material goes every time, because that material is the inflection and the identity is in the consonants.

e

A picture of it

THE PICTURE #
Unwritten vowels in abjads
Unwritten vowels in abjads Read the two boxes on the left as the two ingredients every Semitic word is made of, and the crow's-foot marks as "many words come from this one". One root yields many words, and one vowel pattern applies to many roots -- that many-to-many crossing is what makes the pattern predictable and the root not. Follow the chain to the written form and notice what it contains: only the left-hand ingredient. The final line names who supplies the missing one, which is the reader, using the grammar of the sentence rather than the page. {"generator":"mermaid-svg-renderer@3.2.1","source":"../Socrates/.diagram-cache/_src/unwritten-vowels-in-abjads.md","sourceIndex":1,"sourceLine":4,"sourceHash":"34a1aca7be1fe6603141df149388612103b0dd1df2be0b5a99041002df38ed91","diagramType":"er","layoutVariant":"source","repairedDuplicateIds":[],"motion":"entrance-with-reduced-motion-fallback","presentation":"editorial","attempt":1,"viewBox":{"x":0,"y":0,"width":967,"height":674},"qa":{"passed":true,"findings":[]}} supplies the meaning supplies the grammar is spelled as E01 ROOT string consonants k t b -- written string carries the lexical idea E02 WORD string example kataba, kitab, katib E03 PATTERN string vowels a i u -- usually unwritten string carries tense, voice, derivation E04 WRITTEN_FORM string letters the consonants only string recovered_by syntax and context

How to readRead the two boxes on the left as the two ingredients every Semitic word is made of, and the crow's-foot marks as "many words come from this one". One root yields many words, and one vowel pattern applies to many roots — that many-to-many crossing is what makes the pattern predictable and the root not. Follow the chain to the written form and notice what it contains: only the left-hand ingredient. The final line names who supplies the missing one, which is the reader, using the grammar of the sentence rather than the page.

f

What became clearer

WHAT CLEARED #
WHAT CLEARED

The vowels can go missing because in these languages they carry inflection rather than identity, and inflection is what a sentence's grammar already constrains. That is a fact about the morphology first and the script second, which is why the same omission is workable in Arabic, awkward in Ottoman Turkish, and close to unusable in English. The vowel marks these scripts do possess are not a late admission of failure but a precise map of where the mechanism runs out.

g

Where to go next

ONWARD #
  • Why the Greek adaptation of the Phoenician script produced vowel letters out of consonants Greek did not need.
  • How readers of deep orthographies such as English and Arabic recognise words differently from readers of shallow ones such as Finnish.

Nearby on the shelf

4