Etymologically

Learn etymology · Guide 3 of 3

How proto-languages are reconstructed

Some entries in this dictionary trace a word back to forms like Proto-Indo-European *ph₂tḗr, “father”. Nobody ever wrote that word down: the language it belongs to was spoken some six thousand years ago, long before writing reached its speakers. So how can anyone claim to know it? The answer is a technique called the comparative method, and this guide walks through it.

What is a proto-language?

Languages change, and when the people who speak one language spread out and lose touch, their speech changes in different directions until it becomes several different languages. The Latin spoken across the Roman Empire became Italian, Spanish, Portuguese, French, Romanian and the other Romance languages. Languages that descend from one ancestor make up a language family, and the ancestor is the family’s proto-language.

Sometimes the ancestor is well recorded, as Latin is. Usually it is not. English, German, Dutch, Swedish and Gothic all come from a language that left no texts, called Proto-Germanic. Proto-Germanic, Latin, Greek, Sanskrit, the Celtic, Slavic, Baltic, Iranian and Armenian languages, and several more, all descend in turn from a much older unwritten ancestor, Proto-Indo-European. It is from this family that most of the languages of Europe, and many of those of Iran and India, come.

Reconstructed words are written with an asterisk, as in *ph₂tḗr, to show that they are not recorded but inferred.

Step 1: gather words that may be related

The method begins with lists of words of similar meaning in the languages being compared. Linguists prefer basic vocabulary: numbers, family members, body parts, common animals and everyday actions. These are the words least likely to have been borrowed from a neighbour, which matters because a borrowed word tells us about contact between languages, not about their common ancestry.

Step 2: line up the sounds

The next step is to compare the words sound by sound. Here are some words in four Romance languages. Pretend for the moment that we don’t know Latin; the last column is hidden from us.

Words beginning with c- in four Romance languages (and, for checking later, Latin)
MeaningItalianSpanishPortugueseFrenchLatin
goatcapracabracabrachèvrecapra
fieldcampocampocampochampcampus
to singcantarecantarcantarchantercantāre
dearcarocarocarochercārus
bodycorpocuerpocorpocorpscorpus
shortcortocortocurtocourtcurtus

Each column of matching sounds is called a correspondence set. At the start of these words, Italian, Spanish and Portuguese have a k sound (spelled c), while French has ch, pronounced “sh”, in the first four words but c in the last two. That looks messy, but it isn’t: French has ch where the other languages have an a after the consonant, and keeps c where they have o or u. The difference is caused by the neighbouring sound, so both patterns are one correspondence: k in Italian, Spanish and Portuguese matches French “sh” before a and French k elsewhere. Correspondences must be regular, recurring in word after word; one or two matches could be chance.

The middle of the word goat gives another set: Italian p, Spanish and Portuguese b, French v. Other words show the same pattern, such as “to know”: Italian sapere, Spanish and Portuguese saber, French savoir.

Step 3: work out the ancestral sound

For each correspondence set, the linguist asks: what single ancestral sound could have turned into all of these? Three principles guide the choice.

  • Some changes are common, their reverse is rare. Across the world’s languages, k often turns into “ch” or “sh” before vowels like a, e and i, which are pronounced further forward in the mouth. “Sh” turning into k is much rarer. So the ancestor more likely had k, and French changed it. Likewise, sounds between vowels tend to weaken: p becomes b, and b becomes v, not the other way round. So the middle of goat was probably p.
  • The fewest changes win. An ancestral k needs one change, in French only. An ancestral “sh” would need the same unusual change to have happened separately in Italian, Spanish and Portuguese.
  • Majority is only a tie-breaker. Most languages having a sound is weak evidence by itself, since several languages can make the same change independently.

Putting the sounds back together, we reconstruct something like *kapra for “goat”, *kantare for “to sing”, and so on. Now uncover the Latin column: capra, campus, cantāre. Latin c was always pronounced k. The reconstruction is right. Even the intermediate stage turns up in writing: Old French spelled these words with ch because it was then pronounced “tch”, halfway between k and “sh”.

A reconstruction recovers what people said, not what they wrote. The Romance words for “horse” (Italian cavallo, Spanish caballo, French cheval, Portuguese cavalo) point to *caballus. But the word for horse in Classical Latin literature was equus; caballus was a colloquial word for a workhorse. The comparative method found the everyday spoken Latin, the language the Romance languages actually descend from, rather than the polished Latin of books.

The same method, without the answer key

Now take the same procedure back further, to languages whose common ancestor was never written down. Here are some words in Sanskrit, Greek, Latin, Gothic and Old Irish, the oldest well-recorded members of five branches of the Indo-European family.

Some basic words in five branches of Indo-European
MeaningSanskritGreekLatinGothicOld Irish
fatherpitár-patḗrpaterfadarathair
footpád-poús, pod-pēs, ped-fōtus—
threetráyastreîstrēsþreistrí
tendáśadékadecemtaíhundeich

The correspondences are clear once you look for them:

  • Sanskrit, Greek and Latin p match Gothic f and, in Old Irish, nothing at all. The Celtic languages lost p in general: Old Irish íasc, “fish”, matches Latin piscis. The ancestor had *p.
  • Sanskrit, Greek, Latin and Old Irish t and d match Gothic þ (“th”) and t. These are the changes of Grimm’s law, described in the guide to sound laws, which also explains why Gothic has d in fadar.
  • In “ten”, Greek and Latin k match Sanskrit ś (“sh”) and Gothic h. The ancestor had a k-like sound, written *ḱ, which became “sh” or “s” in Sanskrit and several other branches.

Working through every sound of every word this way, and checking each correspondence against hundreds of other words, gives reconstructions like *tréyes, “three”, and *déḱm̥, “ten”. The table also holds a puzzle: where Greek and Latin have a in “father”, Sanskrit has i. Elsewhere, Greek and Latin a regularly match Sanskrit a. A different correspondence set implies a different ancestral sound, and so the reconstruction *ph₂tḗr contains a sound written h₂. The story of that sound is the best evidence that the method works.

How do we know it works?

A reconstruction is a prediction about something we can’t observe. Like any scientific prediction, it can be tested when new evidence comes to light, and it has been, several times.

Romance and Latin

As above: reconstructing from the Romance languages gives back Latin, and specifically the spoken Latin that the written record only hints at.

The laryngeals and Hittite

In 1879 the young Swiss linguist Ferdinand de Saussure, studying the patterns of vowels in Indo-European words, argued that Proto-Indo-European must have had sounds that had vanished from every language then known, leaving only traces in the vowels around them. For decades most scholars were unconvinced. Then, in 1915, the Czech scholar Bedřich Hrozný deciphered Hittite, an Indo-European language written on clay tablets in what is now Turkey more than 3,000 years ago. In 1927 the Polish linguist Jerzy Kuryłowicz showed that Hittite has a sound, written ḫ, in just the places where Saussure’s lost sounds were predicted: Hittite ḫant-, “front”, beside Latin ante and Greek antí, “before”. These sounds, now called laryngeals and written h₁, h₂ and h₃, are part of every modern reconstruction, including *ph₂tḗr.

Bloomfield and Cree

In the 1920s the American linguist Leonard Bloomfield applied the method to four Algonquian languages of North America. One set of correspondences, found in a single word, did not fit any of the others, so he reconstructed a separate ancestral sound for it. Later he found that Swampy Cree, a language he had not used, has a distinct consonant cluster in exactly that position. The method had predicted a sound in a language its author had not studied.

More than words

The comparative method reconstructs grammar as well as vocabulary: the endings of nouns and verbs, and the way sentences were put together. And the reconstructed vocabulary says something about the speakers. Proto-Indo-European had words for snow (*snóygʷʰos), the horse (*h₁éḱwos, Latin equus, Sanskrit áśva-), and the wheel (*kʷékʷlos, the source of Greek kýklos, Sanskrit cakrá- and English wheel). It also had a sky god, *dyḗws ph₂tḗr, “sky father”, who survives as Greek Zeus patḗr, Latin Jupiter, and Sanskrit Dyaúṣ pitā́. Since wheeled vehicles were invented only in the fourth millennium BC, a shared word for the wheel is one reason most linguists place Proto-Indo-European around that time or later.

What reconstruction can’t do

  • It gives a model, not a recording. We know that h₁, h₂ and h₃ were different sounds, but not exactly how they were pronounced. In 1868 August Schleicher wrote a short fable in his reconstruction of Proto-Indo-European; scholars have rewritten it many times since, and each version looks very different from the last.
  • It smooths over variety. A real language has dialects, and a reconstruction usually flattens them into one.
  • Meanings are harder than sounds. Meaning change follows no laws, so a reconstructed meaning is often just the common ground between the meanings of the descendants.
  • It only reaches so far back. After many thousands of years, sound change and the replacement of words wipe out the evidence. Most linguists think the method stops working somewhere beyond eight to ten thousand years, which is why proposals to link Indo-European with other families into still larger groupings remain unproven.

Reading a reconstruction

A few symbols you will meet in Proto-Indo-European forms:

  • *: the form is reconstructed, not recorded.
  • h₁, h₂, h₃: the laryngeals, three lost sounds made at the back of the mouth or in the throat. h₂ coloured a neighbouring e to a, and h₃ to o.
  • ḱ, ǵ: k and g made further forward in the mouth; kʷ, gʷ: k and g with rounded lips, as in English queen.
  • bʰ, dʰ, gʰ: “breathy” consonants, pronounced with a puff of breath.
  • m̥, n̥, r̥, l̥: the consonants used as vowels, the way the n in button forms a syllable of its own.
  • Accent marks (é, ḗ) show where the stress fell; a macron (ē) marks a long vowel.