The Science · Memory and learning

How a Language Gets Into You

Two thousand words covers eighty-five percent of any English text. A word needs around a dozen encounters in context before it is truly yours. What the applied linguists figured out.

May 29, 2026 · 9 min read

The thing nobody tells you when you start learning a language is that grammar is not the bottleneck. Grammar is finite. You can learn the rough shape of Spanish verb conjugation in a weekend and the polite contours of Japanese in a week. Vocabulary is the work. It is the long, quiet, multi-year accumulation of tens of thousands of small associations between sounds and meanings, and it is the part of the project that decides whether you will ever be able to read a novel, watch a film without subtitles, or carry on a conversation that drifts off the topic you prepared for.

What is strange about vocabulary is that it is wildly uneven. A small handful of words do almost all of the work in any language. Paul Nation, the New Zealand applied linguist who has spent his career counting these things, has shown that the most frequent two thousand word families cover roughly eighty-five percent of running text in English. Five thousand gets you to about ninety-five percent. The remaining tens of thousands of words show up rarely, in narrow contexts, and yet they are the difference between competent and fluent. Language is a power-law distribution. Most of the mass is in the head; the long tail is where the texture lives.

Image slot

The power-law shape of word frequency. A small core does almost all the work.

Image prompt: Editorial data visualization. A long-tail distribution curve plotted on warm cream paper, x-axis labeled 'word rank', y-axis labeled 'frequency'. The first 2000 ranks are shaded a soft amber and labeled '~85% of all running text'. The next 3000 are shaded a paler tan and labeled '~95%'. The tail stretches off to the right, very low, labeled '5000 to 50,000+: the texture of a language'. Hand-drawn axis ticks, serif labels, paper grain. Stone-50 palette.

The power-law shape of word frequency. A small core does almost all the work.

Krashen, and the idea that almost stuck

To understand how a language gets into someone, the unavoidable starting point is Stephen Krashen. In 1982 Krashen published Principles and Practice in Second Language Acquisition, a book that, depending on who you ask, either revolutionized the field or set it back twenty years. The central claim was that adults acquire a second language much the way children acquire a first, through exposure to messages they can roughly understand, and that explicit grammar instruction plays a smaller role than the twentieth-century classroom assumed.

Krashen called the engine comprehensible input, and he gave it a name that has outlived most of the controversy: i+1. The idea is that you acquire language when you are exposed to something just slightly above your current level. Not so far above that you cannot parse it, not so far below that you learn nothing. The sweet spot is the sentence where you understand most of the words and have to lean into one or two. He also drew a sharp line between acquisition, which he described as the subconscious internalization of patterns, and learning, the conscious memorization of rules. His claim that acquisition does most of the heavy lifting was, at the time, heretical.

It is fair to say that Krashen has been criticized inside academic linguistics for overclaiming and for proposing a theory that is hard to falsify. The hypotheses are notoriously slippery to test. Plenty of working linguists will tell you he leaned too hard on intuition and too lightly on data. That criticism is real and worth knowing about. But the core insight, that the primary fuel for language acquisition is large quantities of input you can mostly understand, has held up remarkably well, and it is the foundation underneath every modern language app, every immersion school, and every recommendation to read graded readers and watch dubbed cartoons.

Paul Nation and the two-thousand-word threshold

If Krashen supplied the theory, Paul Nation supplied the receipts. Nation spent decades at Victoria University of Wellington counting words, building corpora, and asking simple empirical questions about how vocabulary works. His canonical reference, Learning Vocabulary in Another Language (2001, second edition 2013), is the book the field quietly agrees on.

Nation's frequency analysis is the result most people remember. The most frequent two thousand word families in English give you coverage of roughly eighty-five percent of any running text. Five thousand gets you to about ninety-five. Researchers sometimes call the two-thousand mark the lexical threshold for comprehension, the rough point below which extensive reading becomes painful and above which it becomes possible. The implication, once you sit with it, is striking. Most of the work of becoming functional in a language is concentrated in a relatively small core. The romantic image of a learner memorizing dictionaries is mostly wrong. The first two thousand words are the cliff. Most of what follows is gentler.

Nation also gave the field a usable program for getting there, called the Four Strands. The idea is that a healthy diet of language study balances four activities in roughly equal measure: meaning-focused input (reading and listening to things you can mostly understand), language-focused learning (deliberate study of vocabulary and grammar), meaning-focused output (speaking and writing), and fluency development (going faster with material already familiar). Most learners, left to their own devices, pick one strand and skip the rest. Most apps do the same. The strongest programs do all four.

The most frequent 2,000 word families of English cover around 85 percent of the running words in most texts. This is a lot of coverage from a relatively small number of words.

Paul Nation, Learning Vocabulary in Another Language, 2001

The dozen-encounters number

A word is not learned in a single sitting. This sounds obvious and yet it is the place where most flashcard regimes quietly fail. The classic empirical study here is Saragi, Nation, and Meister, published in System in 1978. The researchers gave subjects a novel containing carefully tracked low-frequency words and then tested vocabulary retention afterward. The finding that survived into the field's collective memory: words encountered roughly ten or more times in context, across a meaningful reading experience, were the ones reliably retained. Words encountered fewer times mostly were not.

Later work, including replications and refinements by Norbert Schmitt and others, has tightened the number a bit and complicated it. The exact threshold depends on the learner, the word, and the richness of the context. But the order of magnitude is robust. A word needs something like a dozen meaningful encounters before it crosses from a flickering recognition into a part of you. Not twelve glances at a flashcard. Twelve encounters in real, varied, comprehensible context. This is the gap between the words you can name when prompted and the words that come to you unbidden when you need them.

Schmitt's body of work also made another point that the field had been blurring for decades. Knowing a word is not a single binary. There is form (spelling, pronunciation), meaning (the core sense and its near neighbors), and use (collocations, register, the situations where the word fits and the ones where it does not). These dimensions come in incrementally. You might recognize a word in print years before you can use it comfortably in speech. Vocabulary acquisition is not the flipping of a switch. It is a slow brightening, dimension by dimension.

Image slot

One word, twelve encounters, across weeks.

Image prompt: A horizontal timeline sketched on warm cream paper. A single Spanish word, 'amanecer', appears at twelve points along the timeline, each time embedded in a different short sentence (e.g., 'el amanecer fue tranquilo', 'caminamos hasta el amanecer', 'antes del amanecer'). Each appearance is dated with a small day marker showing weeks of spacing. Hand-drawn timeline ticks, serif sentence fragments, light paper grain. Stone-50 palette.

One word, twelve encounters, across weeks.

Pimsleur, and the proprietary truth about spacing

The other ancestor in this story is Paul Pimsleur, the linguist who, in the 1960s, devised what he called graduated interval recall. The premise was that words should be reintroduced into a lesson at expanding intervals, in the moments just before the learner was likely to forget them. The Pimsleur audio courses, still sold today, are built on this insight, and they are one of the rare commercial language products grounded in a real piece of cognitive science rather than marketing.

What Pimsleur described in 1967 is, in modern vocabulary, spaced repetition applied to language. It is the same engine behind Anki, behind Duolingo's review queue, behind every serious vocabulary trainer of the last twenty years. The spacing science has its own article in this series. The point worth keeping here is that spacing and the dozen-encounters threshold are the same problem from two angles. Spacing tells you when the next encounter should happen. The encounters threshold tells you how many are needed in total. Both have to be satisfied.

Why generation and reading both matter

Two more findings round out the picture. The first is the generation effect, which Slamecka and Graf documented in 1978 and which Jan Mondria, among others, later applied specifically to vocabulary in 2003. The finding is that words you have to actively produce, even partially, are remembered better than words you merely recognize. A cloze-deletion exercise where the learner supplies the missing word beats a passive rereading every time. This is why a frame that occasionally shows a sentence with a single word missing is doing more work than a frame that simply displays the word.

The second finding belongs to Stuart Webb and the extensive reading tradition. Webb has shown, across multiple studies in the 2000s and 2010s, that sustained reading at an appropriate level produces substantial incidental vocabulary acquisition. Not as fast per encounter as deliberate flashcard study, but cumulatively enormous, and crucially the kind of acquisition that includes the use dimension Schmitt described. You do not just learn what the word means. You learn the company it keeps.

What this all adds up to

Put the threads together and a picture emerges that contradicts how most people try to learn a language. The picture is not: drill grammar, then memorize vocabulary, then start reading. The picture is: get into comprehensible input as fast as possible, even at the cost of feeling underprepared, and stay there. Meet each new word many times, in many different sentences, spaced out over weeks. Occasionally produce, do not only recognize. Read more than you think you need to. Trust the long arc.

The deep reason classroom language instruction so often fails is that it concentrates encounters. A unit on transportation gives you eight new words on Monday, drills them on Tuesday, tests them on Friday, and never returns to them. The dozen-encounters threshold is never reached. The words flicker into recognition and then back out. Six months later the learner cannot summon them and concludes, wrongly, that they are bad at languages. They are not bad at languages. They were never given the encounters.

Further reading

For the primary sources: Stephen Krashen, Principles and Practice in Second Language Acquisition, 1982 (the original statement of the Input Hypothesis); Paul Nation, Learning Vocabulary in Another Language, 2nd edition, 2013 (the canonical reference, including the frequency analysis and the Four Strands); Saragi, Nation, and Meister, System, 1978 (the ten-encounters study); Norbert Schmitt's work on incremental vocabulary acquisition and the dimensions of word knowledge, summarized in his 2010 book Researching Vocabulary; Paul Pimsleur, "A Memory Schedule", Modern Language Journal, 1967 (graduated interval recall); Slamecka and Graf, Journal of Experimental Psychology, 1978, on the generation effect, and Mondria 2003 for the vocabulary application; and Stuart Webb's work on extensive reading, collected in How Vocabulary is Learned (Webb and Nation, 2017).

Related reads

Bring it home

A quiet device that uses what the research already knows.

QuipCast is a reflective e-paper frame that cues short, meaningful content on the cadence you set. Launches summer 2026. Waitlist signups get founder pricing and a three-year price lock.

Join the waitlist
How a Language Gets Into You - QuipCast