Nine Thousand Words to Read a Newspaper

On the objection that a domain model with six hundred concepts is too complex, and why the number is actually low.

I published a concept map of the AI governance domain recently. ~600 concepts, sixteen neighborhoods, every concept with a generated page and a generated context diagram. It took an hour or so total effort.

The most likely reaction in a corporation would be some version of: that is too complex.

Two different complaints wearing the same clothes

There is a real complaint hiding inside the bad one, and they need separating. The real complaint is about what you are being shown at once. The bad complaint is about the size of the domain, and it is a category error, because the domain does not care what you find comfortable. AI governance has as many concepts as it has. You can decline to name them, and every one of them will still be there, operating on you unnamed.

These get conflated constantly, and the conflation has an origin: for most of the history of this profession, diagrams were drawn by hand. Drawing was expensive, so every diagram had to earn its cost by showing as much as possible, so diagrams were dense, so dense diagrams became what a diagram is. Under that constraint, a large model really was unusable, and “keep it simple” was sound engineering advice rather than an aesthetic preference.

That constraint is gone. When the view is generated, you render one concept and its immediate neighbors, and then you render the next one, and the cost of the second view is zero. You never look at six hundred things. You look at eight things, six hundred times, in whatever order your question takes you.

So the answer to “that is too complex” is not to shrink the model. It is to stop showing people the whole thing.

What a vocabulary costs

Here is where I want to bring in numbers from a field that has studied this properly, because software architecture has strong opinions about complexity and almost no measurements, whereas second language acquisition has spent decades counting.

The question “how many words do you need” has a reasonably settled answer, and it is organized around text coverage: what proportion of the words in a normal text you already know.

Roughly, in word families:

  • 2,000 word families gets you about 80% coverage. This is the functional floor. You survive. You do not read.
  • 3,000 gets you to roughly 95% of spoken language. Conversation becomes possible, with effort.
  • 5,000 puts you near 95% of written text, which sounds excellent and is not. At 95%, one word in twenty is unknown, which is about one per line. You guess continuously and you are often wrong.
  • 8,000 to 9,000 reaches about 98%, and 98% is where unassisted comprehension actually lives.
  • An educated native speaker is usually estimated somewhere between 15,000 and 20,000.

Look at the gap between the fourth line and the third. Going from 95% to 98% coverage, from struggling to comfortable, costs roughly four thousand additional word families. Three percentage points, four thousand words. The curve is savage at the top and there is no shortcut in it.

People have tried to build the shortcut anyway. Ogden’s Basic English compressed the language to 850 words in 1930. Globish and Voice of America’s Special English both settle around 1,500. These are real engineering achievements and they work for their purpose, which is bounded: announcements, instructions, news read slowly, transactions between people with no shared language. What none of them will ever do is let you argue a subtle point, or read a novel, or notice that somebody is being sarcastic. The reduction that makes them learnable is the same reduction that caps what they can say.

That trade is not a flaw in the design. It is the trade. You can have a small vocabulary or you can have a large expressive range, and the exchange rate is not negotiable.

What a hundred words can do

A short detour, because animals make this point more sharply than people do.

Chaser, a border collie trained almost daily for years by the psychologist John Pilley, learned the names of 1,022 individual objects and could fetch any one of them on request. Alex, the African grey parrot who worked with Irene Pepperberg for thirty years, had roughly a hundred labels.

Chaser’s number is ten times larger. Alex was doing something Chaser could not approach.

Alex could look at a tray of objects he had never seen and answer what color, what shape, what material, and how many. He treated same and different as abstractions rather than as trained pairs. He had something close to a concept of zero, which by the evidence he worked out rather than being taught. And he produced speech instead of only responding to it, including asking for things and, unmistakably, refusing.

Chaser had an enormous lexicon of proper nouns. Alex had a small vocabulary with categories, relations and operators in it.

That is the distinction I keep trying to make about domain models, and I have never found a cleaner illustration of it than a parrot outperforming a dog by a factor of ten in the wrong direction. A thousand labels is a list. A hundred terms with real relationships between them is a language. The count is the least interesting number in either case.

Which is also why I am unbothered by ~600, and would be more worried by six hundred concepts with no edges between them than by six thousand that were properly connected.

Thirteen thousand words and still lost in Santiago

Which brings me to the number that keeps me honest.

Duolingo tells me I know about thirteen thousand Spanish words. I went to Chile. I could barely get by, and the conversations that worked mostly worked because the other person spoke some English.

By the table above, thirteen thousand should have been plenty. It was not, and the reasons are worth more than the table.

The units do not match. The research counts word families: help, helps, helped, helping, helper, unhelpful are one family, not six. App counters are closer to individual word forms. My thirteen thousand forms is plausibly four or five thousand families, which puts me right around 95% coverage. Which is to say: exactly the struggling-but-surviving band. The number was not lying to me. I was reading it in the wrong unit.

Recognition is not production. I can recognize far more Spanish than I can produce, and the gap widens under time pressure. Reading gives you as long as you want. A conversation gives you about four hundred milliseconds.

And I picked the hardest room in the house. Chilean Spanish is notorious among learners: /s/ aspirated or dropped entirely, vowels reduced, delivery fast, voseo, and a thick layer of local idiom. Native Spanish speakers from other countries report the same difficulty. This is not my failure of preparation, it is a known property of the dialect.

The conclusion I want to draw from this is the opposite of the one the table suggests on its own.

Vocabulary is necessary and it is not sufficient. Knowing the words is not knowing the language. There is no vocabulary size at which fluency switches on, because fluency is made of things vocabulary does not contain: how words combine, what register signals, what is idiom and what is literal, and the enormous amount of meaning that is carried by what nobody bothered to say.

Back to the six hundred concepts

Both halves of that transfer directly, and I need both or I am selling something.

First half: six hundred is not a large vocabulary. It is a small one. It is below the survival floor for a natural language and it describes a professional domain with regulators and penalties in it. Anybody who finds it excessive is not making a claim about the model, they are making a claim about how much of their own working domain they have agreed to name. Most enterprise glossaries run to a few dozen terms, and that is not evidence of a simple domain. It is evidence of an unnamed one.

There is a reason organizations prefer it that way, and it is not laziness. A named concept can be argued with. A named relationship can be shown to be wrong. The twenty-term glossary and the one-slide architecture survive because vagueness is politically cheaper than precision: nobody can be shown to have been wrong about a concept that was never written down. Simplification is often presented as a service to the reader and is frequently a service to the author.

Second half, and this is the one I would rather people took away: the map is not the fluency. A concept map is a vocabulary. It is the thing you need before you can have a real conversation, and it is nowhere near the whole conversation. Handing someone 600 concepts does not give them judgment, any more than my thirteen thousand words gave me Santiago.

What the map does is make the vocabulary cheap, and that matters because the vocabulary used to be the expensive part. It used to take years of exposure to learn what the words in a domain meant and how they connected, and that time was mostly spent on acquisition rather than on thinking. Generate the vocabulary and that time is freed for the part vocabulary cannot give you.

Which is why the interesting move is what happens after the map exists:

Render one hop at a time. Never show anyone the whole graph. The whole graph is for the machine.

Federate it. One map of an entire enterprise is as useless as one diagram of 600 classes. Split it along the lines the organization actually thinks in, which are usually bounded contexts whether or not anyone calls them that.

Give every context an owner. A concept with no owner is a concept nobody corrects. This is the step that decides whether the thing is alive in a year, and it is the step that gets skipped.

Let the owners refine. Their job becomes correcting a graph rather than writing four hundred pages, and correcting a graph is a job a busy expert will actually do.

Let some graduate. A few of these maps earn their way into real metamodels with real models built on them. Most stay reference material that the metamodels are designed against. Both are good outcomes and you cannot tell in advance which is which, which is precisely why the map needs to be cheap.

Fluency still comes from use. It always did. What has changed is that you no longer have to spend the first three years just learning the words.

Where the hour actually went

I should account for that hour more honestly than I have so far, because the breakdown is not what you would guess.

Almost none of it was thinking. The thinking was done years ago and is encoded in the tower. What the hour actually contained was generation, and generation came in two kinds.

The first kind is the one everybody is currently excited about: a model generated from a description, and prose generated per element. That part is fast, and it is genuinely cheap, which is the argument this whole piece rests on.

The second kind has no intelligence in it whatsoever. First, ~1K Java sources were generated from the model - which proved the model validity. In “real”/“graduated” models these files are used to load and save data. Then a static site generator walked the model and wroted 5.5K files, 1.5Gb total - several pages per concept with context diagrams and graphs. The total size has exceeded the GitHub pages limit and I had to publish the site to my webspace.

This is worth naming, because you will hit it too. When generating knowledge artifacts becomes cheap, the bottleneck does not disappear, it relocates, and it relocates to publishing. Pre-rendering every view of every element is a perfectly reasonable architecture at fifty concepts and a poor one at six hundred, and nothing warns you at the crossover.

The fix is architectural rather than a faster machine. Stop pre-rendering. Ship the model itself alongside a viewer that reads it at runtime and renders the view the reader actually asked for, instead of building every view anybody might ever ask for in advance. One model and one generic viewer, rather than thousands of pre-built files.

That is what the meta model and the React model are for. The meta model is deliberately small, projected to TypeScript, and metacircular by construction: a package containing classes describes package and class, so a runtime bootstraps from that model alone and carries no foreign metamodel with it. Put a reflective React viewer on top of that and the static generation step largely stops existing. The generator still runs for whatever genuinely needs to be static, and the per-concept views become something the browser computes when somebody asks for one.

I am reporting that as pending rather than done. But its shape matters for the argument, because the last expensive thing in this pipeline is also the least interesting thing in it. There is no judgment in writing several thousand files. Nobody’s experience is encoded there.

Chatting with a concept

Once that viewer exists, one more thing becomes nearly free. A reflective viewer already holds the model in the browser. It knows which concept you are looking at, what it specializes, what specializes it, and everything one hop away. That is the expensive input for a language model, and it is already sitting there, structured, at the exact moment somebody would want to ask a question.

So: chat scoped to a concept. Not a chat window bolted onto a documentation site, which is what everybody ships, but a conversation whose context is the neighborhood you are standing in. Ask about meaningful human control and the assistant has the concept, its supertype, its relations, and the handful of concepts it sits between, without retrieving anything. You could scope it tighter still and chat with a single reference or a single attribute, where the question is usually “why is this relationship here, and who decided that.”

The interesting part is not the chat. It is that most retrieval systems spend their engineering effort guessing which context is relevant to a question, and a concept map has already answered that structurally, as a property of the model rather than as a similarity score. The neighborhood is the context window. Somebody drew those edges deliberately, and they turn out to do a second job nobody drew them for.

The commercial consequence follows quickly, inspiered by an AI course I completed some time ago. It answers the question: how does anybody get paid for a well-made concept map?

Mostly they do not. A consultancy builds a genuinely good domain model, delivers it as a deck or a PDF, and the value evaporates, because the artifact is inert, uncitable, and stale within a quarter. A hosted, access-controlled, conversational model is a different object. Put it behind authentication on something like Vercel, sell subscriptions, and what is being sold is not a document. It is ongoing access to a maintained vocabulary that answers questions. A corporation can run the identical mechanism inward, where the access control is the entire point and nobody pays anything.

Now the part I owe you, having spent two thousand words arguing the other side.

An assistant that answers your questions about a map is still delivering vocabulary. Faster, more pleasantly, and in a better order, but vocabulary. It does not hand anybody fluency, because fluency is made of things the map does not contain, and I would be suspicious of anything sold on the promise that it does. The honest claim is narrower and still worth paying for: it compresses the acquisition phase, which used to consume years, so that the part actually requiring judgment starts sooner.

I am not planning to build this, not in the near future. I have written it down so that my own future self finds it here via a web search or by asking an AI assistant.

The sentence I would keep

If the model looks too complex, the model is probably right and the view is probably wrong.

Shrinking the domain to fit the diagram is how we got twenty-term glossaries for domains that fine people into the millions. Generate the views instead, one neighborhood at a time, and let the vocabulary be as large as the thing it has to describe.

Then go and have the conversation, which is the part no map has ever done for anyone.


Sources for the vocabulary figures

  • Nation, I.S.P. (2006). “How large a vocabulary is needed for reading and listening?” Canadian Modern Language Review. Source of the 8,000 to 9,000 word family figure for written text and the lower figure for spoken.
  • Hu, M. and Nation, I.S.P. (2000). On the 98% coverage threshold for unassisted comprehension.
  • Nation, I.S.P. and Waring, R. (1997). “Vocabulary size, text coverage and word lists.” Native speaker estimates.
  • van Ek, J.A. and Trim, J.L.M. Threshold Level, Council of Europe. The concept underlying CEFR B1, and the correct term for a minimum functional vocabulary.
  • Ogden, C.K. (1930). Basic English. The 850 word vocabulary.
  • Pilley, J.W. and Reid, A.K. (2011). “Border collie comprehends object names as verbal referents.” Behavioural Processes. Chaser’s 1,022 proper nouns. Note that Chaser was an exceptional individual under years of daily training; typical dogs are estimated around 165 words.
  • Pepperberg, I.M. The Alex Studies (1999) and subsequent papers. Alex’s roughly one hundred labels, the categorical abilities, and the zero-like concept.

A note on units: these studies count word families, not word forms. Consumer language apps generally count something closer to word forms, so their totals are not directly comparable and will read high by a factor of roughly two to three.