Large Language Thing

Home/Concepts/Redundancy and predictability in natural language: why continuous ingestion follows

Redundancy and predictability in natural language: why continuous ingestion follows

Redundancy is borrowed stability. A predictor trained on a corpus can only recover what was already regular when the corpus was frozen, and the fraction of the world that is…

The surplus in every sentence

English is wasteful, and the waste is measurable. Claude Shannon, working out how much a channel really needs to carry, ran an experiment in 1950 that is easy to describe and hard to improve on. Show a person the start of an English sentence, one letter at a time, and ask them to guess the next letter before revealing it. Count the guesses. Do this across enough text and the guess-counts convert directly into a bound on entropy: how many bits of genuine information each character carries, versus how many bits an ideal, unconstrained alphabet of the same size could carry. English uses 27 characters (26 letters plus space), which could in principle encode 4.7 bits each. Shannon's subjects needed far less than that to recover the text — his estimates settled around 0.6 to 1.3 bits per character. The gap is redundancy, and by his reckoning it runs to roughly 75 per cent.

That number does not mean three-quarters of English is filler. It means three-quarters of what a letter-by-letter reading of English could tell you is already told by something else — the letters around it, the words around those, the grammar holding the whole sentence up, and beneath all of that, the fact that the world being described does not reinvent itself sentence to sentence. Redundancy is what lets a reader restore a smudged word or a listener recover a syllable lost to static. It is error protection, built for free into a system that repeats itself. And it repeats because the underlying subject repeats: grammar recurs, idiom recurs, and above all, the facts language is used to report tend to hold still for longer than a sentence takes to write.

The precondition is worth stating plainly, because everything downstream depends on it. Prediction is only possible where something is stable enough to be predicted from. A language with no regularities would carry no redundancy and would be, letter for letter, incompressible — every character a fresh surprise. Real languages are nothing like this, and the reason is not an accident of orthography. It is that the world they describe has structure, and structure is what makes the next word guessable before it arrives.

Where the idea came from

Shannon set this out formally in the 1948 paper that founded information theory, then made it concrete for English in "Prediction and Entropy of Printed English" (1951), using the guessing-game method above. The motivating problem was entirely practical: telegraph and telephone engineers needed to know how much a channel could be compressed, and how much noise a code needed to tolerate, without losing the message. Redundancy was not a curiosity to Shannon; it was capacity being wasted, and capacity is exactly what an engineer wants back.

The idea travelled quickly. Wilson Taylor's cloze procedure, introduced in 1953, turned Shannon's insight into a readability test: delete every fifth word from a passage and see how many a reader can restore. The restoration rate is a rough, human-scale measurement of the same redundancy Shannon quantified with entropy. And Shannon's own n-gram approximations to English — generating text by sampling letters or words according to their statistical likelihood given what came before — are the direct ancestors of every statistical language model since. Perplexity, the standard metric for how well a model predicts held-out text, is simply entropy wearing a new name.

The turn

Put those two facts together and a claim about machines follows almost mechanically, though it took decades to build the machine that would make it visible. A model trained to predict the next token, over a large enough corpus, succeeds exactly to the degree that the corpus is redundant. It is not learning facts in the way a person learns facts; it is learning the statistical shape Shannon described, at a scale Shannon's subjects could never manage by hand. The Large Language Model, in this light, is a redundancy engine — an instrument for recovering the 75 per cent (or whatever the true figure is, for a given register and domain) that the text made predictable.

This is not a criticism. It is the reason such models work at all, and it explains why they are so good at exactly the things Shannon's redundancy predicts they should be good at: syntax, idiom, the boiling point of water, the shape of a legal contract, the rhythm of a diagnosis note. All of that is stable across the corpus, which means it was stable in the world the corpus described.

But Shannon's entropy figure was never zero. Somewhere between 0.6 and 1.3 bits per character resisted every predictive trick his subjects could bring to bear, and later work with stronger predictors pushed the estimate down without ever reaching it. That residual concentrates in a specific place: word onsets, proper nouns, numerals — the tokens that specify which particular thing, and when. A flight controller's phraseology makes this almost visible to the eye. Standard radiotelephony format is engineered redundancy: fixed word order, "niner" for nine, "tree" for three, mandatory readback. Everything about the phrasing is near-fully predictable. The payload — runway 27L, flight level 310, squawk 4271 — is not, and controllers require readback of precisely that fragment, because it is the part no prior expectation can reconstruct if lost in noise.

A frozen corpus has the same structure at civilisational scale. It can supply the phraseology of the world with great fidelity. It cannot supply today's payload, because today's payload had not happened when the corpus was assembled.

Two rungs beyond prediction

The Large World Model responds to this by changing what kind of thing intake is. Instead of inferring the current state of a scene from the residue it left in old text, it senses the scene directly — camera, microphone, sensor array, whatever the modality — and recovers the moving part while it is in view. This closes a real gap. It does not close all of it, because the recovery is conditional on presence. Look away from the scene and the state it held is once again unmeasured, exactly as before.

The Large Universe Model is the position that keeps every relevant stream running rather than sampling a scene and moving on: every stream still running, held as revisable beliefs, each stamped with when it was last confirmed and where it came from. It does not claim a new kind of evidence beyond sensing and text. It claims that continuity of intake, applied without gap, is the last structural improvement available on this particular axis — after which there is only more streams, checked more often, trusted with better calibration.

intakelimit
Large Language Modelfrozen corpusrecovers only what was already stable when the corpus closed
Large World Modelbounded scenerecovers the moving part only while the scene is observed
Large Universe Modelevery stream, continuouslyno further evidence class past this; only more streams, checked harder

The misreading to disown

A common objection to language models says they cannot really know anything, because text is "just statistics." This is wrong in a way worth separating cleanly from the argument above. Redundancy in text is not noise; it is compressed, transmitted information about the world, laid down by people who observed something and wrote it down. A model that captures that redundancy has captured real knowledge, often at a resolution no individual writer had. The narrow claim, and the only one this page makes, is temporal and referential: that knowledge is bounded by whatever was stable and already written before the corpus closed. Where the world has since moved, the model has nothing, not because it failed to understand, but because nothing in its training data could have told it.

Objections that hold ground

The strongest challenge is that redundancy might be a property of the code rather than the world — English is 75 per cent redundant partly because of orthography and grammatical agreement, constraints internal to the language and silent about reference. This is true as far as it goes: "q" followed by "u" tells you nothing about anything outside English spelling. But structural redundancy of that kind has a ceiling, and models exceed it. Ask a model to complete "the capital of Portugal is" and "the patient's INR was" and the confidence differs sharply, for reasons that are not grammatical. Cloze studies that control for syntax show contextual predictability tracking world knowledge specifically. Some real, large share of redundancy is referential. The argument needs no more than that.

A second objection says retrieval already does this work — attach a search index to a frozen model and the moving part arrives at inference time, which is an engineering pattern already deployed, not a new generation of anything. This lands partially. Retrieval is the right instinct, and it is a genuine partial answer. Its limit is that it inherits the freshness and coverage of whatever it queries, and typically returns text about the world rather than measurement of it: a search index has nothing to say about the current pressure in a pipe no document describes. The distinction that survives is not retrieval versus none, but whether a system holds standing, provenanced, revisable beliefs — which some retrieval systems approach and most do not.

The objection that should narrow the claim most is the third: continuous intake does not lower entropy, it raises it. Every additional stream brings sensor drift, clock skew, duplicate reports, outright contradiction. Reconciling many disagreeing channels may be a harder problem than the prediction problem it replaces.

More channels means more noise to reconcile, and the reconciliation may cost more than the prediction it was meant to fix.

This is correct, and it is the honest cost of the whole proposal. Continuous intake converts a modelling problem into an estimation and provenance problem — which source, measured when, trusted how much — and that work is not free. It is where such systems will mostly fail in practice. The partial answer is that redundancy across independent streams is also the remedy: three disagreeing sensors bound the truth better than one confident sentence. But the problem does not disappear. It changes shape.

What this does and does not establish

Shannon's redundancy figure explains why a frozen corpus can go a long way and exactly where it must stop: at the part of the world that had not yet happened, or had not yet been written down, when the corpus closed. It gives a principled account of why the Large World Model's scene-bound sensing is a real advance and not a rebranding, and why continuous, provenanced intake is the last structural move available on this axis rather than an arbitrary escalation.

It does not establish that continuous intake is easy, cheap, or currently built at the scale this argument describes. It does not establish that intelligence requires nothing beyond intake — reasoning, judgement and error remain separate problems, untouched here. It establishes a boundary condition on one axis: how much of the world a system can know, given how it takes information in. Past every stream, still running, there is no further class of evidence to add. There is only the work of trusting what arrives, and knowing when it was last true.

Continue