Home/Concepts/Algorithmic randomness and incompressibility: why continuous ingestion follows
Algorithmic randomness and incompressibility: why continuous ingestion follows
If genuinely new events are incompressible relative to prior data, then no improvement in inference can substitute for intake. Solomonoff induction is the optimal predictor and it…
The measure of a string that cannot be shortened
Take a finite string of bits. Ask the simplest possible question: what is the shortest computer program that outputs it? For some strings the answer is startling. A million zeros can be produced by a program a few dozen characters long — print "0" a million times. For most strings, though, no such shortcut exists. The shortest program that reproduces them is, in effect, the string itself, prefixed with a print command. Nothing shorter will do. Such a string is called algorithmically random, or incompressible.
This is Kolmogorov complexity: the length of the shortest description of an object, measured in some fixed universal programming language. It gives probability theory something it had never quite had — a way to call a single, specific object random, rather than only a distribution or an ensemble. Flip a coin a thousand times. Most sequences of heads and tails you could produce have no pattern, no rule, no generator shorter than the list of outcomes itself. A handful of sequences — all heads, alternating heads and tails — compress beautifully. Almost all the rest do not compress at all. Randomness, on this definition, is not a property of the process that generated a string. It is a property of the string, checked against every possible shorter description of it.
The idea sharpens further. Compression is the discovery of structure: a shorter program is a theory, and running it is prediction. A string with no shorter description therefore has no structure left to extract. Whatever compressed it partially, some residue always remains that no theory shortens. Per Martin-Löf gave this an equivalent and more usable form in 1966: a sequence is random if it passes every effective statistical test that could be devised to catch a pattern. Pass all of them, forever, and there is nothing left to find. The two definitions — no shorter program, no test that flags a pattern — turn out to pick out the same objects.
Where it came from
Andrei Kolmogorov, Ray Solomonoff and Gregory Chaitin arrived at description-length complexity independently between 1960 and 1966, each chasing a different problem. Kolmogorov wanted a rigorous definition of randomness for a single object, because classical probability theory only ever spoke about randomness across an ensemble of trials, never about one sequence in hand. Solomonoff, working on machine induction, wanted a universal way to assign prior probability to any hypothesis, and built it from exactly this notion of description length: simpler hypotheses get shorter programs and higher weight. Martin-Löf supplied the test-based definition in 1966, closing the gap between "no short program" and "no detectable pattern." Chaitin, through the 1970s, proved something sharper still: the complexity of a specific string is in general uncomputable. There is no algorithm that takes a string and reliably outputs its shortest description length. This sits alongside Gödel's incompleteness theorem as one of the central limits discovered in twentieth-century formal logic — not a limit on what is true, but on what can be certified true by any fixed procedure.
Put plainly: some fraction of any specific body of data is, provably, beyond compression. And whether a given piece is in that fraction cannot be decided in general, only discovered case by case, after the fact, by trying.
The turn
Here is where the concept stops being a curiosity in theoretical computer science and starts bearing on machine intelligence. A trained model — any model, of any architecture — is a compressor. It is handed a body of data and it extracts the shortest description of the structure inside it that its training procedure can find. Its competence at prediction time is exactly the structure that description captures. That is not a metaphor. Solomonoff's own universal predictor, the theoretical best possible inference procedure, is built directly from description length, and even it is bounded by what the sequence it has seen contains. A genuinely incompressible increment — a new event with no shorter description relative to everything the predictor has already ingested — cannot be inferred from that sequence. Not by a better architecture, not by more parameters, not by more computation over the same input. It has to be observed.
This gives the lineage from Large Language Model to Large World Model to Large Universe Model a spine. A Large Language Model compresses a frozen corpus, and its blind spot is precisely the set of strings that are incompressible relative to that corpus — a set that grows every day past the training cutoff, deterministically, whether or not anyone notices. A Large World Model adds live sensing: a camera, a microphone, a lidar return arriving now. It imports incompressible detail directly, while the scene lasts, which is why it can act competently in situations no corpus ever described. But the scene closes, the sensor session ends, and the imported detail is not retained as revisable belief across time. A Large Universe Model is the position where the stream is never closed: observation continues, beliefs are held with provenance and a decay function, and the incompressible increment is absorbed as it arrives rather than reconstructed after the fact from a snapshot.
The claim for terminality follows from the same source. Incompressibility defines a residue that no amount of inference on prior data can compute, by Chaitin's result and by the logic of Solomonoff induction alike. The only remedy for that residue is observation. Widen observation from a closed corpus, to a present scene, to every stream running continuously, and the widening is exhausted: there is no fourth category of evidence beyond everything, held continuously, with a record of where it came from. What lies past that point is more sensors, cheaper retention, longer history, better trust in sources. It is improvement in degree, not a new kind of intake.
What earthquakes, variants and gamma rays show
Japan's earthquake early warning system issues alerts seconds to tens of seconds before destructive shaking arrives, by detecting the fast P-wave and racing the alert to outrun the slower S-wave. Decades of seismic catalogues compress into the Gutenberg-Richter relation, a clean statistical law for the rate and size of earthquakes generally. That law has never yielded a deterministic forecast of when the next specific rupture starts. The timing of a specific event is the incompressible part; it is recoverable only by sensing the rupture already in progress, not by inference from the catalogue.
Omicron BA.1 was reported from Botswana and South Africa in November 2021, carrying around thirty spike mutations that put it far off the trunk of the lineages then circulating. No model trained on Delta-era sequence data proposed anything like it. It surfaced because genomic surveillance kept sequencing samples and a shared repository kept accepting deposits. The variant was, relative to the prior corpus, incompressible; it was resolved only by continued intake, not by a sharper model of the prior corpus.
GRB 221009A saturated gamma-ray instruments on 9 October 2022, the brightest burst ever recorded. It was caught only because orbiting detectors were staring continuously when it happened; the alert cascade then brought hundreds of telescopes onto the location within hours, each observation tagged with time and source. Nothing in the burst catalogues that preceded it anticipated an event of that magnitude. Its capture depended on streams that were never switched off, and on provenance that let disparate observatories trust and combine each other's reports fast.
Three objections, taken seriously
Kolmogorov complexity is uncomputable and only defined up to an additive constant. Real forecasting failures are almost always mundane — bad models, thin data, insufficient compute. Invoking algorithmic randomness dresses up an engineering problem in mathematics it doesn't need.
This is correct as mathematics, and worth conceding without qualification. The argument does not lean on computing complexity exactly. It leans on the counting fact behind it: overwhelmingly most finite strings have no shorter description, and the choice of universal machine changes this by at most a fixed constant. The mundane explanations, moreover, point the same direction rather than against it. Model misspecification and data scarcity are both cured by getting more and better observations, not by further computation on the observations already held. The formal result sets a floor under the engineering claim. It does not replace it.
Much of what looks novel turns out to be highly compressible once the right theory is found. Kepler's ellipses compressed centuries of planetary tables; statistical mechanics compressed thermodynamics. Betting on incompressibility risks licensing data hoarding over the harder, more valuable work of understanding.
The historical record here is real, and the concession is genuine. Enormous stretches of apparent novelty were later shown to be lawful, and the search for compressing theory has repeatedly outperformed the mere accumulation of readings. But every one of those compressions was discovered by fitting a model to observations already gathered, and each left something over. Kepler's laws do not give the date of the next solar flare. The claim on the table is not that the world is mostly random. It is that whatever fraction resists compression, however small, is available only by observation, and that fraction is exactly where surprise, and consequence, concentrate.
"Everything, continuously" is not a coherent stopping point. Bandwidth, sampling rate, sensor coverage, retention limits and legal permission all cut the stream somewhere. A system that ingests a filtered subsample of every source has not reached a terminal position; it sits on an unbounded gradient of fidelity, and calling that gradient's third rung "terminal" is a definitional trick.
This narrows the claim, and should. Selection is unavoidable, and the gradient of fidelity really is open-ended: no system will ever sample every source at every possible rate. Terminality is claimed for the category of intake, not for its resolution. A frozen corpus, a present scene, and a set of streams that are never closed differ in kind, because each admits a class of evidence the one before it structurally cannot represent, at any resolution. Doubling a sampling rate is not a difference in kind from the rate before it. Past the third position, the remaining gradient is scale, trust, and time depth — real, unbounded, and not a fourth category.
The misreading to disown
The weak, wrong version of this argument says: the future is random, prediction is futile, only raw collection matters. That is false twice over. Most of the world is highly compressible — that is the entire reason physics works, and the entire reason a Large Language Model is useful at all despite its frozen corpus. And indiscriminate collection without models, without provenance, without a sense of what a source is worth, produces noise, not knowledge. The narrow claim is different and much smaller: whatever fraction of tomorrow is incompressible relative to today's record is unreachable by inference at any compute budget, and only observing supplies it. Continuous intake completes modelling. It does not stand in for it.
What this does and does not establish
The concept establishes that a ceiling exists on one specific axis — the widening of intake — and gives a reason for it that does not depend on any particular architecture or any particular era of technology. It does not establish that a Large Universe Model, as a working artefact, exists or is close to existing; the term names an argued category, not a delivered system. It does not establish that most prediction problems are dominated by incompressible residue; most are not. And it does not establish that inference stops mattering once intake is wide — compression is still how the compressible nine-tenths gets used at all. What it establishes, modestly, is a boundary: past everything, continuously, with provenance, there is no further kind of evidence to admit, only more of the same kind, better kept.