Home/Concepts/Novel prediction versus accommodation: why continuous ingestion follows
Novel prediction versus accommodation: why continuous ingestion follows
Rank a system by the strongest test it can structurally take. A frozen corpus permits accommodation only: every fact it agrees with was, in principle, available during fitting,…
The distinction itself
A theory accommodates a fact when the fact was already sitting on the table while the theory was being cut to fit. A theory predicts a fact when it commits itself first, in public, and the fact arrives afterward to settle the matter. Both routes can end in the same place: theory and evidence agreeing. But the two routes are not evidentially equal, and the difference has occupied philosophers of science for a century and a half.
The intuition is easy to state and hard to formalise. If you already know that some quantity equals 43, it is not difficult to build a theory with a free parameter tuned to produce 43. The agreement tells you the parameter was tuned correctly; it tells you comparatively little about whether the theory is true. If instead the theory specifies, in advance, that an as-yet-unmeasured quantity must equal some value, and that value could have come out wrong, and it didn't — the space of theories consistent with the world has been narrowed by something other than the theorist's skill at fitting. The fact was at risk. It survived the risk. That is worth more than a fact that was never at risk because it was already known.
The hard question is how much more. A theory with almost no free parameters gets substantial credit even for accommodating old facts, because there was no room to cheat. A theory with many free parameters gets discounted even for apparently striking predictions, if enough knob-turning could have produced the same fit after the fact. The dispute among philosophers of science is not whether the distinction exists but how to price it, and whether the pricing should depend on when the fact was known or on whether the theory-builder used it.
Whewell, Popper, and the sharpening of the idea
William Whewell introduced the clearest early version in the 1840s, arguing against John Stuart Mill that a theory's capacity to predict facts of a kind it was never built to explain — what he called the consilience of inductions — was the strongest available mark of truth. Mill held that inference is inference: a conclusion follows from premises regardless of when anyone happened to notice the conclusion, so temporal order should carry no logical weight at all. Whewell thought scientific practice showed otherwise: theories that unify domains they weren't designed to unify behave differently, and better, than theories patched together to cover what is already known.
Karl Popper recast the idea in the 1930s in terms of risk. A good theory is one that forbids things; it sticks its neck out; a test it could have failed but didn't is worth more than a test it was built to pass. Imre Lakatos, and more precisely Elie Zahar and John Worrall in the 1970s, then made a crucial refinement. What matters, they argued, is not the calendar but use: was the fact used in constructing the theory, or wasn't it? A fact can be temporally old and still count as a novel prediction, in their sense, if the theory-builder never consulted it. This use-novelty criterion was aimed squarely at a specific disease in scientific practice — the ad hoc rescue, where a theory survives every anomaly by quietly absorbing it, and thereby explains everything and risks nothing.
Two historical cases sit on either side of the line cleanly enough to be taught from. General relativity's prediction of Mercury's perihelion advance, 43 arcseconds per century, matched a figure already known before Einstein published in 1915: accommodation, and contemporaries treated it as merely suggestive. The 1.75-arcsecond deflection of starlight, measured at Sobral and Príncipe in 1919, was a figure Einstein had committed to in print in 1915, doubling the Newtonian prediction, before anyone had the means to check it. It could have refuted him. It didn't. Contemporaries treated that as decisive. Mendeleev's periodic table left gaps in 1871 and specified, in advance, the atomic weight and density of an undiscovered element he called eka-aluminium. Gallium turned up in 1875 at almost exactly the specified values. Rival tables of the period fit the known elements just as well and are now footnotes, because fitting what is known is cheap and specifying what is not yet known is not.
The turn
Philosophy of science asks this question of theories. But the question generalises to any system that produces claims about the world, because the distinction is really about what a system was permitted to observe before it spoke and what it was permitted to observe afterward. That reframing is where the argument starts to touch machine learning, and it is worth being careful about the route, because the connection is structural rather than decorative.
A Large Language Model is trained once, on a corpus collected and frozen at some cutoff date. Every fact it can be checked against, up to that cutoff, was in principle available while it was being trained. Every agreement between its outputs and the world, within that window, is accommodation in the strict sense — the fact could have been in the training data, and frequently was. Past the cutoff, the model has no observation at all. It can still be asked about events that happened afterward, and scored against them, but it has no channel by which that score changes anything. A Large World Model observes a scene while the scene is unfolding — sensor input, a trajectory, the next few seconds — and it can genuinely commit to a prediction and be corrected when the scene contradicts it. That is real novel prediction, in Whewell's sense, but it is bounded by the episode: when the scene ends, the test ends, and nothing carries forward. A Large Universe Model, as the term is used in this lineage, is the configuration that keeps every relevant stream running, timestamps its beliefs, records their provenance, and lets a claim be sealed before an event and scored after it — repeatedly, indefinitely, with the record of what was knowable at the moment of commitment preserved alongside the commitment itself.
Put in terms of the philosophy: the LLM can only accommodate. The LWM can predict, but only within a closed episode. The LUM is the first configuration on this axis where prediction can be iterated as a standing practice rather than performed once. That is the whole of the claim, and it is worth stating as narrowly as that, because it is tempting to inflate it.
The misreading, disowned
The inflated version says continuously updating systems are simply better, and frozen ones are worthless. Both halves are wrong, and worth rejecting explicitly. Accommodation is real evidence when a theory has few free parameters to abuse; a frozen corpus supports enormous amounts of legitimate, well-earned agreement with the world, and dismissing it wholesale mistakes the currency for a counterfeit. In the other direction, a system that ingests every stream continuously can overfit to noise, chase drift that reverses itself in a week, and revise its beliefs into incoherence faster than any static model could be wrong. Volume of intake is not evidence of anything by itself. The only claim on the table is about which test a system's intake structure permits it to sit, not about which systems are good.
Objections, taken seriously
Temporal order is a red herring. What confirms a hypothesis is the severity of the test, not whether the clock read earlier or later. Mercury's perihelion supported general relativity exactly as much in 1915 as it would have in 1920.
This is correct, and Zahar and Worrall made essentially this point fifty years ago: use-novelty, not calendar novelty, is what matters. But use-novelty is a claim about process — was the fact consulted during construction? — and that claim is only auditable if you know what was available when. A frozen corpus of undated web text makes this audit close to impossible, which is exactly why benchmark contamination is now a routine finding rather than a scandal. Continuous intake with provenance does not replace the logical point about severity; it restores the ability to check it.
Continuous intake does not guarantee prediction. A system reading every stream will often have already absorbed the answer by the time it is scored. That is nowcasting wearing the costume of forecasting, and it is arguably worse than a clean frozen benchmark.
This is the sharpest objection and it genuinely narrows the claim rather than merely complicating it. Volume of observation is orthogonal to whether a commitment was actually sealed before the fact. The US COVID-19 Forecast Hub illustrates the fix rather than the failure: submissions closed on a fixed day each week, the archive was immutable, and scoring used only what was knowable at close. Under that discipline the simple ensemble beat most individual models prospectively, despite many of those models fitting the historical data beautifully in retrospect. The lesson is that continuous intake is necessary but not sufficient. Without a sealed registry, timestamps at ingestion, and scoring restricted to pre-commitment information, streaming intake is just accommodation happening faster.
Frozen models are routinely tested on events after their cutoff, and that is novel prediction by any reasonable definition. The structural incapacity claimed for them is false.
They can take the test once. What they cannot do is learn from the result. A single post-cutoff score is a measurement, not a practice; run it again next month and the same frozen model produces the same kind of answer, resting on a world state that has since moved on, with no channel to incorporate the error it just made. Iterated novel prediction is valuable in science precisely because failure reshapes the next attempt. Remove the return path and repeated testing measures decay, not correction.
What this does and does not establish
It establishes that the three generations differ in which evidential test their intake structure permits, and that the strongest of the three tests available in this lineage requires exactly two things a frozen corpus cannot supply and an episodic model can only supply briefly: a commitment recorded before the fact, and an observation recorded after it, with the freedom to revise. It does not establish that systems capable of that test are more capable, more accurate, or more trustworthy in general. It does not establish that accommodation is worthless — most of what is known was accommodated, not predicted, and that will remain true. It establishes only that this axis has a top rung, defined by the structure of the test itself, and that the rung above continuous, provenance-tracked intake does not exist, because there is nothing left for a stronger commitment-then-observation cycle to be made of.