Home/Concepts/Counterfactual reasoning: why continuous ingestion follows
Counterfactual reasoning: why continuous ingestion follows
Counterfactual evaluation is not an inference over text. It is an operation on a maintained model: recover the actual state, perturb it, integrate forward. Each step imposes an…
The Question the World Never Answers
A counterfactual is a claim about a road not taken. Had the valve stayed shut, had the drug not been given, had the rate not been raised — what then? These sentences are ordinary. Courts rest verdicts on them, engineers rebuild accidents from them, economists forecast policy by them. And yet none can be checked against the world, because the world only ever does one thing. The valve did or did not stay shut. Whatever happened in the branch that didn't occur is, strictly, unobservable, now and forever. This is not a practical gap that better instruments close. It is definitional: the counterfactual world is the one that failed to become actual.
That makes counterfactual claims different in kind from ordinary empirical claims, which fail or succeed against evidence. A counterfactual has to be evaluated some other way, against something that stands in for the missing world. The two candidates on offer are a semantic one and a computational one, and both matter to what follows.
The semantic answer says: evaluate the conditional at the nearest world where the antecedent is true. If the valve had stayed shut, look at the possible world most like ours except for that one difference, and ask what holds there. The computational answer, arriving decades later, says something sharper: build a model of how the actual causes worked, infer the specific conditions that held at the time, hypothetically alter one variable, and run the model forward again. The first tells you what "true" means for a counterfactual. The second tells you how to compute it. Both treat the counterfactual as parasitic on the actual — the alternative is never seen directly, but it is constrained, sometimes tightly, by what the actual state was and by how the system behind it works.
Where the Analysis Comes From
Nelson Goodman set the problem out formally in 1947. Counterfactual conditionals resist the truth-conditional machinery built for ordinary statements, because their antecedent is false by stipulation — "if the match had been struck" is asserted precisely when the match was not struck — and standard logic makes any conditional with a false antecedent trivially true, which gets the meaning wrong. Robert Stalnaker, in 1968, and David Lewis, in 1973, answered with possible-world semantics: a counterfactual is true if it holds in the world nearest the actual one where the antecedent obtains. This fixed the logic but left "nearest" doing a great deal of unexplained work.
Judea Pearl, from the 1990s, gave the idea computational teeth. Within a structural causal model — a set of variables linked by equations representing mechanism, not mere correlation — a counterfactual is computed in three steps. Abduction: use the observed outcome to infer the values of the background, unobserved factors that were actually in play. Intervention: surgically set the variable in question to the counterfactual value, leaving the rest of the causal structure untouched. Prediction: propagate the altered state forward through the model to see what results. Pearl placed this three-step operation at the top of a ladder of causal reasoning, above mere association and above intervention on its own, because it requires both a working causal model and knowledge of the specific case.
The Intake Requirement Pearl's Recipe Makes Explicit
Once counterfactual reasoning is stated as a procedure rather than a semantics, something becomes visible that talk of "nearest worlds" obscured: each step needs input, and the input has a shape. Abduction needs the actual state, at the moment in question, in enough detail to pin the background variables — not roughly, but at the resolution the model requires. Prediction needs a causal model that has not silently gone out of date. Intervention needs the two to be joined correctly. Miss any one and the recipe still runs, but it runs on guesswork rather than reconstruction.
This is where the lineage from Large Language Model to Large World Model to Large Universe Model turns out to sit on exactly this axis, and the connection was not planned when the ladder was built — it falls out of applying Pearl's own steps to each generation in turn.
A Large Language Model has a corpus frozen at some cutoff. Asked a counterfactual, it can only recall how similar counterfactuals were phrased and resolved in text it has already seen. It has no actual state to abduce from, because it has no "now" — only a fixed date behind it. What looks like counterfactual reasoning is plausible narration dressed in causal syntax.
A Large World Model does better, within limits. It observes an actual scene while that scene lasts — sensor state, positions, a dynamics model good enough to ask what the forklift would have done had it turned left instead of right. Abduction is genuine here: the state is really seen, not recalled from text. But the reach ends at the scene boundary and the session's end. Ask the same model about a state from last month, outside any scene it witnessed, and it has nothing to abduce from, for the same structural reason the language model has nothing.
A Large Universe Model is the position where the actual state is not a scene but a maintained record — every relevant stream still running, held as belief with provenance and a decay function rather than as settled fact. Abduction can then reach any past instant the record covers, not just the present one, because the record carries timestamps and revision history rather than a single frozen snapshot. Intervention and forward integration proceed on a causal structure that is checked against incoming data rather than assumed static. This is Pearl's recipe with a live substrate underneath it, and there is no fourth intake regime beyond it: the actual record is what continuous intake, by definition, is.
Three instances make the pattern concrete rather than abstract. Investigators into the Boeing 737 MAX accidents answered a precise counterfactual — what the aircraft would have done had MCAS not commanded nose-down trim — by re-running recorded flight data at 8 Hz through a validated aerodynamic model with that one input removed: abduction from stored streams, intervention on a single variable, forward integration through mechanism. The Bank of England's fan charts do the same thing for interest rates, with a harder constraint: GDP arrives with a six-week lag and is revised for years, so the abduction step works on a state estimate that is itself still moving. And in Bolitho v City and Hackney, the House of Lords needed to reconstruct what a doctor would have done and whether it would have saved a child, from observation charts recorded only every few hours; the sparseness of that record, not any flaw in the legal reasoning, is what left the counterfactual too thin to support the claim.
Three Objections, Taken Seriously
Counterfactuals are evaluated against a model, not against the world. A correctly specified structural model with decades-old data answers them perfectly well offline. Continuous intake is an aid, not a precondition.
This is true, and it narrows the claim considerably. A counterfactual whose consequent is fixed in the past — what would have happened to the 1980 cohort had exposure been zero — can be answered from a frozen dataset, because the question is frozen too. The requirement bites only for counterfactuals evaluated at or beyond the present: what our position would be now, what a patient's trajectory will be. Those need abduction to reach a live state and a causal model that has not drifted since it was fitted. Historical counterfactuals need good archives. Live ones need continuous intake. The claim holds for the second class, not the first.
The second objection is that continuous observation gives correlation, not causal structure; no volume of passive streaming identifies a causal graph without intervention. This is correct and unresolved by intake alone. What continuous observation buys is not identification itself but more identifying events — outages, staggered rollouts, policy discontinuities — caught with a pre-period baseline because the observatory was already running when they occurred. A frozen corpus catches a handful of these by accident. A live one catches them as a matter of course. Intake multiplies opportunities for identification; it does not manufacture identification from nothing.
The third objection reaches back to Lewis directly: nearness between possible worlds has no principled metric, and no amount of data about the actual world settles which counterfactual world is nearest. This is right as metaphysics, and no amount of intake touches it. What richer intake does is shrink the set of candidate worlds consistent with the known facts, so that the unprincipled similarity ranking has less work left to do. The vagueness survives. Its practical bite shrinks.
The Misreading to Disown
The weak, wrong version of this argument says that a system observing everything can simply see what would have happened. It cannot, and no intake regime changes that, because a counterfactual world is by construction one that did not occur. The defensible claim is narrower and epistemic: intake determines the quality of the abduction step, which fixes the background conditions the counterfactual is computed against. Better observation sharpens the actual state and the model built on it. It gives no window onto the alternative itself.
What This Establishes, and What It Does Not
Continuous intake is necessary for counterfactuals whose answer must hold now, and it is the only intake regime — among the three considered here — that supplies abduction, a live causal model, and forward integration at once. That is what places the Large Universe Model as the terminal rung on this particular axis. It does not establish that such a system reasons correctly, that its causal graph is identified, or that the metaphysics of nearest worlds has been solved. Those problems are older than any model and will outlast this one.