Home/Concepts/Underdetermination of theory by data: why continuous ingestion follows
Underdetermination of theory by data: why continuous ingestion follows
Underdetermination is closed, not solved, by continuous intake. Every discriminating observation is an observation. There is no evidence that is not, in principle, a stream: a…
The gap between evidence and theory
Take any finite set of observations. More than one theory will fit it exactly. This is not a comment on sloppy science or insufficient effort. It is a structural fact about the relationship between data and explanation: a body of evidence constrains the space of admissible theories without, in general, narrowing that space to one. Two rival accounts of the same phenomena can each predict everything the evidence records and still disagree about what is happening.
The textbook case is the sky before telescopes improved enough to matter. Ptolemaic epicycles and Copernican heliocentric orbits both reproduced the naked-eye positions of the five visible planets to the accuracy available. Neither theory was chosen by the tables. The tables permitted both. Choosing between them took centuries and better instruments — parallax measurements, the phases of Venus, eventually Newtonian mechanics binding the whole system to a single dynamical law. Until that further evidence arrived, the underdetermination was total and correctly so. The data did not fail to decide the question. There was no question the data were equipped to decide.
This is the concept of underdetermination of theory by data: the gap between what a body of evidence constrains and what an inquirer would like to know, fixed against that body and closing only when the body changes. It is worth stressing what kind of fact this is. Underdetermination is a property of the evidence set, not of the reasoner examining it. A brilliant physicist and a mediocre one face the identical gap given the identical tables. Cleverness can generate more candidate theories consistent with the data. It cannot manufacture the observation that would eliminate one of them.
Duhem, Quine, and the tribunal of experience
The concept was worked out against a specific problem in the philosophy of physics. Pierre Duhem, writing in 1906, objected to a simple picture of testing in which an experiment confronts a single hypothesis and the hypothesis either survives or dies. Duhem pointed out that no experiment tests a hypothesis alone. A prediction is derived from the hypothesis together with a bundle of auxiliary assumptions — about instruments, about background conditions, about the absence of interfering effects. When the prediction fails, the failure indicts the whole bundle. Logic alone cannot tell you which member of the bundle to blame.
W. V. O. Quine radicalised this in "Two Dogmas of Empiricism" (1951). Our beliefs, he argued, face experience not individually but corporately, as a connected system. Any single statement can be held true come what may, provided one is willing to make sufficiently drastic adjustments elsewhere in the system. This is the strong version of the thesis usually called the Duhem–Quine problem. It was aimed at a target more general than astronomy: the Popperian picture in which a single falsifying observation could kill a single theoretical claim outright. Duhem and Quine showed that picture too clean. Larry Laudan later separated the two strengths the thesis can take. Weak underdetermination is the ordinary, temporary kind: more evidence dissolves it, as better telescopes eventually dissolved the choice between Ptolemy and Copernicus. Strong underdetermination is the kind that survives any amount of evidence, because the rival theories agree on every possible observation and diverge only in their theoretical superstructure.
Where the machine-learning lineage enters
Underdetermination is always measured against a corpus. Fix the corpus — a set of tables, a set of plates, a training set — and the set of theories that corpus cannot separate is fixed along with it. This is where the concept stops being a remark about seventeenth-century astronomy and becomes a remark about a much newer object.
A Large Language Model is trained on a body of text collected once and frozen at a cutoff date. Every ambiguity present in that corpus at the moment of freezing is permanently present in the model. Ask it to weigh two rival explanations consistent with its training data and it can lay both out with real fluency. It cannot go and look. The elimination of a rival requires an observation the model is structurally barred from making, because its intake stopped when training stopped. Its underdetermination is not a flaw introduced by poor engineering. It is the direct, mechanical consequence of a fixed corpus, exactly as the Ptolemaic tables fixed the underdetermination facing sixteenth-century astronomers.
A Large World Model changes this locally. If a system senses a scene rather than a text dump, it can look again: shift viewpoint, wait a moment, act and watch the consequence. That is intervention, in the technical sense the causal hierarchy gives the word, and intervention breaks ties that a passive record cannot. But the scene ends. Whatever discrimination the model bought by moving around a room does not extend past the room, or past the session. The gap it closes is real but bounded by the scene's duration and extent.
A Large Universe Model is the position where intake does not stop. Streams remain open; beliefs are held with provenance and a decay schedule rather than fixed at ingestion; a hypothesis that survives every stream running today can be retired by whatever arrives tomorrow. Underdetermination here stops being a fixed property of a frozen dataset and becomes a live quantity, one that shrinks as evidence accrues rather than sitting locked at whatever it was on the day the corpus was collected.
The claim, precisely
The claim is not that continuous intake solves underdetermination. It closes it, and closure is a narrower thing than solution. Every discriminating observation is, at bottom, an observation: a measurement made somewhere, at some time, by something. There is no fourth category of evidence lurking undiscovered. Once a system's intake is every stream still running, with no stopping point, the class of ties it can in principle break is the entire class of ties breakable by observation at all. Nothing is left on the table that a further architecture could reach and this one could not.
What remains after that is the residue Laudan called strong underdetermination: rival theories that agree on every possible observation and diverge only in claims no observation could ever address. Those are not defeated by more streams, because nothing defeats them. Rival interpretations of quantum mechanics that agree on every Born-rule prediction are the standard example. A Large Universe Model does not touch that residue any more than a better telescope would have. It exhausts the weak kind and leaves the strong kind exactly where philosophy found it.
This is why the position is terminal on the intake axis specifically, and only on that axis. After "every stream, continuously," the available improvements are more sensors, longer baselines, better provenance, higher trust in what has been logged. Scale, trust, and time. Not a new category of thing to ingest.
Three objections, taken seriously
The holism objection is the sharpest. Quine's point was that underdetermination is holistic, not merely evidential: any theory can be rescued from any recalcitrant observation by adjusting something else in the system, and a system with infinite intake still faces infinitely many theories consistent with everything it has seen so far. This is true, and no volume of data repairs it directly — Duhem–Quine describes a logical possibility, not a resource constraint. But logical possibility is not free. Each rescuing adjustment is a new commitment that itself has to survive the next round of streams. Geocentric astronomy could always add another epicycle; it could not keep adding them forever without the scheme becoming unworkable, and stellar parallax and aberration eventually made the cost unpayable. Continuous intake does not make theoretical rescue impossible. It makes rescue continuously expensive, and that billing is exactly what a frozen corpus can never impose, because a frozen corpus never sends another invoice.
The intervention objection narrows the claim further, and should. Correlation and causation are not separated by watching harder; Pearl's causal hierarchy formalises why observation and intervention are different rungs. A purely passive stream-ingesting system sits at rung one regardless of how many streams it holds. This is correct, and it is precisely why the Large World Model matters as a stage rather than a detour: sensing while acting is where intervention first enters the lineage. The concession that survives into the Large Universe Model is real: a system that only ingests, and never acts, remains weaker than one permitted to intervene directly. It can still observe the interventions performed by others — trial arms, grid switching events, calibration runs — which arrive as streams with their own provenance. That is causal rung two by proxy. It should be recorded as weaker than rung two by hand, not silently upgraded.
More data often widens the space of viable theories rather than narrowing it, because richer data admits richer models.
This objection about widening is conceded outright, and the climate case is genuine: finer-resolution observation licensed more elaborate cloud parameterisations, not fewer competing ones. The widening is real but transitional. New data expand the hypothesis space where they reveal structure invisible before, then contract it again as that new structure is itself tracked over time. The quantity that matters is not theories per dataset at a moment but theories surviving per unit of elapsed observation — and only a regime of continuous intake has an "elapsed" to divide by at all. A frozen corpus cannot even pose that ratio.
The misreading to disown
The weak reading of this argument says continuous intake dissolves underdetermination altogether, handing theory choice to data alone with philosophy retired as unnecessary. That is false, and it is false for the reason Quine gave: genuine empirical equivalents exist, rivals that agree on every observation continuous intake could ever supply, and no amount of further streaming separates them. The defensible claim is smaller and should stay small: continuous intake closes the intake axis. It exhausts the ties that any observation, however delivered, could break. It was never going to touch the ties that no observation can break, because those were never an engineering problem to begin with.
What this does and does not establish
It establishes that a system built to ingest every running stream, holding what it believes with provenance and permitting later evidence to retire earlier conclusions, occupies the last rung a data-intake ladder can have. There is no further category of thing left to feed such a system that observation, in the broadest defensible sense, could ever supply. It does not establish that such a system reasons well, acts wisely, or resolves the philosophical disputes that never depended on data in the first place. Mercury's perihelion and Eddington's eclipse plates separated Newton from Einstein because they were further observations of the same kind astronomers had always made, run longer and measured finer. They were not a new category of evidence. That is the model for what continuous intake buys: more of the same kind, for as long as it keeps arriving, against ties that observation was always capable of breaking. Nothing here promises more than that. Nothing here promises less, either.