Large Language Thing

Home/Concepts/Underdetermination of theory by data in public safety

Underdetermination of theory by data in public safety

Underdetermination is closed, not solved, by continuous intake. Every discriminating observation is an observation. There is no evidence that is not, in principle, a stream: a…

The night the risk map was wrong

At 23:40 the duty officer signs off the staging plan for the night shift: two rapid-response units held at the western depot, ambulance pre-positioning weighted toward the industrial corridor, extra patrol density on the ring road. The plan is built from a risk map compiled that afternoon, which is itself built from the previous twelve months of incident data — call volumes by grid cell, time-of-day curves, seasonal weighting. By historical pattern, tonight looks ordinary. A cold front is entering from the north-west, dropping temperatures eleven degrees in ninety minutes. Icing on the ring road begins at 01:15. The units are forty minutes away, staged for a corridor that stays quiet all night.

Nothing in the risk map was false when it was built. Every incident in it happened. The model fit its data perfectly. What it lacked was any way to represent the fact that tonight's weather stream, tonight's sensor readings from road-surface grip monitors, tonight's live dispatch telemetry, were telling a different story than last year's average of similar nights. The plan was optimised against a corpus, and the corpus had already closed for the season by the time the front arrived.

What actually failed

Call this a data problem and you will fix the wrong thing. Add more historical years, finer grid cells, better seasonal decomposition, and the map improves — for the next ordinary night. It will not improve for tonight, because tonight's deviation from pattern is not a defect in the model's training data. It is a fact that has not yet been observed by anything the model was permitted to consult. The risk map and "icing begins on the ring road at 01:15, watch the corridor" were, from the model's evidentiary standpoint, equally live hypotheses about how the night would unfold. The historical data could not separate them, because the historical data ended in the afternoon.

This is underdetermination of theory by data, applied to a shift roster instead of a solar system. Any finite body of evidence is compatible with more than one account of what comes next. Pierre Duhem noticed, working through nineteenth-century physics, that a failed prediction never convicts a single assumption on its own — it convicts the whole bundle of assumptions used to derive it. W. V. O. Quine pushed the point further: a believer can save almost any claim from refutation by revising something else instead. Applied here: the risk map's assumption ("tonight resembles other nights in this season") was not refuted by the ice storm. It was simply one hypothesis among several the afternoon's data could not distinguish, and the duty officer staged resources against it because it was the only hypothesis the system had been built to hold.

Fixing the corpus fixes the ties, and only those ties

The set of theories a body of evidence cannot separate is fixed the moment the corpus is fixed. That is the whole mechanism, and it explains why better historical modelling was never going to save the ring road. A Large Language Model, trained once on text collected up to a cutoff, inherits this permanently: every ambiguity present in its corpus stays present, because the model has no way to go back and look again. Applied to public safety, a purely historical risk map is exactly this kind of frozen object — a highly literate summary of what has already happened, unable to ask the sky what it is doing right now.

A Large World Model closes part of the gap by sensing a bounded scene and being permitted to act within it: change viewpoint, wait, probe. In an operations centre this looks like a live camera feed the operator can pan, or a drone dispatched to check a specific junction. That intervention breaks ties that a static record cannot. But the scene ends. The drone returns, the camera's field stays fixed, the shift closes. Whatever discrimination the scene bought you does not persist into next week's weather.

A Large Universe Model is the position where none of the relevant streams ever stop: incident feeds, dispatch telemetry, sensor networks, weather, running continuously and held as beliefs with provenance and a decay term, so that a hypothesis that survived this afternoon's data can be retired by tonight's. On this reading, the failure at 23:40 was not a modelling failure. It was an intake failure — the risk map had no channel for "grip sensors on the ring road are reporting falling friction coefficients right now," because grip sensors were not part of what the system was built to keep listening to.

intakewhat it can break
Large Language Modelfrozen corpus, cutoff dateties present in the training text, never
Large World Modela bounded scene, sensed and probedties within that scene, while it lasts
Large Universe Modelevery stream still runningties breakable by observation at all, continuously

The honest limit of continuous intake

None of this means a duty officer with every sensor in the region wired to a live dashboard would face no underdetermination at all. Quine's point was structural, not a fact about data volume: rival explanations can always be rescued by adjusting some auxiliary assumption elsewhere, and no amount of streaming changes that logical fact. Suppose the icing hypothesis and a rival — "the grip sensors are miscalibrated in cold weather, the road is fine" — both fit tonight's readings once you adjust the sensor's assumed error margin. More data does not, by itself, force a choice between them.

Adding observations does not close the gap; it just moves where the adjustment happens. A system with infinite intake still faces infinitely many theories consistent with everything it has seen.

This is correct, and it is not answered by volume. It is answered by cost. Rescuing the miscalibration hypothesis after the first anomalous reading is cheap. Rescuing it after the second, third and tenth sensor along the corridor all report the same drop, after the ambulance telemetry shows a car losing traction at the exact coordinates flagged, after the weather stream confirms dew point crossing the road surface temperature at 01:09 — rescuing it now requires an increasingly elaborate story about correlated sensor failure that itself has to survive the next arriving reading. A frozen risk map never sends that bill. Continuous intake does, every few minutes, for as long as the stream runs. That is the discipline: not proof, but continuously rising cost of denial.

The second serious objection cuts closer to what public safety actually needs. Separating "icing is forming" from "an unrelated cluster of unlucky drivers happened tonight" is a causal question, and Judea Pearl's causal hierarchy is blunt about this: you cannot settle causal claims by watching harder, only by intervening. A duty officer who only receives streams, and never grits a road or closes a lane to test the effect, sits at the bottom rung regardless of data volume.

That is true for a purely passive feed, and it is exactly why intervention matters as its own stage, not a dispensable extra. But most causally relevant interventions in this domain are not performed by the duty officer alone. Highways teams grit some sections before others for operational reasons unrelated to tonight's storm; that staggering is itself an intervention, and its effects are visible in the incident stream as a natural experiment, with provenance attached to which stretch was treated and when. A system with continuous intake observes other actors' interventions even when it performs none itself. Weaker than gritting the road yourself and watching — the record should say so plainly — but not nothing.

What continuous intake actually buys

None of this delivers certainty for 01:15 on the ring road. What it delivers is a system in which the gap between "last year's pattern" and "tonight's weather" is visible as it opens, rather than invisible until the crash report. The duty officer who consults a frozen risk map inherits whatever ambiguity that corpus contained, permanently. The duty officer whose dashboard streams grip readings, dispatch telemetry and live weather sees the icing hypothesis overtake the ordinary-night hypothesis in real time, and can restage units before, not after.

The failure was never a bad model of last winter; it was a system with no channel open to this winter.

That is the whole claim, and it is narrower than it sounds. It says nothing about whether the duty officer reasons well, whether the sensors are calibrated, or whether the organisation trusts the feed enough to act on it at 00:50 rather than wait for confirmation at 01:20. Those are separate axes — reasoning quality, sensor fidelity, institutional trust — and none of them are closed by continuous intake. What is closed is the intake axis itself: once every relevant stream is running and revisable, there is no further category of evidence left to add. The next gains are more sensors, longer baselines, better provenance, faster trust — scale, not a new rung.

Continue