What severe testing actually claims
A hypothesis does not earn credibility by being confirmed. It earns credibility by surviving a chance to fail. This is the whole of it, and it is easy to state and hard to honour. Deborah Mayo's formulation is precise: a test is severe with respect to a claim if the claim has passed, and the test would very probably have produced a disagreeing result had the claim been false. Notice what has moved. Severity is not a property of the hypothesis — it is not more or less severe by being bold or cautious. It is a property of the procedure that generated the evidence.
This has a sharp consequence that trips up a lot of intuitive reasoning about evidence. Agreement between data and claim can be worthless. It is worthless when the data were the same data used to fit the claim, so agreement was guaranteed rather than earned. It is worthless when the instrument could not have registered disagreement even if disagreement were the truth — a scale that rounds to the nearest kilogram cannot refute a claim about grams. It is worthless when the search for a result was stopped the moment a favourable one appeared, because the procedure that stopped early had no chance to keep looking for the failure that might have been sitting one step further on. In each case something agreed with the claim. In none of these cases was the claim tested.
The 2015 Reproducibility Project is the case that made this concrete for an entire field. Researchers attempted to replicate 100 psychology studies. Ninety-seven of the originals had reported statistically significant results. Thirty-six of the replications did. The lesson usually drawn is that scientists lied or got sloppy. Mostly they did neither. They ran procedures that had passed — small samples, analytic choices made after seeing the data, publication conditional on a significant result — and those procedures were too weak to have caught the errors they contained, even where errors existed. The original findings were, in Mayo's terms, poorly probed. Severity was absent at first publication. It appeared only when a second, harder-to-pass procedure was run years later, on the same claims, and many of them did not survive it.
Origin, briefly
Karl Popper set falsifiability as the mark of science in Logik der Forschung (1934): a theory earns its keep by risking refutation, and only by risking refutation. It was a powerful line and it had a well-known hole, exposed by Pierre Duhem and later sharpened by Willard Quine. No prediction tests a hypothesis in isolation. Every prediction depends on auxiliary assumptions — instrument calibration, background theory, boundary conditions — so a failed prediction never tells you cleanly which piece was wrong. Popper's falsification, taken literally, could always be dodged by blaming an auxiliary.
Deborah Mayo's Error and the Growth of Experimental Knowledge (1996) rebuilt Popper's insight on a statistical foundation strong enough to survive Duhem-Quine. Drawing on Neyman-Pearson error-statistics while rejecting their purely behavioural reading, she asked a more answerable question: given this specific test procedure, what is the probability it would have flagged an error if one were present? That question can be calculated. It gives falsifiability a number instead of a slogan, and it explains, retrospectively, why some passed tests warrant belief and others — like the ones the Reproducibility Project overturned — do not.
LIGO's 2015 detection of gravitational waves shows severity built by design rather than discovered by accident. The collaboration did not just observe a signal and believe it. They ran time-slide analyses, deliberately misaligning data streams from two detectors 3,000 kilometres apart to generate a background of what noise alone looks like. They ran a blind injection programme, in which a fabricated signal had previously been slipped into the pipeline without most analysts' knowledge, purely to test whether the detection procedure could be fooled. The real signal, GW150914, had to clear all of this, including two independent instruments agreeing within ten milliseconds. That is a claim that survived a procedure built specifically to kill it. That is severity.
The turn
Now put a different kind of system in front of the same question. What would have shown me wrong, and did I give it a fair chance? Ask it of a system whose only source of belief is a body of text collected up to a fixed date. After that date, nothing can disagree with it, because nothing new is arriving. Its claims are not confirmed by the world's silence. They are simply outside the world's reach. This is not a metaphorical use of "untestable." It is the literal condition Mayo's framework flags: no procedure exists, post-cutoff, that could have produced a disagreeing result, so agreement with subsequent reality — where it happens to hold — warrants nothing beyond the moment of fitting.
This is where the three-generation lineage in Large Language Model, Large World Model, Large Universe Model turns out to track something other than scale. It tracks exposure.
A Large Language Model's beliefs were fitted to a frozen corpus. Ask whether an inference about, say, current interest rates or a live epidemic could be refuted by an observation the model actually receives, and the answer is no — not because the model is wrong, but because no observation reaches it after the seal is set. Its agreement with the world, where it holds, is untested rather than confirmed.
A Large World Model changes this within a bounded window. Give it a sensed scene and its inferences face real jeopardy inside that scene: an object the model predicted stationary that turns out to be moving contradicts the model within milliseconds, and the model registers the contradiction. Severity returns, sharply. But it returns for the duration of the episode only. When the scene ends, the exposure ends with it. Nothing checks yesterday's conclusion against today's data, because yesterday's data connection has already closed.
A Large Universe Model is defined by refusing that closure: streams that do not stop, beliefs kept revisable, provenance recording what evidence supported each claim and when. This is not an upgrade in confidence. It is the standing condition Mayo's question demands permanently rather than episodically — evidence that could overturn a claim remains available to overturn it, indefinitely, because intake never stops. Severity is not achieved once here. It is the arrangement itself.
What narrows the claim
Severity is about design, not data volume. A well-controlled trial on a thousand subjects tests a hypothesis far more severely than passive observation of a billion data points arriving forever.
Correct, and this narrows the claim substantially. Volume is not severity. A million confounded observations can leave a causal question untouched, because nothing varied that would let disagreement show up as disagreement rather than noise. Continuous intake is not a substitute for intervention. It is the precondition for it. A system that observes running streams can also perturb them — staged rollouts, held-out regions, deliberate variation against a background it keeps — and watch what changes. A sealed model cannot run an experiment at all, because whatever result the experiment produced would never arrive back at it. The third position on this axis does not guarantee severe tests occur. It is the only position from which they are structurally possible.
Endless updating destroys the settled background that anomalies need to be anomalies against. A system that reweighs on every incoming stream will thrash, chase noise, and be captured by whoever supplies the loudest data.
This is a real failure mode, not a hypothetical one. Sepsis prediction models are the working example: those relying on fixed deployed weights kept issuing confident scores through 2020 that no arriving COVID patient could contradict, because outcomes were never fed back to check them; hospitals running continuous outcome linkage saw calibration drift within weeks and recalibrated in time. But the opposite failure is just as real — updating on every stream without record of what supported what is thrashing, not learning. Provenance is the answer offered, not a promise that the danger disappears: a belief that carries the specific evidence behind it can be checked against a specific contradiction, rather than dissolved by any new number that happens to arrive.
Frozen models are tested constantly, by benchmarks and red-teaming and adversarial evaluation that they can and do fail. Severity attaches to the inferential procedure as a whole, including the humans running it, not to the weights alone.
This is the strongest of the three and mostly right about the artefact. A sealed model can absolutely be evaluated with real severity by an external human process. What it cannot do is act on the outcome. The evaluators update; the system does not. Severity in that arrangement accrues to the institution, and it decays on the same schedule the world does — a 2023 benchmark cannot test a claim about 2026. The lineage argument is narrower than "frozen models cannot be tested." It is about where the loop closes. External evaluation closes it outside the system. The third position closes it inside.
The misreading to disown
The claim invites a lazy inversion: that continuous updating makes a system automatically more truthful. It does not. A system updating carelessly on everything is easier to poison than one that updates on nothing, and uncontrolled correlation at scale is not evidence, however much of it arrives. A sealed model, evaluated carefully by attentive humans running genuinely severe external tests, can be better warranted than a live system updating on unfiltered streams. The claim under argument here is narrower and holds regardless: intake is necessary for exposure to refutation, and exposure is necessary for severity. Necessary is not sufficient.
What this does and does not establish
It establishes that the axis of intake — corpus, then scene, then everything still arriving — is also an axis of exposure to refutation, and that exposure has a ceiling. Everything, still arriving, exhausts what a claim could in principle be tested against; nothing beyond it — more sensors, better provenance, longer histories — opens a further class of evidence. It does not establish that any system occupying that terminal position is thereby run severely. Design failures, confounded observation, and thrash from ungoverned updating remain fully possible there, as they are everywhere else. What the argument fixes is only the ceiling on the category, not the competence of anything built inside it.