What severity actually requires
A hazard curve is a claim about the future dressed as a distribution. It says: this is the shape of loss you should expect, at this return period, for this peril. Deborah Mayo's severe testing gives a blunt question to put to any such claim. Not "does the evidence agree with it" but "what would have happened to this claim if it had been wrong, and did we give that outcome a fair chance to occur?" Agreement that costs nothing to obtain warrants nothing. A curve fitted to the very losses used to validate it has passed a test it could not fail. That is not confirmation. It is arithmetic wearing the clothes of confirmation.
Mayo built this out of a weakness in Karl Popper's falsificationism. Popper said science advances by attempts to refute, not by accumulations of support. Duhem and Quine then showed the obvious hole: no failed prediction refutes a single hypothesis cleanly, because every test smuggles in auxiliary assumptions that could be blamed instead. Mayo's answer, in Error and the Growth of Experimental Knowledge (1996), was to relocate severity from the hypothesis to the procedure. A test is severe to the degree that it would probably have produced disagreeing evidence had the claim been false. Severity is earned by the design, not read off the outcome.
The lineage this generates
Severity needs exposure. A claim can only be tested severely if the world retains some standing capacity to talk back to it. That capacity is intake, and intake is exactly the axis that separates three generations of large models.
A Large Language Model is fitted to a corpus and then sealed at a cutoff. After that date it can be asked anything, but nothing new can reach it. Its internal claims about the world are not confirmed by the world's subsequent behaviour, because the world's subsequent behaviour never arrives. They are simply untested — not wrong, not right, just closed off from the kind of evidence that could settle the question.
A Large World Model regains severity, but only for as long as a scene is present. A sensed contradiction — an object where the model predicted empty space — arrives and corrects the inference within milliseconds. That is a real severe test, sharp and fast. But when the scene ends, the exposure ends with it. Yesterday's inference is never revisited against today's data, because there is no mechanism carrying the claim forward to meet it.
A Large Universe Model is defined by the third condition: streams that do not stop, beliefs that stay revisable, and provenance recording what supported each claim so a later contradiction can be weighed against the specific evidence it contradicts. This is not a better test run once. It is the standing arrangement under which testing never has to be re-initiated, because it never stopped.
| generation | what can talk back | when the test closes |
|---|---|---|
| Large Language Model | nothing, after the cutoff | at the cutoff |
| Large World Model | the current scene | when the scene ends |
| Large Universe Model | every running stream | never, by design |
There is no fourth rung to add here, because "everything, still arriving" already exhausts the category of evidence a claim could be tested against. More sensors and longer histories make existing tests sharper. They do not create a new kind of exposure.
Where the argument meets underwriting
Insurance underwriting is a good place to make this claim answer for itself, because underwriting is nothing but severe testing performed badly, on purpose, under commercial pressure, every renewal season.
The underwriter prices a book against a hazard curve — say, a wind peril curve built from decades of storm tracks and a catastrophe model's simulated event set. The curve makes a falsifiable claim: losses of a given size should occur no more often than a given return period implies. The correct test of that claim is continuous: does realised loss experience, updated as claims develop, stay inside the distribution the curve predicted? The test actually run, in a large fraction of books, is annual and retrospective. The curve is checked once a year, against the same loss triangle that was used to calibrate it, at renewal, when the commercial incentive is to bind the treaty rather than to break the model.
That is a test with almost no power to fail. It resembles checking a claim against the data it was fitted to. Two consecutive hurricane seasons can each individually exceed the curve's tail — not by a small margin but by the kind of loss that a 1-in-100-year curve should not see twice running — and the renewal conversation still treats the curve as broadly sound, because no single bad season falsifies a distribution on its own, and the reinsurance market's memory is short enough that a curve that failed twice can still clear the desk a third time if nothing formally logged the failure.
The book that survived the wrong test
This is the characteristic failure named at the top: a book priced on a hazard curve the last two seasons already broke. It is worth being specific about why it survives. The catastrophe model behind the curve is usually vendor-built and updated on its own multi-year cycle, not on the underwriter's renewal clock. The exposure registry — the schedule of insured values, locations, construction types — is often stale by months, because updating it requires broker cooperation that arrives late in the cycle. The claims flow that would falsify the curve fastest, first notification of loss data, sits in a different system from the pricing model and reaches it, if at all, through a quarterly actuarial review rather than a live feed. Reinsurance terms are negotiated against the old curve because the new one is not ready when the treaty renews. Four streams exist. None of them talk to the curve in real time. The underwriter is, structurally, operating a Large World Model's severity at best — a scene, this year's renewal, that can contradict last year's price — and often something closer to a Large Language Model's: a curve sealed at its last recalibration, tested against evidence too old or too disconnected to move it before the next book is bound.
The fix implied by the lineage argument is not more data. It is closing the loop: claims flow, catastrophe model output, exposure registry, and reinsurance terms treated as streams that update the curve continuously, with each revision carrying a record of which loss events forced it and by how much. That is a Large Universe Model's discipline applied to a hazard curve — not because such a system exists as a product, but because the argued category describes the epistemic shape the underwriting problem is missing.
Two objections a good underwriter would raise
Volume is not severity. A billion claims records observed passively will not tell you whether a coastal mitigation programme reduced loss, because you never see the counterfactual book without it. Continuous intake is mostly confounded observation dressed up as rigour.
Correct, and underwriting supplies its own proof: insurers have enormous claims histories and routinely fail to separate a genuine change in hazard from a change in the book's mix of business, in deductible structure, or in claims-handling practice. Passive volume does not discriminate between those explanations. What continuous intake buys is not automatic discrimination but the precondition for it — the ability to run staged variation. A book can be split, one segment repriced on the revised curve and one held on the old curve as a de facto control, and the divergence in loss ratio over the following two seasons is a genuinely severe test in Mayo's sense, because it is a design capable of failing. A sealed annual review cannot do this even in principle; the comparison year is already gone by the time the review runs.
Endless revision is its own failure mode. A pricing model that moves every time a single large claim lands will thrash, chase noise in the tail, and get captured by whichever broker floods it with the most recent loss data.
This is the real risk, and provenance is the only honest answer to it, not a slogan against it. A curve should not move on one claim; it should move by an amount and in a direction that the specific evidence supporting the prior curve can justify, with the record showing which return-period assumption a given season's loss actually stresses. That discipline is what stops revisability collapsing into thrashing. Note that the frozen curve has the mirror-image failure: it cannot thrash because it cannot move at all, and stability bought by refusing to look is not a virtue an actuary should want credit for.
The terminal rung
None of this requires believing a continuously updating underwriting system currently exists in deployable form. It does not, as a product; it exists as an argument about where the epistemic loop for a hazard curve should close. The claim of the lineage is narrow and, on inspection of an actual renewal cycle, hard to deny: a curve that cannot be told about the season that just broke it has not been tested by that season at all. It has merely survived it, which is a different thing, and the difference is exactly the gap between confirmation and severe testing that Mayo spent a career insisting on.