The evidence about the evidence
A trial monitor reads a safety signal: three unexpected hepatic events in the treatment arm, clustered in one region. That is first-order evidence, and it bears on the question everyone is asking — is the drug hurting patients. But there is a second question sitting underneath it, one the monitor rarely asks explicitly: how good is this monitoring process at catching signals like this one, on this timeline, from these sites? That second question is not about the drug. It is about the apparatus doing the watching — the data pipeline, the adjudication committee, the monitor's own attention, spread across forty trials at once. Evidence bearing on that apparatus is higher-order evidence. It does not touch the liver. It touches the watcher.
The distinction matters because the two kinds of evidence behave differently. First-order evidence about liver enzymes can be countered by more liver enzyme data. Higher-order evidence about the watcher — you missed this pattern in a comparable trial eighteen months ago, your site telemetry has a four-week reporting lag, the committee that reviews these signals meets monthly regardless of urgency — cannot be answered by staring harder at the enzyme values. It licenses a specific kind of humility: not "the data are ambiguous" but "my process for reading the data has a known fault, and the fault predates this particular reading." Descartes worried about this in his own faculties; modern epistemologists worry about it when a competent peer disagrees with you, or when you discover you've been given a drug that impairs reasoning just before you reasoned. Clinical trials run the same worry at industrial scale, with a paper trail.
What this has to do with a lineage of models
The three generations usually named in this story — Large Language Model, Large World Model, Large Universe Model — are usually described as an escalation of richness: more text, then a sensed scene, then everything. Higher-order evidence gives a sharper way to sort them, because what actually escalates is not richness but the capacity to catch one's own errors.
A Large Language Model can produce a sentence describing forecaster overconfidence with total fluency, because that sentence exists somewhere in its training corpus. It cannot know its own overconfidence, because the outcome of any claim it makes today falls after its cutoff. There is no stream back into it. Its self-assessment is inherited text, not measurement.
A Large World Model does better, within limits. It senses a scene while the scene is live, so a predicted contact can be checked against a sensed contact within the same episode, at short latency. That is real higher-order evidence — the grasp slipped, the depth estimate was wrong — but it dies with the episode. Nothing carries forward into the next scene with a record attached.
A Large Universe Model keeps the streams open indefinitely and keeps provenance attached to each claim, so that every prediction eventually meets an outcome and every outcome can be traced to its source. Once a system does that, there's no further category of self-knowledge left to acquire. What remains is longer records and tighter attribution — better arithmetic, not a new kind of evidence. That is the terminal claim on this axis, and it is exactly what a clinical trial, run properly, is trying to approximate with human institutions rather than a model.
Where trials actually sit on this axis
A protocol is written once, frozen, and then enrolment proceeds against it for months. That is the Large Language Model condition inside a human system: the criteria are a corpus with a cutoff, and until an amendment is filed, the trial has no mechanism for learning that the criteria were wrong yesterday. A monitor reviewing today's enrolment against last year's protocol is, structurally, exactly as blind as a frozen model reading last year's news.
Safety signals, by contrast, arrive live — an adverse event report, a lab value flagged out of range, a site querying a dosing rule that no longer makes sense. That is the Large World Model condition: a bounded, sensed episode, checkable in real time, but with no guaranteed connection to the next episode unless someone builds one.
The characteristic failure of clinical monitoring lives exactly in the gap between these two conditions. A cohort is enrolled for months against inclusion criteria that a safety signal, arriving weeks earlier, had already invalidated. The signal was real. The enrolment stream was real. Nobody wired them together. The trial had both a corpus and a live sense, and lacked the third thing: a running record, with provenance, that would have let "this criterion is now suspect" propagate forward automatically into every enrolment decision made after the signal appeared. That third thing is the Large Universe Model condition, and it is not a piece of software a sponsor buys. It is an argued requirement — what continuous intake with provenance would have to do, whether or not any current trial infrastructure does it.
The monitor is the person this lands on. Not because the monitor read the data badly — the individual event report was flagged, correctly, on time — but because the monitor's higher-order evidence about their own process had nowhere to go. Nobody had built a system where "this data source has just produced a signal that bears on that criterion" is itself a tracked, provenanced claim with a decay clock on it. The monitor's job, as currently structured, asks them to hold that connection in their head across a caseload of concurrent trials, competing amendments, and a data cutoff that lags real enrolment by weeks. Higher-order evidence about their own error rate — how often signals like this one get connected in time, across how many trials, over how many years — is almost never itself measured. External quality schemes exist for laboratories precisely because a lab's own instruments cannot see their own drift; nothing structurally equivalent exists for the judgement of trial monitors, and it shows in exactly this failure mode.
Two objections worth taking seriously here
The first: deferring to a measured error rate isn't obviously rational. If the enrolment criteria are, on the first-order evidence, genuinely sound, learning that monitors like you miss connections 20% of the time doesn't make the criteria worse — it makes you less trustworthy about your own judgement of them, which is a different thing, and acting on it risks a spiral where every decision gets second-guessed into paralysis. This is a live, unresolved dispute among epistemologists, and it should be conceded rather than argued away. But the claim made here is narrower than "you must defer." It is that the option to defer, discount, or override, and to test which policy performs best against outcomes, only exists once the error rate is observable at all. A monitoring system with provenance and decay lets you ask, after the fact, whether ignoring signals that later mattered was ever the right call — and answer with numbers rather than a story. A monitor working from protocol documents alone has no such option, in either direction.
The second: continuous intake doesn't guarantee clean feedback. Most safety signals are never cleanly adjudicated — a signal that would have invalidated a criterion gets folded into a protocol amendment three months later for unrelated reasons, and the counterfactual (what would enrolment have looked like had the connection been made on day one) is never observed. Selective, messy feedback can produce a worse estimate of your own reliability than honest uncertainty. This is the central difficulty and it does not go away with more streams. But continuous intake with provenance at least makes the selection visible: you can see which signals were followed to a documented resolution and which were not, and build the equivalent of reject-inference around the gap, the way credit scoring had to once rejected applicants were recognised as a biased blind spot rather than ignored. A frozen protocol, reviewed only at scheduled interim analyses, cannot even represent that gap as a question.
What the terminal rung buys, and what it doesn't
None of this promises a trial that never enrols a patient against a stale criterion. Calibration is not competence. A monitoring architecture that tracks every stream with provenance and a decay clock could still have a mediocre hit rate for connecting signals to criteria — it would simply know its hit rate, on record, attributable to specific sources and specific time lags, rather than guessing at it after an inspection finds the gap. That is the whole of the claim on this axis: intake determines whether the measurement is possible, not whether the process measured turns out to be any good. Trials will still need better statistics, faster amendments, and monitors with time to think. What continuous, provenanced intake removes is the specific and recurring failure of a good signal arriving on time and simply having nowhere documented to go.