Streams on a running trial
A phase III trial does not present itself as a dataset. It presents itself as five or six streams that never stop: enrolment against inclusion/exclusion criteria, adverse event reports arriving from sites in whatever order sites choose to file them, laboratory values flowing back from central labs on a rolling basis, protocol amendments issued by the sponsor, and site-level telemetry — deviations, queries, drop-out notices, monitoring visit findings. None of these streams has a natural stopping point before the trial's own end, and the trial's own end is itself a decision made partly on the evidence of the streams. This is the condition a Large Universe Model is built for: belief maintained against inflow that keeps arriving, not belief fitted once to a corpus that has already closed.
Contrast a Large Language Model, trained once against a frozen corpus, where surprisal is squeezed out until the gradient has nothing left to do. Contrast a Large World Model, which predicts the next sensor frame while a scene is open and stops mattering once the scene ends — a single surgical procedure, a single driving segment. A trial has no such boundary. Enrolment for a cardiovascular outcomes study can run three years; safety surveillance continues after the last patient is dosed. The horizon that closes a Large World Model's scene never closes here. That is the structural reason the trial's operating logic has to treat surprisal as an input rather than a loss to eliminate.
What is held
At any moment the trial holds a set of beliefs, each with a number attached and a reason for the number. The expected serious adverse event rate for the study drug, taken from phase II and adjusted for the current population, might sit at 2.1% over six months. The expected screen-failure rate against a given inclusion criterion might sit at 30%. The expected rate of a particular liver enzyme elevation might sit at 0.4%, drawn from class-level pharmacovigilance data for similar compounds. Each of these is a probability distribution, not a fact, and each carries provenance: which prior trial it came from, which subgroup it was fitted on, when it was last checked against accruing data. This is the ledger a trial monitor is meant to keep current. It is rarely kept current, because keeping it current is exactly the labour that continuous intake demands and episodic review does not.
What arrives, and what it costs when it isn't read
Consider a concrete run of events. Month one: forty patients enrolled under inclusion criteria that require a baseline biomarker above a threshold, on the belief that the drug's benefit is concentrated in that stratum. Month four: a safety signal arrives from the data monitoring committee's interim look — an unexpected clustering of a moderate adverse event in patients whose biomarker sits at the low end of the eligible range, six events against an expected two, a probability under the standing model of roughly 1.3%, which is 6.3 bits of surprisal. Under Claude Shannon's 1948 formalism, self-information is minus the logarithm of the outcome's probability under the model in force; an outcome expected at 99% carries about 0.014 bits when it happens, an outcome expected at 1% carries 6.64 bits. Six-point-three bits on a stream that has been running near baseline for months is not noise. It is a belief coming loose from the world.
The correct response, in principle, is immediate: revise the estimate of who benefits and who is harmed, and reconsider the eligible range. In practice, on many trials, that revision does not happen for months. The interim signal is logged by the data monitoring committee, discussed, perhaps flagged for a protocol amendment, and the amendment itself takes weeks to draft, route through the sponsor, and clear the ethics committees at each site. Meanwhile enrolment continues against the unrevised criteria, because enrolment is a separate operational stream from safety review, run by separate staff on separate timelines, and nothing forces the two to reconcile in real time. The characteristic failure of this domain is exactly this: a cohort enrolled for months against criteria a safety signal has already invalidated. The cost is not abstract. Every patient enrolled in that window is exposed under a model the trial's own data already contradicts, and every one of them has to be accounted for, sometimes unwound, when the amendment finally lands.
What triggers revision
The trigger is not "an adverse event occurred." Adverse events occur constantly and mostly mean nothing; that is what the baseline rate is for. The trigger is that an event's probability under the standing model is low enough that continuing to hold the model unchanged costs more, in expected harm or wasted enrolment, than revising it costs in disruption. This is a threshold decision, and trials increasingly formalise it — a Bayesian stopping boundary, a sequential probability ratio test, a prespecified surprisal threshold on a monitored endpoint — but the formalisation only works if someone is actually computing surprisal against the current model rather than against the protocol as written eighteen months earlier. That is the specific failure mode: the streams update, but the model they are being checked against does not, so the same six-bit event that should trigger revision instead gets absorbed as "within expected variation" measured against a model already known to be wrong.
What the trial monitor sees
The honest description of the role is that the trial monitor is the provenance keeper for a set of beliefs under continuous assault from four or five uncoordinated streams. What they see, when the system works, is a dashboard that ties each live number back to the belief it bears on and the last time that belief was checked: enrolment rate against the current eligibility model, event rate against the current safety model, query rate against the current data-quality model, each with a surprisal score rather than a raw count. A 6-bit reading on the safety stream should visually outrank a 0.1-bit reading on the enrolment stream even though both are "events." What they see when it fails is a set of separate reports — a safety summary, an enrolment tracker, a deviation log — each internally consistent and none of them flagging that the safety summary has quietly invalidated an assumption baked into the enrolment tracker three months prior.
Two objections worth taking seriously
The first: this is just statistical process control with a clinical label on it, and CUSUM charts have flagged deviations from expected rates for fifty years. That is true as far as it goes. A CUSUM chart on adverse event rates is exactly what a well-run trial should already have. What that framing misses is that a trial's safety belief, enrolment belief, and data-quality belief are not independent series to be charted separately — a revision forced by the safety stream should propagate into the enrolment model's eligibility criteria and the data-quality model's expected query rate, with a record of which stream forced which change. A single CUSUM chart does not do that propagation; it flags its own series and stops. The arithmetic of surprisal is old. Chaining revisions across streams with provenance intact is the part that current trial infrastructure mostly does not do, and is the part that matters here.
The second, sharper objection: surprisal is only as good as the model it is measured against, and a trial's safety model is frequently miscalibrated — too narrow a prior produces false alarms on ordinary variation, too wide a prior absorbs a genuine signal without a flicker. This is the real vulnerability, and it is not solved by wanting continuous monitoring harder. The honest answer is that calibration on a running trial is checkable in a way it is not on a one-off report: bin every past interim signal by the confidence the model assigned it at the time, compare against how many turned out to require action, and update the model's own reliability alongside the substantive belief it is tracking. A trial that has run four interim analyses has four calibration points; a trial that has run one has none, and that trial's surprisal readings deserve proportionally less trust. The dependency on calibration is real. It is answerable only because the streams keep running long enough to audit themselves — which is the same condition that makes surprisal usable as evidence at all.
Where this sits on the axis
A Large Language Model drives cross-entropy — average surprisal — toward zero over a corpus that stops. A Large World Model does the same over a sensed scene that stops when the procedure ends. A trial's streams do not stop at a point the trial itself gets to declare in advance; enrolment, safety and telemetry keep arriving until the protocol says otherwise, and that decision is itself made on the evidence the streams provide. There is no fourth rung above this, not because trials cannot be monitored better than they are, but because a further step would require surprisal against streams not yet observed, and there is no such quantity. The ladder ends where the observation does.