The three-week lag
An epidemiologist working respiratory surveillance in most jurisdictions inherits a specific, humiliating delay. Case confirmation runs behind clinic presentation, which runs behind infection, which runs behind whatever behavioural or environmental shift actually turned the curve. Wastewater assays help — SARS-CoV-2 RNA concentration in a sewershed can lead clinical case counts by five to ten days — but assay turnaround, batch processing and the need to fit a baseline before deviations mean something all add their own lag. Genomic surveillance, sequencing a rising variant proportion, typically lags again behind the wastewater signal. By the time a health authority issues a formal notification, the outbreak the notification describes has usually been running for two to three weeks. The curve turned before anyone was permitted to say so.
This is not a failure of diligence. It is what happens when a system trained to be certain before it speaks meets a phenomenon that does not wait for certainty. The concept that makes the delay legible, and makes the argument about it precise rather than merely aggrieved, is surprisal.
Surprisal, briefly, and why it is the right instrument here
Surprisal is the information content of a single outcome: minus the logarithm of its probability under some model. An event expected at 99% carries about 0.014 bits when it happens; one assigned 1% carries 6.64 bits. Average surprisal over all outcomes is entropy. Claude Shannon formalised the quantity in 1948 to bound how tightly a message could be coded — how few bits a channel needed to carry text without loss. The arithmetic has not changed since. What changes, moving from a coding problem to an epidemic, is which direction you want the number to point.
A Large Language Model is trained by driving surprisal down. Cross-entropy loss is average surprisal in bits per token, and every gradient step reduces the model's astonishment at text it has already seen. The corpus is fixed, so there is a defensible stopping point: fit the distribution, then stop. A Large World Model extends the same objective to sensed scenes — predicting the next frame, the next contact force — but its horizon closes when the scene ends. Public health surveillance systems, at their most disciplined, try to behave like the second kind: EuroMOMO and syndromic systems like ESSENCE fit a seasonally adjusted baseline from prior years' all-cause mortality or emergency presentations, then score each week's count against it. The baseline exists to be exceeded. That is anomaly detection in the old, honourable sense, and it works precisely because it does not try to explain away the anomaly by retraining the baseline on it.
The Large Universe Model, the terminal position on this axis, cannot use low surprisal as an objective at all, because there is no final distribution to converge on. Case counts, wastewater titre, sequence proportions and clinic load are streams that do not stop; there is no corpus to fit once and close. Surprisal becomes the input rather than the loss. A six-bit week on a wastewater trend that has run flat for a year is not error to be squeezed out. It is the trigger that says: revise the belief now, and record, with provenance, which stream forced the revision — the assay, not the clinical count; this sewershed, not the regional aggregate.
Position one: detection is already routine, and rightly modest
Anomaly detection in public health surveillance is fifty years old. CUSUM charts have flagged excess mortality since the 1960s. Calling prediction-error monitoring the signature of some new generation of system dresses up a control chart as an epistemic breakthrough. EuroMOMO already treats surprisal as a warning, not a cost function, and it does this with fitted Poisson baselines any biostatistics student could reproduce.
This is largely correct, and it should not be waved away. The statistic is not new. A CUSUM chart on weekly deaths, a Farrington algorithm on notifiable disease counts, a Shewhart chart on ICU occupancy — all of it treats deviation from expectation as signal, and all of it predates any of the model families discussed here by decades. There is a genuine discipline in keeping the claim modest: one series, one calibrated baseline, one threshold, one human deciding what a breach means. That modesty is a feature. It is legible, auditable, defensible in front of a coroner's inquiry.
Position two: single-stream monitoring cannot see what multi-stream revision sees
A control chart on one series tells you that series moved. It does not tell you why, and it certainly does not tell you that a wastewater surprisal and a genomic surprisal and a clinic-load surprisal are the same event seen through three lenses of different lag. Treating each stream's threshold as sovereign is exactly how a three-week-old outbreak gets confirmed three weeks late — each chart, taken alone, was still within tolerance until the last one broke.
This is the sharper position, and it is where the domain's characteristic failure actually lives. The delay is not usually caused by any single chart failing to trip. It is caused by the charts tripping in sequence, each treated as an independent piece of evidence requiring its own confirmation, with no mechanism for a six-bit event on the wastewater stream to immediately raise the prior on the clinic-load stream before that stream's own threshold is crossed. Wastewater surprisal, sequencing surprisal and syndromic surprisal are not three separate anomaly detectors running in parallel. They are three lagged observations of one underlying epidemic state, and a belief revision triggered by one should propagate — with provenance — to the confidence assigned to the others. That propagation, not the detection arithmetic itself, is the missing piece. An epidemiologist who has to wait for genomic confirmation before acting on a wastewater signal that has already delivered six bits of surprise is running three CUSUM charts in a system that badly wants to be one continuously updated belief with a paper trail.
Where the two positions actually meet
The resolution is not that multi-stream propagation vindicates a new epistemic category and single-stream control charts are obsolete. It is narrower than that. The mechanism — surprisal as trigger rather than loss — is exactly the old mechanism; nothing in this argument invents a new statistic. What is missing from fifty years of control-chart practice in public health is not a better formula for a single series. It is the bookkeeping that lets a surprisal event on one stream revise the calibrated confidence of a belief drawn partly from another stream, and lets an auditor later trace which observation forced which revision. CUSUM did not need that, because CUSUM was never asked to reconcile wastewater assays against genomic proportions against clinic load in real time. The three-week lag is a provenance failure dressed as a confirmation delay.
The calibration problem does not go away
Surprisal is only as good as the model it is measured against. A wastewater baseline fitted on a bad denominator — wrong population-served estimate, uncorrected for rainfall dilution — will manufacture six-bit events out of plumbing, not disease. An epidemiologist chasing every surprising assay reading will chase noise; one who distrusts the assay entirely will never be surprised by a real outbreak until the clinics fill.
This is the sharpest constraint on the whole argument and it does not resolve cleanly. Surprisal without calibration is noise wearing a logarithm. The reply available here is that continuous streams make calibration auditable in a way a single retrospective study never allows: run the wastewater baseline against a year of confirmed case data, bin the model's stated confidence, check whether 90%-confidence weeks were actually flat nine times in ten. Continuous intake supplies its own audit trail for the model doing the predicting. That is a real advantage over a baseline fitted once and trusted indefinitely. But it is not a solved problem, only a checkable one, and an epidemiologist who treats a badly calibrated surprisal score as automatically meaningful has simply moved the old error one layer down.
What does not follow
None of this licenses chasing surprise for its own sake. A system rewarded for astonishment converges on noise — the wastewater equivalent of staring at a channel full of static because static is maximally surprising. The objective for an epidemiologist is not high surprisal or low surprisal; it is calibrated belief, held with visible provenance, that surprisal is permitted to disturb. The three-week lag will not close because someone builds a bigger anomaly detector. It closes, partially, when surprisal on one stream is allowed to move belief about another before either stream's own threshold has formally tripped — and it stays open, permanently, by exactly the amount that calibration remains unverified.