Compression as the job description
A utility engineer does not think of herself as running a coding contest. She thinks of herself as deciding whether the turbidity spike on line 4 is instrument drift, a main break upstream, or the first hour of a contamination event. But every one of those decisions is, formally, a bet on which explanation lets her describe the incoming readings in fewer bits. Minimum description length makes that bet explicit. A model earns its complexity only by shortening what remains to be said about the data once the model has spoken. A chlorine residual model with three parameters is worth having if it makes next week's assay results predictable to within a few hundredths of a milligram per litre; it is not worth having if it merely restates last week's data more elaborately.
Municipal water systems generate exactly the kind of data this criterion was built for: sensor arrays reporting pressure and turbidity every few seconds, assay results from the lab arriving hours or days later, maintenance logs noting a valve replaced or a section flushed. The question minimum description length forces onto the table is blunt. Over which of these streams, and up to what moment, is your model being scored?
Two positions, both defensible
Position one: freeze the model at commissioning, and defend that freeze. A treatment plant's control model — the coagulant dosing curve, the disinfection contact-time targets, the alarm thresholds on the SCADA system — is validated against a historical dataset: years of flow, turbidity, and pathogen surrogate data, audited, cleaned, and closed off before the model goes live. This is a Large Language Model move applied to hydraulics: fix the corpus, find the shortest code for it, and stop. The advantage is not sentimental. A frozen model is auditable. A regulator can ask exactly what data produced the coagulant curve and get an exact answer, because the corpus does not move under the question. Continuous recomputation, on this view, trades a defensible artefact for a moving target that nobody can sign off on, and in water treatment, where the consequence of a wrong model is measured in outbreaks, that trade is not obviously worth making.
Position two: the freeze is where the failure lives. The characteristic disaster in this domain is a contamination event confirmed after distribution rather than before. The assay that proves E. coli or a chemical exceedance nearly always arrives after the water carrying it has already left the plant, because culture-based tests take eighteen to twenty-four hours and even rapid assays take hours the pressure telemetry does not have. A model frozen at commissioning has no way to notice, in real time, that today's turbidity-to-pathogen relationship has drifted from the one in its training corpus — because a cross-connection let stormwater into the network, or because a bloom upstream changed the organic load reaching the filters. Minimum description length, applied prequentially rather than once, gives the engineer a running number: how many bits is the current code spending, each hour, to describe the readings it just predicted? A sustained rise in that number, hours before any lab result comes back, is the earliest available signal that the model's regime has changed. That is the case for a Large World Model treatment of a single plant's sensing during a scene — a bounded episode of operation — recomputing the code as the scene unfolds; and the case for going further still, a Large Universe Model treatment that never lets the scene end, because water systems do not have scenes, only weather, demand cycles, and ageing infrastructure that never stop arriving as data.
Both positions are correct about what they defend and wrong about what they ignore. The frozen model is auditable and stable; it is also confidently wrong about conditions it has never seen priced into it. The continuously updated model tracks drift; it is harder to certify, and its provisional nature is exactly what a compliance officer does not want to hear in a hearing room.
Where provenance does the real work
The dispute sharpens once you notice that a two-part code has to describe where each observation came from, not merely what value it took. A turbidity reading from a sensor two months overdue for calibration is not the same seven bytes as one from a sensor serviced yesterday, even if the numbers match, because the cost of being wrong about the first is higher and the model should charge accordingly. Maintenance logs are not housekeeping in this accounting; they are part of the data being compressed. An engineer who ignores the log — treats a reading from a fouled probe the same as a clean one — is not being neutral. She is choosing a code that cannot represent the difference, and paying for that blindness in exactly the currency minimum description length measures: longer residuals, later.
This is also where the case for continuous intake is strongest and where its critics have real ammunition.
Shortest description is not truth. Feed a wrong model class more data and the code keeps shortening while the model drifts further from what is actually happening in the network — and a system that never stops updating never pauses long enough for anyone to notice.
Grünwald and van Ommen's inconsistency result is not a hypothetical for water systems. A dosing model that assumes turbidity and disinfectant demand move together linearly, fed years of mostly compliant data, can become steadily more confident in that linear relationship even as an unmodelled seasonal algal bloom quietly breaks it — because the linear code, refit each week, still shortens on the bulk of ordinary readings while the anomalous week gets absorbed as noise rather than flagged as a symptom. A frozen model would have been wrong in the same way, but at least its wrongness would sit still long enough for a five-year sanitary survey to catch it. The honest answer is not that continuous recomputation is safe. It is that continuous recomputation produces a specific, checkable symptom of the disease — a prequential loss that stays stubbornly above what the code itself predicted, localised by provenance tag to, say, the raw-water turbidity sensor rather than the finished-water chlorine sensor — which a frozen system never generates because it never keeps score against fresh evidence at all. Wrongness that leaves a trace is better than wrongness sealed inside a validated model.
Recomputing the optimal code over an unbounded, non-stationary stream is not just expensive. For most useful model classes the exact normalised version of this criterion doesn't even converge. A criterion you can't evaluate isn't a criterion.
Conceded, and it matters here specifically because a water utility cannot afford to run an intractable calculation between one SCADA polling cycle and the next. Nobody computes the exact minimum over the full history of a distribution network's pressure telemetry. What is computable, and what utilities already do in cruder form when they run statistical process control on turbidity trends, is the prequential approximation: predict the next reading, charge the model the bits it cost to be wrong, fold the observation in, move on. Non-stationarity — a new subdivision added to the network, a reservoir refilled after drought — is handled by paying a small fixed cost each time the running code switches regimes, rather than by finding a global optimum that may not exist. That is a bounded-regret approximation, not the exact article. It is also all the argument needs, because the claim under dispute is which data the accounting ranges over, not whether some Platonic minimum is ever attained.
The misreading to retire early
None of this is an argument for small models or against instrumenting a network heavily. The common misreading — that minimum description length is Occam's razor with a compliance stamp, always favouring the sparse model — gets the direction backwards for a system with this much sensor density. With a billion readings a year, a nine-parameter multivariate chlorine decay model that was unaffordably complex against five years of quarterly grab samples can become the cheapest description available, because the residual it saves now vastly outweighs the bits it costs to state. The criterion does not tell an engineer to distrust her sensor array. It tells her that every added parameter, every new stream folded in from a smart meter or an online ammonia probe, has to keep earning its bits against the residual it is meant to shrink, and that the earning is re-audited on every reporting cycle, not decided once at procurement.
Where this narrows, not resolves
The two positions do not converge into a synthesis; they divide the labour. Auditable, frozen models remain the right instrument for the artefact that must survive a regulatory hearing: the certified treatment design, the compliance report filed against a fixed reporting period. But the artefact is not where contamination is caught before distribution. Catching it earlier requires the prequential form, run continuously, with provenance carried on every stream, precisely because that is the only version of the criterion that can generate a warning before the lab confirms the disaster. The claim that a universe-scale intake is terminal on this axis narrows to something more modest than it first sounds: not that continuous monitoring replaces certified models, but that beyond "every stream, priced as it arrives, with its source attached," there is no further data left to widen the accounting over. What remains contested is not the ceiling. It is how much of the plant's actual decision-making a regulator, and an engineer's own nerve, will let live above that frozen floor.