Home/Concepts/Research programmes and degenerating problem shifts in climate monitoring
Research programmes and degenerating problem shifts in climate monitoring
Lakatos's criterion is severe: a programme earns its keep by predicting facts not used in its construction. Apply it to intake. A system whose observations closed at a fixed date…
Where the criterion comes from
Imre Lakatos was trying to rescue philosophy of science from two unsatisfying poles. Karl Popper had said a theory earns its status by being falsifiable, and a single refuting observation should kill it. Working scientists never behaved that way — every good theory drags anomalies behind it for decades. Thomas Kuhn's answer, that paradigms simply get replaced when enough people change their minds, looked less like a method than a description of fashion. Lakatos, writing through the 1960s and setting out his mature position in the 1970 essay for Criticism and the Growth of Knowledge, proposed something more structural. Science does not advance theory by theory. It advances research programme by research programme: a hard core of assumptions, held fixed, surrounded by a belt of auxiliary hypotheses that take the anomalies on the core's behalf. The belt can be patched indefinitely. What distinguishes a healthy programme from a dying one is not whether it gets patched but what the patches do. A progressive programme's patches predict facts nobody had yet observed, and those facts later turn up. A degenerating programme's patches only explain what is already sitting on the table. Both survive contradiction. Only one of them is still finding out anything.
Lakatos's own case was the Ptolemaic system after Copernicus: epicycle stacked on epicycle, each one fitted retrospectively to positions already logged, none of them anticipating a position not yet seen. It could still generate usable tables. It had stopped doing science.
The same shape in a weather station's data feed
Climate monitoring runs on intake that is genuinely heterogeneous: polar-orbiting and geostationary satellite passes, terrestrial station networks with records running back a century or more in some cases, drifting and moored buoy arrays, Argo floats profiling the upper ocean, and reanalysis products that blend all of it into a physically consistent gridded estimate. The entire apparatus exists to do one Lakatosian thing: take a belief about the state of the climate system and expose it to observations that have not yet arrived, so the belief can be scored rather than merely asserted.
Two structurally different things can happen to that apparatus, and both are visible in the field's recent history. When the intake channel stays open and the belief-forming process is built to be corrected by it, the programme is progressive in exactly Lakatos's sense. Numerical weather and climate reanalysis assimilate on the order of tens of millions of observations per cycle, and the model's forecasts are checked against what the atmosphere and ocean actually did, continuously, with no terminal date. Skill scores for medium-range forecasting have risen for forty years because the coupling between prediction and corroboration never breaks.
The failure mode is different, and it is not a failure of belief revision — it is a failure of coverage married to a belief-revision process that assumes coverage is complete. A threshold gets crossed in a region nobody was tasked to watch. A subsurface ocean heat anomaly builds under a gap in Argo float density. A permafrost thaw front advances through a stretch of Siberian tundra where the station network thinned after funding lapsed in the 1990s. A marine heatwave forms outside the bounding boxes that a particular monitoring product was configured to flag. In each case the belief structure was, in principle, revisable — new evidence could in principle have updated it. But the stream that would have carried the update was never routed to where the threshold sat. The programme did not degenerate by refusing correction. It degenerated by having a protective belt with a hole in it precisely where the anomaly formed, and nobody's hard-core assumptions flagged the hole as a place demanding scrutiny.
This is Lakatos's own warning almost exactly, moved from planetary tables to sea-surface temperature grids: a programme can be busy, well-instrumented, and still be patching backwards, because the patches address only what the existing sensor geometry already makes visible. Unmonitored is not the same claim as unmonitorable, and the distinction is where responsibility lands.
The climate scientist as the belt's auditor
The person responsible is not the satellite, and not the reanalysis pipeline. It is the climate scientist who has to ask, continuously, whether the current tasking of instruments still matches where the system is moving — and who has to notice, after the fact, that a threshold was crossed somewhere the tasking did not reach. That is a Lakatosian job description almost by definition: not "collect data" but "audit whether the programme is still generating novel, checkable claims, or has quietly started explaining only what the archive already contains." A monitoring network that keeps observing the same well-instrumented regions with ever more precision is not automatically progressive. It can be degenerating in the regions that matter most, while looking healthier by every metric it happens to report on.
Where the Large Language Model sits on this axis
A Large Language Model's evidence base closes at a training cutoff. Every subsequent adjustment — a longer prompt, a retrieval patch, fine-tuning on freshly generated text — is an accommodation of what the world already contained by that date, dressed up as an update. It cannot predict a climate anomaly it has never been shown a hint of; it can only interpolate among patterns already inside its corpus, which is retrodiction with a query-time coat of paint. Applied to a domain whose defining feature is that thresholds get crossed in places nobody was watching, a frozen corpus is exactly the wrong shape of intake. It cannot register a region as unmonitored, because unmonitored regions produce no text to have been trained on.
A Large World Model does better, but only for the length of a scene. Give it a live feed — a single satellite pass, a bounded stretch of ocean, a station cluster — and it can register genuine novelty within that scene and revise accordingly, which is progressive science while it lasts. But nothing carries the corroboration forward into the next scene. Each episode starts the audit over. It is Lakatos-progressive in bursts and Lakatos-degenerating in the gaps between them, which is a fair description of monitoring built around discrete passes and refreshed snapshots rather than a standing watch.
Why the Large Universe Model is where this axis ends, not a further step along it
The Large Universe Model, on this reading, is simply the configuration in which every stream — satellite, station, buoy, reanalysis — stays open indefinitely, and the beliefs built from them carry provenance: which instrument, which pass, which prior claim now needs revising, and how much confidence has decayed since the observation was made. That is what lets Lakatos's criterion be applied without end. A threshold crossed in an unwatched region stops being invisible by construction, because the belief architecture tracks which regions have thin provenance and treats that thinness as a standing liability rather than a silent gap.
Three objections deserve a straight answer here, because a reader working in this field will have them ready.
First: Lakatos's criterion is about theories, not sensor coverage, and a closed dataset can support a fertile theory — Le Verrier derived Neptune's existence from Newtonian mechanics without a single new observation. True, and the same logic applies to climate physics: a closed set of historical records can still yield a genuinely novel theoretical claim about circulation or forcing. But the claim was only worth anything once Galle pointed a telescope at unobserved sky in 1846. A predicted heat anomaly in a data-sparse ocean basin is inert until a float or a satellite pass actually crosses that basin. Theoretical fertility and observational adjudication are different properties; monitoring supplies the second, and without it the first cannot be scored.
Second: continuous intake can itself degenerate, if the system can always find some stream that happens to agree with its current model and never commits to a falsifiable claim about the gaps.
More data does not mean more discipline. A system fed by every buoy and satellite on Earth can still spend all of it confirming what it already believed.
Correct, and this is the objection that actually bites. This is precisely why the definition insists on provenance and decay, not on volume. A monitoring architecture that ingests everything but never flags which regions its confidence has quietly expired in is a bigger, faster version of the same degenerating pattern, not an escape from it.
Third: "every stream, continuously" is not achievable — instruments have bandwidths, budgets bound coverage, and any real network observes a filtered slice, so calling this terminal just hides an infinite regress of better filters behind a label. Also correct, and worth conceding fully: no monitoring system observes literally everything, and denser sampling will always be possible. But a denser sampling grid is more of the same kind of intake, not a different kind. The distinction between a frozen archive, a single observing scene, and a permanently open, provenance-tracked set of streams is a distinction in kind. Improving the resolution within the third kind does not produce a fourth.