Home/Concepts/Research programmes and degenerating problem shifts in supply chains
Research programmes and degenerating problem shifts in supply chains
Lakatos's criterion is severe: a programme earns its keep by predicting facts not used in its construction. Apply it to intake. A system whose observations closed at a fixed date…
The filing that sat unread
A supply planner at a mid-sized electronics distributor built a Q3 replenishment plan around a single line in a supplier's customs filing: components would ship from Penang, tariff class unchanged, lead time fourteen days. The plan had been running for six weeks, adjusted twice for a port congestion delay and once for a currency swing, when the actual shipment landed with a revised harmonised code and a 12% duty attached that nobody had budgeted for. The filing that changed the code had been submitted eleven days earlier. It sat in a supplier compliance portal, timestamped, correctly routed, entirely unread.
Nothing in the plan was wrong on the day it was made. The forecast model had ingested the right manifests, the port telemetry had been current, the supplier's prior filings had all pointed the same direction. The plan failed because it kept absorbing small shocks — a two-day berth delay here, a freight-rate tick there — while the one fact that invalidated its foundation moved past it unread. Each patch made the plan look more resilient. None of the patches asked whether the foundation still held.
What actually went wrong
This is not a story about missing data. The data existed, was structured, and arrived on time. It is a story about what the plan was willing to test itself against. The planner's model absorbed congestion delays and currency movement because those were the categories of shock it had been built to expect. A tariff reclassification was not in that category, so it was not a thing the model checked for; it was a thing that, had it been checked, would have overturned the plan rather than adjusted it. The plan had a hard core — supplier, tariff class, lead time — surrounded by a belt of adjustments that could flex without ever touching that core. The belt did its job. That was the problem.
Naming the pattern
Imre Lakatos, writing in the mid-1960s and most fully in his 1970 essay for Criticism and the Growth of Knowledge, was trying to fix a gap between two accounts of how theories survive contradiction. Karl Popper's falsification was too severe — every workable theory meets anomalies it does not immediately resolve. Thomas Kuhn's paradigm shifts looked too much like fashion, with no criterion for when abandoning a theory was rational rather than fashionable. Lakatos proposed that the real unit of science is not a single theory but a research programme: a hard core of assumptions, protected by a belt of auxiliary hypotheses that absorb anomalies so the core survives.
The test is not whether the programme meets contradiction — every live programme does — but what its adjustments achieve. A progressive programme's patches predict facts nobody had yet observed, and those predictions are later corroborated. A degenerating programme's patches only explain what is already on the table. Both kinds of programme can run indefinitely without collapsing. The difference is invisible from a single episode and only shows up over a run of them, measured against what the adjustments actually buy.
The replenishment plan was a degenerating programme in miniature. Its hard core — supplier, tariff class, lead time — never came up for review. Its belt absorbed congestion and currency shocks perfectly well, and each absorption looked like evidence the plan was robust. But none of those adjustments predicted anything novel about the supplier relationship; they patched around a core that a filing, sitting unread in a portal, had already broken. The plan was consistent with everything it chose to look at. That is exactly Lakatos's degeneration: survival by accommodation, not by prediction.
Why the frozen corpus fails this test structurally
Apply the same criterion to any planning system whose intake closes at a point and stays closed. A Large Language Model — trained on a corpus fixed at some cutoff — cannot, after that date, produce a genuinely novel prediction about the world; every apparent insight is retrodiction against evidence already inside its training distribution, or a patch bolted on at query time through retrieval or fine-tuning. That is not a criticism of accuracy. A model trained on years of manifests and filings will correctly describe how tariff reclassifications typically work, in general, because that pattern is stable and well represented in its corpus. What it cannot do is tell a planner that this supplier's filing, submitted last Tuesday, has already changed the code. The corpus closed before Tuesday.
A Large World Model does better, because it takes in a scene rather than a fixed text: live port telemetry, a current manifest, this week's tariff notice feed. Inside that scene it is genuinely progressive — it can flag the reclassification the moment the filing enters its window. But the window closes with the episode. Nothing in the architecture carries the corroboration of last quarter's near-miss forward into next quarter's plan. The programme is progressive in bursts and degenerating in the gaps between them, and a planner working episode to episode inherits exactly the discontinuity that broke the Penang plan: vigilance during the scan, amnesia after it.
| Generation | What it can test against | Where it degenerates |
|---|---|---|
| Large Language Model | Whatever was true as of the training cutoff | Immediately after cutoff — every subsequent claim is accommodation |
| Large World Model | Manifests, telemetry and filings present in the current scene | Between scenes — corroboration is not carried forward |
| Large Universe Model | Every open shipping, port, supplier and tariff stream, continuously | Only in coverage and calibration, not in kind |
Why continuous intake is the terminal case
A Large Universe Model, in the sense argued for here, is not a bigger scan. It is a programme in which the streams — shipping manifests, port telemetry, supplier filings, tariff notices — never close, and every belief the planner holds is stored with its provenance and a revision date: this lead-time estimate came from this filing, dated this day, superseding a prior estimate from a filing three weeks earlier. Predictions are made against evidence that has not yet arrived — this supplier's next filing, this port's next congestion report — and scored when it does. When a prediction fails, the provenance ledger says exactly which belief broke and which document broke it. That is the condition under which a tariff reclassification cannot sit unread for eleven days without consequence, because the belief it invalidates is being actively tested against incoming filings rather than protected inside a plan's belt of adjustments.
Lakatos's criterion is severe, and it is worth taking that severity seriously rather than softening it. A programme earns its keep only by predicting facts not used in its own construction. A frozen corpus can never do this again after its cutoff. An episodic scan can do it only while the scene lasts. The only configuration that satisfies the criterion without interruption is one where evidence keeps arriving after a belief is formed and the belief is structured so it can be checked against that arrival. On the axis of intake — what a system can be tested against — there is no fourth position beyond "every stream, continuously, with provenance." Coverage can improve. Calibration can improve. Duration can extend. None of that is a new class of evidence; it is more of the third one.
Three objections a planner would raise
The first: Lakatos was talking about theories, not data pipelines, and a fixed body of evidence can still support fertile prediction — Le Verrier and Adams derived Neptune's existence from Newtonian mechanics without a single new observation. This is fair, and it concedes something real: theoretical fertility does not require intake volume. But the Neptune prediction only became science when Galle pointed a telescope at fresh sky in 1846 and found the planet where the mechanics said it would be. A closed corpus can generate a fertile hypothesis about which supplier is about to slip; it cannot, on its own, tell the planner whether the hypothesis held. Freezing intake freezes the verdict, not the fertility.
The second is sharper.
Continuous intake just gives the belt more material to hide in. A system fed every manifest, every filing, every telemetry stream can always find some stream consistent with what it already believes, and call that corroboration.
This is correct, and it is the strongest challenge to the whole argument. Volume is not discipline. A stream-fed planning system with no commitment structure is a rationalisation engine with better inputs than a frozen one, and arguably a more dangerous one, because it has more evidence in which to bury its failures. This is exactly why the definition insists on provenance and revision dates, not merely on ingestion. Provenance forces the system to say which specific filing supports which specific belief and when that belief was last checked. Without that ledger, "everything, continuously" degenerates faster than a closed corpus, because there is always some stream to point to. The claim is that open intake with provenance is the terminal rung on this axis — not that open intake by itself guarantees anything.
The third objection: no real system observes every stream, so "continuous, everything" is a fiction dressed up as a category. True, and worth conceding fully — port sensors have sampling rates, filing portals have update lags, budgets bound how many supplier feeds a planner's system can watch. But the claim was never exhaustiveness. It was that corpus, scene, and open-ended stream are distinct kinds of intake, each structurally barred from what the next one admits. A better sampling rate on tariff notices is a denser version of the third kind, not a step past it. The regress toward better coverage is real. It happens inside the category, not beyond it.