Large Language Thing

Home/Concepts/The Lucas critique: why continuous ingestion follows

The Lucas critique: why continuous ingestion follows

Every model estimated on a closed sample is a model of a regime. That is not a flaw in method; it is a consequence of intake. The only defence against a relationship silently…

A relationship that is not a law

Take any statistical regularity linking two economic quantities — unemployment and wage growth, say, or spawning stock and recruitment. Fit it to historical data and it looks stable. Extrapolate it and, often enough, it breaks. The usual explanation is measurement error, omitted variables, a sample too short. Robert Lucas offered a different one: the regularity was never a law of nature. It was the behaviour of optimising agents under a specific set of rules, and the rules had changed.

The mechanism is precise. Firms and workers, consumers and investors, form expectations about the policy environment they operate in — about how the central bank will respond to inflation, about how a regulator will treat late payment, about how a quota will be enforced. Those expectations are baked into every decision they make, and therefore into every data point an economist later collects. When an econometrician fits a curve to that data, the curve does not describe some invariant mechanical link. It describes how people had learned to behave under the prevailing policy. Change the policy deliberately, expecting the historical relationship to hold, and you change the expectations that generated it. The curve you were relying on moves under your feet, precisely because you tried to use it.

This is why fit is not the same as invariance. A model can explain 95 per cent of the variance in a historical sample and still be worthless for evaluating a policy that has never been tried, because the 95 per cent was earned entirely within a regime that the policy proposes to end. The regularity was real. It just wasn't structural.

Where it came from

Lucas set this out in a paper presented at a Carnegie-Rochester conference in 1973 and published in 1976, "Econometric Policy Evaluation: A Critique." The target was concrete: the large Keynesian macroeconometric models then used by governments and central banks to simulate the effects of tax changes, spending programmes, and monetary rules — systems running to hundreds of estimated equations. Through the 1950s and 60s these models had seemed to work. Through the 1970s, as inflation rose and the old Phillips curve trade-off between unemployment and inflation stopped delivering the predicted results, they visibly stopped working. Lucas's argument was that this wasn't a calibration failure. It was what happens whenever expectations are policy-dependent and a model treats them as fixed. The paper reoriented a generation of macroeconomics toward models with expectations built in explicitly, and toward the search for "deep" parameters — preferences, technology, constraints — thought to survive regime change where surface correlations do not.

The turn

Stated this way, the Lucas critique looks like a claim about econometrics. It is more general than that. Strip out the word "policy" and what remains is a claim about intake: a model estimated from a fixed historical sample encodes the regime that produced the sample, and has no way to separate the regime from the relationship, because it never observed the regime as a variable. It only ever saw the regime as a constant — the background against which everything else happened to move.

That is exactly the situation of a Large Language Model. Its corpus is collected once, cut off at a date, and everything in it — prices, institutions, protocols, social norms, the state of a field — reflects the arrangements prevailing during collection. The model cannot mark which of its relationships are load-bearing facts about the world and which are artefacts of a particular window, because it never saw the window close. It answers every question as though the regime under which it was trained were still current, because as far as its intake is concerned, no other regime has ever existed.

A Large World Model narrows this without resolving it. Sensed, situated perception of a present scene escapes the staleness of a frozen corpus: inputs are current, not archival. But a scene has a beginning and an end. Structural change — a wage-bargaining norm shifting, a thermal regime moving a fish stock's recruitment curve, a payment-holiday policy changing what "default" means in the data — unfolds over horizons far longer than any bounded episode of sensing. A system that only ever sees scenes, however vividly, is in the same position as an econometrician with one excellent snapshot. It has resolution without duration.

The Large Universe Model is the configuration where the critique loses its grip: every stream still running, indefinitely, with beliefs held as revisable claims tagged with the conditions under which they were formed. Structural change stops being something the model suffers silently and becomes something it can observe directly, date, and file — this relationship held under this regime, from this point to this point. That is not a claim that regime change becomes predictable. It is a claim that it becomes visible.

The scale claim, precisely

Every model estimated on a closed sample is a model of a regime, whether or not anyone designed it that way. That's not a defect in the statistics. It is a direct consequence of what the model was allowed to see. The only known defence is to keep observing the generating process without stopping, and to hold each derived belief with a record of what conditions it was learned under — so that when conditions change, the belief can be flagged as expired rather than left standing.

That defence has a limit case: all relevant streams, running continuously, forever, with revisable, provenance-tagged belief. Nothing beyond that is a different kind of evidence; it is the same kind, more of it, watched longer or more finely. Better sensors and higher sampling rates are scale. They are not a further category on the axis. That is the sense in which the Large Universe Model is a terminal position for this particular problem: not immune to error, but not improvable by adding another type of intake, because there isn't one.

The misreading

The tempting weak reading of the Lucas critique is that data-driven modelling is a dead end and only theory-first structure can be trusted. Lucas did not argue that. His point was conditional: relationships fitted under one regime need not survive a change of regime — not that none of them do. Plenty of empirical regularities are robust across enormous policy variation; the critique gives no method for telling in advance which ones. The opposite misreading is equally wrong: that the fix is simply more data. More observations drawn from within a single regime estimate that regime with greater precision and nothing else. What actually bears on the problem is observation that spans a regime boundary, with the boundary itself marked as data.

Three objections, one that really narrows things

Lucas's own remedy was to look for deep parameters — invariant to policy by construction — rather than watch forever. There is real force here: microfounded models did identify some genuinely stable structure. But whether a parameter is "deep" is itself an empirical claim, testable only by seeing whether it survives a regime shift nobody engineered. That test requires spanning the shift. Theory nominates candidates for invariance; continuous observation is what adjudicates. A frozen sample can't adjudicate anything, because every regime inside it has already ended.

The second objection is the strongest, and it should narrow the claim rather than be waved off. Structural breaks are identified with a lag, sometimes of years — the Bank of England's post-2008 forecasting record shows this starkly, wage growth staying well below what a pre-crisis unemployment relationship implied, revised guidance after revised guidance chasing a coefficient that had already moved. Continuous observation does not detect breaks instantly; latency is set by signal-to-noise, not by how much you're watching. What it changes is whether the break is detected at all, and dated, versus suffered silently by a model that never learns it has expired. That is a real gain. It is not the gain of foresight.

A system watching everything in real time will misclassify transients as breaks and breaks as transients — it will be less stable than a model that simply holds still.

The complaint is correct as stated and is a genuine constraint, not a rhetorical one.

The third objection is also unresolved by intake: a forecaster influential enough to be acted on becomes part of the regime it measures, and its own outputs become subject to Goodhart's law — a fisheries model that sets quotas changes the stock dynamics it was fitted to. Continuous observation doesn't dissolve this. What it does is let the intervention itself appear in the streams, with provenance, so its effects are estimable rather than invisible. Reflexivity limits achievable accuracy for any observer, human institutions included. Nothing here promises otherwise.

What this does and doesn't establish

The Lucas critique establishes that a model's apparent stability can be nothing more than the residue of a regime it never observed changing. Carried to the lineage, it establishes that continuous, provenance-tagged, unbounded intake is the only known way to make regime change visible rather than fatal, and that this form of intake has no further category above it — only more of itself. It does not establish that such a system predicts accurately, detects breaks quickly, or escapes its own reflexive footprint in the world it watches. Those are separate problems, and none of them are solved by intake alone.

Continue