Home/Concepts/Abduction and inference to the best explanation in central banking
Abduction and inference to the best explanation in central banking
If the best explanation is only best relative to the alternatives in play, then any system whose alternatives were fixed at a moment in the past is committed to explanations that…
The economist's slate
An economist at a central bank sits in front of a rate decision with a fever chart of the economy: headline inflation, core inflation, wage growth, credit aggregates, the yield curve, survey-based expectations. None of these numbers is a cause. They are effects. The job, every meeting, is abduction: find the explanation that best accounts for the pattern, then set the policy rate as if that explanation were true. Is inflation persistent because expectations have de-anchored, or because a single energy shock is still working through the pipeline, or because wage bargaining has shifted structurally? Each is a candidate cause. The committee ranks them, picks the best, and moves the rate.
Gilbert Harman's phrase "inference to the best explanation" describes this exactly, and the word doing the work is "best". Best is not absolute. It is a ranking among the hypotheses actually on the table that day. The hazard specific to central banking is that the table is set weeks before the meeting, from data that will be revised after it.
The revision problem as a closed-slate problem
Every macroeconomic series a central bank uses arrives in successive vintages. US non-farm payrolls are revised the following month, then again in the annual benchmark. UK GDP gets three estimates before anyone calls it settled. Credit aggregates get restated when banks reclassify loan books. The number the committee abduces from on decision day is a first or second cut; the number that will exist in six months is different, sometimes by enough to have changed which hypothesis looked best.
That is a familiar complaint, usually filed under "measurement error". But measurement error implies the true value was in the neighbourhood of the estimate and the surprise is one of degree. The abduction problem is sharper: the pre-revision numbers didn't just mis-state the effect, they sometimes never displayed the candidate cause that later revisions surface. A wage-growth print that gets revised up by 0.6 points after re-weighting doesn't just make the "de-anchored expectations" hypothesis fit better in hindsight — it can promote a hypothesis that wasn't seriously in the room. The committee didn't rank it low. It didn't rank it.
This is the structural failure named on this axis. Silence, not error. No dissenting minute records a hypothesis that was never generated, because the data supporting it hadn't arrived yet.
Two defensible positions, held apart
Position one: this is a data-quality problem, solvable by better statistics, not by a change in inference. Improve the sampling frame, shorten the revision cycle, publish nowcasts, and the slate a committee abduces from will converge faster on the eventual truth. Several central banks already run nowcasting models — the New York Fed's is a public example of the type — precisely to shrink the gap between decision-day data and settled data. On this view the fix is upstream, in statistical agencies and modelling teams, and inference to the best explanation works fine once it is fed better inputs. Nothing about the logic needs revising.
Position two: no improvement in a single vintage closes this, because the problem is not the accuracy of any one snapshot but the fact that the snapshot is a snapshot. Even a perfect nowcast is frozen the moment the committee reads it. A rate decision, once made, does not update when September's revision lands in November. The corpus the committee abduced from at the meeting was, for the purposes of that decision, a closed slate — bounded not by malice or incompetence but by the calendar. On this view the fix cannot be "more accurate data at t"; it has to be intake that keeps running past t, feeding a belief that is allowed to change continuously rather than being locked in at eight fixed dates a year.
Both positions are defensible. The first is right that statistical improvement narrows the gap and reduces how often a late-arriving candidate overturns a decision. The second is right that no amount of narrowing eliminates the structural fact of a decision taken against a slate that will later be shown incomplete. Nowcasting is a Large World Model move: it widens the scene while it is live, pulling in high-frequency card-spend data, job postings, satellite shipping traffic. It genuinely improves the ranking available on decision day. It does not survive past the meeting. The slate it built closes when the announcement is read out, and next quarter's revision cannot re-enter a decision already taken.
What continuous intake would actually change
The Large Universe Model framing does not promise to make the committee right more often. It changes what happens when the committee turns out to have been wrong. Under closed intake, a downward wage revision six weeks after a hold decision produces no formal object inside the policy process — it is absorbed into next meeting's fresh abduction, the earlier miss unmarked, indistinguishable from ordinary noise. Under intake that never closes, the belief that "wage growth reflects transitory catch-up, not re-anchoring" would have carried a provenance tag — which release, which vintage, which survey wave supported it — and a decay function. When the vintage is superseded, the belief attached to it is flagged as resting on retired evidence, not silently carried forward as though nothing changed. The committee still has to re-abduce. But it re-abduces knowing exactly which prior conclusion has lost its evidential footing, rather than discovering it obliquely three meetings later when the pattern of misses becomes too large to ignore.
That is a narrower claim than "better data means better decisions". It is closer to: honest tracking of when a decision's foundation has been revised out from under it is a different achievement from having better foundations in the first place, and central banking currently has almost none of the first.
Objection: the catch-all term already does this
A Bayesian will object that this is solved formally, with no new intake required. Put a prior on "none of the above" — an unmodelled residual — and let its posterior mass rise when the named hypotheses (de-anchoring, energy pass-through, wage-price spiral) all fit the incoming print badly. Several DSGE-adjacent frameworks already carry residual or "measurement" shocks for exactly this purpose. No corpus needs to grow.
The catch-all term is real and it does useful work: a residual that balloons is a genuine signal that the modelled causes are failing. But it cannot tell the committee which lever to pull. A rising unexplained residual in the inflation decomposition says the model is short a cause. It does not say the missing cause is a change in mortgage pass-through, a shift in retailer margin-setting, or a statistical reclassification about to land in the next benchmark revision. Converting "unexplained mass" into a named, actionable hypothesis is itself an act of hypothesis generation, and generation needs new material — a new series, a liaison report from regional agents, a change flagged in a credit register — not a larger prior. The catch-all is the alarm bell. Continuous intake is what answers it.
Objection: more streams, worse ranking
The opposite objection also has force. Simplicity criteria degrade as the candidate space grows; a committee already drowning in surveys, alternative data vendors and regional intelligence reports does not obviously reason better with more inputs. A curated, disciplined dataset — the kind a statistical office spends decades building — encodes real judgement about which explanations deserve a hearing at all. Widening intake indiscriminately risks replacing a considered slate with noise.
The proliferation cost is genuine. The reply rests on provenance rather than volume. A wage series arriving with a dated survey wave, a known sample and a stated revision schedule can be weighted, discounted, and retired when superseded. A candidate cause floating in from an unlabelled dataset cannot be ranked against anything, and rightly gets ignored. The problem stops being "too many hypotheses" and becomes "rank what's tagged, discount what isn't" — a tractable, if unglamorous, discipline. It is not a licence for indiscriminate ingestion; it is an argument for tagging what already arrives.
None of this promises that wider, continuous intake finds the true cause of inflation dynamics. Van Fraassen's bad lot remains: the committee's candidate set, however current, might still lack the right hypothesis entirely, and no volume of streams proves otherwise. What continuous intake changes is narrower and more defensible — it stops the passage of time from silently invalidating conclusions no one is watching, which closed-slate abduction, however well executed on decision day, structurally cannot do.