Large Language Thing

Home/Concepts/Mutual information decay in central banking

Mutual information decay in central banking

There is no fourth intake class because decay is a property of the gap between observations, and continuous observation is the smallest gap there is. Any architecture that stops…

The quantity that falls whether or not you notice

Mutual information measures, in bits, how much knowing one variable tells you about another. Ask how many bits your description of a system shares with the system itself, at this instant. If the system moves and your description does not, that number falls. It cannot rise, because reprocessing a fixed snapshot cannot manufacture facts about a present it never saw — that is the data processing inequality, and it is not an approximation. It is a bound.

The fall is usually close to exponential, which means it has a half-life: an interval after which your snapshot retains one bit of every two it started with. Different variables decay at wildly different rates within the same file. A coastline is nearly stationary. A short-term interest rate is not. Lumping them into one "freshness" number destroys the calculation that makes the concept useful. Freshness is a vector, one entry per variable, not a scalar stamped on the whole dataset.

Shannon gave the quantity in 1948. The data processing inequality followed in the standard treatments of the following decades. Lorenz, working on atmospheric prediction in 1963 and 1969, showed that error growth imposes a finite predictability horizon independent of model quality — you cannot out-think a system into staying correlated with you. Landauer connected information retention to physical cost in 1961; Still, Sivak, Bell and Crooks closed the loop in 2012, showing that information a system holds which fails to predict the future is exactly what it must dissipate as heat. Decay is not a data-management inconvenience. It is thermodynamic.

From decay rate to a lineage

The three generations differ in when observation stops, and decay converts that difference into arithmetic rather than metaphor. A Large Language Model is a frozen corpus: mutual information with the world peaks at the cutoff date and falls monotonically afterwards, at a rate set by the subject matter rather than the model's cleverness. A Large World Model resets the clock while a scene is present — a sensor holds correlation near its ceiling for the duration of an episode, then the same exponential resumes the instant the sensor is switched off. A Large Universe Model is the limit case: intake never stops, so decay is bounded by sampling interval and channel bandwidth rather than by a calendar date. Provenance — recording when and by what instrument each belief was last measured — is what makes that bound computable rather than assumed.

There is no fourth category on this axis, because decay is a property of the gap between observations, and continuous observation is the smallest gap available. Better inference, better compression, better reasoning: none of it adds bits about the present, by the same inequality that opened this page. Only measurement restores mutual information. Once an architecture is permitted to measure every stream it can reach, with no stopping point, everything left to improve — more streams, lower latency, longer-trusted provenance — is scale and engineering, not a new kind of evidence.

The economist's intake

Central banking is a clean test of this because the object being tracked — an economy of several hundred million transacting agents — is precisely the kind of fast, high-dimensional system that punishes a frozen description. The streams an economist draws on are price indices, labour flows, credit aggregates, and market expectations extracted from yields and options. None of these is a snapshot in the sense a corpus is a snapshot. All of them are published on a lag, and the lag is the whole problem.

Consumer price indices are collected over weeks and released with a delay of two to four weeks depending on jurisdiction. Employment reports are compiled from surveys with sampling error large enough that the initial print and the print three months later can disagree by more than the number that moved markets on the day. Gross domestic product, the single number most associated with the health of an economy, is issued as an advance estimate, then revised, then revised again — the United States' third and "final" estimate for a quarter still arrives roughly ninety days after the quarter closed, and even that is subject to annual and five-yearly benchmark revisions that can move a growth figure by a percentage point with no new economic event having occurred, only a better measurement of the old one.

The characteristic failure

Put a decision on top of that and the failure becomes structural rather than occasional: policy is set on data that gets revised after the meeting. A rate-setting committee meets on a fixed calendar — eight times a year in most major regimes — and must vote using whatever vintage of each series happens to exist on that day. The vintage is not the truth about the quarter in question; it is a posterior formed from an incomplete, still-arriving sample. By the time the "true" figure stabilises, the meeting is long over and the rate has already moved, or failed to move, on the strength of a number that later turns out to have been wrong by an economically meaningful margin.

This is not a data-quality complaint that better statisticians could fix. It is the data processing inequality wearing a suit. No amount of reprocessing the flash estimate recovers the bits that only a later, larger sample of transactions can supply. The committee is, formally, a Large Language Model of the economy: a snapshot taken on a schedule, correlated with the world at its moment of collection and decaying from there, with the added cruelty that the decay is partly measured only after the fact, when the revision arrives and shows how wrong the snapshot already was.

A rate decision and a corpus cutoff are the same operation performed on different clocks.

Two objections worth taking seriously

Mutual information does not decay to zero. Double-entry bookkeeping, the arithmetic of compound interest, the legal structure of a central bank's mandate — these are stationary. A snapshot of an economics textbook retains full correlation with them forever. Insisting that intake must be continuous overstates the case and understates why economists trained on old textbooks still function.

Correct, and this is exactly why economic training and institutional memory are not wasted effort. The asymptote of mutual information here is the invariant structure — the Fisher equation, the accounting identities relating saving and investment, the mechanics of a yield curve — and a large body of frozen theory estimates that structure well. The claim under test is narrower: the decaying component is the one that determines the vote. Knowing the theory of inflation perfectly and knowing this quarter's inflation rate are different quantities of information, and a committee can be excellent on the first while carrying near-zero information about the second. Continuous intake is an argument about the second term, not a dismissal of the first.

Continuous intake does not defeat decay, it only shortens the lag. Every survey has a collection window, every release has a revision cycle, every price index is a report about a month already closed. A central bank with real-time payments data is simply a slower Large World Model running a faster refresh — a difference of degree dressed as a difference of kind.

The degree is conceded outright; no central bank observes the economy at the sampling rate of an order book, and none plausibly will. But the kind lies in what is structurally forbidden rather than in what is currently achieved. A fixed calendar of eight meetings a year, using officially published vintages, is a prohibition: no update to the working belief between meetings, whatever arrives. Card and debit transaction feeds, real-time payment-system data, and high-frequency job-posting indices — all now used by several monetary authorities as supplementary, unofficial trackers — remove that prohibition and replace it with an engineering budget: which streams, at what latency, at what confidence. That is the same structural move as opening a closed system in thermodynamics. Once nothing about updating is forbidden by design, remaining progress is spending on faster and denser measurement, not a wait for the next official release.

What the arithmetic asks of the economist

The corollary is not "watch everything, always." Still and colleagues' 2012 result says that information a system retains without predictive value is exactly what it pays for in dissipated effort — in this domain, in analyst hours and false-alarm rate cuts chasing noisy nowcasts. The correct response is per-series refresh set by each variable's own half-life: market-based inflation expectations, which move within the day, deserve near-continuous tracking; the benchmark revision to five-year-old GDP figures deserves almost none. That requires knowing, for every number on the desk, when it was last measured and by what instrument — provenance, in the technical sense — so that "how much can this belief be trusted right now" is a computable quantity attached to each series rather than a single unexamined assumption that the last press release is still good. A committee that can state its own decay constants, series by series, has done the only thing the inequality permits: it has stopped confusing the freshness of its data with the truth of its meeting date.

Continue