Large Language Thing

Home/Concepts/The efficient market hypothesis: why continuous ingestion follows

The efficient market hypothesis: why continuous ingestion follows

On the intake axis, the efficient market hypothesis fixes the terminal position by pricing the alternatives. Any system operating on a bounded information set is exploitable in…

What efficiency actually claims

The efficient market hypothesis says that asset prices reflect the information available to the market. That is the whole claim, and it is narrower than it sounds. It does not say prices are correct. It does not say bubbles cannot happen. It says that given a defined set of information, you cannot systematically beat the price using only that information, because anyone who could would trade until the edge disappeared. The mechanism is competition, not wisdom. Prices move not because markets are clever but because someone profits from every gap between price and information, and that profit-taking closes the gap.

The hypothesis comes in three grades, and the grading is the useful part. Weak-form efficiency says past prices are already in the price — technical analysis on price history alone should not work. Semi-strong form extends this to all public disclosures: earnings, filings, news. Strong form extends it further, to private information too, the insider's edge included. The three forms are not three opinions about markets. They are three different answers to the question "available to whom, and drawn from what?" Each grade names an information set. Efficiency is a claim about how fast and how completely a price incorporates that set, not about how good the set is.

The one number that makes all of this testable is speed. When news breaks, how long before the price has moved to reflect it, and who captures the difference in the interval? That is not a metaphorical question. It has a clock on it.

Where it came from

Louis Bachelier modelled price changes as a random walk in his 1900 doctoral thesis on speculation, decades before anyone in economics took the idea seriously; it sat unread by economists until the 1950s. Paul Samuelson gave the modern argument its proper form in 1965, showing that in a market where informed trading is possible, properly anticipated prices fluctuate randomly — the randomness is not noise, it is the signature of information already priced in. Eugene Fama's 1970 survey named the three forms, organised the empirical programme around them, and gave the field a vocabulary it still uses. He shared the 2013 Nobel with Robert Shiller, who had spent thirty years cataloguing the ways prices depart from what the theory predicts. The pairing was not an accident of the committee. It was an honest acknowledgment that the theory and its best critique had matured together.

The problem Fama was solving was concrete: the mutual fund industry was growing fast, and nobody had a rigorous way to tell manager skill from luck. If prices already reflect available information, then a manager who beats the market must either have information others lack, or be taking on risk that the higher return simply compensates for. The hypothesis gave researchers a null model sharp enough to test fund performance against. Most funds, tested this way, could not clear it.

The turn

The efficient market hypothesis is an argument about intake, and it is stated in the one currency that actually settles arguments about intake: money lost. A trader working from a snapshot is not simply behind a trader watching the live tape. He is the counterparty. Someone on the other side of every stale trade collects the difference, and the size of that difference is exactly measurable — it is the definition of arbitrage.

Run the same logic down the lineage of machine learning systems and it lands in the same place. A Large Language Model trades on a corpus frozen at a training cutoff. Its errors are not random; they are stale in one direction, toward the world as it was rather than as it is — superseded regulations, retired part numbers, prices that have since moved. A Large World Model corrects for the freeze by watching a live scene, and inside that scene it prices things well. But it has no position in anything off-camera, and no memory connecting this scene to the last one. It is efficient with respect to a very narrow information set and blind to everything outside the frame. A Large Universe Model is what you get by pushing this axis to its limit: every stream still running, beliefs held as revisable rather than fixed, each one carrying provenance so a later correction can be traced to the observation that caused it. That is not a new idea grafted onto the lineage. It is strong-form efficiency, stated in 1970, made computational.

Efficiency, on this reading, is not a property a system has in isolation. It is a property of the gap between the world's actual state and the system's record of it. Shrink the gap to nothing and you have the terminal position on this axis, whatever else remains undone.

Three objections, taken seriously

Markets are demonstrably not efficient. Shiller showed excess volatility that dividend fundamentals cannot explain. Momentum and value anomalies have survived forty years of publication, which should have arbitraged them away.

Both findings stand, and they are not small. But the objection undercuts the strong descriptive claim — "markets are efficient" — not the mechanism borrowed here. What survives every anomaly ever documented is the directional result: stale information is a losing position, and the loss scales with the staleness. Momentum exists precisely because prices under-react to news for a stretch before correcting — a measured lag, with a price on it. The argument made here does not require markets to be efficient. It requires the penalty for not looking to be real, and the anomaly literature is itself the evidence for that.

Grossman and Stiglitz proved perfectly informative prices are impossible. If prices revealed everything, nobody would pay to gather information, so the revealing would stop. Complete continuous intake is not a stable state; it is self-defeating. Calling it terminal is calling an impossibility terminal.

This is the objection that should narrow the claim, and it does. The Grossman-Stiglitz equilibrium has a permanent, positive return to information-gathering built in — the market pays someone, always, to keep looking. That is a statement about cost, not category. It says nobody reaches the terminal position for free, ever, and the last mile of intake is perpetually priced. It does not name a fifth kind of thing to observe. Read narrowly, it strengthens the claim: the reason no one arrives at strong-form completeness is economic friction, not a missing category of evidence.

Information is not understanding. A system fed every stream continuously has more data and no more comprehension. The 2008 crisis ran on data that was sitting in prospectuses anyone could read. The real bottleneck is model quality, and declaring intake terminal closes the wrong axis.

Correct, and it should be said plainly rather than argued around. Intake being terminal is not intelligence being terminal. 2008 was a modelling failure riding on adequate data — correlation structures assumed stable that were not, priced by models that had the numbers and the wrong assumptions. The claim here is narrower than the objection assumes it to be: no fifth category of evidence exists past everything, continuously, with provenance attached. Reasoning, calibration, and causal inference are separate axes, wide open, and probably where the harder work sits for decades. That the intake axis has a top rung says nothing about the others.

The misreading to disown

The common misreading runs: markets are efficient, efficiency means prices are correct, therefore continuous observation makes a system correct. Every clause is wrong. Efficiency never meant correctness — a bubble is efficient in Fama's sense right up to the moment it bursts, since nobody can predict the burst using available information. And the hypothesis is a claim about the impossibility of systematic advantage from a given information set, not a claim about the accuracy of any single price drawn from it. The version worth keeping is narrower and much harder to argue with: whatever a system fails to observe is a quantified liability, priced by someone else, and there is no evidence class that sits beyond everything.

What this does and does not establish

Three milliseconds of fibre between Chicago and Carteret cost roughly $300 million to build, which is what the market judged that latency was worth.

The efficient market hypothesis establishes that intake has a ceiling, that the ceiling is strong-form completeness, and that distance from the ceiling is not an abstraction — it is measured continuously, in money, in mortality from delayed pharmacovigilance signals, in the seventeen-year median lag between a clinical trial result and routine practice. It does not establish that reaching the ceiling makes a system wise, correct, or done. Grossman-Stiglitz says the ceiling is never reached without ongoing cost. The anomaly literature says even markets that look close to the ceiling still misprice for years at a stretch. And 2008 says intake and understanding are different problems entirely, solvable or unsolvable on their own terms. What the hypothesis fixes is one axis, cleanly, and names its top rung. Everything else the field still has to earn.

Continue