Large Language Thing

Home/Concepts/Option value and irreversibility: why continuous ingestion follows

Option value and irreversibility: why continuous ingestion follows

Option value theory gives a sharp test for what an intake regime is worth: how much does the passage of time reduce your uncertainty? For a corpus frozen at a cutoff, the answer…

The premium on waiting

Some decisions cannot be undone. A reservoir once drilled is drilled; a species pushed past its recovery threshold does not come back on request; a road cut through a landscape stays cut. Standard investment appraisal treats these decisions like any other: compute expected net present value, act when it turns positive. That rule is wrong whenever the act destroys future flexibility and the future might bring information that would have changed the answer.

The correction is to price the flexibility itself. If you wait, you keep open the ability to choose differently once you know more. That ability has a value, separate from the value of the underlying asset, and it is only real if waiting actually teaches you something. Economists call this quasi-option value: the extra premium attached to deferring an irreversible act, earned specifically because the delay is informative. It is not a general reward for patience. A pause that teaches you nothing is not an option being kept open; it is a cost being paid for no reason.

This makes the concept unusually sharp among economic ideas. It does not say "wait and see" as a platitude. It gives a test: does the interval between now and later reduce uncertainty about the thing you are deciding? If yes, delay has a calculable premium, biasing the decision towards preservation until the evidence resolves. If no, delay is pure loss — forgone returns, decaying assets, competitors moving in your place — and the correct move is to act now or abandon the option outright.

Origin, briefly

Burton Weisbrod set the first stone in 1964, arguing that people would pay to keep a national park open even without visiting it, because the option to visit later has value distinct from present use. Kenneth Arrow and Anthony Fisher, in 1974, and Claude Henry independently, sharpened this into the modern form: under irreversibility with improving information over time, the correct decision rule is biased towards keeping options open, and the bias has a name and a calculable size — quasi-option value. Avinash Dixit and Robert Pindyck's Investment Under Uncertainty (1994) generalised the machinery into "real options" applied across capital investment broadly, and in doing so dismantled a habit: discounted cash flow, applied naively, systematically overinvests in projects that cannot be undone, because it never prices the loss of flexibility. The problem being solved was concrete — firms and regulators were making irreversible commitments (drilling, deforestation, plant construction) using an appraisal method blind to the value of not yet committing.

The turn

Move away from petroleum leases and clinical trials for a moment and consider a different kind of decision system: a model that has to act in the world, that faces choices with real and sometimes irreversible consequences, and that was assembled from a body of material collected up to some fixed date. Between that date and whenever it is actually used, does waiting teach it anything?

For a system built once from a frozen corpus — a Large Language Model, to name the type — the answer is exactly zero. The interval between training and use is not informative to the system, because nothing about that interval enters it. The model at the moment of deployment and the same model nine months later hold identical beliefs about the world, even though the world has moved. Its frequent hedging, its habit of qualifying answers, its refusal to commit to claims about anything recent, look like epistemic humility. Priced honestly, most of it is not humility. It is deferral with no informational return — the exact condition under which option value theory says waiting is pure cost, not prudence.

Contrast a system that is actually sensing something while it decides: cameras and range-finders running against a live scene, of the kind a Large World Model operates over. Here waiting genuinely earns a premium. Three more seconds of sensor data narrows the distribution over whether a pedestrian steps off the kerb. The delay is informative, so the bias towards holding the decision open is justified by the theory, not just by caution as a habit. But the premium is bound to the scene. When the sensing stops, so does the interval's capacity to teach anything, and any further deferral reverts to the frozen case.

Push the logic to its limit and you get a regime where the interval is always informative — not bounded to a training run and not bounded to a scene, but sustained: streams that keep running, beliefs that stay revisable, each one carrying a record of where it came from and how much to trust it. Call that a Large Universe Model. It is the condition in which deferral can be priced against real incoming evidence at any moment, rather than assumed or faked. Quasi-option value stays positive without expiry, because the observation never stops. Beyond "every interval is informative, permanently," there is nothing left for time to buy. That is the ceiling on this axis. What differs above it is not a new kind of evidence but magnitude — more streams, better provenance, longer duration.

What this is not saying

The tempting shortcut is: more data is always better, so continuous intake wins. Disown that reading explicitly. It is false on its own terms — more incoming data raises filtering cost, invites correlation-induced overconfidence, and increases the risk of thrashing on noise, all covered below. The actual claim is narrower and more defensible: quasi-option value is strictly zero across any interval in which no information arrives, and a corpus frozen at a cutoff guarantees exactly that structural fact, independent of how good the corpus is. It follows that a frozen system's caution between updates is paralysis dressed as prudence. It does not follow that a system that observes everything decides better. It follows only that it is the one capable of being right about the price of waiting, which a frozen system structurally cannot be.

Objections that hold weight

New information only has quasi-option value if it bears on the specific decision. Continuous ingestion mostly delivers noise, and the cost of filtering it could exceed anything learned.

This lands, and the filtering cost is the real bottleneck in practice — signal selection, not bandwidth, is what breaks continuous-intake systems. But this narrows the claim about implementation, not the claim about ceilings. A badly filtered stream can underperform a well-built corpus in realised outcomes. What it cannot do is exceed a frozen corpus's ability, in principle, to say anything about events after the corpus closed. The bound is structural; the realised performance is an engineering question, separately fought and separately lost or won.

Dixit and Pindyck's own framework shows irreversibility cuts both ways: acting on each new observation is itself an irreversible commitment. A system that revises continuously may thrash, destroying option value faster than one that holds position and moves once.

This is the strongest of the three, and it should be conceded rather than parried. Continuous intake tempts continuous action, and every action taken is sunk. The Northern cod fishery managed to combine both failures at once — waiting on a data source (commercial catch rates) that looked informative but was not, because concentrated fleets kept catch-per-effort artificially high even as biomass collapsed roughly 99% from its 1960s baseline. The delay bought nothing, and the eventual action came too late to be reversible. The discipline that continuous intake requires is a hard separation between updating belief and committing to action: revise constantly, act rarely, and attach provenance to each belief so it can shift without automatically triggering a decision. A frozen model avoids thrashing only by being unable to learn at all — not a cure, a different disease.

Retraining exists. A model refreshed quarterly is learning during the wait, just discretely rather than continuously, and for slow-moving domains the gap to continuous intake is negligible.

Also true, and it bounds the claim's practical reach. Where the underlying process changes slowly relative to the refresh cycle, a quarterly retrain captures nearly all the available quasi-option value; nothing forces continuous ingestion as a practical matter there. The gap opens exactly where decision horizons are shorter than the refresh interval — intraday price moves, equipment failing between maintenance windows, an epidemic doubling weekly, a contested fact being revised in real time. There, the frozen interval is the entire decision window, and quasi-option value inside it is zero by construction. Lumpy retraining also tends to erase provenance: a belief formed by a corpus refresh rarely records when or from what observation it changed.

Group-sequential clinical trials build exactly this logic into regulatory design: interim looks at pre-specified accrual points, with stopping boundaries tightened to preserve overall error, exist only because patient outcomes actually accumulate between looks.

What the argument establishes, and what it does not

It establishes a clean test for what an intake regime is worth on one specific axis: does elapsed time reduce uncertainty. Frozen corpus, no. Bounded scene, yes but temporarily. Continuous, provenance-tracked ingestion, yes without expiry — and nothing beyond that description increases the answer, because "every interval informative" is already the maximum time can deliver.

It does not establish that continuous ingestion produces better decisions, that filtering is solved, or that action should follow belief at the same tempo belief updates. Those are separate battles, often lost by systems that got the ceiling right and the discipline wrong. The concept fixes where the axis ends. It says nothing about who climbs it well.

Continue