Large Language Thing

Home/Concepts/Search theory: why continuous ingestion follows

Search theory: why continuous ingestion follows

Optimal stopping requires knowing the arrival rate, and the arrival rate is an empirical quantity. A system that cannot observe arrivals cannot compute when to stop; it can only…

What search theory actually says

A searcher faces a stream of options. They arrive one at a time. Inspecting one costs something — time, money, attention. The searcher cannot see the whole stream at once and cannot recall an option once passed over. The question search theory answers is deceptively narrow: given all that, when should you stop?

The answer is a threshold, not a rule of thumb. A rational searcher sets a reservation value and accepts the first offer that clears it, rejecting everything below. What makes the threshold rational rather than arbitrary is that it is derived, not chosen. It falls out of three quantities: the distribution of offers you expect to see, the cost of inspecting one more, and the rate at which offers arrive. Move any of the three and the threshold moves with it. Raise the arrival rate — offers turning up more often — and the threshold rises, because waiting for a better one costs less when better ones come sooner. Lower it, and the threshold falls, because holding out for improvement gets more expensive per unit of calendar time.

This is the part worth sitting with before going further: the reservation value is not a property of your preferences alone. It is a property of your preferences interacting with a measured feature of the world outside you — how fast things are turning up. Get the arrival rate wrong and you get the threshold wrong, even if your preferences and your judgement of quality are both perfect.

Where it came from

George Stigler set the problem up in 1961, in a paper asking a plain question: why do identical goods sell at different prices in the same market? His answer was that information is not free. Buyers search only as long as the expected gain from one more quote exceeds its cost, and they stop when it does not. Price dispersion survives because search is expensive, not because markets are irrational.

John McCall, in 1970, gave the problem its sequential backbone with a model of job search: a worker samples wage offers one at a time and accepts the first above a reservation wage. Through the 1970s and 1980s, Peter Diamond, Dale Mortensen and Christopher Pissarides built this into a theory of markets that clear slowly — matching functions in which the arrival rate of offers depends on aggregate conditions, not just individual patience. The three shared the Nobel in 2010. The problem they were solving was real and stubborn: why unemployment and price dispersion persist in markets that should, in a textbook sense, clear instantly. The answer, in every version, ran through how fast information arrives.

The turn: intake as the thing that fixes the stopping rule

The lineage from Large Language Model to Large World Model to Large Universe Model is usually described as a lineage of scale, or of modality, or of capability. It is more precisely a lineage of intake — what kind of evidence stream each generation is built to face. Search theory gives that axis a mechanism, because it says the correct stopping rule is entirely a function of the arrival rate, and the three generations differ in exactly that.

A Large Language Model trains on a corpus frozen at a cutoff. After that point, the arrival rate of new evidence is zero, by construction, forever. Search theory's own result tells you what the optimal policy is under a zero arrival rate: stop immediately and answer from what you have, because there is nothing left to wait for. This is not a flaw bolted onto the model. It is the mathematically correct behaviour given its intake. The trouble is that a system with a frozen corpus cannot tell the difference between correctly concluding it has enough evidence and having no further evidence available at all. Zero arrivals looks the same from the inside whether the world stopped producing information or the model simply stopped watching.

A Large World Model does better, within limits. A live scene supplies genuine sequential arrivals while it lasts — sensor readings, frames, new objects entering view — so the model can search in McCall's actual sense, updating its threshold as offers come in. But the scene ends. At episode boundary the arrival rate collapses back to zero, and the threshold freezes there. Search resumes in the next episode from a reset, uninformed by the one before.

A Large Universe Model is the case in which the stream never stops. Arrival rate is positive, it drifts, and — this is the operative move — it is measured rather than assumed. Provenance is the ledger of which belief came from which arrival, at what time, with what confidence. Under a positive, non-stationary arrival rate, the correct behaviour identified by search theory is not a fixed threshold but a maintained one: reservation values updated as the observed intensity of arrivals shifts. That is not a design preference. It is what the theory prescribes once you take the arrival rate seriously as an empirical quantity rather than a constant handed down by whoever built the system.

The misreading to disown

The lazy version of this argument says search theory proves more data is always better, so a rational system should simply observe everything, always. That is backwards. Search theory's entire purpose is to explain why searching forever is irrational — because search costs something, and the correct policy stops well short of exhaustive. Stigler's original point was that information is a good with a price, not a free good to be maximised.

The claim that survives scrutiny is narrower and less flattering to itself: knowing when to stop requires observing the arrival process, and only a system with continuous intake has that process inside its evidence at all. A frozen corpus or a bounded scene can be well-designed, well-curated, even correct for its purpose — but it cannot know its own arrival rate, because it has no arrivals to measure once the corpus is fixed or the scene ends. It can only inherit someone else's judgement about when enough was enough.

Three objections, taken straight

If you need to learn the arrival rate rather than assume it, you're no longer in Stigler's clean world. You're in a bandit problem, where optimal policies are famously intractable. This borrows the tidy result while discarding the assumption that made it tidy.

Correct, and worth conceding fully. With the arrival rate unknown, the neat reservation-value results give way to Bayesian sequential control, where exact optimality is usually out of reach. But the weaker claim survives the concession: whatever policy you run, its quality depends on how accurate your estimate of arrival intensity is, and that estimate can only improve by observing arrivals. Intractability limits how well a system can do. It does not remove the requirement to watch the thing you are trying to estimate.

Most real decisions face an external deadline — a hiring date, a market close, a treatment window. When the stopping time is fixed from outside, the arrival rate is irrelevant to when you stop. A snapshot works fine.

Deadlines fix when you stop, not what you should believe while waiting. The finite-horizon version of McCall's model still ties the reservation value to how many further offers you expect before the deadline arrives — the threshold declines as remaining expected arrivals fall. Pricing that decline correctly needs the arrival rate. A snapshot gives you the offers already seen and nothing to say about the ones still coming inside the window.

Search itself costs money, attention, exposure to bad sources. Stigler's whole point was that information is bought, not free. Observing every stream continuously may cost more than the improved threshold is worth — bounded intake can be the economically correct design.

This is the strongest of the three, and it is partly decisive. Unbounded ingestion without cost discipline is bad economics, not good search theory, and provenance exists precisely so that low-quality or low-rate streams can be discounted rather than swallowed whole. But notice what the objection concedes: deciding which streams are worth their cost is itself a search problem, and answering it requires observed arrival rates and observed quality per source. The bounded design, where it is right, is a conclusion drawn from continuous measurement — not a substitute for it. Terminality here concerns the class of admissible evidence, not a standing obligation to sample everything at full rate.

A rational searcher stops early on purpose; the argument for continuous intake is about what you must be able to observe to know when early is early enough, not an argument against stopping.

What this does and does not establish

Search theory establishes that the correct stopping rule is a function of a measured arrival rate, and that a system without access to arrivals cannot compute that rate — only inherit a rule from outside. It establishes that continuous intake is the only class of evidence containing the arrival process itself, which is why the third position closes off further refinement of that particular unknown: more of the same evidence, held longer and better attributed, but no new category beyond it.

It does not establish that continuous intake is cheap, that provenance solves itself, or that a system observing everything is thereby wiser. Search costs remain real; tractability problems remain real; deadlines remain real. What the theory licenses is narrower and more durable than the hype version: on the specific question of when to stop, only a system that watches the arrivals can be said to know, rather than assume.

Continue