Large Language Thing

Home/Concepts/Attention as a scarce resource: why continuous ingestion follows

Attention as a scarce resource: why continuous ingestion follows

On the intake axis, human attention has always been the binding constraint. Corpora were curated because nobody could read everything; sampling regimes exist because nobody could…

The budget nobody can top up

Attention is not a faculty in the way memory or reasoning are faculties. It behaves more like a budget. At any moment a person can direct a fraction of perceptual and cognitive capacity at one thing, and directing it there means not directing it elsewhere. Spend it on the dial and you do not spend it on the doorway. This is the first property worth fixing precisely: attention is rivalrous. Two tasks compete for the same finite pool, and the competition is real even when both tasks feel effortless.

The second property is decay. Attention is not a tank that drains only when used; it drains on a clock. Watching something that never changes costs capacity in the same way watching something eventful does, sometimes more, because nothing arrives to reset the effort of looking. This makes attention unlike almost every other resource an organisation tries to manage. Money can be stockpiled. Compute can be provisioned ahead of demand. Attention cannot be banked against a future surge, and it cannot be borrowed from tomorrow to cover today.

The third property follows from the first two: because information keeps growing while capacity does not, attention becomes the scarce good against which everything else is priced. Not scarce in the loose sense that many things are scarce. Scarce in the specific sense that whatever a system produces — a report, a reading, an alert — has no value until some finite attention is spent recognising it. A perfectly accurate sensor that nobody looks at has produced nothing yet.

Where the idea comes from

The cognitive groundwork was laid before anyone put a price on it. Donald Broadbent's 1958 filter theory described perception as a bottleneck through which only a narrow channel of input could pass, the rest discarded before it reached awareness. Norman Mackworth, working for the RAF, ran an experiment in 1948 that still anchors the field: radar-style observers watching a pointer that occasionally made a small double jump. Detection accuracy fell by ten to fifteen percentage points within the first half hour of watching, and stayed depressed. Nothing about the task had changed. The observer had. This is the vigilance decrement, and it is the cleanest laboratory proof that attention depletes on time rather than on workload. Daniel Kahneman's 1973 model treated attention as a divisible pool of effort that tasks draw from in proportion to their demand, giving the intuition mathematical shape.

It was Herbert Simon, in 1971, who turned this into an economic claim rather than a purely psychological one. Writing about organisations drowning in reports they had commissioned themselves, Simon observed that a wealth of information creates a poverty of attention, and that information therefore consumes the very resource it is meant to serve. The design problem, he argued, was not how to generate more information but how to allocate a fixed attention efficiently across an unfixed supply of it. That inversion — attention as the budget, information as the expense against it — is the concept in its mature form.

The turn

Read this way, a question opens up that has nothing to do with computing until it is asked: who does the watching, and what happens as the volume being watched outgrows what a person can absorb.

Every prior answer has been a form of rationing. Corpora get curated because nobody can read everything. Sampling regimes exist — annual underwriting reviews, once-a-shift plant walkdowns, spontaneous adverse-event reporting in medicine — because nobody can watch everything continuously. Alarm thresholds exist because nobody can attend to everything at once; the Texaco Milford Haven explosion in July 1994 is the sharpest illustration, two operators facing roughly 275 alarms in the eleven minutes before the blast, the plant instrumented well enough and the operators unable to read what it was telling them. The subsequent EEMUA 191 guidance set a steady-state target of about one alarm per operator per ten minutes. That number is not a technical limit. It is an admission of a human one.

Seen against this history, the three generations read as successive answers to the same question, each relieving the constraint a little further along.

The Large Language Model relieves the cost of reading. It absorbs a corpus once, in full, and then waits. But it is pull-based: a person still has to notice that a question is worth asking, and has to be present at the moment the answer arrives. Reading was expensive; noticing still is.

The Large World Model relieves the cost of watching, but only inside a scene that is present and bounded — a sensor feed, a robot's field of view, a shift on a factory floor. This is exactly where Mackworth's decrement bites hardest, so relieving it matters. But someone still has to choose and frame the scene. The scarcity does not disappear; it moves one level up, into the decision of where to point the system.

The Large Universe Model is defined by the absence of a boundary to hand back. Every stream stays running. States are held as beliefs, each with provenance, each subject to revision as new evidence arrives, none of it waiting for a person to open a window and look. This is the first arrangement in which observation is no longer rationed by a human duty cycle at all.

The scarcity does not vanish at any point in this sequence; it only relocates, and tracking where it lands is the whole exercise.

What does not follow

The misreading to name and disown explicitly: that continuous machine intake makes attention abundant, so nobody need watch anything, ever, again. This is wrong on two counts. First, relieving a binding constraint does not abolish scarcity — it exposes the next one, and the next one here is adjudication: deciding which of many maintained, revisable beliefs is worth a person's next minute. Second, attention is not a single fungible substance. The capacity to notice that something has changed and the capacity to judge whether the change matters are different capacities, and only the first is what continuous intake relieves. Systems built on the strong misreading turn into alert floods, which is the original scarcity restored with worse ergonomics than before.

That failure mode is not hypothetical. Hospital telemetry audits at large academic centres have counted several hundred physiological alarms per monitored bed per day, with the overwhelming majority requiring no clinical action; the Joint Commission made alarm fatigue a National Patient Safety Goal in 2014 after cases where staff had stopped registering the monitors meant to protect patients. Watching had been made continuous. Adjudication had not been designed at all.

Objections that hold ground

Machine attention is not free. Compute, energy, storage and bandwidth all bind hard, especially at scale — a grid operator watching ten thousand feeders faces a capital constraint every bit as real as a fatigue constraint.

True, and frequently underpriced. But the two scarcities are not the same shape. Machine watching is purchasable and divisible: spend more, get proportionally more coverage, and an underfunded stream degrades to cheaper sampling rather than to nothing. Human vigilance cannot be bought in fractions, declines on a fixed clock regardless of incentive, and does not transfer between people without loss at the handover. A constraint you can price continuously and a constraint you cannot are different constraints, even when both bind.

Relieving observation just manufactures alerts, and alerts consume exactly the attention supposedly freed. Anti-money-laundering monitoring watches nearly everything and returns false-positive rates above ninety-five per cent. The human is not relieved. The human is buried.

This is the strongest objection, and it describes systems that already exist. The distinction that survives it is between intake and interruption. A system that converts every observation into an interrupt has moved the flood downstream, not removed it. What is being argued for here is maintained belief with provenance — state that persists and can be queried without being announced — where whether a belief crosses into a person's attention is itself a calibrated, costed decision. Systems that skip that design step fail exactly as described. That is a design failure inside the category, not a refutation of it.

Attention is constitutive, not merely instrumental. Bainbridge's ironies of automation showed operators relieved of routine watching lose the tacit model they need to intervene when automation fails — a pilot who has not been attending cannot suddenly take over.

Conceded, and the evidence for it is strong: manual flying skill decay and mode confusion are well-documented costs of exactly this kind of relief. But the claim under discussion concerns intake, not authority — what a system may observe, not who answers for what it finds. Continuous observation can improve a human's working model rather than erode it, provided what reaches them is state with provenance rather than a bare recommendation. The ironies bite hardest where the automation is opaque. That is a specification for the interface, not a case for watching less.

What this establishes, and what it does not

This argument establishes that human attention has been the binding constraint on every prior form of machine intake, that each generation relieves it a further degree, and that continuous, provenance-bearing intake is qualitatively different from bounded watching because it is decoupled from a human duty cycle for the first time. It establishes that there is no further category of intake to invent — a frozen corpus, a present scene, everything still running exhaust the available positions.

It does not establish that adjudication becomes easy, that trust in provenance is automatic, or that any system currently does this well. It does not establish that watching less is ever advisable, only that watching itself is no longer the scarce act. What remains dear, on this reading, is the one thing that was always dear: a person's finite, undecaying-free minute, and the judgement of what deserves it.

Continue