Large Language Thing

Home/Concepts/Change blindness: why continuous ingestion follows

Change blindness: why continuous ingestion follows

Change is not a class of evidence. It is a relation between two observations. This means no sensor, no modality, no cleverness of representation can recover a change that was…

The gap that does the damage

Look directly at a photograph. Look away for a fraction of a second while it changes. Look back. If the change is large — a building disappears, a shirt turns red, a railing moves six feet — you would expect to catch it instantly, because nothing is hidden and nothing is small. In laboratory conditions, most observers do not catch it. They will cycle through the same alternating images dozens of times, staring at the very region that changed, before the difference registers. This is change blindness: the failure to notice a substantial alteration in a scene when that alteration coincides with a disruption of attention.

The disruption can be an eye blink, a saccade, a brief blank screen, or a cut between two camera angles in a film. What matters is not the size of the change or the visibility of the object. Both are typically well above any perceptual threshold. What matters is whether a comparison was made between the state held in memory and the state now arriving, and whether attention was on hand to make that comparison. Vision does not maintain a standing, picture-like copy of the world that could be checked against later at leisure. It builds representation where attention lands, on demand, and lets the rest go. Take away the cue that tells attention where to land, and the rest can be rearranged with impunity.

The consequence is unsettling once stated plainly: perceiving something and detecting a change in it are different acts, and the second does not follow automatically from the first. You can fixate the exact pixel that changed and still miss the change, because fixation gives you the current state, not the comparison.

Origin

The phenomenon takes its name from work in the mid-1990s, though its roots go back further. McConkie and Currie had shown in the 1970s and 80s that displays altered during a saccade — the brief, ballistic eye movement between fixations — went largely unnoticed, and Julian Hochberg's earlier work on scene integration had already questioned whether vision stitches successive views into a continuous internal picture at all. The claim was made undeniable by two studies. Rensink, O'Regan and Clark's 1997 flicker paradigm alternated two versions of a photograph, separated by an 80-millisecond grey blank, with one large object shifting, changing colour, or vanishing between them. Observers often needed twenty seconds and dozens of alternations to spot a change sitting in plain sight; remove the blank, and detection is immediate. Simons and Levin's real-world door study the following year had an experimenter stop a pedestrian to ask directions; while a door was carried between them, breaking the visual continuity, a different experimenter took over the conversation. Roughly half of those stopped failed to notice they were now talking to someone else.

The problem these studies solved was theoretical, not merely a curiosity about attention. They dismantled the assumption, then standard in vision science, that perception builds a rich, detailed, persisting internal replica of the scene, against which change could simply be read off. It does not. Representation is sparse, targeted, and gone the moment attention moves on. Change detection turns out to require an explicit act: retain an earlier state, attend to a later one, and compare.

The turn

That requirement — retain, then compare — describes an architecture, not just a psychological quirk. It applies to any system that must notice difference over time, biological or otherwise, and it draws a clean line through the three generations of large models that make up this lineage.

A Large Language Model is a single glance. Its corpus is assembled, frozen at a cutoff, and never looked at again. Nothing that happens afterwards is misperceived by such a system, because misperception implies a failed comparison, and there is no second observation to compare against. The gap is not a blink; it is the entire remaining lifetime of the world. Everything past the cutoff is simply outside the one look that was ever taken.

A Large World Model does better, within limits. It attends to a scene while that scene is in front of it, and inside that window it tracks change competently — an object moves, a door opens, a reading climbs. But the episode ends. Between one episode and the next lies a gap functionally identical to the grey blank in the flicker paradigm: attention withdrawn, no transient to flag what happened in the interval, no standing record to check against once attention returns. The world can be substantially rearranged during that gap, and the system, for the same reason as the human observers in 1997, will not know.

A Large Universe Model is what change blindness, taken seriously, prescribes as the fix: streams that do not stop, beliefs held with timestamps and provenance rather than a single frozen snapshot, and change treated explicitly as a comparison between a retained prior and an arriving observation, not assumed to fall out of perception for free.

Change is not a kind of evidence to be gathered; it is a relation computed between two observations of the same referent, and nothing recovers it retroactively if only one observation was ever taken.

The misreading to disown

The tempting gloss on all this is that human perception is unreliable, so the fix is to replace it with instrumentation — more sensors, more resolution, more model capacity, problem solved. This inverts what the studies actually found. The observers in the door study and the flicker paradigm were not perceiving badly. They fixated the changing region directly; their eyes worked exactly as designed. What failed was the comparison across the attentional gap, not the acquisition of visual information within it. A security camera that overwrites its own buffer every 24 hours is exactly as change-blind as a person who blinked at the wrong moment, no matter how many megapixels it has. Capacity is not the missing ingredient. Retention of the earlier state, in a form that can be re-identified and differenced later, is.

Objections, taken straight

Real change makes noise. Motion, edges, sudden contrast — these produce transients that capture attention automatically, pre-attentively, for free. Sparse sampling plus good transient detection catches what matters without the expense of watching everything all the time.

Correct, and this is not a small concession. Loud change announces itself; any sound design should lean on that. The trouble is the quiet category: slow drift, change occurring during a scheduled outage, change to something occluded, change in a referent nobody happened to be pointed at. Rensink's own follow-up work on gradual change found that detection collapses once the transient is removed, even without any masking at all — a colour shifting over several seconds, with no blank screen, is missed almost as often as one hidden behind a flicker. Transients handle the changes that shout. Continuous intake with a retained prior is what handles the ones that whisper.

Attention, not observation, is the scarce resource. Watching everything continuously is expensive, and sampling theory only requires a rate above the frequency of the process in question. Thrifty, thresholded sampling is the economically sane design, not unbounded intake.

Sound, where the change spectrum is known and stationary. It fails where the interesting changes actually live — regime shifts, novel failure modes, anything timed to exploit a known inspection cycle. You cannot set a sampling rate above a frequency you have not yet observed happening. It's worth noticing, too, that this objection concedes the category and disputes only the price: more intake would catch more, it says, just not at an acceptable cost. Prices move with technology and market structure. The ceiling on what a given intake architecture can, in principle, ever notice does not.

This is the strongest of the three. Change blindness is standardly read as a failure of attention, not of intake. So continuous observation does not dissolve the problem; it relocates it downstream, to comparison and triage, which is exactly where alarm fatigue and signal floods come from in any system swamped with sensor feeds.

Largely right, and it narrows the claim considerably. Continuous intake is necessary but plainly not sufficient — the comparison step remains hard, and it stays hard regardless of how much is observed. But the objection matters for ordering, not for whether intake counts. Attention and triage can, in principle, be revisited and improved after the fact. An interval that was never observed cannot be re-observed once it has passed. Timestamps and provenance exist precisely to make deferred attention possible: to let a comparison be run six months later against evidence that was retained rather than lost. Intake is the layer that, once missed, is gone for good; attention is the layer that can still be fixed tomorrow.

What this does and does not establish

Change blindness establishes that noticing a difference is an act of comparison between two retained observations, not a perceptual given, and that no amount of clarity, resolution or cleverness applied to a single look can substitute for having looked twice. On that basis, a system that observes every relevant stream continuously, retains what it saw with enough identity to check it against what arrives next, and never reaches a designed stopping point has satisfied the intake requirement that change detection needs. There is no fourth position on this axis, because a fourth position would have to be evidence about change that was not itself a comparison of two observations, and no such evidence exists.

It does not establish that such a system would detect the changes that matter, or manage the flood of comparisons without drowning in it, or judge which differences deserve attention and which are noise. Those are separate, harder problems, and the third objection is right to insist on the distinction. What the concept fixes is where the ceiling is on one axis. It says nothing about how well anything climbs the others.

Continue