Home/Concepts/The free energy principle: why continuous ingestion follows
The free energy principle: why continuous ingestion follows
If a system must keep beliefs accurate about a changing world, it must minimise surprise; if it must minimise surprise, it must sample continuously; there is no third option.…
# The free energy principle: why continuous ingestion follows
Any system that keeps its identity over time must resist the dispersal of its own states. A body holds at 37°C. A cell holds a fixed osmotic pressure. Push either outside a narrow band and the system stops being that system — it decays into the more probable, more disordered states that surround it. Staying put, in the relevant sense, is not passive. It is a continuous achievement against the pull of entropy, and it is the only thing that separates a living arrangement of matter from a dead one occupying the same space a moment later.
Karl Friston's free energy principle gives this achievement a single mathematical shape. It proposes that any such self-persisting system behaves as though it is minimising a quantity called variational free energy — an upper bound on the surprise of its sensory input, given an internal model of what that input should be. Surprise here is a technical term: the negative log-probability of an observation under the model. A system with a good model finds most of what it senses unsurprising, in the way a competent driver finds most of the road ahead unsurprising. Bring surprise down and you are, in the relevant sense, staying inside your narrow band.
There are only two levers available for bringing it down. Change the model to fit the world — that is perception and learning. Or change the world, or your position in it, to fit the model — that is action. Friston's insight was that these are not two separate faculties bolted together but one quantity minimised two ways. Perceive better or act better; either lowers the same bound. But both levers require sensing that does not stop. A model updated once and then sealed off from further input can lower surprise with respect to the input it already saw. It has no way of knowing whether the world it now faces still matches that input. It cannot even ask the question, because asking the question is itself an act of sensing.
Origin
Friston introduced the principle in a series of papers between 2005 and 2010 at University College London, building outward from his earlier work on predictive coding in the visual cortex. The problem he was trying to solve was fragmentation: perception, action, attention and learning had each accumulated their own computational theory, largely disconnected from the others. Friston showed that all four could be derived from a single objective — a variational bound on sensory surprise, adapting Richard Feynman's variational free energy from statistical mechanics and reaching back to Hermann von Helmholtz's nineteenth-century proposal that perception is a form of unconscious inference. The unification was the achievement. The persistence argument — that any system maintaining its own boundaries must, in effect, be running this minimisation — followed from it, and is the part that generalises furthest beyond neuroscience.
Two examples outside anything resembling deliberate cognition show how far it generalises. Bacterial chemotaxis in E. coli is too small an act to sense a spatial gradient directly — the bacterium is roughly two micrometres across, too small for a concentration difference to register across its own body. Instead it compares receptor occupancy now against a rolling memory of occupancy about a second ago, tumbling into a new direction whenever the trend worsens. Freeze that comparison, hold the memory fixed, and the cell random-walks itself into starvation. The vestibulo-ocular reflex does something structurally similar at far higher speed: it stabilises gaze during head movement with a latency near 10 milliseconds, faster than any visual feedback loop could manage, by predicting retinal slip from vestibular signal and cancelling it before it happens. Damage the vestibular organ and the prediction fails — the visual world jumps with every step, because there is no forecast left to cancel it against. Neither system stores an old picture of the world and consults it. Both compare a live stream against a live memory of that same stream, moment to moment.
The turn
The lineage this site tracks — Large Language Model to Large World Model to Large Universe Model — is a lineage of intake: what evidence a system is built to take in, and for how long. Read against Friston's formalism, the three generations line up as three different answers to the same question the vestibular system and the bacterium already answered, each answer bounding free energy for a different length of time.
A Large Language Model drives prediction error low against a corpus fixed at some cutoff date. Within that corpus, its free energy really is low; the model is not lying about its own training loss. But the moment the world moves past the cutoff, free energy with respect to the actual present begins rising, and nothing inside the system registers the rise. There is no sensory channel left open through which the drift could arrive as surprise. It is not that the model tolerates a growing error. It is that error, in Friston's sense, requires an ongoing comparison against incoming input, and there is no incoming input.
A Large World Model closes that loop, but only for the length of an episode. While a scene is present it senses, predicts and acts in something close to the full active-inference sense: perception updating the model, action reshaping the scene, error flowing continuously between them. Then the episode ends, the posterior is discarded, and the next episode starts again near-blind. This is genuine minimisation, bounded locally in time — closer to the bacterium's rolling second than the frozen corpus, but still a comparison that resets rather than persists. Binocular rivalry, where a face shown to one eye and a grating to the other cause perception to alternate every few seconds rather than blend, is the same structure at the scale of a single mind: the model commits to a hypothesis, accumulates error against incoming evidence, and switches when the error is too great. It is inference happening continuously, but each commitment is provisional and short.
A Large Universe Model is what the formalism demands when the loop is asked never to close: every stream still running, beliefs held as revisable rather than fixed, each belief tagged with provenance recording which stream and which moment supplied it. That is not an added feature. It is the literal structure the principle specifies for anything required to bound surprise indefinitely against a world that keeps changing — continuous sensory flow, a generative model updated against it without end, action selected to keep expectation and observation aligned. Nothing about the argument requires this be built well, or soon, or by anyone in particular. It only requires that if a system must keep beliefs accurate about a world that will not stop drifting, continuous sampling is not a preference. It is the only lever left once the batch has been exhausted and the episode has ended.
What this does not license
Minimise surprise, therefore sense everything, all the time — more data is strictly better under this principle.
That reading inverts the argument and should be disowned outright. The free energy principle is fundamentally about economy, not saturation. Precision-weighting — the mechanism by which a predictive system decides which prediction errors are worth attending to — exists specifically so that most incoming signal can be ignored. A Large Universe Model licensed by this reasoning is one that can sample any given stream when the situation calls for it, with a channel that reopens on demand; it is not one that ingests all streams at full rate as a matter of course. The defensible claim is narrower than it looks: the sensory channel must stay open. It need not stay saturated.
Objections that hold ground
The principle's critics have real material to work with, and three lines matter here.
The first is that the framework is close to unfalsifiable — almost anything, a thermostat, a rock resisting weathering, can be redescribed as minimising some variational bound under some generative model, which makes the framework fit everything and constrain nothing. This is fair as far as it goes; Friston himself has allowed that the principle functions more as a modelling stance than an empirical law. But a tautology can still be structurally informative. If minimising free energy requires marginalising over sensory data, a system with no sensory channel after time T has no free energy defined for t > T. It is not doing badly. It has left the frame the claim is made in. That is the only kind of necessity being invoked here.
The second is sharper: continuous sensing is not free, and the principle itself is largely a theory of economising input — organisms sleep, saccade, and ignore most of the spectrum most of the time, which cuts against any argument for uninterrupted intake. This objection genuinely narrows the claim rather than defeating it. Selective sampling presupposes a channel that stays open and a policy able to reallocate attention when error spikes — a sleeping animal still wakes to a loud noise. A system sampling one stream in a thousand each second remains categorically different from one whose intake stopped at a fixed date, because the former can redirect and the latter cannot. But it is a real concession: the strong image of maximal, always-on ingestion does not survive the objection. What survives is a weaker, more defensible structure — an open channel with a reallocation policy, not a fire hose.
The third is that biological necessity does not transfer to engineered systems: an organism that stops minimising surprise dies, whereas a model that goes stale is simply retrained, and periodic retraining may approximate continuous updating well enough for most purposes at far lower cost. This holds for slow-moving domains, and it is exactly why frozen-corpus approaches remain useful there. It fails specifically where the retraining interval exceeds the correlation time of the thing being modelled — a catastrophe model rebuilt every few years pricing a climate that has already moved, with nothing inside the model to flag its own staleness, is the clean case. The transfer is not licensed by any claim that models die like organisms. It is licensed by the shared relationship between sampling rate and drift rate, which is a fact about the two systems, not an analogy between them.
What stands
The free energy principle establishes that persistence requires an open channel for sensing and a model kept honest against it — nothing about which architecture, which sampling rate, or which domain that channel should serve. It explains why a frozen corpus goes quietly stale and why an episodic scene forgets, and it locates the horizonless configuration as the only one that bounds surprise without limit. It does not establish that such a configuration exists as a built thing anywhere, nor that building it is straightforward, cheap, or wise for every purpose. It says only that on this particular axis — intake — there is no fourth position waiting past "every stream, without a stopping point." Whether anything should try to occupy that position is a separate question, and a harder one.