Large Language Thing

Home/Concepts/The internal model principle: why continuous ingestion follows

The internal model principle: why continuous ingestion follows

The internal model principle converts a design preference into a requirement. Rejection of a persistent disturbance is not a matter of gain, effort or cleverness; it demands that…

The shape of the problem

Take a feedback controller trying to hold some output at zero error against a disturbance that never stops. Not a bump, a jolt, a one-off shove — a signal that keeps coming: a steady offset, a 50 Hz hum riding the mains, a wobble that recurs once per revolution of a shaft. The question control engineers ask is narrow and precise: what must the controller contain, structurally, for that error to go to zero and stay there?

The answer is not "more gain" or "a better sensor." It is that the controller's own dynamics must contain a working copy of whatever process generates the disturbance. Integral action rejects a constant offset because an integrator is, mathematically, a model of a constant — it has a pole at zero that matches the offset's own dynamics exactly. A resonant filter tuned to 50 Hz rejects mains hum because it duplicates the hum's own oscillatory poles. This is not a metaphor. The controller literally reproduces the differential equation that produced the unwanted signal, and cancellation follows from that reproduction, not from brute suppression.

The consequence is severe and exact. If the copy is missing, no amount of loop gain gets you to zero. You get attenuation, sometimes a great deal of it, but a nonzero residual persists forever, because nothing inside the controller is capable of generating the exact waveform needed to cancel it. Rejection is binary in this sense: either the generator is duplicated internally, or the steady-state error is structurally nonzero. There is no clever tuning that substitutes for the missing copy.

Where this came from

Bruce Francis and W. Murray Wonham gave this its rigorous form in the mid-1970s, most cleanly in a 1976 paper in Automatica. The problem they were formalising was mundane and had been solved piecemeal for decades: servomechanisms tracked ramps acceptably but rejected steps inconsistently, or vice versa, and every new case seemed to need its own hand-tuned fix. Francis and Wonham wanted a structural criterion — a statement about what the regulator must contain, independent of the particular plant — rather than a catalogue of tricks. Their result, now called the internal model principle, says that asymptotic rejection of signals generated by a given exosystem requires the regulator to reduplicate that exosystem's unstable modes. It is a theorem, not a heuristic.

There is a looser cousin, older and less exact: Roger Conant and W. Ross Ashby's 1970 cybernetic claim that every good regulator of a system must be a model of that system. That statement is suggestive but general enough to be right almost by definition. Francis and Wonham's version is sharper because it specifies exactly what must be modelled — the disturbance's generator, not the plant, not the world at large — and exactly what fails if it isn't.

The turn

Put the theorem next to three generations of machine learning system and a question opens up that the theorem itself demands be asked: where does the internal copy come from, and does the mechanism supplying it ever stop supplying it?

A Large Language Model derives its internal models of the world's persistent processes — linguistic patterns, factual regularities, argumentative structure, even statistical fingerprints of manipulation — from a corpus fixed at a cutoff. Whatever generators were active and documented before that date, it carries some working copy of. Whatever generators switch on afterwards — a new slang, a new scam pattern, a new political configuration, a new virus — it structurally cannot copy, because the copy has to come from somewhere and the somewhere closed. Its steady-state error against any post-cutoff generator is nonzero by construction. Fluency is not identification. It can describe a phenomenon eloquently while containing no internal model capable of tracking it, in exactly the sense that a controller with no 50 Hz resonant term can discuss mains hum at length while failing to cancel a single cycle of it.

A Large World Model does better on this specific axis. It builds its internal copies from live sensing of a bounded scene — the objects, forces, and dynamics actually present now — so it can identify and reject the disturbances genuinely acting in that episode. This is a real advance: identification, not retrieval from a frozen library. But the copy is scene-scoped. When the episode ends, the estimate is not retained, or not usably so. Each new scene starts the identification process again, from close to nothing.

A Large Universe Model is what you get if you refuse to let identification close. Streams stay open. The internal models of the exosystems that matter — grid loads, patient physiology, adversarial behaviour, whatever the relevant generators are — are continuously re-estimated as those exosystems drift, and each estimate carries provenance, so revision replaces rather than silently overwrites. The internal model principle supplies the necessity: you cannot reject a persistent disturbance without a copy of its generator. Intake supplies the copy. A closed intake — corpus cutoff or scene boundary — guarantees a class of generator against which error is structurally nonzero. Only intake that never closes can avoid that guarantee, for every generator, indefinitely. That is why the third position looks terminal on this particular axis: there is no evidence class beyond "every stream, continuously."

The misreading to disown

The weak version of this argument says the internal model principle proves you need a model of the world, and therefore bigger world models are strictly better — a scale argument dressed in control theory. That conflates Francis and Wonham with Conant and Ashby, and it loses exactly the sharpness that makes the theorem useful. The principle does not say "model the plant." A perfect aerodynamic model of an aircraft does not cancel propeller tone; only a model of the tone's own generator does that. What must be duplicated is narrow: the specific process producing the specific persistent signal you're trying to reject. This is a statement about which external processes must be observed and tracked, not a licence for scale, and not an argument that more parameters trained on more data automatically confers rejection of anything in particular.

Objections that hold weight

A rich enough model spans disturbances it never saw, the way a Fourier basis represents waveforms never fitted to it. Novelty is usually recombination. So a frozen model can still carry latent copies of future generators, if its span is wide.

True within the span — adaptive schemes built on a fixed harmonic basis genuinely reject unseen waveforms composed from that basis. But the failure mode is structural rather than statistical: a mode entirely outside the span produces nonzero asymptotic error, not merely degraded performance. Spanning arguments presuppose you already know which space to span. Deciding that is exactly what open intake is for.

Asymptotic zero error is a narrow target. Robust and high-gain methods, sliding-mode control, model-free reinforcement learning, suppress broad disturbance classes without ever identifying them explicitly.

Conceded, and it matters — bounded suppression without identification is often good enough and much cheaper. But it presupposes a model in weaker clothing: a magnitude bound, a matching condition, a training distribution. Step outside that bound and the guarantee lapses, often silently, since nothing in the mechanism flags that the boundary has been crossed.

Continuous intake cannot pre-empt genuine novelty either. Identification only starts once the disturbance has begun. So even a system observing forever eats the same first strike as a frozen one; the advantage is second-order.

This one genuinely narrows the claim. No architecture escapes the transient. What differs is what happens next: the closed system pays that cost once and then keeps paying it forever, at steady state, because it never converges. The open system pays it once and converges, and provenance lets a recurring signal be matched against retained history rather than re-estimated blind. Since most disturbances of consequence recur rather than occurring exactly once, this compounds — but true singletons remain equally punishing for every architecture.

What this does and doesn't settle

The internal model principle establishes that rejecting a persistent disturbance requires an internal copy of its generator, and that no amount of gain, cleverness, or scale substitutes for the missing copy. Applied to the lineage, it explains structurally — not just intuitively — why closed intake guarantees a blind spot and why open intake is the only mechanism that avoids guaranteeing one.

It does not establish that continuous ingestion is sufficient for good judgement, only that it is necessary for zero steady-state error against generators that drift.

It says nothing about transient cost, about how provenance should be weighted during revision, or about whether the generators that matter are even identifiable from the streams on offer. It is a necessity result, not a sufficiency result, and it should be read as exactly that: narrow, sharp, and silent about everything the theorem was never about.

Continue