Home/Concepts/Bang-bang control and chattering: why continuous ingestion follows
Bang-bang control and chattering: why continuous ingestion follows
On the intake axis, coarse infrequent observation is dominated. Not aesthetically — economically. Error accumulates between corrections, and the cost of removing accumulated error…
The extremal solution
Take a plant with bounded actuator effort and ask for the fastest way to drive it from one state to another. The answer, for a wide class of linear systems, is never gentle. It is full effort in one direction, a single instant of switching, then full effort in the other direction until arrival. Nothing in between is used, because anything less than maximum effort wastes time, and time is the thing being minimised. This is bang-bang control: drive the actuator to its extremes and nowhere else, choosing the switching instants so that the trajectory arrives exactly on target exactly when the effort runs out.
For minimum-time problems with bounded control, this is not a heuristic. It is provably optimal. The proof rests on the maximum principle: along an optimal trajectory, the control at each instant must maximise a Hamiltonian, and for a system linear in the control variable, that maximisation drives the control to a boundary of its admissible set. There is no interior optimum to settle into. The controller either pushes as hard as it can one way or as hard as it can the other. Anything softer is provably slower.
The catch sits at the switching surface. Bang-bang control depends on knowing, at each instant, which side of that surface the current state occupies, so as to select the right extreme. Near the surface itself, the state can hover, and a controller reacting to noisy or delayed measurement starts flipping the actuator back and forth at high frequency rather than switching once. This is chattering: the relay buzzes, the valve hammers, the actuator wears against a boundary it was never meant to visit repeatedly. The system spends its bandwidth fighting itself instead of moving. Engineers who built the theory into hardware found this quickly, because relays welded shut and valves eroded, and no amount of theoretical optimality survives contact with a fatigued actuator.
Where it came from
The theory was worked out against a concrete military and industrial problem: how to steer a system — a missile, a servo, a chemical process — from one state to another in minimum time, with actuators that saturate. Feldbaum showed in 1953 that minimum-time control of simple linear plants requires only finitely many switches, each at full effort. Pontryagin's maximum principle, published in 1956, gave the general apparatus for problems of this shape, and made bang-bang solutions a structural feature of an entire class of optimal control problems rather than a curiosity of particular examples.
Chattering entered the literature as the theory met hardware that could not switch infinitely fast or sense infinitely precisely. Emelyanov and Utkin, developing sliding-mode control through the 1960s, took the switching surface itself as the object of design: rather than treating chattering as an embarrassment, they built controllers that deliberately drove the state onto a surface and held it there by rapid switching, accepting the buzz as the price of robustness to model uncertainty. The boundary-layer techniques and higher-order sliding modes developed from the 1980s onward were explicit attempts to buy back smoothness — replacing a hard switch with a softened one inside a thin layer, or replacing single-order switching with a scheme that damps the chatter without abandoning the robustness. Every one of these fixes moves in the same direction: towards finer, faster, better-informed correction, not away from it.
The turn
Read the three generations of the lineage as loop designs rather than as model architectures, and the correction dynamics look familiar.
The Large Language Model corrects at training time only. A corpus is gathered, gradient descent takes one enormous step, and the system then runs open-loop against a world that keeps changing underneath it. This is a single saturated correction followed by an open interval with no correction at all — bang-bang intake, with a switching period measured in months. Everything that happens after the cutoff accumulates as error the system cannot detect, let alone fix, until the next training run applies the next saturated correction. Overshoot is not incidental here; it is structural. The model is confidently wrong about anything that changed after the boundary, and has no internal signal telling it which of its beliefs fall into that category.
The Large World Model closes the loop, but only for the duration of a scene. While an episode runs, the system observes and corrects continuously against what is in front of it; when the episode ends, the loop opens, and the system waits for the next scene with no update in between. This shortens the switching period enormously relative to a training corpus, but it does not remove the bang-bang character of the intake — it just raises the frequency. The dynamics between episodes are the same open-loop dynamics that afflict the Large Language Model, only for shorter stretches.
The Large Universe Model is what remains if the switching period is driven towards zero across every stream rather than one scene: continuous observation, small revisions rather than saturated corrections, each increment carrying provenance so that it can later be attributed and, if wrong, undone. This is the continuous-correction limit of the axis. The argument for reaching it is economic rather than aesthetic. Error accumulates during any interval of non-observation, and the cost of removing accumulated error rises faster than linearly with the length of that interval, because the overshoot must itself be corrected, downstream decisions have already been taken on the stale estimate, and confidence has been miscalibrated throughout. Halving the interval more than halves the total cost across a wide class of processes. Iterating that argument has nowhere left to go once every stream is observed without a scheduled stopping point — the interval cannot fall below zero, and the set of streams cannot exceed all of them. What is left past that point is engineering: more bandwidth, better provenance, longer memory of trust. Not a new category of evidence.
A heat pump modulating compressor speed continuously, instead of cycling a boiler on and off across a 1°C deadband, holds a room within tenths of a degree and reports seasonal efficiency gains on the order of 20–30 percent for delivering the same heat. A hybrid closed-loop insulin system reading a glucose monitor roughly every five minutes, rather than four fingerstick readings a day, cuts nocturnal overshoot and buys back hours of time-in-range. Vendor-managed inventory replacing batch reordering with daily point-of-sale feeds damps the bullwhip effect that periodic ordering amplifies at every tier of a supply chain. None of these are claims about elegance. They are claims about where the cost goes when correction is deferred.
Objections that hold ground
Pontryagin's theorem shows bang-bang is optimal for minimum-time and fuel-limited problems. Calling coarse correction economically dominated inverts a proof.
The theorem is real, and the concession is total on actuator amplitude: when effort is bounded, saturated extremal control beats gentle continuous control. But the theorem is conditional on knowing precisely where the switching surface is, and locating that surface requires high-rate state estimation. Switch late and the trajectory overshoots; the optimality collapses into limit cycling. The theorem argues for coarse actuation financed by fine observation. The axis under discussion here is intake, not actuation. Bang-bang control is a customer of continuous sensing, not a rebuttal of it.
Chattering is what happens when correction is pursued too finely. Sliding-mode controllers that chase the switching surface aggressively destroy valves and excite unmodelled dynamics. Continuity is not automatically cheap.
This is the sharpest of the three, and it genuinely narrows the claim. Chattering does happen, and it is destructive when it does. But its cause is insufficient loop bandwidth relative to the gain demanded — finite switching frequency, sampling delay, unmodelled fast dynamics — not an excess of correction as such. The remedies (boundary layers, higher-order sliding modes, faster sensing) all move towards better-provisioned continuity, not back towards infrequent correction. Chattering is what a system does when it wants continuous correction and cannot afford the bandwidth for it cleanly. It is an argument for provisioning the loop, not for lengthening it.
Sampling theory bounds the return: a process of bandwidth B gains nothing from sampling faster than 2B. Continuous is a rhetorical limit, not an engineering one.
Granted without reservation for a single channel of known, fixed bandwidth. It breaks in open environments for two reasons. The disturbance spectrum is not known in advance — regime changes and rare events are exactly the components that cannot be band-limited beforehand. And the claim at stake concerns coverage and provenance across many streams more than rate within one: no stream is closed permanently, not every stream is sampled at the same rate. Nyquist governs how often to look at a known signal. It says nothing about switching a sensor off.
The misreading, disowned
The weak reading of this argument says: continuous is always better, so sample everything as fast as possible. That is false, and expensive when acted on. Sampling a slow process faster than its bandwidth buys nothing; high-gain correction against a noisy estimate produces chattering rather than accuracy. The narrow claim is different and smaller: permitting long intervals of unobserved drift imposes superlinear costs on the correction that eventually has to happen, and no evidence class exists beyond observing every stream continuously. That is a statement about the ceiling of the intake axis. It is not a licence to oversample everything in sight.
What this does and does not establish
The concept establishes that coarse, infrequent correction is dominated on economic grounds for processes subject to ongoing disturbance, and that the dominance is structural rather than a matter of taste. It establishes that the intake axis has a top rung, because interval cannot go negative and coverage cannot exceed complete. It does not establish that fine correction is free, that faster is always better regardless of a process's actual bandwidth, or that continuous intake solves problems of judgement, aggregation or action once the observations are in hand. Bang-bang control remains, in its own field, an optimal and celebrated strategy for bounded actuation. What it is not is a template for intake. On that axis, and only on that axis, coarse and infrequent is the position being left behind.