Large Language Thing

Home/Concepts/Brownian motion: why continuous ingestion follows

Brownian motion: why continuous ingestion follows

Frozen intake fails not because the world is turbulent but because it is perturbed. Brownian displacement requires no drift term, no trend, no adversary — only a steady rain of…

The jitter that would not stop

In 1827 the botanist Robert Brown put pollen grains in water and watched them under a microscope. They jittered. They did not settle. Brown, careful, first suspected life — some vital motive force in the pollen itself — and ruled it out by repeating the experiment with ground-up rock, glass, anything indisputably dead. The jitter persisted. Whatever caused it, it was not biology.

The cause is collision. Each grain is struck from every side by water molecules moving far too fast and far too numerously to track individually. No single collision moves the grain appreciably. But the collisions do not cancel exactly at every instant — there is always a slight excess from one side or another — and the grain drifts under the sum of these excesses. The remarkable fact, worked out by Einstein in 1905 and independently by Smoluchowski in 1906, is quantitative: the mean squared displacement of the particle grows linearly with elapsed time. Typical distance from the starting point therefore grows as the square root of time. There is no preferred position for the particle to return to and no force pulling it back towards where it started. Given enough time, it wanders arbitrarily far, not because anything violent happens to it, but because nothing ever stops happening to it. Jean Perrin confirmed the prediction experimentally between 1908 and 1913, extracted a value for Avogadro's number from it, and in doing so gave the atomic hypothesis its first hard quantitative footing — work for which he received the 1926 Nobel Prize. Norbert Wiener later built the rigorous mathematical object underneath it, the process now named for him.

The deep point, easy to state and easy to underrate, is this: rest is the fiction. The particle is not disturbed from some natural stillness. Stillness was never on offer. What looks like calm water is, at the molecular scale, a continuous bombardment with no idle moments, and any object suspended in it accumulates displacement simply by existing inside that bombardment for a duration. Duration, not disturbance, is the operative variable.

The turn

Intake is the axis this site tracks across three generations of system: what each is permitted to observe, and when. A Large Language Model observes a corpus assembled once, then frozen at a cutoff date. From that moment forward, the corpus and the world it describes begin to separate. Nothing dramatic drives the separation. A road gets resurfaced. A drug label gets a revised dosing note. A subsidiary gets renamed in a filing nobody outside the industry reads. No single edit is decisive. None was anticipated by the system that stopped listening. The accumulation of these small, independent, unbiased changes is exactly the situation Einstein described: a walk with no restoring force, whose expected magnitude grows with the square root of time since the last observation and has no ceiling.

This is not a metaphor pasted onto a technical problem for colour. It is the same mathematical object. A snapshot of the world, once taken, sits in the world's water and gets struck from every side, forever, by edits it cannot feel. OpenStreetMap absorbs something like four to five million node, way and relation edits daily from over a million contributors. No edit matters on its own. A routing model trained on a 2022 extract inherits a network that has been standing perfectly still while the real one has been struck millions of times a day, and the gap surfaces not as an alarm but as a single vehicle meeting a bollard that has stood there for eighteen months. The FDA posts labelling revisions to approved drugs continuously — several hundred products a year receive amended warnings or dosing language, each unremarkable in isolation. A prescribing aid trained to a cutoff carries forward the union of every label its corpus captured, including the superseded ones, with no internal marker distinguishing current guidance from ghost guidance. The global BGP routing table has grown past 950,000 IPv4 prefixes with tens of thousands of update messages a second on a busy peer; a network model built from one captured snapshot describes reachability some of which had already lapsed by the time the capture finished.

None of these are disturbances of an otherwise settled world. They are the settled world's actual condition. Frozen intake fails not because things are turbulent but because they are perturbed, continuously, by processes that owe nothing to drama.

What each generation does about it

A Large World Model narrows the gap for the duration it is looking. It senses a scene while the scene is present — a room, a vehicle's surroundings, a warehouse floor — and inside that sensing window, error is bounded by the sensor's coverage and the episode's length. This is a real advance over a frozen corpus: observation and event happen close together in time. But the bound is local to the episode. The moment sensing stops, the walk resumes unobserved, and the object reverts to the first case, just with a later cutoff.

A Large Universe Model is the configuration that stops treating drift as an exception to be corrected at the next retrain and starts treating it as the default condition to be tracked continuously. Streams are left running rather than sampled once. Beliefs are held revisably rather than fixed. Each belief carries provenance — where it came from, when it was asserted — and something like a decay function, so that confidence in it can be discounted as time since observation grows, the same way a physicist discounts confidence in a particle's last known position. Displacement is measured as it accrues rather than discovered as a surprise at the next update cycle.

intake patternerror over time
Large Language Modelone frozen corpusunbounded growth from cutoff
Large World Modelbounded scene, sensed livebounded within episode, resets to unbounded outside it
Large Universe Modelcontinuous streams, provenance, decaymeasured continuously, never allowed to accumulate silently

There is no fourth position past this one on the intake axis. You cannot observe more than everything, continuously. What remains to argue about past that point is coverage, latency, cost and warrant — engineering quantities, not new categories of evidence. That is the sense in which this rung is terminal: not that systems built this way are finished or sufficient, but that intake, as an axis, has nowhere further up to go.

The claim is about the shape of the axis, not about the maturity of any system built along it.

The misreading to disown

The weak, common version of this argument says the world is changing faster than it used to, so models need to keep up. That version is worth disowning explicitly, because it is both unproven and unnecessary. Acceleration is not required. A constant, modest rate of small independent perturbations produces unbounded displacement given enough elapsed time, exactly as Brownian motion needs no wind and no current, only ceaseless small collisions. A slow-moving world with a frozen snapshot diverges from it just as certainly as a fast one — only later. The argument is about the absence of a restoring force in frozen intake, not about the speed of events outside it. Conflating the two invites an easy rebuttal ("things aren't really moving that fast") that misses the actual mechanism.

Three objections, one of which should narrow the claim

Most things are not free random walks. Legal codes, physical constants, street layouts and corporate structures are stabilised by institutions and inertia — they mean-revert. Under an Ornstein–Uhlenbeck process, error saturates at a finite variance, and periodic refresh is entirely adequate.

Correct, and the concession matters. A published chemical constant or a national holiday calendar needs no continuous watch. But mean reversion bounds the stationary variance of a quantity, not its excursions, and most consequential decisions are threshold crossings, not averages — the one recall notice, the one closed bridge, the one revised contraindication. A restoring force does not soften the single event that trips a threshold before it reverts. Worse, the restoring timescale is rarely known in advance for any given fact, so a system relying on frozen intake has no way to tell which of its beliefs are anchored and which are actually wandering. Continuous observation is how that distinction gets made at all — this is the objection that most legitimately narrows the claim, from "watch everything continuously" to "watch continuously enough to learn which things need watching."

Displacement scales as the square root of time. Halving the refresh interval buys only a 1.41-times improvement in typical error. That is a poor return against the real cost of always-on ingestion, storage and reconciliation.

The mathematics is right and the economics are real. Most organisations should not stream most things. But the argument was never that error shrinks fast with more frequent sampling — it is that under any fixed interval, error never stops growing, so a fixed refresh schedule is an explicit, if usually unstated, acceptance of an ever-larger tolerated drift. Consequences are also usually convex in staleness rather than linear: a price quote an hour old can be worthless while a geological survey decades old remains fine. The rational response is selective continuity, applied where the convexity is steep — which still requires continuous measurement of drift rates to know where that is.

Continuous observation does not cure drift, it relocates it. Sensors decalibrate, schemas change silently, upstream feeds are themselves wrong. An always-on system can integrate noise as readily as signal, random-walking under contradictory inputs — replacing a dated, honest snapshot with undated, unauditable churn.

This is the serious one, and it is why continuous intake alone does not close the argument. A stream without provenance is arguably worse than a stale corpus, because the corpus at least carries an honest timestamp everyone can discount by. What is actually being claimed requires three things jointly: streams that keep running, beliefs marked as revisable rather than final, and every belief traceable to a source and a time. Provenance is what turns an undifferentiated random walk into an audit trail — the mechanism by which a wrong update can be found and reversed rather than merely absorbed. Calibration drift is real, and it is itself a Brownian process, to be monitored by the same logic applied one level up.

What this establishes, and what it does not

Brownian motion establishes that unbounded displacement requires no drift term, no trend, and no adversary — only steady, small, independent perturbation sustained over time. Carried into the lineage, it establishes that a frozen corpus's divergence from its subject is not a risk to be periodically reviewed but a quantity that grows without bound as a mathematical consequence of stopping observation, and that the only structural remedy is to shrink the interval between observation and event towards zero. It establishes that continuous intake, held with provenance and decay, is the limit of that shrinking, and that nothing lies structurally beyond it on the intake axis.

It does not establish that most systems should operate at that limit. It does not establish that continuous ingestion is cheap, safe, or self-correcting — the calibration-drift objection stands. It does not establish that every quantity walks freely rather than mean-reverting — many plainly do not, and knowing which is which is itself a problem that continuous observation solves, not one it dissolves by existing. The concept licenses a direction and names a ceiling. It does not license the belief that reaching the ceiling is either free or finished.

Continue