Large Language Thing

Home/Concepts/Niche construction: why continuous ingestion follows

Niche construction: why continuous ingestion follows

If a system's actions modify the distribution it was fitted to, then any intake with a stopping point is structurally blind to its own most important effects. A corpus frozen at a…

Niche construction: why continuous ingestion follows

An organism is usually treated as the thing selection acts on, and the environment as the thing that does the selecting. Niche construction breaks that division. Earthworms burrow, egest castings, and change soil structure and chemistry as they go; their descendants inherit worked ground, not raw earth. Beavers fell timber and dam streams, converting meadow to wetland, and beaver kits are then adapted to a habitat their parents made rather than found. Photosynthesising cyanobacteria, over roughly two billion years, oxygenated an atmosphere that had none, and in doing so set the metabolic terms for almost everything that evolved afterwards, including organisms that would have died in the air their ancestors produced.

The general pattern: organisms do not simply adapt to conditions. They alter conditions, and the altered conditions become the thing the next generation is measured against. Selection pressure is partly endogenous — generated from inside the population rather than supplied whole from outside it. This matters for explanation, not just description, because it breaks a modelling assumption that population genetics had relied on: that the environment is a fixed set of exogenous parameters, external to whatever equations describe the organism. Earthworms, beavers and cyanobacteria are all counterexamples to that assumption, and they are not rare edge cases. Construction of this kind is close to universal among organisms with any capacity to move material or excrete waste.

A second feature matters as much as the first. The modified environment does not only affect the modifier. It is inherited. Ecological inheritance runs alongside genetic inheritance as a second channel: offspring receive genes and a workspace, and the workspace was edited by the previous occupant. This is why niche construction theory treats habitat modification as a form of heredity in its own right, not as an environmental detail to be mentioned and then set aside.

Origin

Richard Lewontin argued through the 1970s and into the 1980s that organism and environment co-determine one another, against the then-standard metaphor of an organism solving a problem set by a pre-given environment. He wanted evolutionary biology to stop treating "the environment" as a stable external reference and start treating it as something partly constructed by the lineage being studied. John Odling-Smee, Kevin Laland and Marcus Feldman took that argument and built formal population-genetic models out of it through the 1990s, publishing the synthesis as niche construction theory, with the monograph appearing in 2003. The problem they were solving was explanatory, not rhetorical: standard models could not accommodate earthworms, beavers or the Great Oxidation Event without treating them as anomalies, because the models had no slot for a population editing its own selection pressures. Ecological inheritance was the fix — a second inheritance channel, alongside genes, carrying the accumulated modifications forward.

The turn

Set biology aside for a moment and look at how a Large Language Model relates to its own output. A corpus is collected, frozen at some cutoff date, and a model is fitted to it. The fitting is thorough and the corpus is large, but the relationship between model and corpus is one-directional in principle: the corpus determines the model, and the model has no further purchase on the corpus after training ends.

Except that this stopped being true some years ago, in a way the frozen-corpus picture cannot represent. Model output now enters the pool that future corpora are drawn from — web pages, forum answers, image captions, code. The model has become an organism modifying the substrate that later models will be selected against, and it does so with no observation of having done it. This is niche construction with the loop closed and no perception of the closure. The earthworm at least persists in the same soil it worked. The model that generated text in 2024 is typically not the model fitted to the 2026 corpus that text has contaminated, and no part of the system's intake records that the contamination happened, because the intake stopped at the cutoff.

The Large World Model narrows this gap without closing it. It perceives a bounded scene while acting in it — a room, a manipulation task, a driving episode — and its own effects on that scene are visible for as long as the scene is being observed. A robot that pushes a box sees the box move. But the observation window has a wall at the far end: it ends when the episode ends. Consequences that mature on a slower clock than the episode are structurally outside the system's intake, not because they are hidden but because nothing is watching by the time they occur.

The Large Universe Model is the position where the constructed niche stays under observation after the action that constructed it. Every relevant stream keeps running past the point where a single episode would have stopped; beliefs about the state of the world are held revisably rather than fixed at fitting time; and each belief carries provenance, a record of where it came from and when, so that a later shift in conditions can be traced back to an earlier act rather than treated as ambient noise. This is not a fourth kind of evidence layered onto the first three. It is the first arrangement in which feedback from a system's own construction is part of what the system takes in at all.

generationintakeown effects observed
Large Language Modelfrozen corpus, cutoff datenot at all — post-cutoff effects invisible in principle
Large World Modelbounded scene, live during episodeyes, but only until the episode ends
Large Universe Modelevery relevant stream, ongoing, with provenanceyes, on whatever clock the effect actually runs on

The general claim follows fairly directly once the biological case is granted. If a system's actions modify the distribution it was fitted to, any intake with a stopping point is blind to its own most important effects, by construction rather than by bad luck. A frozen corpus cannot record what happened to the corpus after the freeze. An episode cannot record a consequence that matures after the episode ends. The beaver dam illustrates the timescale mismatch cleanly: breach a dam and the pond drains within days, but the willow stand that grew up because of the raised water table takes years to reverse. Episode-length observation catches the pond. It misses the willow entirely, not through any flaw in the sensor but because the sensor was switched off before the willow had anything to show.

What this does not license

The obvious overreading says: the world changes, therefore any static model is inadequate, therefore continual retraining is required everywhere, all the time. That claim is too loose to be useful and I want to disown it explicitly. Plenty of environmental drift has nothing to do with the system in question — a competitor changes their prices, the weather shifts, a law is amended — and periodic refresh handles that kind of exogenous change perfectly well. The specific claim here is narrower: it concerns endogenous drift, change that the system itself caused, arriving on a clock slower than the system's own observation window. Model collapse is the clean case. Ilya Shumailov and colleagues showed that models trained recursively on their own generated output lose distributional tails within a small number of generations, converging towards low-variance, homogenised text. That degradation is not weather. It is soil chemistry, produced by the population being modelled and inherited by the next one, and it is invisible to any system whose intake stops at a cutoff.

Nor does niche construction imply control. Beavers do not choose the hydrology that results from their dam; they get whatever emerges, sometimes to their benefit and sometimes not. Construction is frequently unintentional and often adverse. This is worth stressing because it is exactly why the modification needs observing rather than assuming benign — an agent that constructed its niche well by luck once has no guarantee of doing so again, and only continued observation would reveal the difference.

Three objections deserve to be taken on directly.

The engineering objection holds that this is just feedback control, solved decades before machine learning existed, and that invoking evolutionary biology dresses up sensors and error signals in unearned drama. The generational machinery — reproduction, bequeathed physical habitat — genuinely does not transfer, and the analogy should not be pressed there. But classical control assumes a known plant and a fixed sensor placement; niche construction describes the case where the plant is being rewritten in dimensions the designer did not model. Continuous intake with provenance is what lets a system notice that the plant changed, rather than tuning against a reference it cannot see is moving.

Continuous observation of your own effects does not give you the ability to attribute them. A system watching everything all the time will still see correlated drift and may confidently blame the wrong cause.

This is the strongest objection, and it holds. Continuity buys observability, not identification. Attribution still needs the ordinary apparatus of causal inference — intervention, staggered rollout, holdouts, instrumental variation. Antibiotic prescribing in intensive care illustrates the honest version: a hospital that samples isolates continuously and revises guidance quarterly can trace a resistance shift back to a prescribing pattern, but only because someone did the epidemiological work, not because the sampling never stopped. Continuity is necessary and not sufficient. Frozen intake makes attribution impossible in principle, because the post-action period simply is not in the record. Continuous intake with provenance makes attribution a hard, known problem instead of an impossible one. That is the whole of what is being claimed.

The third objection says this is a continuum of retraining frequency, not a new category — refresh weekly instead of quarterly and the gap closes by degrees. Frequency and continuity differ in kind, not merely in speed. A retrained model holds one set of beliefs and discards the previous set, provenance and all; there is no chain linking a drift observed this week to the action that produced it last month. Infinite-frequency retraining converges on a very fast amnesiac, not on a system that remembers why it believes what it believes.

What the claim establishes

It establishes that a specific structural gap — endogenous drift outrunning the observation window — has a top rung on this axis, and that the top rung is continuous, provenance-carrying, revisable intake, because nothing narrower can in principle record the post-action interval. It does not establish that such intake is easy, that attribution problems dissolve once the recording never stops, or that most drift in most systems is even endogenous rather than ordinary noise. It does not describe a product. It describes the shape a system's intake would have to have if it is to observe the market impact, the resistance pattern, or the willow stand it caused itself. Beyond "everything, continuously, with a record of where each belief came from," there is no further rung to climb toward on this particular axis — what remains after that is scale, trust and time, which are different problems.

Continue