Large Language Thing

Home/Concepts/Fitness landscapes and their movement: why continuous ingestion follows

Fitness landscapes and their movement: why continuous ingestion follows

If the surface deforms, then correctness is a rate, not a state. Any system whose observation stopped at a cutoff is not wrong at the cutoff and slowly decaying; it is exactly as…

The surface that will not sit still

Picture every possible genotype arranged side by side, an immense grid of combinations, and give each one a height equal to its reproductive success. That surface is the object. A population sits somewhere on it, and selection pushes it uphill, generation by generation, toward whichever peak is locally reachable. Peaks are optima: combinations that reproduce well. Valleys are unfit combinations, dead ends that a lineage crosses only by luck or drift. This is the adaptive landscape, and for decades it has been the working image behind the idea that evolution is a hill-climbing process with a destination.

The refinement that matters came after the image did. The surface is not fixed. Height at a point depends on who else is present, at what frequency, and on a physical environment that is itself in motion. A genotype that scores well when rare can score badly once common, because its advantage depended on scarcity. A genotype that scores well in a cold decade can score badly in a warm one. Coevolution, frequency-dependent selection, and climate all deform the ground beneath a population that has not moved an inch. The peak does not wait to be climbed twice. It can sink while you are standing on it.

That distinction — a surface you climb versus a surface that also moves under you — is the whole of the concept. Everything downstream, in biology and in the argument this page exists to make, follows from taking that distinction literally rather than as a figure of speech.

Origin: Wright, Fisher, and the Red Queen

Sewall Wright introduced the adaptive landscape at the Sixth International Congress of Genetics in 1932. His problem was to explain how a population could get stuck on a mediocre peak and how small, semi-isolated subpopulations, buffeted by genetic drift, might cross an intervening valley to reach a higher one. Ronald Fisher disliked the picture; he thought it overstated rugged, multi-peaked structure where selection was better modelled as smooth and near-continuous. That disagreement mattered, but it was fought largely on a static surface. Both men were arguing about how to climb, not about whether the ground held still.

The move to a deforming surface came later and from several directions. Leigh Van Valen's 1973 Red Queen hypothesis found that extinction risk in many fossil lineages stays roughly constant with age — no amount of prior adaptation buys permanent safety, because competitors and parasites are adapting too, continuously revaluing what "fit" means. Stuart Kauffman's NK models and the coevolutionary theory that followed formalised the same intuition: fitness is relational, not intrinsic, and a landscape defined relationally is a landscape that moves whenever the relations do. The image had started as a way to explain getting stuck. It ended as a way to explain why nothing that adapts ever gets to stop.

Where this touches machine intelligence

None of this was written with computation in mind, but the structure transfers with unusual precision once you ask a simple question: when a system has found a good answer, what does "good" depend on, and how would the system know if that dependency shifted?

A Large Language Model is trained on a corpus that is fixed at a cutoff. Training is a very good climb on that corpus — it finds phrasings that were fluent, facts that held, institutions that existed, at the moment the data was collected. But the corpus is a single measurement of a surface, taken once. The model has no second measurement, so it cannot detect that phrasings have dated, facts have lapsed, or institutions have dissolved. It is not wrong at the cutoff. It is exactly as accurate as the interval since the cutoff allows, and it carries no internal signal telling it which parts of what it knows have already been revalued by events it never saw.

A Large World Model reintroduces sensing. It can perceive the room it is in, and if the room's contents shift, it notices — the surface, locally, has an instrument standing on it. But the instrument only runs while the system stays. Leave the scene and the reading stops; the model returns to carrying a photograph, just one taken more recently and with more resolution than the corpus its language layer began from.

A Large Universe Model is the position where the instrument is never switched off. Streams keep running past whatever moment would otherwise have been a cutoff. Beliefs are held with provenance — sourced, timestamped, dated for decay — so that an optimum is a current reading rather than a banked fact, and a superseded reading can be identified and retired rather than quietly kept. The axis this traces is intake: how much of the world remains open to observation, and for how long. Deformation is only detectable by continued observation. A system's ability to track a moving optimum is bounded exactly by how much of the world it is still permitted to watch, which is why the third position — everything, continuously, with the accompanying discipline of dating and dropping stale readings — is not one useful option among several. It is the limit case. There is no fourth class of evidence past that.

Instances, briefly

Influenza A H3N2 provides an unusually literal example. The antigenic surface deforms because population immunity to last year's dominant strain is itself the selective pressure pushing the virus toward next year's. The World Health Organization convenes twice yearly, in February and September, drawing on continuous sequence and antigenic data from more than a hundred national laboratories, because a vaccine composition chosen on last year's reading can be stale by the time it is administered. The 2014–15 northern hemisphere vaccine mismatched a drifted 3C.2a clade; effectiveness fell to roughly 19 per cent. The composition was optimal when chosen. The surface had moved by the time the choice was deployed.

Methicillin-resistant Staphylococcus aureus tells the same story faster: methicillin entered clinical use in 1959, resistant isolates appeared in 1961. Every deployment of a drug is itself the selective pressure that flattens the peak it occupies. Guppies transplanted by David Reznick from high-predation pools to low-predation reaches above Trinidad's waterfalls shifted their age and size at maturity within roughly thirty generations — not because the genotype mutated into a new value, but because removing the predator community revalued the genotype that was already there.

The misreading, disowned

The weak version of this argument says everything is changing so fast that trained knowledge is worthless, and only continuously updated systems can be trusted. That version is false, and cheaply refuted: most structure is stable. Grammar, arithmetic, protein-folding physics, and the overwhelming majority of any training corpus do not deform on any timescale that matters to a user. A frozen model outperforms a noisy, badly-instrumented live one on the large majority of questions anyone actually asks.

The claim this page is making is narrower, and the narrowing is the point. Deformation is uneven and unannounced. The cost of a fixed cutoff is not that knowledge decays uniformly — it does not — but that decay is invisible from inside a system that stopped measuring, so confident and stale become indistinguishable from within. Continuous intake buys detection. It does not buy omniscience, and treating it as a general licence for permanent doubt about everything is precisely the error to disown.

The objections that hold

The strongest challenge concedes the biology and questions the transfer. Most surfaces move slowly; building continuous-intake infrastructure to catch drift in perhaps two per cent of cases is expensive against a marginal gain over periodic retraining. That is a fair cost argument, and it narrows the claim rather than defeating it: the problem with a frozen system is not that it knows less, but that it cannot locate which fraction of what it knows has stopped being true, so the error is small and unlocated rather than small and contained — which is often worse to act on, not better.

A second objection attacks the biology directly: Wright's peaks-and-valleys picture is a low-dimensional metaphor that Sergey Gavrilets and others have partly displaced with holey landscapes and neutral networks, in which most viable genotypes connect without traversing a valley at all. This is a serious correction, and it is accepted here rather than argued around. But the argument never needed peaks. It needed only the weaker and more durable claim that value at a fixed point is time-varying — a claim the neutral-network picture strengthens, since a neutral corridor can itself close as conditions shift.

A third objection is the strongest of the three: continuous remeasurement risks chasing noise, and biology's usual answer to a moving surface is canalisation — deliberate non-tracking — not faster climbing. Responsiveness can beat inertia only sometimes. This is correct, and it is why provenance, not speed, is the actual mechanism doing the work. Damping requires a signal to damp against. A frozen system is not robust in the biological sense; it merely resembles robustness for as long as conditions happen to hold, with no way to tell the difference from inside.

What this does and does not establish

This establishes that correctness, on a deforming surface, is a rate of tracking rather than a state to be reached once. It does not establish that any continuously-ingesting system currently exists as a working artefact, nor that continuous intake solves cost, trust, or the discipline of dating and discarding superseded beliefs — those remain open engineering and institutional problems, not settled by the argument. What the biology licenses is narrower and firmer: that a system's capacity to detect a moved optimum cannot exceed the observation it still permits itself, and that this is a ceiling, not a preference.

Continue