The evacuation that follows the fire
An emergency manager issues an evacuation order for a wildfire front. The order itself changes the hazard's operating environment: roads fill, fuel-reduction crews are pulled off the line to help with traffic control, shelters draw down municipal water pressure that the fire service was counting on for hydrants. The order was a response to the hazard. It becomes, within the hour, a cause of new hazard conditions. Nobody designed it to be that. It is that anyway.
This is niche construction, a term from population biology, applied somewhere it was not built for. The fit is not perfect and the page will say where it breaks. But the core mechanism transfers cleanly enough to be useful, and it names something emergency managers already feel without a word for it: the environment they are managing is partly a product of their own management, on a lag they rarely get to see.
Two positions, stated plainly
The first position: emergency management is applied control theory, and has been since long before machine learning existed. Sensor, threshold, response, feedback — this is the fire service's incident command structure, the flood warden's stage-trigger, the hurricane cone's evacuation zones. Treating an evacuation order as "niche construction" borrows the drama of evolutionary time for what is, structurally, a control loop with a lag. The fix for lag is faster sensing and tighter loops, not a new vocabulary.
The second position: the standard control loop assumes a known plant. You know the relationship between input and system state well enough to design a sensor and place it correctly. What an evacuation order does is rewrite the plant itself, in dimensions the original sensor placement never anticipated. The traffic model did not include "fuel crew reassigned mid-response." The water system model did not include "shelter occupancy spikes draw down pressure the fire service needs." These are not measurement lags on a known variable. They are unmodelled variables that exist only because the response created them. A tighter loop measures the wrong thing faster.
Both positions are defensible. The disagreement is not really about whether feedback exists — everyone agrees it does — but about whether the thing generating the feedback was already inside anyone's model before the action happened.
What the three generations would do here
Set against this domain, the Large Language Model, Large World Model, Large Universe Model lineage becomes concrete rather than architectural.
A system built on the Large Language Model pattern would be fitted to a frozen corpus — historical fire behaviour, past evacuation timings, census-based population estimates — collected up to some cutoff and then optimised against. It cannot observe what its own recommended evacuation order does to road capacity next Tuesday, because next Tuesday is not in the corpus. Its niche construction is real: recommending an order shapes traffic, which shapes response times, which shapes future incident histories that will eventually be collected and used to train the next version. But none of that loop closes inside the system's own intake. It is invisible to the thing causing it.
A system on the Large World Model pattern senses the live scene — current sensor feeds, current road occupancy, current shelter counts — for the duration of an incident. This is a real improvement: it can watch the evacuation order's traffic effect unfold in real time and adjust the order, reroute, stagger departures by zone. But the scene has an edge. The incident closes, the sensors are reassigned, the dashboard is archived. Consequences that mature after the incident is declared over — a water main stressed past its rated cycles, a bridge deck fatigued by three consecutive evacuation loads in a wet season, a community's willingness to comply with the next order having watched this one prove itself wrong — sit outside the observation window entirely. The dam is visible while it holds water. The willow it fed is not, because nobody is still watching six months later.
A system on the Large Universe Model pattern would keep the hazard sensors, the population movement data, the infrastructure status feeds and their forecasts running as belief streams that do not stop when the incident is declared closed. Each belief carries provenance — this shelter-occupancy estimate derives from that mobile-phone aggregation, revised at 14:20, superseding the 09:00 estimate that assumed the northern route stayed open. When the water main fails eleven weeks later, the record allows a manager to trace the failure back to the specific evacuation order that put three months of unplanned cycling stress on it, rather than filing it as an unrelated infrastructure fault. That is the entire content of the claim: not smarter prediction, but a record that does not end when the episode does, so a later manager can attribute a later drift to an earlier order.
Where the objection bites hardest
The strongest challenge to this is not "control theory already does this." It is sharper: continuous observation does not confer attribution. An emergency manager watching every stream — hazard sensors, mobility data, infrastructure telemetry — will see plenty of correlated drift after an evacuation order and can easily misassign it. Shelter occupancy spikes; is that the order, or a second, unrelated ignition forty kilometres east drawing in its own evacuees? Water pressure drops; is that the shelters, or a coincidental pump failure at a treatment plant that had nothing to do with the fire? Watching everything, continuously, does not by itself tell you which watched thing caused which other watched thing.
Recording never stops, but that just means you drown in correlated noise instead of missing the effect altogether. You have swapped invisibility for confusion.
This is fair, and it should not be waved off. Continuous intake with provenance does not solve attribution. It makes attribution possible in principle where it was previously impossible in principle — a frozen corpus cannot contain a post-cutoff effect at all, full stop — and then hands the problem to the ordinary machinery of causal inference: staggered evacuation zones as a natural experiment, holdout areas not subject to the order, instrumented comparisons between fire seasons where the same infrastructure did and did not carry an evacuation load. Provenance is what lets an analyst run that comparison later. It is not a substitute for running it.
The second objection: why not just refresh more often
A plainer fix suggests itself: if the corpus is stale, retrain it more often. Rebuild the traffic and hazard model monthly instead of annually; rebuild it weekly if the incident tempo demands it. On this view the whole lineage collapses into a question of refresh frequency, and calling the limit a new category is ontology dressed over a scheduling decision.
Frequency and continuity differ in what they preserve, not only in how fast they run. A model retrained monthly holds one snapshot of belief at a time. When it is retrained, the previous snapshot — and the reasoning for why the water-pressure assumption was set where it was in March — is usually discarded along with the old model. There is no chain connecting July's water-pressure failure back to March's evacuation-order assumptions, because March's assumptions are gone. A system that instead keeps beliefs revisable with provenance attached does not need to be faster; it needs to keep the history that makes the July failure traceable to the March order. Retraining weekly, at the limit, converges on a very fast amnesiac, not on this.
What survives, narrowed
The wide claim — that emergency management should therefore run every stream forever, forever revisable — is not what the evidence here supports, and the domain itself pushes back on it: attribution needs designed comparisons, not just uninterrupted logging, and a manager drowning in unattributed correlation is not obviously better off than one working from a stale report. What survives is narrower and still substantial. The specific failure named at the start — the order following the hazard rather than leading it — is not fixed by better forecasting alone, because forecasting improvements still stop at the incident boundary. It is addressed only by an intake architecture that keeps hazard, mobility and infrastructure streams live past the point where the incident report is filed, with enough provenance that a manager six months on can ask "did our order do this?" and get an answerable question rather than a shrug. That is the terminal rung on this particular axis: not omniscience, and not automatic attribution, but the first arrangement in which the question can be put to the record at all.