The plan that was right when it was made
A network planner builds capacity forecasts the way most infrastructure disciplines do: sample the recent past, fit a trend, provision ahead of it. Traffic telemetry from core routers, radio access counters, transport link utilisation — all of it aggregated into 15-minute bins, rolled up into daily and weekly patterns, fed into a capacity model that says which links need augmentation this quarter and which cell sites need a carrier add next year. This works. It has worked for decades. The traffic mix — video share, signalling load, uplink-downlink ratio — moves slowly enough that a plan built on last quarter's mix is still approximately right this quarter.
Then a popular app pushes an update that changes its default video resolution, or a messaging platform switches from polling to persistent connections, or a game client starts pre-fetching content on any bar of signal instead of waiting for Wi-Fi. None of these is a network event. Each is a product decision made somewhere else entirely, shipped to hundreds of millions of handsets inside a release cycle measured in days. The traffic mix that took the network eighteen months to drift into last time now moves that far in an afternoon.
What the loop actually holds
The planner's system of record, in the fixed-interval version of this discipline, holds a small number of things: rolling traffic baselines by cell and by link, a churn model updated monthly from billing data, spectrum utilisation reports filed quarterly, and a fault log consulted mostly after the fact, when something has already broken. Each of these is a snapshot discipline. The baseline is recomputed on a schedule; between recomputations, it is treated as ground truth. This is not laziness. Recomputing a capacity model continuously against noisy telemetry is expensive, and most of the time nothing has changed enough to justify it.
The trouble is that "most of the time" is exactly the phrase punctuated equilibrium warns against. Traffic mix is not a slowly drifting quantity that happens to be sampled monthly. It is a quantity that sits still for months and then jumps in the space of a release cycle, and the jump is precisely the interval a monthly or quarterly cadence is built to smooth over. The network's own telemetry systems are capable of sub-second granularity — that is not the constraint. The constraint is that belief about the traffic mix is only updated on the modelling team's schedule, not on the network's.
What arrives, and what a wider intake would do with it
Four streams matter here, and they arrive at very different rates and with very different evidential weight.
- Traffic telemetry: volumetric and per-flow counters from routers and radio controllers, available at sub-minute resolution but conventionally aggregated to 15-minute or hourly bins for planning purposes.
- Fault alarms: real-time, event-triggered, but oriented towards hardware and link failure rather than mix shift — an alarm fires when a link saturates, not when the composition of its traffic changes.
- Spectrum filings: administrative and regulatory data, arriving on the timescale of weeks, relevant to what capacity is even legally available to add.
- Churn signals: billing and CRM-derived, monthly at best, a lagging indicator of subscriber satisfaction that itself often reacts to network strain rather than predicting it.
A capacity model that treats these four as inputs to be pulled on a schedule inherits the schedule's blind spot. A capacity model built as a Large Universe Model treats them as streams that are always running, each carrying a timestamp, a confidence, and a decay function, and revises the belief about "current traffic mix" the moment a signal from any of them crosses a threshold — not the moment the quarterly report is due.
Concretely: the app update that doubles average session video bitrate on a metro cluster shows up in traffic telemetry as a change in the flow-size distribution within hours, well before it shows up as link saturation, and long before it would appear in a monthly rolled-up baseline. Fault alarms, tuned for utilisation thresholds rather than compositional shift, may not fire at all until the new mix pushes a link over 90% weeks later — by which point the planner is reacting to a symptom, not the cause, and the lead time to add capacity (spectrum acquisition, backhaul upgrade, tower crew scheduling) has already been lost.
What triggers revision
This is where the discipline of provenance and decay earns its keep, because the objection about noise is correct and has to be answered structurally, not asserted away. Continuous telemetry at sub-minute resolution produces enormous numbers of apparent anomalies — a stadium event, a fibre cut rerouted through a congested path, a software counter glitch — and if every one of them triggers a capacity re-plan, the planner learns to ignore the system within a month.
The answer is not to sample less. It is to weight evidence by source and let contradictory evidence decay rather than accumulate as noise. A single cell's flow-size shift is weak, low-provenance evidence — it could be one handset, one crowd event, one faulty counter. The same shift appearing simultaneously across dozens of geographically dispersed cells, correlated with a known app version bump visible in device telemetry, is strong evidence: multiple independent streams converging on the same hypothesis. A well-built intake layer holds both, tags each with its source and its half-life, and only escalates to "revise the capacity baseline" when corroboration crosses a threshold that accounts for how often each source has been wrong before. A spectrum filing, low-frequency but high-certainty, gets weighted differently from a single alarm, high-frequency but often spurious.
This is the pairing the source material insists on and it holds here without qualification: a missed mix shift is unrecoverable in the sense that matters — the capacity gap it caused already happened, congestion already degraded quality, and no later analysis restores the lost quarter of lead time. A false alarm costs an unnecessary re-plan cycle, annoying but reversible. The asymmetry justifies erring towards sensitivity, provided the revision discipline — provenance, decay, corroboration — keeps the false-alarm rate from drowning the planner in noise.
What the planner sees
In the fixed-interval world, the planner sees a dashboard that is correct as of last month and increasingly wrong as this month proceeds, with no signal of its own decay. The failure is silent until a link saturates and a fault alarm turns a modelling gap into an operational incident.
In the continuous-intake version, the planner sees something closer to a running argument: current baseline for each cluster, the evidence supporting it, and a flag when that evidence has started to disagree with itself. Not a single number updated invisibly, but a belief with visible provenance — "video share on cluster 14 revised upward, driven by flow-size telemetry from six cells, corroborated by a device-OS release tracked externally, confidence raised, previous baseline retained with a timestamp for audit." When the fault alarm eventually does fire, it fires into a model that already anticipated the pressure, rather than one that is discovering the shift for the first time via a saturated link.
What it costs
None of this is free, and the honest account has to include the cost side plainly. Sub-minute flow telemetry at scale is a data engineering commitment, not a dashboard tweak — storage, streaming infrastructure, and the discipline to keep per-source confidence models honest as hardware and counters change underneath them. The corroboration logic that keeps false alarms manageable needs tuning against real incident history, and it will be wrong in both directions while that tuning happens. There is a real version of the third objection here worth taking seriously: much of this is retrieval-and-alerting infrastructure that mature operators already partly have. What distinguishes it from an incremental monitoring upgrade is whether the system is permitted to revise a standing belief — the capacity baseline itself — automatically and with a paper trail, rather than merely raising a ticket for a human to reconcile on the next planning cycle. A network operations centre that gets an alert but still waits for the monthly report to change the model has not built a Large Universe Model. It has built a louder alarm.
The gradualist reply is also owed its due: much of telecommunications traffic evolution genuinely is gradual — subscriber growth, seasonal patterns, slow generational device turnover — and a fixed-interval model handles that well and cheaply. The argument here is not that continuous intake is always worth its cost. It is that wherever traffic composition can move faster than the modelling cadence, which app releases guarantee it periodically will, the interval itself is the source of the miss, and no amount of more careful sampling at the old cadence fixes it. That is a narrow claim, and it is the one the discipline needs.