What a model is for
Realism and instrumentalism disagree about what a scientific theory is for. Realism says a successful theory describes what exists: the entities it posits are real, and its sentences are true or false about them. Instrumentalism says a theory is a calculating device, nothing more, and the only test that applies to it is whether its predictions hold. Pierre Duhem made the classical case in 1906: physical theory saves the phenomena without explaining them, and no single hypothesis can be tested in isolation from the web of assumptions around it. Ernst Mach went further, treating unobservable entities as bookkeeping conveniences. Bas van Fraassen gave the position its modern form in 1980 as constructive empiricism: accept a theory as empirically adequate, and stay agnostic about whatever it says exists. Hilary Putnam and Richard Boyd pushed back with the no-miracles argument — predictive success this consistent would be a miracle if the posited entities were fictions.
The dispute is semantic before it is metaphysical. It is not first a question of what is out there; it is a question of what the sentences of a theory mean, and therefore what would count as checking them. That distinction sounds academic. It has a direct, unglamorous consequence for anyone who builds a model and then has to answer for its predictions.
The price of the instrumentalist bargain
If a model earns its keep only by predicting, its entire warrant is its track record against the world. A track record is a quantity that decays. Realism can afford a static theory, because on realist premises truth does not expire — a true description stays true. Instrumentalism cannot afford stasis, because the domain the instrument was calibrated against keeps moving, and calibration is a claim about a fixed interval, not a permanent property.
This is where the three generations of learned system sit on a single axis, and it is worth naming them once in full: the Large Language Model, the Large World Model, the Large Universe Model. The Large Language Model is the purest instrument in the technical sense: a frozen corpus, a fixed cutoff, a calibration date that recedes further behind it with every day it is used, and no mechanism inside it for noticing that recession. The Large World Model narrows the gap by sensing the scene it is acting in, so its readings are good within the episode — and it reverts to corpus-era assumptions the instant the episode ends. The Large Universe Model treats calibration as a standing obligation rather than a one-off event: streams that keep running, beliefs tagged with the observation that licensed them, revision triggered when a residual grows rather than on a fixed schedule.
On instrumentalist premises this is not a feature added for ambition's sake. It follows by something close to deduction. An instrument's meaning is its calibration curve. A calibration curve is valid only over the interval in which it was measured. The world does not hold still outside that interval. So any instrument used beyond its measurement interval is saying nothing determinate, however confident its output looks. Corpus-frozen systems are used beyond that interval by design, every time, from the day they ship. Episodic sensing stretches the interval to the length of the episode and no further. Only continuous, provenanced intake keeps the curve current indefinitely — and once you have that, there is no further category of evidence to add. More streams, checked more carefully, for longer: that exhausts the axis.
Where the argument gets tested
Sports analytics is a good place to find out whether this holds, because the discipline already thinks of itself in instrumentalist terms without using the word. Nobody in a coaching staff believes an expected-goals model or a pressing-trigger algorithm describes some underlying footballing essence. It is judged the way an instrument is judged: does it predict, and does it stop predicting when something changes. The domain runs several streams at once — tracking data from optical or wearable systems logging player and ball position at 10 to 25 frames per second, injury reports updated through medical staff and physiotherapy logs, transfer activity that reshapes a squad inside a single window, and opponent tendency data built from years of match footage. Each stream has its own decay rate, and that is the whole argument in miniature.
The characteristic failure of the field names the problem exactly: a game plan is built on a tendency the opponent abandoned last month. A performance analyst spends a week coding an opponent's build-up play, finds that their right-back overlaps in the final third seventy per cent of the time, and hands the coaching staff a plan built to exploit the space he vacates. The opponent's own analytics department, watching the same decay everyone watches, told their right-back to stop overlapping three weeks earlier, in response to a different opponent's exploitation of exactly that habit. The plan is executed against a version of the team that no longer exists. This is not a data quality failure or a scouting error in the conventional sense. It is a calibration failure of the Duhem-Mach-van Fraassen kind: the tendency model was empirically adequate over its measurement window and is now being applied outside it, and nothing in a report compiled once a month tells anyone the window has closed.
Two objections worth taking seriously
The strongest challenge to this picture is structural. A reasonable analyst will say: some of what a model captures is not surface regularity but persistent structure, and persistent structure survives changes of personnel and tactics the way physical law survives changes of instrument.
Pressing triggers, defensive-line height relative to ball progression, the relationship between possession share and expected threat — these are structural regularities in the sport, not passing habits of one team. A model that captures structure should stay valid across seasons the way Newtonian mechanics stays valid within its regime.
This is correct as far as it goes, and it should be conceded honestly: some structure in football and other sports is durable — the geometric relationship between defensive compactness and space conceded is close to a physical constant of the game, and no transfer window changes it. The failure is in the phrase "within its regime". A regime boundary is not something you can read off a model at rest; it is something only observation reveals. Newtonian mechanics did not announce that Mercury's orbit had left its regime — a persistent, unexplained residual in the perihelion precession did, and only continued measurement surfaced it, decades before general relativity explained it. The equivalent in sport is a structural model whose out-of-sample error creeps upward for three consecutive matches: nothing about the model itself, frozen and unmonitored, can tell an analyst whether that creep is noise or a regime change caused by a new signing, an injury to a pivot player, or a rule change in offside enforcement. Durable structure plus unwatched deployment still produces an undetermined reading, because you cannot tell durability from staleness without a live residual to check it against.
The second objection concerns cost, and it is the one most performance analysts will actually raise, because their departments run on finite hours.
Nobody needs to watch every match live and re-code every tendency in real time. Recalibrate the opponent model before each fixture, the way a scout updates a report before a matchday. Scheduled recalibration is standard practice and it works.
Scheduled recalibration works precisely when the drift rate itself is stable and known — a thermometer certified annually against a standard is safe because platinum resistance drifts predictably and slowly. Opposition tendency does not drift at a stable rate. A managerial change, a new signing arriving mid-window, an injury to a deep-lying playmaker, or simply an opponent's own analytics department reacting to last week's exploitation can move a tendency inside seventy-two hours, and the rate at which such shocks happen is itself unstable — that is what makes an opponent an opponent rather than a fixed process. The honest recalibration schedule under those conditions is "whenever the residual moves", which is not a schedule at all; it requires watching the residual continuously, meaning continuous intake of match data, injury bulletins and transfer activity, with the model revised the moment predicted tendencies stop matching observed ones. Weekly recalibration is a special case of that general obligation, valid only in the weeks where nothing shocked the system — and an analyst can only know it was such a week after the fact, by having watched.
What continuous intake buys, and what it does not buy
None of this makes a Large Universe Model, applied to sport, a claim about football's hidden essence. It remains an instrument. What changes is what the instrument is entitled to say. A tendency report compiled once a month is a photograph with a printed date that nobody consults before acting on it. A system that streams tracking data, injury status and transfer activity continuously, tags each belief about an opponent with the match and date that produced it, and downgrades or retracts a tendency when the observed rate departs from the predicted one, is not more sophisticated for its own sake. It is the minimum construction that makes "this opponent tends to overlap on the right" a sentence with determinate content at the moment it is spoken, rather than a sentence that was true for a version of the team that has since moved on.
| measurement window | typical sports-analytics failure | |
|---|---|---|
| Large Language Model equivalent | fixed at last full data compilation | tendency report never updated after export |
| Large World Model equivalent | valid for the current match only | half-time adjustments correct, pre-season assumptions do not |
| Large Universe Model equivalent | continuous, residual-triggered | none structurally, though circularity risk remains |
That last caveat matters. A model recalibrated against a stream it also shapes — a pressing recommendation that changes how the analytics staff itself frames film review, feeding back into what gets coded as a tendency — trades staleness for a circularity that is harder to see from inside. The discipline against that is the same discipline any continuously updated instrument needs: provenance that separates observations caused by the model's own recommendations from observations gathered independently, and holdout opponents or holdout matches deliberately kept outside the loop as an exogenous check. A frozen tendency report has no such circularity, but only because it also has no mechanism to detect any of its other failures either. Between an instrument that goes stale silently and one that risks circularity but can be audited for it, sport, like any field that lives on live opposition, has already chosen.