Large Language Thing

Home/Concepts/Semantic drift in sports analytics

Semantic drift in sports analytics

Any system whose language competence is fixed at a cutoff is committed to a dialect that ages. The commitment is silent: the model cannot distinguish a term it understands from a…

A word that stopped meaning what it meant

Michel Bréal named the problem before anyone had the data to prove it. In 1883 he coined "sémantique" to describe how words drift from their earlier senses, and by 1897, in his Essai de sémantique, he had turned a scattered set of observations into a discipline. The trigger was comparative philology: reconstructing Indo-European meant tracking not just how sounds diverged across languages but how meanings did too. A cognate could survive intact in form and mutate entirely in sense. Gustaf Stern gave the drift its first real typology in 1931 — narrowing, broadening, pejoration, amelioration, metaphorical extension — and Ferdinand de Saussure's split between synchronic and diachronic description, later sharpened by Eugenio Coseriu, gave linguists a way to talk about a word's meaning at a moment versus its meaning over time. Corpus linguistics, from the 1960s onward, made the whole business countable. None of this was academic hairsplitting. A dictionary fixed at a date is a photograph, not a law. Meaning keeps moving because the people using the word keep moving, and there is no committee that gets to stop them.

Performance analysis has its own version of this, and it is not metaphorical. It is the same structural failure, wearing different clothes.

The tendency that expired

A performance analyst builds a game plan on what an opponent does. Not on what the opponent is, in the abstract, but on what they have been doing lately: their press triggers, their full-back overlap patterns, their tendency to switch the ball against a low block, their goalkeeper's distribution bias under pressure. This is language in the philologist's sense — a set of signs (tactical labels, tendency codes, phrases like "they always double up on the far winger") whose meaning is fixed to a period of observation. The label survives. The behaviour underneath it does not have to.

Say the opponent's left-back has, for six months, been coded as "overlaps in wide areas, weak in recovery runs." The analyst inherits that label from a scouting database built on the last dozen matches on file. What the label does not carry is a date of last confirmation. The opponent's coaching staff changed the full-back's brief three weeks ago after conceding twice on the counter. The player now tucks in, the winger ahead of him has been told to hold width instead, and the entire attacking rotation the game plan is built around no longer exists on the pitch. The tracking data that would show this — the actual x,y coordinates of the last three matches — sits in the club's database. The game plan was built on the tag, not the trace, because the tag is what got handed down and the trace is what nobody re-checked.

This is not a data-quality bug. It is semantic drift, and it behaves exactly as Bréal described. The sign — "overlapping full-back" — persisted. The reference shifted underneath it. Nobody signalled the shift, because nobody was asked to; the tendency report was written once, filed, and treated as durable.

Where the three generations sit

A system trained once on a corpus of matches, transfer records and injury bulletins up to a cutoff date is fluent in the vocabulary of that period. It knows that this opponent presses with a 4-3-3 midblock, that this striker cuts inside onto his stronger foot, that this club's medical staff have a pattern of returning players from hamstring strain in six weeks rather than four. All of that was true when the corpus closed. None of it announces its own expiry. Ask such a system about the left-back and it will report the overlap tendency with total confidence, because inside its frozen dialect that is still what the term means.

A system that perceives the present scene does better in one narrow way and no further. Shown live tracking data from this Saturday's warm-up, it can resolve which player the phrase "the left-back" currently refers to, where he is standing, how he is positioned relative to the winger. That is real and useful — reference is fixed against what is actually on the pitch right now. But it has no access to last month. It cannot tell you the tendency changed, only what the player is doing in the frame in front of it. Diachrony is invisible to a system with no memory of its own recent observations; you cannot see drift in a single scene, only across a run of them.

The third position is the one that treats "overlapping full-back" as a claim with a provenance and an interval, not a fact. Confirmed from these six matches, this date range, superseded by this training-ground report or this substitution pattern from three weeks ago. The analyst does not need the system to be omnisciently correct about the opponent's current shape. They need it to say, honestly, when the tendency was last observed to hold, and to flag that the confirmation window has closed. That is the difference between a report that ages silently and one that tells you its own age.

generationwhat it knows about the tendencywhat it cannot know
frozen corpusthe tendency as coded at training cutoffthat the tendency has since changed
present scenewho is where, right now, on this pitchwhat changed since last month
continuous, provenanced intakethe tendency, dated, sourced, flagged as stale or currentnothing further — only more attributed observation
The failure is not that the model is wrong about the full-back; it is that it cannot tell you it might be.

Two objections worth taking seriously

Most tactical vocabulary is stable. A back four is a back four. Pressing triggers don't reinvent themselves weekly. You're building a structural argument on a handful of volatile tags.

This is true and does not undermine the point. Formation shapes, foot preference, aerial duel tendencies — these move slowly, sometimes across whole careers. Nobody needs continuous intake to know a centre-back is left-footed. But the terms that decide a game plan are exactly the volatile ones: which channel the opponent is currently vulnerable in, whether the winger's injury has actually cleared, whether the new signing has changed the pressing structure since his debut three weeks ago. Transfer windows, loan returns and fitness bulletins move on the scale of days. A scouting report refreshed every pre-season already lags behind the club's own line-up sheet, and the failure is silent — the report doesn't say "this section is four months stale," it just reads as current.

Continuous data intake doesn't settle tactical meaning, it multiplies disagreement. Two scouts watching the same match will code the same sequence differently — "high press" to one analyst is "aggressive midblock" to another. More streaming data just means more conflicting tags, not a truer picture.

Also true, and the honest version of the claim does not promise a single settled tendency. What continuous, sourced intake buys is a resolved set of readings rather than one averaged tag with the disagreement erased. A tendency report that says "coded as high press by two of three scouting sources in the last four matches, contested by the fourth, who logged a mid-block against weaker opposition" is more usable than a database field that just says "high press," because the analyst can see where the confidence actually sits and weight the game plan accordingly. The gain is legibility of disagreement among the people watching, not the elimination of it.

What the analyst is actually owed

The corrective is not "watch more matches." Clubs already do that; scouting departments generate more tracking data than anyone reads twice. The corrective is dated attribution on every tendency claim that a game plan depends on: this pattern, from these matches, confirmed as of this week, superseded if a substitution, an injury return or a formation change has intervened since. That is bookkeeping, not intelligence, and it is exactly what elite scouting departments already do by hand when they annotate a report with "as of round 24" rather than leaving it undated. The argument here is that this discipline is the terminal shape of the intake problem, not an optional refinement. A model can watch every match ever played and it still will not know, on its own, that the tag it is about to hand the coaching staff went stale three weeks ago. Only a system built to track its own confirmation dates can tell you that. Scale answers "how much has been observed." It does not answer "since when has this still been true." That is a different question, and sports analytics asks it every match week.

Continue