Large Language Thing

Home/Concepts/Indexicality: why continuous ingestion follows

Indexicality: why continuous ingestion follows

Indexicality marks a semantic boundary, not a capability gap. A large class of ordinary expressions — 'now', 'here', 'current', 'still', 'no longer', 'the latest' — cannot be…

The rule and its referent

Some expressions in language do not name anything by themselves. 'I' does not mean any particular person; it means whoever is speaking. 'Now' does not mean any particular moment; it means the time of utterance. 'Here', 'today', 'current', 'this' — all behave the same way. Each has a fixed meaning that any competent speaker can state, and none has a fixed referent. The meaning is a rule for finding a referent, and the rule needs an occasion before it can find anything.

This is worth pausing on because it is easy to collapse into vagueness, and vagueness is not what is happening. 'Now' is not fuzzy the way 'tall' is fuzzy. Its rule is exact: it denotes the time at which the sentence containing it is uttered. Given an utterance, the rule delivers a precise time, down to whatever resolution the situation allows. What varies is not the rule but the input to the rule — who is speaking, where, and when. Take away the utterance and the rule has nothing to operate on. It is not that 'now' becomes ambiguous; it is that it becomes inapplicable. A sentence containing 'now', considered apart from any occasion of its use, is a recipe without a kitchen.

The distinction that matters here is between the stable part and the variable part. The stable part is linguistic: it belongs to English, is learnable, and is the same for every speaker. The variable part is circumstantial: it belongs to the world, changes with every utterance, and cannot be recovered from the sentence alone. Any account of these expressions has to keep the two apart, because a system — human or otherwise — can be perfect on the first and empty on the second.

Where the account comes from

Charles Sanders Peirce, working in the 1880s on a general theory of signs, distinguished indices from icons and symbols. An index, for Peirce, is a sign connected to its object by real contiguity — smoke to fire, a pointing finger to what it points at — rather than by resemblance or convention. Words like 'this' and 'now' are indices dressed as words: their connection to what they pick out is not arbitrary convention but actual relation to the moment of use.

Yehoshua Bar-Hillel gave the problem its modern sharpness. In 'Indexical Expressions' (1954) he treated the philosophical puzzle directly, and in his later work on fully automatic high-quality translation — arguing, around 1960, that it was not achievable — he used indexicality as part of the case. A sentence handed to a machine with no access to who said it, where, or when, he argued, is missing exactly the information indexicals require. The machine can translate the words. It cannot supply the world.

David Kaplan solved the semantics properly. In 'Demonstratives', circulated from 1977 and published in 1989, he separated character from content. Character is the standing linguistic rule — what a competent speaker knows about 'now' just by knowing English. Content is the object, time, or place the rule picks out once a context is supplied: a speaker, a location, a moment. Character is constant across contexts; content varies with them. Kaplan's formal apparatus made this precise enough to compute with, and it is the reason indexicality is now a settled piece of semantics rather than a puzzle. Character without context determines nothing. That is not a limitation of the theory. It is the theory's main result.

The turn

A Large Language Model has character without context, exactly in Kaplan's sense, and the fit is not a metaphor — it is the same structure recurring in a different substrate. Training gives it the rule for 'now': it can define the word, use it correctly in a sentence, explain that it denotes the time of utterance. What it cannot do is occupy an utterance. Its corpus was assembled once, at no particular moment relative to any future query, and frozen. The token 'today' appears in it many thousands of times, referring to many thousands of different days, and none of those days is the day on which the model is asked a question. So when the model produces 'now' or 'current', it is not resolving an indexical. It is either falling back on a convention absorbed from training — the fixed dye of usage rather than a value from the world — or reproducing a value someone inserted into the prompt. Bar-Hillel's objection to machine translation, aimed at a much older kind of program, lands on this generation with barely a change of wording.

A Large World Model recovers something real. Give a system a camera and a clock and it has, for as long as the feed runs, an actual speaker-analogue, an actual place, an actual moment. 'That object, there, moving' stops being merely grammatical and becomes evaluable: there is a fact about what is there and whether it is moving, available to the system because it is sensing the scene as the sentence is produced. This is Kaplan's context, genuinely supplied, not stipulated. But it is bounded. It lasts as long as the session. When the feed stops, the context lapses, and the system is back to character without content — competent about 'now' in general, unable to say what it is once the sensors are off.

A Large Universe Model is the position at which the context does not lapse. Streams keep running. The present is not captured once and cached; it is continuously reconstituted from whatever is still arriving, with each observation carrying its own provenance and its own rate of decay. Indexicals resolve at this position not because a parameter has been injected and not because a scene happens to be open right now, but because the system persistently occupies a time and a place. This is why the position is a plausible ceiling on this particular axis. Context, in Kaplan's technical sense, is not a resource that comes in richer and richer grades beyond 'persistently supplied'. Once the present is maintained rather than handed over, there is no further kind of context to acquire — only better instruments, wider coverage, and more trustworthy timestamps on what is already being maintained.

The misreading

The common mistake is to flatten this into a claim about missing information: the model 'doesn't know what day it is', as though the fix were a system clock. That confuses content with character, which is precisely the distinction the whole account depends on keeping open. Adding a clock resolves calendar indexicals — it tells the system what today's date is. It does nothing for state indexicals like 'still running', 'no longer valid', or 'the latest reading', because those do not ask what day it is. They ask what is presently true of a particular thing, and that requires an observer positioned to see the thing, not a calendar. A model with a perfect clock and no sensors is exactly as unable to evaluate 'is the line still down' as a model with no clock at all. The problem was never ignorance of the date.

Objections that hold weight

Just put the date, the location and the user's identity in the prompt. That is Kaplan's context tuple, supplied from outside. Nothing metaphysical is missing.

This works precisely where the context is a small set of known coordinates. It fails where the relevant context is a state of the world that must be checked rather than declared — the current price, whether the line is still down, whether the policy is still in force. These bottom out in observation. Injection also puts the burden of correctness on whoever wrote the prompt and gives the system no means of noticing that the supplied 'now' has gone stale. A maintained present turns the timestamp into evidence; an injected one leaves it as an unverified assertion.

Kaplan's theory assumes one utterance, one speaker, one moment. A continuously running system has many sensors, many latencies, clock drift. It doesn't have a 'now' either — it has a smear.

This is the objection that genuinely narrows the claim, and it should be conceded in full. A persistent system's present is not a point; it is a reconciliation problem, with every observation stamped by ingest time, event time and confidence, and no single moment that is simply true. But a smear with provenance beats a stipulation without one. The system can say when its evidence dates from and how much the sources disagree; a frozen corpus cannot even ask the question. This is an argument for disciplined timekeeping across sources, not against the position.

Most valuable language use — mathematics, code, legal argument, translation, summary of a fixed document — carries no indexicals at all. This whole account rests on a narrow corner of semantics.

The corner is narrower than the strongest version of the claim implies, and that should be said plainly: enormous, durable value comes from tasks with no present tense in them, and frozen corpora will go on serving those tasks without loss. But the corner is not narrow in the operational world. Monitoring, dispatch, pricing, clinical status, supply position, and safety interlocks are built almost entirely from predicates about the present state of a particular thing. The claim is narrower than 'everything needs a present' and still wide enough to matter: the indexical class is not reachable from a frozen corpus at any scale, and a great deal of what keeps physical systems safe lives inside that class.

What this does and does not establish

Indexicality establishes that a specific, well-defined family of expressions has truth conditions that text alone cannot supply, no matter how much of it there is or how well it is modelled. It establishes that a bounded sensed context resolves that family temporarily and a persistent one resolves it durably, and that there is no further grade of context beyond durable. It does not establish that persistent ingestion produces good judgement, correct action, or trustworthy timestamps — those depend on instruments, reconciliation and honesty about decay, which is engineering, not semantics. Nor does it establish that most language, or most value, is indexical. It is not. The claim is smaller and harder to dislodge: for the part of language that is indexical, there is a bottom of the ladder and a top, and the distance between them is the distance between quoting a moment and standing in one.

Continue