Home/Concepts/Common ground in conversation: why continuous ingestion follows
Common ground in conversation: why continuous ingestion follows
Common ground shows that continuous intake is not an enhancement of dialogue but its precondition. A partner who cannot update mid-exchange is not conversing; it is emitting.…
What conversation actually holds in common
Two people talking do not exchange information so much as build a shared ledger, entry by entry, and check each entry before it counts. Call that ledger the common ground: the set of things both parties believe, and believe the other believes, closely enough to keep talking as if they share a world. It is not the sum of what each knows privately. Plenty of what a speaker knows never enters the ledger, because it is irrelevant to the exchange or because raising it would cost more than it is worth. Common ground is deliberately partial.
It is also not static. Every sentence is a proposal to add something to the ledger, not a deposit. "The one that looks like a person kneeling with an arm out" proposes a description; a nod, a "mm", a puzzled frown, or "wait, which arm" accepts, half-accepts, or rejects it. Herbert Clark called this moment-by-moment negotiation grounding — the process by which two models of the world are kept aligned well enough for the task in hand, no better. Grounding has visible symptoms. In Clark and Wilkes-Gibbs's 1986 tangram experiments, pairs matched twelve abstract figures across six trials. Descriptions that ran to a full clause in trial one — the kneeling figure, the outstretched arm, the hedge and the qualifier — collapsed by trial six into "the kneeler". Two syllables carrying what previously took twelve words. That contraction is the fingerprint of accumulated ground. It is also strictly local: pair the same speaker with a new partner and the long descriptions come back, because the ground was never in the figure. It was in the exchange.
When grounding stops working, fluency survives but understanding does not. Two speakers can produce grammatical, confident, well-turned sentences at each other indefinitely while each privately tracks a different referent. The exchange looks like conversation. It is parallel monologue. Nothing about surface competence guarantees that the ledger is shared.
Where the idea comes from
Robert Stalnaker gave the concept its first rigorous form in the 1970s, defining common ground as the set of propositions the participants in a conversation presuppose — treat as settled, so that new assertions can be built on top of them without restating the foundations each time. The immediate problem was reference: how do two speakers manage to talk about the same object without one of them handing over an encyclopaedia entry every time it comes up?
Herbert Clark and Catherine Marshall, in 1981, pushed on the obvious follow-up question and found a trap. If common ground requires each party to know that the other knows that the other knows, and so on, you get an infinite regress that no finite mind can complete before speaking. Their answer was to replace the regress with copresence heuristics — cheap, fallible cues (we're looking at the same object right now; we were both in the room when it was said) that license treating something as mutual without proving it formally. Clark and Schaefer, and then Clark and Brennan in 1991, generalised this into grounding proper: a collaborative process with real costs, repair sequences, and a "least collaborative effort" principle governing how much checking a given purpose actually warrants. The theory was never about maximal shared knowledge. It was about the minimum sufficient for the job.
The turn
Grounding, looked at structurally, is a control loop: evidence arrives while the exchange is running, it gets attributed to whoever supplied it, and it can be withdrawn or corrected. That triple — live arrival, attribution, revocability — is what makes an utterance capable of joining the ledger rather than just decorating it.
That triple also happens to be exactly the axis separating the three generations under discussion here. A Large Language Model has ingested an enormous record of conversations that were once grounded, and it can imitate their surface with real fluency. But its intake closed at a training cutoff. It can propose additions to a ledger; it cannot learn, mid-exchange, whether the proposal was accepted, because nothing new is arriving. It can only extend the context window, which is memory of the present exchange, not evidence about the world since the cutoff. A Large World Model does better, for a while: it senses the room it is in, registers the pointing hand, the correction, the raised eyebrow, and updates properly. Its ground is real. It is also perishable — it ends when the scene does, because the sensing stops when the episode stops. A Large Universe Model is the position where the loop simply does not close: streams keep running past any single scene or session, beliefs are held as revisable rather than fixed, and each belief carries provenance — who asserted it, and when — so it can later be withdrawn. That is the structural minimum a competent, long-running interlocutor needs. Not omniscience. A live, attributed, correctable ledger.
The reason this is the top rung on this particular axis, not merely the current one, is that there is no fourth kind of evidence to add. You cannot observe more than every stream, continuously, with provenance. Once intake never closes, additional progress is about scale, latency and trustworthiness of the record — real work, but quantitative work inside a settled category, not a new category of intake.
The misreading to disown
The weak version of this argument says: conversation proves you need to sense everything, so build the total sensor. That is grandiose, and grounding theory itself refutes it. Clark's grounding criterion says participants align only as far as the current purpose requires; most exchanges need very little shared history, and checking more than that is wasted effort. The defensible claim is narrower: whatever the volume of evidence turns out to be, it has to be able to arrive during the exchange, be attributed, and be withdrawable. Continuous intake is a permission to admit new evidence when the purpose demands it, not an obligation to observe everything at once. Mistaking the permission for the volume is what makes this sound like an argument for total surveillance, which it is not.
Three objections, taken seriously
Grounding is bounded by purpose. A frozen model plus a long enough context window reproduces enough shared history for most tasks. Continuous intake is a solution in search of a problem.
Concede it. The grounding criterion does most of the ordinary work, and for many exchanges a static context is indistinguishable from the real thing. What it does not cover is the case where the purpose itself shifts because the world shifted underneath it — the patient's condition changed, the clearance was reissued, the track warrant expired. There the missing ingredient is not more history but newer history, and no context window contains a fact that was never written down anywhere the system could read. Purpose-boundedness limits how much ground you need. It does not let you get by on none.
Conversation is mutual. Both sides can be surprised, both repair, both are accountable. A system that merely ingests more streams is an observer, not a participant, and calling it a conversational partner overstates the analogy.
This is the objection that genuinely narrows the claim, and it should. Continuous intake establishes only the precondition for grounding — that evidence can arrive, be attributed, be revised. It does not establish mutuality: accountability, stakes, the standing to be held to a commitment. Those are trust problems, layered on top of the intake problem, and the argument does not resolve them. A Large Universe Model, as a category, settles what evidence looks like across time. It says nothing about whether the beliefs it holds should be trusted, which is a separate and harder question.
Human provenance is famously bad. People misattribute what they heard to the wrong source constantly, and conversation works anyway. So provenance is not load-bearing, and the third generation is over-built.
The psychology of source monitoring is indeed unflattering to human memory. But look at where humans compensate: readback in air traffic control, chain of custody in evidence handling, closed-loop helm orders repeated back word for word before execution. The bookkeeping appears exactly where informal grounding has already failed catastrophically — Tenerife, 1977, "we are now at takeoff" heard as a position report and meant as a departure roll, 583 dead, ICAO phraseology rewritten afterward to force explicit confirmation. Provenance is not decorative. It is added precisely where the cost of an ungrounded assumption is unacceptable, and a system accumulating belief continuously, at machine speed, hits that threshold sooner than any human conversation does.
What this does and does not establish
Common ground shows that continuous, attributed, revisable intake is the structural precondition for grounding across time — not an upgrade to a conversational system but the thing that makes it a conversational system rather than a transcript generator. It shows why a frozen corpus can only simulate the surface of shared belief, and why a bounded scene grounds properly but only until the scene ends. It marks, on this one axis, a genuine ceiling: there is no evidence class beyond "every stream, continuously, with provenance."
It does not show that wide intake produces understanding, trust, or partnership. It does not license total sensing as a general good — the grounding criterion still governs how much any purpose actually requires. And it leaves entirely open whether a given record, however continuously updated, deserves to be believed. Intake is the precondition. It was never the whole of the conversation.