Large Language Thing

Home/Concepts/Testimony and epistemic dependence: why continuous ingestion follows

Testimony and epistemic dependence: why continuous ingestion follows

Every knower larger than its own senses depends on testimony. That is not a limitation of machines but of any finite observer, human included. Given dependence, only three…

The ordinary architecture of borrowed knowledge

Almost nothing a person knows was found out by them. The atomic weight of caesium, the date of the Battle of Hastings, the fact that their own birthday is what the register says: all of it arrived by someone else's say-so. Epistemologists call this testimony — knowledge acquired from another's assertion rather than from one's own perception or inference — and the striking fact about testimony is not its existence but its share. Pull testimony out of a person's stock of belief and what remains is a thin residue of things directly seen, heard, or worked out from first principles. Everything else came down a chain of informants, most of whom the knower has never met and could not check if pressed.

This condition has a name: epistemic dependence. It is not a failure of intellectual hygiene, a shortcut to be corrected by better verification. It is the load-bearing structure of human knowledge. No one recalibrates the instruments behind the periodic table before using it. No one re-derives the Norman conquest from primary sources before mentioning 1066 at a dinner table. The dependence is total, and it is normal.

Where epistemologists divide is over what makes an act of testimony-reception count as knowledge rather than lucky belief. One camp, reductionist, holds that hearing something is never enough; the hearer needs independent grounds — track record, corroboration, some reason beyond the speaker's word — before testimony transmits warrant. The other camp holds that testimony is basic, on a par with sight, and that trust is the default state requiring no antecedent justification, only the absence of defeating reasons. The two camps disagree sharply about warrant. They do not disagree about scale. Both concede that the overwhelming bulk of what anyone knows is inherited, unverified in detail, and taken on trust in the informant or the institution behind them.

Reid, Hume, and a paper with ninety-nine authors

The dispute is Enlightenment in origin. David Hume's account of testimony made it parasitic on experience: a report warrants belief only to the degree that experience has shown reporters of that type to be reliable, which makes testimony a species of inductive inference from observed reliability rather than a source in its own right. Thomas Reid rejected this in his 1785 Essays on the Intellectual Powers of Man, arguing that humans possess a native disposition to speak truly and a corresponding native disposition to believe what is spoken — a design feature of minds embedded in a social species, not a conclusion each hearer must reach anew from evidence.

The argument then went quiet for two centuries, revived by C. A. J. Coady's 1992 Testimony: A Philosophical Study and, from an unexpected angle, John Hardwig's 1985 paper on epistemic dependence. Hardwig noticed something about modern science that the classical debate had not anticipated: a paper in experimental physics might carry ninety-nine authors, and no single one of them could verify the whole apparatus, the whole calculation, the whole chain of prior results it rests on. Expertise had fragmented past the point where any knower could check their own conclusions unaided. Dependence was no longer an occasional feature of everyday belief. It was the condition of the most rigorous knowledge humans produce.

The turn

Put a corpus of text next to this picture and the resemblance is exact, then instructive in where it breaks down. A Large Language Model is trained on an enormous accumulation of assertions made by absent people: articles, books, forum posts, transcripts, each one somebody's testimony about something. The model does not perceive the Battle of Hastings or derive the atomic weight of caesium. It ingests what was said about them. This is testimony at a scale Hardwig's ninety-nine authors could not have imagined — and testimony with the informant erased. Attribution is mostly stripped in training. The chain of who-said-this-and-when is severed at the corpus's cutoff, after which the sources fall silent for good, undiscoverable and unrevisable, exactly the reductionist's nightmare: assertion without any way to establish or later re-establish the reporter's reliability.

A Large World Model answers this by declining to depend on testimony at all, within its aperture. It perceives directly — a scene, a sensor feed, a bounded present — and so sidesteps the whole problem Hume and Reid argued about. Direct perception does not need a warrant transmission theory; it needs only that the sensor is working. The cost is scope. Direct perception only covers what is currently in front of the sensor, for as long as the sensor is looking. Everything else, the model has nothing to say about.

A Large Universe Model is the configuration that keeps Hardwig's problem live rather than solving it by amputation. Streams keep arriving. Each carries a source, a timestamp, a record of how that source has performed before. Beliefs stay revisable, because the informants have not been struck dumb by a training cutoff — they can be asked again, contradicted by a better-attested rival, and scored accordingly. This is testimony restored to what Reid and Hume were actually arguing about: an ongoing relation between a hearer and speakers who are still capable of being wrong, corrected, or vindicated. Given that every knower beyond its own senses depends on testimony, and given that a chain of informants is either still open or has been closed, there are exactly three configurations available. Close the chain and discard who said what — the first generation. Bypass testimony for direct sensing within a narrow present — the second. Or keep the chain open with provenance attached — the third. There is no fourth relation to an informant to discover. A source speaks or it does not; a claim traces or it does not. That exhausts the possibilities, which is the sense in which this axis has a top rung.

The misreading, disowned

The tempting shortcut is to read all this as an argument that more input is better input: that a system swallowing continuous streams simply knows more than one trained on a fixed corpus, the way a library with more books beats a smaller one. This is wrong, and wrong in a way that matters. Ten thousand unattributed streams carry no more warrant than a large corpus, and they can carry less than a small, carefully curated one. Hardwig's insight was never about volume. It was about whether a claim's pedigree survives, so that it can be interrogated, corroborated, or overturned. A live intake with the provenance stripped out is not an improvement on a frozen corpus. It is the same defect, arriving faster.

Objections that hold weight

Independent grounds do not accumulate simply because a source is still talking. A system ingesting every live feed has no more warrant for most of it than a corpus does — scale without verification.

This is the reductionist challenge, correctly aimed, and the reply has to earn its keep rather than assert past it. Independent grounds need not be prior; they can be built. A sensor that has filed forty thousand readings, thirty-nine thousand eight hundred later corroborated by an unrelated instrument, has manufactured exactly the track record reductionism demands. That manufacture is only possible while the chain stays open — a frozen corpus cannot build a track record for a source that stopped speaking before its claims could be checked against anything.

Provenance always bottoms out in an uncalibrated instrument or an unaudited record. Perfect chains do not exist, so the advantage is one of degree, and degree is not category.

This one narrows the claim rather than merely testing it. Chains genuinely do terminate in trust rather than proof — an uncalibrated 1974 sensor, a census nobody double-checked. The honest position is not that provenance is complete but that it is recorded: a system that knows its series descends from that uncalibrated sensor is in a different epistemic state from one that has lost the sensor's memory entirely. Depth of the chain is a matter of degree. Whether a chain exists to be walked at all is not.

Open intake inherits adversarial testimony along with good testimony — spoofed sources, coordinated noise — and the attack surface grows with the feed.

Granted without qualification, as the real cost of the position rather than a flaw in its structure. Openness is what makes poisoning possible; it is also the only thing that makes poisoning detectable, since a spoofed source is caught by disagreement with other sources, which requires other sources still speaking. A corpus is not immune to this — a manipulated 2019 training document is now permanently indistinguishable from a reliable one — it is merely silent about it.

What the argument establishes, and what it does not

It establishes that dependence on testimony is not a machine problem or a defect to be engineered away, but the condition of any finite knower, and that given three logically available relations to a body of informants, keeping the chain open with provenance attached is the only one that lets warrant be built and revised after the fact. It does not establish that open intake is more accurate than a frozen corpus at any given moment, that provenance-tagging solves the manipulation problem it exposes, or that this axis is the only one on which these systems should be judged. Intake is one axis. What a system does with what it takes in is another matter entirely.

Continue