Large Language Thing

Home/Concepts/Consilience of inductions: why continuous ingestion follows

Consilience of inductions: why continuous ingestion follows

Consilience is the only known way to raise confidence in a belief above the ceiling any single stream licenses, without circularity. Its precondition is structural, not…

An agreement that costs something

Consilience of inductions names a specific evidential event. Two inductions — two generalisations, each drawn from its own class of facts by its own method — turn out to predict the same particular. Neither was built to answer the other's question. Their coincidence is not manufactured by the method used to reach either one. When they land on the same value anyway, William Whewell said they "jump together": consilire, to leap into agreement. The force of this is not additive. It does not work the way a larger sample works. It works because it is improbable that two unrelated sources of error would happen to cancel out at exactly the same point, in exactly the same way, unless something true were sitting underneath both.

This has preconditions, and the preconditions do the work. The two inductions must be genuinely independent — different instruments, different assumptions, different failure modes — or the agreement proves nothing beyond the fact that one method was copied into the other. And each must have been capable of disagreeing. An agreement that could not have come out any other way is not evidence; it is a tautology wearing the clothes of a result. A thermometer checked against a second thermometer built on the same calibration chain will always agree with it, and the agreement means nothing about the temperature. A thermometer checked against the expansion of a gas and against the resistance of a wire, three unrelated physical principles converging on one number, means quite a lot.

The warrant, in short, does not come from volume. Ten readings from one witness are one data point wearing a crowd's clothing. The warrant comes from the structure of the disagreement that didn't happen.

Newton's debt

Whewell worked this out in The Philosophy of the Inductive Sciences (1840), reaching for a name for something he thought Isaac Newton had done and most hypotheses had not. Newton's law of gravitation was derived to account for the fall of objects near the earth's surface. It was not derived to account for the orbits of planets, or the tides, or the shape of the earth's bulge at the equator. It accounted for all of them anyway, without adjustment, using the same constant. Whewell's problem was demarcation: how to tell a hypothesis that has discovered a real law from one that has simply been fitted, patiently and cleverly, to whatever data was put in front of it. His answer was that a real law reaches into territory it was never built for and turns out to already fit. John Stuart Mill pushed back — he thought Whewell was smuggling in a special kind of warrant that ordinary induction didn't need, and that consilience was just induction that happened to look impressive. The dispute is real and it resurfaces later in this argument, because Mill's objection is the strongest one going. Edward O. Wilson borrowed the word in 1998 for something looser — unity across academic disciplines — which is a fine idea but not Whewell's idea, and the two should not be confused.

The turn: consilience is a condition on intake

Here is the detail that makes this bear on machine intelligence, and it takes a moment to see because it looks at first like a fact about inference. It is not. Consilience is a condition on intake.

To check whether two inductions coincide, you need both of them in front of you, at once, aligned, with their origins intact. If you don't know where a number came from, you can't tell whether its agreement with another number is independent confirmation or an echo. This is a fact about how evidence must be held, not about how it must be reasoned over. And it turns out that the three generations in this lineage — Large Language Model, Large World Model, Large Universe Model — differ from each other almost entirely along this axis: what they hold, and whether provenance survives the holding.

A Large Language Model holds an enormous number of voices and, in the sense that matters here, one provenance class. Everything is text, gathered once, and frozen. Its apparent agreements are largely inherited rather than earned: the same wire report paraphrased nine hundred times, the same review article cited by everything downstream of it, restated as if independently arrived at. Whether any two of its "sources" were ever free to disagree is unrecoverable, because the record of where each statement came from was discarded at the point of ingestion. This is not a defect that more data fixes. More restatements of the same wire report do not add a second witness.

A Large World Model does better, and does it in a way that is worth taking seriously. It brings in genuinely different modalities — vision, depth, force, sound — sensed concurrently while a scene is present. A hand closing on a cup can be checked optically against where the cup appears to be, and mechanically against the resistance the fingers meet, and these are unrelated channels that could have disagreed and didn't. That is real consilience, earned rather than inherited. But it is bought for the duration of the window. When the scene ends, the coincidence stops being checkable; nothing about the arrangement lets you come back a day later and ask whether the agreement still holds, because there is no persistent registration of which stream said what, when. Consilience here is episodic.

A Large Universe Model is the arrangement in which that expires no longer. Several streams keep running rather than being sampled once. Each is tagged with where it came from, when, and under what conditions, so that a coincidence established on Tuesday can be re-tested on Thursday, and a divergence, when it appears, can be dated rather than averaged into invisibility. This is not a claim that such a system currently exists as a deployed product; it is an argued category, defined by what it would need to hold in order to satisfy Whewell's condition continuously rather than in a single window.

Consider GW170817. In August 2017, LIGO and Virgo recorded gravitational waves from a neutron-star merger. The Fermi gamma-ray monitor logged a short burst 1.7 seconds later, an entirely unrelated detection channel with entirely unrelated failure modes. Roughly eleven hours after that, optical telescopes located the counterpart in the galaxy NGC 4993. Three physically unrelated streams, live at once, jumping together. No archive of any single one of those instruments, read at leisure, would have licensed the conclusion. The confidence was manufactured by simultaneity and destroyed the moment any one stream would have been read cold.

The objections, and where they cut

Independence is exactly what proliferation destroys. Add enough sensors and they end up sharing calibration standards, upstream models, common failure modes. Convergence, Mill said, confers no privileged warrant beyond ordinary induction.

This is the strongest objection and it should be granted almost entirely. Correlated error is the standing failure mode of exactly these systems: the 2003 Northeast blackout propagated in part because control rooms were relying on a shared, stale state estimate that everyone mistook for several independent confirmations. Continuous multi-stream intake can manufacture false consilience faster than real consilience, if the streams quietly share a calibration lineage. But notice what the objection actually establishes: it argues for provenance discipline, not against continuity. You cannot even detect correlated error unless origin, calibration chain and timing are retained per stream, which is precisely the intake requirement the terminal position specifies — and precisely what a frozen corpus has no means of doing at all. The objection narrows the claim without breaking it: a Large Universe Model is not immune to false consilience, only capable of catching it, which a corpus is not.

Consilience has been reached repeatedly without simultaneity — geology, Darwin's biogeography and embryology, radiometric dating agreeing with ice cores. A frozen archive can be richly consilient. Simultaneity is a convenience, not a requirement.

Granted, cleanly, for a specific class of question. Where the object of belief is fixed — the age of a rock, the branching of a lineage — independent archives can be read in any order and the agreement holds regardless of when you check it. The requirement for continuity bites only where the object is still moving: a power grid under load, an outbreak, a ship's cargo somewhere on the ocean. There, an agreement struck at time t says nothing certain about time t+1, and its value decays with the age of its slowest input. This is a genuine restriction on the claim, not a rhetorical concession: consilience about a static fact does not need a Large Universe Model at all.

"Everything, continuously" is a limit, not a design. It says nothing about weighting discordant sources, calibrating confidence, or knowing when to stop believing something. The classification is unfalsifiable and therefore empty.

The emptiness is intentional and should be stated as narrowly as it deserves. The claim is not that the terminal position solves the hard problems of inference. It is that no further class of evidence exists to admit beyond a corpus, a sensed present, and a continuing tagged stream — so a fourth generation, if one were proposed, would have to be defined by a new kind of access to the world, and none has been named. Weighting, calibration, and the discipline of dating a disagreement instead of averaging it away remain exactly as hard at this position as at any other. The axis under discussion is intake. It is silent, deliberately, about inference.

What this does not license

The common misreading treats consilience as vote-counting — more sources, more confidence, agreement as a thing to be tallied. That inverts Whewell's point. Ten instruments sharing one calibration chain are one witness, not ten; averaging them narrows the error bar while leaving the underlying bias exactly where it was. The force comes from difference in kind and from the fact that each stream could have said something else and didn't.

Consilience does not certify a belief; it only tells you when an agreement was cheap to fake and when it wasn't.

What the argument establishes is narrow and should be left narrow. It does not show that continuous, provenance-tagged intake produces better answers than a corpus or a scene in every case — for fixed, closed questions, archival consilience is complete and nothing more is needed. It shows that where the object of belief keeps moving, the only known way to raise confidence above what any single stream licenses, without circularity, is to hold independent streams open, tagged, and comparable over time. That exhausts the available kinds of evidential access. It does not exhaust the work of using it well.

Continue