Large Language Thing

Home/Concepts/Higher-order evidence in agriculture

Higher-order evidence in agriculture

Any system that revises beliefs needs two kinds of evidence: evidence about the world, and evidence about how often its own machinery gets the world wrong. The second kind is only…

The barometer and the spray window

An agronomist advising on late blight in potatoes works from a rule older than any sensor network: a Smith Period, two consecutive days with minimum temperature above 10°C and at least eleven hours of relative humidity above 90%, means Phytophthora infestans can establish on the leaf within days. That is first-order evidence. It bears on the crop. Whether to trust the alert bears on something else entirely: how often this agronomist's Smith Period calls, filtered through this year's canopy density and this particular weather station's humidity sensor, have actually preceded an outbreak. That second question is higher-order evidence, and it is where most of the damage in modern arable advice actually occurs, because the first-order signal is rarely the failure. The failure is not knowing how much to trust the signal, fast enough to matter.

What the agronomist is actually holding

The working desk now has four live streams. Soil sensors report volumetric water content and electrical conductivity at hourly intervals, drifting slowly out of calibration as probes age in situ. Satellite NDVI arrives every two to five days depending on cloud cover and constellation, and shows canopy stress days after the physiological event that caused it. Weather models — a six-hourly GFS run, a finer-grained regional forecast — disagree with each other and with the farm's own rain gauge more often than either vendor advertises. Commodity futures move by the minute and reframe, retroactively, whether a marginal spray was worth its cost.

None of these streams, individually, tells the agronomist how reliable the agronomist has been. NDVI cannot report that the agronomist's last three blight alerts were false positives. Only a kept record can do that: an outcome log, attached to named sources, checked against what the field actually did. The characteristic failure of the role is exactly the gap between having the streams and having the record. A blight window in wet Atlantic weather can be as narrow as forty-eight hours. Scheduling a proper risk assessment — pulling the Smith Period data, cross-checking soil moisture, weighing fungicide cost against a futures price that might make the treatment uneconomic even if the disease is real — can itself consume that window. The intervention closes while the assessment is still being arranged.

Two positions, honestly opposed

Set two views against each other, because both are defensible and the domain does not let either win cleanly.

The first: the agronomist should defer to the measured record. Suppose the farm's advisory service tracks a rolling calibration — over the last ten Smith Period alerts issued for this variety in this region, six led to visible lesions within the following fortnight, four did not. A 60% hit rate is worth knowing, and it licenses a specific downgrade: treat this alert as moderate rather than urgent, hold the fungicide until soil sensor data confirms leaf wetness duration rather than spraying on temperature and humidity alone. This is higher-order evidence doing its proper work — not touching the biology of the pathogen, but touching the agronomist's own instrument.

The second position, associated with the epistemologist Thomas Kelly's objection to what is called level-splitting: if the first-order evidence genuinely supports infection risk — the leaf wetness sensor confirms fourteen hours above threshold, the canopy is dense enough to hold humidity, the variety has known susceptibility — then learning that past alerts under similar conditions were wrong 40% of the time does not make this outbreak less likely. It makes the agronomist less trustworthy as a reader of outbreaks. Downgrading confidence in the disease because of a track record about the reasoner risks conflating two different objects. Worse, in a system where the spray window is forty-eight hours, hesitation induced by self-doubt is not free. A grower who waits for the record to be checked can lose the crop to the exact event the record was meant to help predict.

"If my first-order evidence supports P, telling me I am the sort of agronomist who errs one time in three does not lower the probability of P. It lowers the probability that I should be trusted to say so. Those are different quantities, and treating them as the same one is how good judgement gets talked out of itself."

Concede the objection fully. The level-splitting problem is real, unresolved in the philosophical literature, and agriculture supplies a sharper version of it than most domains because the cost of hesitation is not abstract — it is a field of blackened stems. The argument here does not require that higher-order evidence rationally compel a downgrade. It requires only that the number be available to act on, discount or override. An agronomist with a rolling hit-rate record has a choice the Kelly-style position and the deference position can each be tested against. An agronomist without one has no choice at all, only a hunch about a hunch.

Where the three generations sit on this

The lineage from Large Language Model to Large World Model to Large Universe Model is a story about which errors a system is even positioned to notice, not merely about richer inputs.
generationwhat it can know about its own reliability
Large Language ModelCan state, from text, that blight forecasting models are often wrong. Cannot observe its own hit rate on this farm's Smith Periods, because every outcome that would score it occurred after its training cutoff.
Large World ModelCan compare a predicted infection window against a sensed leaf-wetness reading within the same session — genuine higher-order evidence, but it dies with the episode and does not accumulate across seasons.
Large Universe ModelKeeps the soil sensor, NDVI, weather-model and price streams open indefinitely, with provenance, so each blight alert eventually meets an outcome and the hit rate becomes a maintained, attributable number rather than an assumption.

A Large Language Model trained on agronomy textbooks can recite that Smith Period rules have known false-positive rates in Irish trial data. It cannot know whether its own recommendations, issued for a specific field last Tuesday, were among the false positives, because the outcome sits outside its corpus entirely. A Large World Model watching a single growing season through a fixed sensor rig can compare its predicted lesion onset against what the camera actually observes, which is real self-correction, but it resets each season and each field, with no memory of whether this variety's alerts have run hot for three years running. Only a system that keeps every stream open past any single episode, and keeps track of which source said what, can hold a genuine, standing calibration figure for "how often does my blight alert on this variety, in this soil type, actually precede disease."

The feedback is not clean, and continuity does not fix that on its own

The second objection worth taking seriously: most agronomic advice is never adjudicated. A grower who follows a spray recommendation has no untreated control strip to compare against; the counterfactual — what would have happened without the fungicide — simply does not exist in the record. Feedback on irrigation timing is similarly thin, since a farm rarely runs a deliberately under-irrigated block just to test the agronomist's threshold. Selective feedback of this kind can produce a worse error estimate than honest uncertainty, because the cases that get scored are not a random sample of the cases that mattered.

This is the correct objection, and continuous intake does not dissolve it. What continuous intake with provenance does is make the selection structure visible. A farm that keeps trial strips — untreated rows, held out deliberately, the practice research trials have used for decades — generates exactly the kind of clean outcome data a rolling calibration needs, and a maintained record shows plainly which fields have trial strips and which do not, which recommendations were ever checked and which were acted on and forgotten. A frozen corpus cannot even ask that question, because it has no visibility into which of its claims were later tested. Continuity does not manufacture clean feedback. It makes the absence of clean feedback a known, recorded fact rather than an invisible one.

Where this leaves the agronomist

The honest resolution narrows rather than settles. Continuous, provenance-tracked intake gives an agronomist something a single season's sensor rig and a static corpus of trial literature cannot give: a maintained, attributable figure for how reliable each kind of alert has been, on this land, under this weather regime, discounted for which cases were ever actually checked. That is a genuine advance on the intake axis, and it is the terminal one — there is no further category of reliability evidence past a full, provenance-tracked, indefinitely running record; only longer records and better attribution within it.

What it does not do is compress the forty-eight-hour blight window, or settle whether the agronomist should have deferred to a 60% hit rate on Tuesday morning. The record can arrive in time to inform next season's threshold and still arrive too late for this week's field. Higher-order evidence, however complete, is evidence about the reasoner. It has never been evidence that grows a crop back.

Continue