Large Language Thing

Home/Concepts/Granger causality in agriculture

Granger causality in agriculture

Granger causality cannot be computed over an unordered archive, and cannot be kept true by a frozen one. The criterion needs three things: order, continuation, and breadth of the…

The blight that arrived on schedule

On a 340-hectare wheat block in early May, the soil moisture probes at 10cm depth logged three consecutive nights above 90% relative humidity in the canopy microclimate — the threshold at which Zymoseptoria tritici spores germinate on leaf surfaces. The satellite pass two days later showed NDVI easing back half a point across the block's northern third, the kind of dip that precedes visible lesion by four to six days. The regional weather model, updated that morning, showed a 48-hour dry window opening the following afternoon — the only spray window before a week of rain that would ground every sprayer in the county.

The agronomist responsible for the block had all three readings by Tuesday lunchtime. The fungicide recommendation went to a scheduled Wednesday review, the usual slot for signing off variable-rate applications across the client's holdings. By Wednesday afternoon the dry window had opened and closed. It rained for six days. The infection that had been contained to the northern third at the point of decision had run through the canopy by the time spraying was physically possible again. Yield loss on that block came in at 14%, roughly four times what a Tuesday application would have cost in disease.

Nothing about the data was wrong. The moisture threshold was real, the NDVI dip was real, the forecast window was real and it closed exactly on schedule. What failed was the interval between evidence and action — a gap the size of a calendar entry.

What the agronomist actually had

Four streams were running: point sensors reporting soil and canopy conditions at fifteen-minute intervals, a satellite product refreshed every two to five days depending on cloud cover, a numerical weather model reissuing its forecast six-hourly, and a commodity feed pricing wheat futures that would ultimately decide whether the fungicide spend was justified against the crop's value. Each stream, on its own, told a partial story. The moisture readings said infection risk was rising. The NDVI said something had already begun to change physiologically. The weather model said there was one door and it would shut. None of these streams, alone, said act now. The claim that mattered was a claim about order: did the rising humidity precede the NDVI dip, and did both precede a window that would close before intervention could follow?

That is not a question about mechanism. Nobody needed a differential equation for spore germination to make the spray decision correctly. It was a question about which series moved first, by how much, and whether that lead was reliable enough to act on before the next observation arrived. That question has a name.

Predictive precedence, not mechanism

Clive Granger made this precise in 1969, formalising a suggestion Norbert Wiener had floated in 1956: one series is said to Granger-cause another if the past of the first improves prediction of the second beyond what the second series' own past already supplies. Granger was working on macroeconomic data — arguing whether money supply moved output — where controlled experiment was impossible and observational inference was the only tool available. He was explicit, and later regretted how often people forgot, that the test is defined relative to an information set: the collection of series available to the observer up to the present moment, which he wrote as Ω_t. Change Ω_t and the verdict can change. A confounder sitting outside every instrumented channel will make two effects look causally related to each other when both are driven by a third thing nobody is watching.

In the wheat block, the relevant information set was: soil moisture, canopy humidity, NDVI, six-hourly forecast updates, and nothing else formally logged. Did canopy humidity Granger-cause the NDVI dip, conditional on everything else in Ω_t? Almost certainly — the agronomic literature on septoria progression is built on exactly that kind of lagged relationship, established over decades of trial data. But the test that mattered on that Tuesday was not whether the relationship existed. It was whether the assessment of it could be completed before the weather model's own forecast, itself part of Ω_t, expired. The information set was adequate. The institution consuming it was not open on the same clock.

The information set is the whole argument

This is where the failure stops being a scheduling story and becomes an intake story. Granger's criterion contains, in its own definition, the axis this site is built on. Causality-as-precedence is not a property of spores and leaves alone; it is a property of the world relative to what an observer is permitted to see, and how promptly.

A frozen archive of agronomic trial data — the kind that trains a model on last decade's disease-progression studies — can recover that septoria-conducive humidity precedes NDVI decline, because that finding is stable across seasons and regions. What it cannot do is tell you that this window, on this block, closes at 2pm Wednesday. The general relationship survives in a corpus; the current instance does not, because currency is not a property a frozen corpus can hold. A bounded scene — a single field walked and sensed for a season — can track the local lead-lag as it happens, but only for the channels pointed at that field and only for as long as the sensors keep running. Neither reaches the case the agronomist actually faced: many streams, still arriving, needing to be conditioned against each other continuously, with the verdict revised the moment the forecast reissues.

conditioning sethorizonfailure mode on this block
Large Language Modeltrial literature, frozen at training cutoffnone beyond cutoffknows septoria precedence in general, cannot see this week's forecast
Large World Modelthis season's sensors and imagery, livebounded to the monitored field and seasontracks the local lead-lag correctly but the institution reading it is not continuously open either
Large Universe Modelsoil, canopy, satellite, weather, price, held open with timestampscontinuing, revised on each reissuethe case where the Wednesday review is replaced by a standing test that fires when Ω_t changes
The failure was not a bad model of septoria; it was a conditioning set that updated faster than the institution consuming it.

Two objections from the field

Granger causality is not causality. It sits at the observational rung of Pearl's hierarchy. Two series driven by an unmeasured third will pass the test convincingly. Adding more sensors to a system that only watches does not lift it out of that rung.

Correct, and agronomy has its own version of the unmeasured confounder: a subsoil compaction pattern from three seasons back, never mapped, driving both moisture retention and canopy stress independently. No amount of NDVI and weather data removes that risk by watching harder. But widening the conditioning set does remove specific, nameable confounds once they are identified — soil compaction, once mapped, either enters Ω_t or the finding stays suspect and is flagged as such. And continuous intake turns accidental interventions into evidence: a neighbouring block sprayed a day early because of a machinery slot, an irrigation line that failed for six hours, a fertiliser trial with a randomised control strip. These are not designed experiments, but a system holding every stream open will notice them as they happen and read causal structure off them. That is narrower than the objection assumes it can be answered, and it should be stated that narrowly.

Sampling defeats the criterion regardless of breadth. Canopy microclimate changes on the scale of hours; a satellite product refreshed every four days aggregates over exactly the window where the causal action happens, and can manufacture or reverse the apparent lead-lag.

This is the sharper problem, and it is real. NDVI is not a five-minute stream and never will be under current satellite revisit rates; treating a four-day composite as if it moved at the pace of the disease it is meant to flag is a category error that has produced false precedence claims in the agronomic literature before now. The corrective is not to force everything onto one grid, which is what a frozen archive has usually already done by the time it is published. It is to let each stream stay at its native cadence — moisture at fifteen minutes, canopy humidity likewise, satellite at its real revisit interval, forecast at six hours — and to condition the disease-progression model on lags appropriate to each, rather than resampling the fast series down to the slow one's rate. That does not fix satellite latency. It stops the latency from silently contaminating the moisture data's much finer resolution.

Where the ladder ends

Granger's Ω_t is the intake axis written out in 1969, decades before anyone thought to organise machine intelligence by what it is allowed to see. A frozen corpus fixes Ω_t at a cutoff and strips most of the ordering that made the criterion computable in the first place. A bounded scene restores order and rate for as long as the scene runs, which is real progress and still not enough for a decision that must beat a forecast's own expiry. Holding every relevant stream open, timestamped, with provenance attached and revised the moment a new value arrives, is what the criterion actually asks for. There is no further intake improvement past that point — no fourth channel, no finer grid, that changes the shape of the problem rather than filling in its detail. What is left, once the streams are open and current, is not more data. It is whether the agronomist's Wednesday meeting can be replaced by a standing test that fires on Tuesday, before the door closes.

Continue