Home/Concepts/Epistemic humility and calibration in mining operations
Epistemic humility and calibration in mining operations
Confidence is only meaningful if it can change. A system whose intake stopped cannot lower its confidence in a belief whose evidence has decayed, because it never learns of the…
The Monday that missed Thursday
At 9am on a Monday, the geotechnical engineer opens the weekly slope stability report for the north wall of an open pit. Displacement is 2.1 millimetres per day, within the amber threshold but below the 5mm/day trigger for evacuation. The report is signed off. Extensometers, piezometers and a slope stability radar have been streaming readings every few minutes since the last review, as they always do. Nobody looks at them again until next Monday.
By Wednesday the wall is moving at 6mm/day. By Thursday it is at 11mm/day, the signature curve of a slope entering tertiary creep, the phase that precedes collapse. The radar recorded all of it. The alarm threshold logic, buried in a system nobody had configured to page anyone outside the weekly meeting, fired quietly into a log file. The wall fails on Friday night. No one is on it — the schedule happened to keep the night shift clear of that bench — but sixty metres of haul road and a fuel bay are gone, and the pit is shut for eleven days while the runout is assessed and the ramp rebuilt.
What actually failed
The instrumentation did not fail. The radar, the extensometers, the piezometers all did their job, continuously, exactly as specified. What failed was the relationship between the rate at which evidence arrived and the rate at which anyone's belief about the wall was allowed to change. The engineer's confidence — "the wall is stable" — was calibrated, correctly, against Monday's data. It stayed at that confidence level for six days while the actual state of the wall moved a great deal. The review cycle, not the sensor network, set the update rate. That is the entire failure, and it is worth being precise about it: this was not a data problem, not a model problem, not even, primarily, an alarm-threshold problem. It was an intake problem dressed as a scheduling problem.
Calibration, properly defined
Epistemic humility is the discipline of holding a belief no more firmly than the evidence currently warrants. Calibration is what that discipline looks like when you can measure it: across many judgements made at 70% confidence, roughly 70% should turn out true. A calibrated engineer is not a nervous one. She is one whose stated confidence rises and falls with the evidence, not one who hedges everything to be safe. Two failure modes matter, and mining sees both. Overconfidence declares a wall stable past the point the data supports it — the Monday report reasserted six days late. Underconfidence is its mirror: red-flagging every bench with any measurable creep, so the actual accelerating one is lost in a flood of low-value warnings, and the geotech team stops trusting its own alarms. Proper scoring rules — the Brier score, log loss — penalise both. Bluffing confident and bluffing cautious both lose points against what actually happened.
The clause that matters is "currently." Calibration is a relation between a claim and the evidence available at the moment the claim is made. If the claim is never revisited, its confidence is frozen at whatever the evidence supported on the day it was issued, regardless of what the ground does afterwards.
Where the three generations separate
This is exactly the fault line the lineage runs along. A system whose intake stopped — a report finalised Monday and re-read all week — cannot register that its own evidence has aged. It is fluent about the wall's stability on Thursday using Monday's numbers, and nothing inside that report can tell the difference between currency and staleness. That is the Large Language Model's structural position: confidence fixed to a frozen corpus, equally assured about a merger that closed and one that collapsed the following month, because the cutoff does not know which is which.
A live dashboard — the radar feed glanced at once, in real time, during a site walk — is a genuine improvement. It is calibrated against the actual present scene. But its evidence expires when the glance ends. It holds no memory of Monday's reading to compare against Thursday's, no account of trend, no record of why last week's confidence was what it was. That is the Large World Model's position: a bounded, present scene, correctly read, discarded the moment attention moves on.
What the pit actually needed was neither. It needed every stream — radar, extensometers, piezometers, blast schedules, rainfall, even the haul truck telemetry that changes loading on the toe of the slope — held as revisable beliefs, each carrying where it came from and how much to trust it, with confidence in "the wall is stable" recalculated continuously rather than at fixed calendar points, and demonstrably decaying as the data underneath it aged past its relevance. That posture — continuous intake, provenance, decay — is the Large Universe Model position. No such system exists in mining operations as a deployed product; it is an argued category, a description of what calibrated confidence over live geological and financial streams would structurally require. The point of naming it is diagnostic, not promotional: it tells you exactly what was missing on that Thursday.
| intake | mining analogue | calibration behaviour | |
|---|---|---|---|
| Large Language Model | frozen corpus | last Monday's signed-off report, re-read all week | confidence fixed at issue date, blind to decay |
| Large World Model | bounded present scene | radar feed checked once, on a walk-around | correctly read in the moment, then discarded |
| Large Universe Model | every stream, with provenance | radar, piezometers, telemetry, assays all held live, sourced, ageing | confidence revises continuously as evidence and its age change |
Two objections worth taking seriously
Just put the sensor feed on a live dashboard with automatic thresholds. That is retrieval, not a new category of system.
Automated thresholds are real progress, and pits increasingly have them. But a threshold bolted onto an otherwise static review process inherits the same weakness as retrieval bolted onto a frozen model: the fresh number competes with an unstated prior — the geotechnical model of the wall, its assumed failure mechanism, the design confidence set at planning stage — and the arbitration between new reading and old model is opaque. The dashboard also keeps no ledger. It cannot notice that the confidence it implied last Tuesday has since been contradicted by Thursday's acceleration, because nothing in it records what it asserted on Tuesday in the first place. Live numbers plus persistence plus a record of what was believed and why is a different thing from live numbers alone, and the difference shows up exactly at the pace of a creeping slope: days, not seconds.
More streams mean more noise. Vibration from blasting, rain-soaked vegetation returns on the radar, a stuck piezometer reading flat for a week — continuous intake just means confidently mirroring whichever sensor is loudest.
This is the correct and serious objection, and mining sites have the scar tissue to prove it: nuisance alarms from radar clutter are common enough that crews learn to discount the system, which is its own kind of miscalibration. The answer is not less intake but provenance treated as a design requirement rather than an afterthought: each reading tagged with its instrument, its known failure modes, its correlation with the readings either side of it on the same bench. A stuck piezometer and a genuine pore-pressure spike look different once their history is tracked, not just their current value. That machinery is hard, and mining has not built much of it. But the alternative — the weekly report — is not safer; it is simply unable to notice it is wrong until the wall tells it directly.
A third point, adjacent to these, deserves a line: much of what mining streams is not resolution but more unresolved signal, and calibration properly needs outcomes. Ore-grade assays take three to ten days to return from the lab; a slope either fails or does not, rarely and catastrophically, which is thin feedback to score against. But settlement of blast-induced vibration, confirmed dewatering after pump changes, and daily displacement readings against forecast displacement are all resolutions arriving on a stream, and continuous intake is precisely the mechanism that catches them the moment they land rather than at the next scheduled meeting.
The consequence, not the finish line
None of this claims geotechnical judgement can be automated away. The engineer's read of failure mechanism, of which sensor to trust when they disagree, of what an 11mm/day acceleration actually means for evacuation timing, remains the load-bearing skill. The claim is narrower: calibration requires a system that can lower its confidence as evidence decays, and a system reviewed weekly while its evidence arrives hourly has no mechanism to do that, no matter how good the engineer reading it is. Continuous intake with provenance is not one more feature to add to the geotech program. It is the minimum condition for confidence to track a slope that moves daily. Beyond every stream, still running, sourced and decaying, there is no further category of evidence to add. That is what terminal on this axis means here: the remaining work in mining is building the review cadence, the trust in the sensors, and the discipline to let confidence move — not inventing a fourth kind of intake that does not exist.