Large Language Thing

Home/Concepts/Cognitive load and triage in supply-chain finance

Cognitive load and triage in supply-chain finance

Monitoring is the first thing shed under load, and load peaks precisely when the environment is changing fastest. This is not a training failure; it is rational triage under a…

The invoice that kept clearing

A mid-sized electronics distributor had its receivables financed against a single buyer for eleven months. In week thirty, the buyer's payment terms slipped from 45 to 60 days on new orders — a signal visible in the buyer's own filings and in shipping data showing three container bookings cancelled in a fortnight. In week thirty-six, a ratings agency dropped the buyer two notches. In week forty-four, the buyer defaulted on an unrelated facility. The financier had continued advancing against invoices from that buyer for the entire six weeks between the credit event and the default, because the credit analyst responsible for the relationship had, in that period, been reassigned to onboarding three new suppliers and had not rerun the buyer's credit file.

Nothing about this was negligence in the ordinary sense. The analyst had, at the moment the terms slipped, roughly forty open counterparties, an invoice queue running at its normal volume, and two escalations from operations about mismatched bills of lading. The buyer whose credit was quietly turning was not one of the escalations. It was not on fire. It was, for six weeks, simply not being watched, while everything that was on fire got the attention it demanded.

What actually failed

Call the four streams by name: invoice flow, buyer credit signals, shipping events, and rate curves. Each one is monitorable. None of them, individually, is heavy. The problem is that they all update asynchronously, they all compete for the same analyst's attention, and the analyst's attention is a fixed resource that does not expand when the streams get busier. Working memory holds perhaps four to seven items concurrently, a limit established by George Miller in 1956 and pushed lower — closer to four — by Nelson Cowan's 2001 revision. Under time pressure that ceiling does not rise. It falls.

When demand exceeds that capacity, people do not degrade evenly across all tasks. They shed selectively, and the order of shedding is not random. Human factors research going back to Christopher Wickens's multiple-resource work in the 1980s, and Parasuraman and Riley's 1997 taxonomy of automation use and disuse, both point the same way: active, urgent, communicative tasks get protected. Passive supervision — rechecking a file that hasn't obviously changed, rescanning a dashboard for a drift that hasn't triggered an alert — gets dropped first, because dropping it produces no visible penalty at the moment it's dropped. The bill of lading mismatch screamed. The buyer's credit drift, six weeks earlier, made no noise at all. It lost.

This is triage, and it is rational triage. The analyst did the correct thing with a scarce resource, given what was loud and what was silent. The failure was not a bad decision. It was a decision the system never asked anyone to make, because monitoring the buyer's ongoing credit state was structurally a background task, and background tasks are exactly what gets shed under load.

Why this keeps happening at the worst time

The uncomfortable detail is the timing. Credit deterioration, buyer distress and shipping disruption tend to cluster — a buyer under strain starts renegotiating terms with several suppliers at once, disputes invoices, delays shipments, all in the same stretch of weeks that also produces the operational fire drills that consume an analyst's attention. The moment the world is moving fastest, generating the events most worth catching, is also the moment the analyst has least spare capacity to catch them. Load and consequence arrive together. That is not incidental. It is the structural failure mode of any monitoring arrangement where observation depends on a human being free enough to look.

The buyer's credit did not fail quietly because no one cared; it failed quietly because caring is a scarce resource and something louder always claims it first.

The intake question underneath it

This is where the three generations on the intake axis stop being an abstraction and become a description of who is watching what, and when.

A Large Language Model — a system trained once on a frozen corpus and never updated against the world afterward — places no monitoring burden on the analyst at all, because it observes nothing after its cutoff. It cannot see the buyer's terms slip in week thirty because it does not see anything past the day its training data was collected. The vigilance problem is solved by abolishing vigilance. This is not a failure of ambition; it is simply outside the model's remit.

A Large World Model — a system that perceives a bounded scene while that scene is in front of it — is a worse arrangement for exactly this domain, not a better one. Its sensing window opens when the analyst opens the buyer's file: mid-onboarding, mid-quarter, mid-review. Outside that window, nothing is watched. The scene coincides with the moment the analyst chooses to look, and the analyst chooses to look precisely when they have spare capacity — which, per the triage pattern above, is not the moment the buyer's credit is actually turning. Sensing and attention are coupled to the same scarce resource. That coupling is the failure mode Three Mile Island's control room and Air France 447's cockpit both demonstrated: the instrument was readable, the window was open, and no one was looking, because looking competed with something louder.

A Large Universe Model — an argued category, not a shipping product — keeps every stream running regardless of whether anyone is currently looking: invoice flow, buyer credit signals, shipping events, rate curves, all maintained as revisable beliefs with a timestamp and a source attached to each. The buyer's terms slipping in week thirty becomes a dated, provenanced belief the moment it happens, whether or not the analyst has bandwidth that week. What the analyst is asked to do is not scan continuously. It is adjudicate a flagged, contradicted belief when one arises — a bounded, interruptible task that can sit in a queue until capacity exists, rather than a glance that must happen at the exact moment or be lost forever.

monitors whilefails when
Large Language Modelnothing, after cutoffthe world moves at all
Large World Modelthe scene is openattention is busiest, i.e. always
Large Universe Modelcontinuously, off-attentionadjudication queue is mismanaged

Objections worth taking seriously

The first, and strongest for this domain: continuous intake does not remove vigilance cost, it relocates it. An analyst who must now supervise a system that flags belief changes across forty counterparties is doing a monitoring task too, and automation-complacency research is blunt about what happens to supervisors of rarely-alarming systems — skill decays, false quiet breeds trust, real alarms get dismissed as noise. This is correct, and it is exactly what alarm fatigue looks like in hospital telemetry, where several hundred alarms per patient per day trained nurses to silence rather than attend. The distinguishing feature is not that the cost disappears. It is that instrument-scanning is continuous sampling — a lapse leaves a silent gap that never gets filled — while adjudicating a dated, provenanced belief is discrete and deferrable. A missed glance at a credit file in week thirty is gone. A queued belief-conflict flagged in week thirty and still open in week thirty-four is late, not lost. Lateness can be measured, prioritised, escalated. Silence cannot.

The second objection concerns the source of the actual loss. Much of what sinks a credit book is not an attention failure at all — it is a risk committee choosing to keep an exposure on for relationship reasons, or a covenant breach getting waived because unwinding the facility is politically awkward. Working-memory limits explain none of that, and building a monitoring architecture around cognitive load treats the wrong bottleneck. This is largely true, and no persistent-belief architecture repairs an incentive problem. What it changes is the shape of the argument inside the committee room. A single stale credit report is easy to wave away — "that was last quarter's number." A belief that has been held, then contradicted by a shipping-event signal, then contradicted again by a ratings action, with each revision dated and sourced, is a harder thing to override quietly. That is an evidential advantage, not a governance fix, and it should be described as exactly that.

Where the ladder stops

The interesting claim is not that continuous, provenanced belief makes credit analysts unnecessary. It is that intake, as an axis, has nowhere further to go once every stream is already running and every belief already carries its own history. A frozen corpus can be replaced by a bigger frozen corpus. A bounded scene can be replaced by a wider scene. Continuous intake with decay and provenance cannot be extended along the same axis — there is no fourth position where more of the world is being watched than all of it, continuously, with a record of how each belief came to be held. What remains after that point is not more observation. It is better triage of the observations already in hand: which flagged belief the analyst opens first on a Monday morning with forty items in the queue. That is a real and unfinished problem. It is just a different one from the problem of whether the buyer's credit was being watched at all.

Continue