Large Language Thing

Home/Concepts/Goodhart's law in retail operations

Goodhart's law in retail operations

If measures decay under optimisation, then any system whose evidence stops arriving is on a decay curve from the moment it is deployed. Its accuracy is highest at the cutoff and…

The demand curve that already moved

A category manager building a spring assortment in November is not looking at spring demand. There is no spring demand yet to look at. She is looking at last spring's sell-through rate, this autumn's early-read POS data, a supplier notice about a cotton price rise, and a forecast model trained on eighteen months of history. She commits an open-to-buy budget against that composite picture, locks in purchase orders with twelve-to-sixteen-week lead times, and waits. By the time the stock lands on the shop floor, the demand curve she planned against has already been overtaken by whatever actually happened in the intervening months — a competitor's promotion, a shift in weather, a TikTok trend that moved a category three weeks before the buy but wasn't yet visible in the tool she used.

This is not a forecasting failure in the ordinary sense. The forecast was reasonable given what it saw. The problem is what it saw: a slice of history, frozen at the point the buying decision had to be made, evaluated against a benchmark — last season's sell-through — that was itself a fossil of an earlier, different distribution. The assortment gets planned against a target that has already moved by the time the plan is executed. That is Goodhart's law, running quietly through a category management process that has never heard the name.

Goodhart's law, stated plainly

Charles Goodhart made the observation in 1975 while chief adviser at the Bank of England. The authorities had been targeting a measure of the money supply because it correlated well with inflation. Once they targeted it, banks restructured deposits to sit outside the definition, and the correlation that had justified the target broke down. Marilyn Strathern gave the compressed version in 1997, writing about British university research assessment: when a measure becomes a target, it ceases to be a good measure. Donald Campbell had said much the same about social indicators a year after Goodhart, from a different discipline, independently.

The mechanism is selection, not conspiracy. A proxy — a money supply aggregate, a research output count, a weeks-of-supply ratio — correlates with the thing you actually care about across some historical distribution of behaviour. The moment you optimise against the proxy, you change the distribution that produced the correlation. The proxy detaches. The target quietly stops being served while the proxy keeps reporting success. Nobody has to be lying for this to happen. The system just stops observing the part of reality that used to keep the proxy honest.

Two positions on the buying cycle

Retail operations have built an entire discipline around managing this decay, largely by trying to slow it down rather than solve it. Call this the disciplined-cycle position. A category manager reviews a category on a fixed cadence — often quarterly, sometimes monthly for fast fashion — because reacting to every nightly POS fluctuation would introduce more instability than it removes. Order quantities that chase daily sell-through swings amplify upstream: a small retail-level demand wobble becomes a large swing in orders to distribution centres, which becomes a larger swing in orders to suppliers, who over-produce or under-produce in response. This is the bullwhip effect, documented by Jay Forrester in 1961 and confirmed repeatedly in retail supply chains since: variance in orders placed with suppliers routinely runs several times higher than variance in the underlying consumer demand that supposedly justified them. On this view, a frozen review cycle is not a failure to observe reality quickly enough. It is a deliberate low-pass filter, and it is doing its job precisely by not reacting to everything the streams report.

Chase the live number and you get a supply chain that oscillates. The four-week cycle exists because the four-week cycle is calmer than the truth. A category manager who re-forecasts nightly is not more accurate. She is more anxious, and she passes that anxiety up the supply chain as wasted capacity and expedited freight.

The opposing position takes the same facts and draws the opposite conclusion. The disciplined cycle is calm because it has stopped listening, not because it has found the right cadence. Between reviews, the streams that would have told the category manager her demand curve had moved — POS data updating nightly, inventory telemetry flagging unexpected sell-through velocity on a handful of SKUs, supplier notices about input shortages weeks before they hit shelf — sit unread. The quarterly cycle converts a continuous signal into four discrete snapshots a year, each one already stale by the time the buy commits. Under this reading, the calm the disciplined position prizes is the calm of a system deliberately declining to know things it could know, and the cost of that calm is paid later, in markdowns on stock that missed the curve and stockouts on SKUs that moved faster than the model expected. Retailers on the losing end of this trade-off tend to discover it as a bimodal season: heavy discounting on one shelf, empty pegs on another, both caused by the same frozen assortment plan.

Neither position is straightforwardly wrong. The disciplined cycle correctly identifies that fast reaction has a cost, paid in supply chain volatility. The continuous-intake position correctly identifies that the cost of not reacting is also real, paid in margin, and that it is systematically hidden because nobody attributes a markdown to a forecast that was three months out of date by design.

Two objections worth taking seriously

The first objection is that faster intake does not defeat Goodhart's law, it just gives the gaming a shorter feedback loop. Retail has its own version of the rat-tail bounty. A category manager who ties supplier scorecards to on-time-in-full delivery gets suppliers who ship early and incomplete rather than late, because early counts and late is penalised even when early creates receiving-dock backlogs the metric never sees. A sales team incentivised on weeks-of-supply near quarter close has an incentive to over-order in week eleven so the ratio looks healthy in week twelve, then work off the excess in week thirteen when nobody is measuring. Continuous, granular intake does not remove this; it can sharpen it, because a supplier watching the retailer's live inventory telemetry can time a forecast inflation to exactly the window the retailer is most likely to act on it.

This is a fair hit, and the honest answer is that it is only partly answerable. What continuous intake changes is not whether gaming happens but whether it can be traced. A frozen quarterly review has no record of which upstream signal justified a given buy; when the buy turns out wrong, there is no way to localise the failure to a specific supplier notice or a specific demand spike that was misread. A system that retains provenance — this order quantity rested on this POS trend plus this supplier lead-time notice, timestamped — lets the category manager find the moment a proxy detached and withdraw the belief that rested on it. Gaming under fast feedback is faster and more visible. Gaming under a frozen cycle is slower and, because nothing was recording the reasoning at the time, effectively permanent once discovered three seasons later, if it is discovered at all.

The second objection is about physics, not incentives. Even granting that continuous intake is the right target in principle, a garment still takes twelve weeks to arrive from a factory in a different hemisphere. No amount of nightly POS granularity shortens a container ship's transit time or a supplier's minimum production run. Real-time demand sensing meets a supply chain that can only respond on a cycle measured in weeks. The category manager's constraint is not how fast she can know; it is how fast anyone downstream can act on what she knows.

This is exactly the objection that separates knowing sooner from acting sooner, and retail operations live entirely inside that gap.

What narrows

Both objections land, and both narrow the claim rather than break it. Continuous intake with provenance is the only configuration in which a category manager can trace a demand assumption back to the observation that supported it and retract it when that observation stops holding — that much survives. What it does not deliver is instantaneous correction of the physical assortment, because lead times, minimum order quantities and container schedules impose a floor on reaction speed that no amount of streaming data removes. The realistic version of the argument is that assortment planning should shorten its re-forecast interval and retain provenance on every buy, catching a moving demand curve within weeks rather than a season, while accepting that the buy itself still executes on the supply chain's clock, not the data's. That is a narrower claim than "watch everything, act instantly." It is also the only version that survives contact with a shipping schedule.

Continue