Large Language Thing

Home/Concepts/Reproducibility and the replication crisis in venture capital

Reproducibility and the replication crisis in venture capital

Every claim drawn from published science has a half-life. Retraction Watch lists over 50,000 retractions; papers are often cited hundreds of times after withdrawal. Any system…

The Memo That Outlived Its Market

In March 2021 a partner at a mid-size growth fund wrote a sixteen-page investment memo backing a vertical SaaS company selling scheduling software to freight brokers. The memo cited three things: a hiring signal (the target had gone from 40 to 140 engineers in nine months, evidence of a founder who could recruit against a hot market), a filing signal (two competitors had just closed Series Bs at valuations implying a addressable market north of $4bn), and early product telemetry (weekly active broker accounts up 30% quarter on quarter). The fund led the round.

By the second quarter of 2023 all three signals had inverted. The hiring hadn't just slowed, it had reversed — the company had cut 60% of that engineering headcount. The competitors' filings told the same story in aggregate: freight rates had collapsed, brokers were consolidating, and the addressable market the whole cohort had priced against had shrunk by more than half. Product telemetry showed weekly actives flat for five straight quarters, the growth curve dead since the freight recession began. None of this was secret. It was sitting in the company's own board decks, in public rate indices, in the hiring pages of three direct competitors. The partner defended the position at three consecutive investment committee meetings anyway, using language from the original memo almost verbatim, and the fund did not write the position down until a bridge round forced the issue in late 2023 — roughly a year after the market the thesis assumed had stopped existing.

What Actually Failed

The proximate failure looks like stubbornness, and there was some of that. But the structural failure is more interesting and more common: the memo was written once, as a conclusion, not as a belief with an expiry condition attached. Nothing in the fund's process asked the question a good research design asks by default — under what future evidence does this claim get downgraded? The hiring signal, the filing signal and the telemetry signal were treated as diligence inputs consumed at the moment of the decision, not as streams to keep watching. Once the round closed, the memo went into a drive folder and the partner's attention moved to sourcing the next deal. The three data streams kept running. Nobody was reading them against the original thesis, because the thesis had already served its institutional purpose: it had gotten the deal approved.

This is exactly the shape of the replication crisis in published science, transposed. A landmark finding is published, cited, and built upon; the failed replication, when it eventually surfaces, is a smaller, later, less-cited document, weakly linked to the original, and most readers of the original never see it. Amgen's oncology group, trying to build on 53 "landmark" preclinical cancer papers, could confirm the findings of only 6. The Reproducibility Project's re-run of 100 psychology studies reproduced 36 of the original effects. The gap between publication and correction is not an accident of any one paper; it is structural, produced by incentives that reward the first confident claim and barely reward the quiet work of checking it later.

The Venture Version of a Retraction

In venture, the "retraction notice" almost never arrives as a retraction. It arrives as a filing nobody cross-referenced, a hiring page nobody rechecked, a telemetry dashboard nobody reopened. The Reinhart-Rogoff debt paper had an identifiable, dateable correction — a graduate student found a spreadsheet error in 2013 and the 90% threshold claim collapsed publicly. Venture theses rarely get that clean a moment. The freight-tech thesis didn't get retracted; it got quietly outlived by three separate data streams that kept moving after the decision was made, none of which anyone had a standing obligation to keep checking.

That is the diagnosis, and it generalises. A thesis in venture is a claim built from evidence with a half-life, exactly like a claim built from a psychology experiment or a preclinical cancer study. The filing was true when it was filed. The hiring signal was real when it was measured. The telemetry was accurate on the day it was pulled. None of that implies durability. The mistake is treating a true-at-T claim as a true claim, full stop, and then acting on it eighteen months past T without asking whether the streams that produced it still support it.

A diligence memo is a hypothesis with a publication date, not a fact with a shelf life.

Why a Frozen Corpus and a Bounded Scene Both Miss This

A fund's version of a Large Language Model is its accumulated memo archive, plus whatever market data was pulled at diligence and frozen into the deck. It has the advantage of every prior deal's reasoning available in one place, and the fatal weakness that the archive is silent about anything that happened after the memo was filed. It contains the freight-tech thesis exactly as written in March 2021, with no annotation that the market it assumed had come apart by 2023. Worse, it contains the confident version — the version written to get the deal approved — not the corrected version, because no one writes a follow-up memo titled "our March 2021 thesis was wrong" unless a write-down forces it.

A fund's version of a Large World Model is the quarterly portfolio review: a bounded scene, sensed once per quarter, board deck by board deck. This is real progress over a static archive — it catches problems a frozen memo never would, because someone is at least looking at the current state of the company. But it resolves nothing about provenance. A quarterly review tells you the freight-tech company's actives are flat this quarter. It does not tell you whether the original market-sizing filing the thesis relied on has since been undercut by two competitors' collapse, because that is a fact about a different company, in a different filing, on a different clock, and nobody's quarterly review process is built to reopen a three-year-old comparable set.

what it holdswhat it misses
memo archive (LLM-equivalent)the thesis as written at closeeverything that happened after
quarterly review (LWM-equivalent)the portfolio company's current statethe provenance of the thesis it was measured against
continuous streams with provenance (LUM-equivalent)filings, hiring, telemetry, market structure, all still running, tagged to the claims they supportnothing structurally, though coverage and latency remain real limits

What the failure actually demands is a position where filings, hiring signals, product telemetry and market structure keep running past the decision date, where every clause of the thesis carries the evidence that produced it, and where a demotion — the hiring reversal, the rate collapse, the flat actives — automatically downgrades the belief rather than waiting for a write-down to force a rewrite. That is not a faster memo. It is a different relationship between claim and evidence, in which "our thesis" is a live, tagged, decaying object rather than a document filed once and cited from memory.

Two Objections Worth Taking Seriously

The first: retraction is a lookup problem, not an architecture problem. Crunchbase already flags a company as shut down; PitchBook already flags a down round. Why not simply check the flag at query time against a frozen thesis, rather than rebuild the whole intake pipeline?

"You don't need the fund to watch everything continuously. You need a dashboard that pings you when the company you backed goes bankrupt. That's a status check, not an epistemology."

This is fair, and it is exactly right for the sharp cases — bankruptcy, a down round, a public shutdown. Those are the retractions: dateable, discrete, cheap to flag. But most thesis decay in venture is not a bankruptcy. It is the freight-tech pattern: hiring quietly reversing, telemetry quietly flattening, a competitor's filing quietly revealing that the market is smaller than assumed. Nothing gets flagged, because nothing crosses a bright line. The claim degrades without being formally withdrawn, and a status-check dashboard built to catch discrete events will not notice a slow bleed. Catching that requires watching the underlying streams, which is a continuous intake system by another name — the objection concedes the architecture while trying to outsource it to someone else's flag.

The second objection cuts the other way: continuous ingestion of hiring boards, product dashboards and scraped filings does not solve the problem, it multiplies it. A quarterly memo has a stable, known error profile. A system pulling in LinkedIn scrapes, Twitter sentiment and press releases in real time absorbs noise faster than correction, and early signals — a founder's optimistic tweet, a single strong month of telemetry — are the least reliable evidence there is. This is the strongest objection, and it is correct about naive versions of the argument. It is exactly why the position that actually solves the freight-tech problem is not "ingest more, faster" but ingest continuously while tagging every belief with the evidence and confidence that produced it, so a single strong telemetry month raises a weakly held belief rather than a confirmed one, and a hiring reversal three quarters later can demote it cleanly. Throughput without provenance is just faster staleness.

The Limit, Not the Failure

None of this means every venture thesis rots at the same rate. Infrastructure bets on durable primitives — payments rails, identity, compute — decay slower than vertical bets on a single market structure that a rate cycle can dissolve in a year. The heterogeneity is real. But it is itself the argument for the third position rather than against it: a static memo archive cannot tell you which regime a given thesis belongs to, because that classification is a property of the subsequent evidence stream, not of the memo's text. Filings, hiring signals, telemetry and market structure are not settled once at diligence. They are still running. A fund's intake either keeps running with them, carrying provenance and allowing demotion, or it freezes a confident document and calls it a thesis. There is no fourth kind of intake left to invent past "every stream, continuously, tagged to what it supports." What remains is discipline in using it.

Continue