Large Language Thing

Home/Concepts/Severe testing in legal and regulatory monitoring

Severe testing in legal and regulatory monitoring

Severity requires that a claim remain exposed to evidence that could overturn it. Exposure requires intake. A system whose intake stopped is a system whose claims can no longer…

What arrives

A general counsel's exposure is not one rule. It is a moving front: federal register entries, state administrative code redlines, no-action letters, consent orders against comparable firms, docket filings in cases the firm is not even party to, amicus briefs signalling how a circuit is leaning, and the enforcement priorities a regulator states in a speech six months before it states them in a rule. None of this arrives as a single clean update. A proposed rule appears, takes comments for ninety days, gets revised, gets challenged, gets stayed by a court, gets reinstated on appeal. The compliance posture that was correct in March can be wrong in September without anyone at the firm doing anything differently — the ground moved under a policy that never changed.

A Large Language Model's grip on this world is fixed at its training cutoff. Ask it whether a specific disclosure obligation still applies and it will answer with total confidence, because it has no mechanism for knowing the rule was superseded eight weeks after the corpus was frozen. It is not wrong because it is stupid. It is wrong because nothing arrives. Its claim about the rule was never exposed to the rulemaking that overturned it, so agreement with the old text was never tested against the possibility of amendment — the amendment simply could not reach the model to disagree.

What is held

The severe-testing version of legal monitoring does not hold "the current rule." It holds a claim, its provenance, and a decay clock. For any compliance position — say, a data-retention practice justified under a specific state privacy statute — the system records which docket produced the current reading, which agency guidance it relied on, the date that guidance was issued, and whether it has been cited approvingly or distinguished in subsequent enforcement actions. The claim is not "we comply." It is "we comply, per interpretation X, sourced from filing Y, last checked against docket activity through date Z."

This matters because the characteristic failure in this domain is not ignorance of the rule. It is staleness dressed as knowledge. A posture built on a rule superseded two quarters ago looks, from inside the firm, identical to a posture built on current law. The memo reads the same. The training deck for staff reads the same. Nothing internal signals decay. Decay is only visible from outside, in the stream that superseded it — and only if something is still watching that stream.

What triggers revision

A severe test here is not "did a new document mention our topic." That test is too weak to fail; keyword matching on regulatory text returns noise at a volume no one reads, which is functionally the same as reading nothing. A real trigger has to be built to be capable of contradicting the specific claim on file, not just adjacent to it.

Concretely: the claim "our retention practice satisfies §X" is tested by watching for (a) amendments to §X itself, (b) enforcement actions citing §X against firms with comparable practices, (c) court decisions narrowing or widening the statute's scope, and (d) agency guidance documents that reinterpret §X without amending it — the quiet mechanism by which rules change without a formal rulemaking. Each of these is a distinct channel with a distinct latency. Statutory amendment is slow and visible. Guidance reinterpretation is fast and easy to miss, because it often arrives as an FAQ update or a speech, not a filing with a docket number.

A test with real severity treats the guidance-reinterpretation channel with more suspicion, not less, precisely because it is the channel that has previously undone compliance postures quietly. The system does not wait for agreement. It looks for the update that would kill the claim and checks that channel harder than the ones that only ever seem to confirm it.

Continuous total intake is mostly uncontrolled, confounded observation. A firm scanning every stream is not running an experiment. It is drowning in correlation and calling it vigilance.

This is fair, and it is close to the actual failure mode of naive monitoring tools that flag every mention of a keyword and produce three hundred alerts a week that no one reads past the subject line. Volume is not severity. What makes the arrangement severe is not that everything is watched, but that watching is structured around what would specifically overturn each held claim, and that some claims are tested by deliberate probing rather than passive scanning — a firm can, for instance, file a declaratory ruling request or a comment specifically designed to surface whether the agency's current reading matches the firm's assumption, rather than waiting for the ambiguity to resolve itself in an enforcement action against someone else.

What the operator sees

General counsel does not see a feed. A feed is where severity dies, because a feed rewards the appearance of coverage over the accuracy of any given claim. What general counsel sees is a claim ledger with three columns that matter: the position currently relied upon, the strength of the test it has survived, and the date the clock resets.

A position that has survived a hard test — say, an enforcement action in which the agency had the opportunity to challenge exactly this reading and did not — is marked accordingly, and marked differently from a position that has simply not yet been contradicted because no one has looked hard at the relevant docket in four months. These are not the same epistemic state, and collapsing them is the mechanism by which stale postures get carried forward as if they were current. "No news" is not evidence of stability. It is evidence that the test window closed without a result, and the ledger has to say so rather than defaulting to green.

When something in a watched stream does threaten a held position — a circuit split, a new consent order, a guidance FAQ that reads the statute more broadly than the firm's memo assumed — the operator sees the contradiction attached to the specific claim it threatens, with the provenance of both the old claim and the new evidence laid side by side. Not a summary. A comparison built to be adjudicated.

What it costs

This is not free, and pretending otherwise is where these systems earn their bad reputation. Structuring intake around falsification rather than keyword coverage requires someone — usually outside counsel or a specialist monitoring desk — to define, per major compliance position, what would count as contradicting it. That is slow, unglamorous work, and it does not scale by throwing more scraped text at the problem. A firm with two hundred active compliance positions across a dozen jurisdictions faces two hundred distinct falsification conditions, not one dashboard.

There is also a thrashing cost. Regulatory streams contain genuine noise: proposed rules that die in committee, guidance that gets withdrawn a week later, enforcement actions later vacated on appeal. A ledger that revises a compliance posture on every signal will whipsaw the firm into rebuilding processes for changes that never took effect.

A system that updates on every stream will chase noise and thrash between conclusions. Endless updating means there is never a settled background against which anything counts as an anomaly.

The answer is not less updating but better-anchored updating: revision is gated on provenance strength, so a withdrawn guidance document does not trigger a rebuild, but a final rule published in the Federal Register does. Stability comes from weighting evidence by what kind of thing it is — proposed versus final, dicta versus holding, speech versus rulemaking — not from refusing to look.

The failure a general counsel fears most is not a rule they never heard of; it is a rule they heard of once and stopped checking.

What this rules out

A benchmark of "compliance current as of last audit" is the legal equivalent of a frozen corpus, and audits, like Large Language Model training cutoffs, are tested severely by the auditors at the moment they happen and by no one afterward. That severity belongs to the audit firm's process, not to the compliance posture itself, which sits unexamined until the next audit cycle — often twelve months, sometimes longer, an eternity against a regulatory environment that issues guidance weekly.

A live dashboard tracking today's docket activity against today's posture, with no memory of yesterday's assumptions and no mechanism for checking whether last quarter's cleared position has since been undercut by a slow-moving amendment, has the opposite problem: sharp in the moment, blind to drift. What the domain actually needs, and what the claim ledger with provenance and decay is built to be, is neither snapshot: a standing arrangement in which every compliance position remains attached to the evidence that supports it and stays open to the specific document, filing, or ruling that would show it wrong — for as long as the position is relied upon, which in regulatory practice is usually longer than anyone remembers to check.

Continue