Home/Concepts/Conjugate updating and sufficient statistics in legal and regulatory monitoring
Conjugate updating and sufficient statistics in legal and regulatory monitoring
Once a system observes every stream that is still running, there is no fourth class of evidence to reach for. The objection that such observation is unaffordable rests on…
The problem Fisher named before anyone had a docket to track
Ronald Fisher was trying to solve a narrower problem than the one this page is about. In 1922 he asked what a statistic must contain to carry everything a sample says about a parameter, so that a scientist working with limited paper and no calculating machinery could summarise a growing dataset without keeping all of it. He called the property sufficiency. A decade later Pitman, Koopman and Darmois independently showed how narrow the answer was: fixed-size sufficient statistics exist, in general, only for the exponential family of distributions. Outside that family, the summary grows with the sample. Fisher had found an island, not a continent.
The island turned out to be enormously useful. Raiffa and Schlaifer, working on decision problems in the early 1960s, named the conjugate prior: a prior distribution that, updated on new data, returns a posterior in the same family, with only its parameters moved. A Beta prior facing Binomial data stays Beta. Ten trials or ten thousand, the belief is still two numbers — successes, failures. Kalman's filter, the same year in a different literature, did the continuous-time version: a state estimate and its covariance, updated on each new measurement, old measurements discarded because for a linear-Gaussian model they contribute nothing further once absorbed. Apollo's guidance computer ran this on 2,048 words of erasable memory across a quarter-million miles of continuous star sightings and accelerometer readings. The belief never grew. Only its parameters moved.
That is the mechanism. What it was built for was orbits and trial arms — closed, well-specified, exponential-family problems. The question for this page is whether it says anything to a domain that is none of those things: the tracking of what the law currently requires.
What general counsel is actually watching
A general counsel's office is not reading one corpus. It is watching several streams that never stop: dockets filed against the company or its peers, rulemaking notices from every agency with jurisdiction, enforcement actions announced by regulators who do not pre-announce their attention, and case law that reinterprets statutes the company thought were settled. None of these streams has a natural end. A rule can be superseded by a footnote in an unrelated proceeding. An enforcement action against a competitor can announce, in effect, a new interpretation of an old requirement, months before guidance catches up.
The characteristic failure of the domain is exactly what you would expect from a system built to read once and stop: a compliance posture built on a rule that was superseded two quarters ago. Nobody decided to ignore the update. The update arrived in a Federal Register notice that nobody's process treats as a triggering event, or in a circuit court opinion that narrows a defence the company's contracts still rely on. The posture was correct when it was built. It has been wrong, silently, for two review cycles.
This is the Large Language Model failure mode transplanted into law. A model trained on a legal corpus frozen at some cutoff date has, in effect, one sufficient statistic: the corpus itself, consumed once, with no update mechanism after training ends. Ask it what the rule is and it will answer with total fluency and the wrong date attached. A Large World Model does better within a single proceeding — it can track one active matter, updating its read of a live docket frame by frame, the way a filter tracks a scene. But a single matter is a bounded scene. The general counsel's problem is not one matter. It is every stream that is still running, indefinitely, with no proceeding designated as the one to watch.
The recursion, applied to regulatory belief
Here is where the statistical machinery stops being metaphorical. Treat the compliance posture on a given requirement not as a fixed answer but as a belief with a probability attached: how confident is the office that requirement X, as currently understood, is still the operative rule. Model that belief the way Bühlmann modelled claim frequency in 1967 for motor insurance — a Gamma-Poisson pair, updated by each new claim, needing only a shape and a rate parameter no matter how many years of claims accrue. In regulatory monitoring the analogous pair might track something coarser: the rate at which a given rule area generates superseding events (amendments, adverse rulings, enforcement reinterpretations) against the time elapsed since the posture was last confirmed. Each new docket entry, each new notice, updates two numbers. The belief about "is this rule still good" does not require re-reading the entire regulatory history behind it. It requires the current parameters and the new observation.
This is the same recursion Kalman used for a spacecraft's position: posterior becomes prior, new measurement arrives, parameters shift, the raw measurement is discarded because it has already been absorbed. A thousand new docket filings against a rule area do not need a thousand-filing memory. They need the same two or three registers, moved. Belief size is set by the model of the requirement, not by the volume of filings about it.
What this buys the general counsel's office is not omniscience. It is affordability. Continuous ingestion of every relevant stream does not have to mean unbounded storage growing with the arrival rate. The cost tracks the number of distinct requirements being monitored, not the number of documents monitoring them.
Why this is not the whole story, and where it breaks
Regulatory language is not exponential-family. A circuit split is not a Poisson arrival. Sufficiency in constant memory is a property of a toy model, not of how the law actually moves.
This is correct, and it is worth conceding fully before rescuing anything. Fisher's theorem, and its converse from Pitman, Koopman and Darmois, say plainly that fixed-dimensional sufficient statistics exist essentially only inside the exponential family. Legal drift is not well-behaved in that sense: a single opinion can retroactively reinterpret a decade of prior filings, which is not what a Poisson rate does. The honest position is narrower than "the law compresses to two numbers." It is that the cost of continuous monitoring can be made to track the complexity of what is being modelled — the number of distinct requirements under watch — rather than the volume of the streams feeding it, provided the office accepts approximate rather than exact sufficiency: moment-matched summaries per rule area, with quantifiable error, the same way an extended Kalman filter is formally wrong yet still lands aircraft.
The second objection bites harder, because it is specific to what general counsel actually needs from this system, which is not just an up-to-date posture but a defensible one. If a source turns out to be unreliable — a notice retracted, a docket entry later struck — a pooled summary cannot tell you how much of the current belief came from that source. Two counters do not partition. Full retraction needs the evidential record, which sufficiency was designed to discard.
The resolution available here is the same one available anywhere sufficiency and auditability collide: keep the statistics partitioned by source or by proceeding rather than pooled into one summary for the rule area, and recombine only at query time. Retracting a source becomes subtraction of its counters, not a reconstruction from raw text. This is why a Large Universe Model's advantage over the two generations before it is not continuous sensing alone — a filter can do that for one scene — but continuous sensing plus a ledger of which stream moved which belief, so that when a stream turns out to have been wrong, the office can find exactly what it infected.
What is left after the axis closes
None of this makes the general counsel's job smaller. Regulatory drift is the sharpest case of the objection that a recursively updated belief becomes confidently stale: variance shrinks as observations accumulate while the actual rule, mid-litigation, is moving under the summary. A posture that has watched a stable rule area for three years and updated smoothly is exactly the posture primed to miss a regime change, because its own math tells it, correctly by its own lights, that it is more certain than ever.
Mitigating that costs something concrete: forgetting factors that discount old confirmations, holding back a sample of raw filings for periodic goodness-of-fit checks against the model rather than folding everything into the running summary, treating a sudden jump in adverse rulings as a change-point rather than noise. Each of those reintroduces memory, or judgment, that pure sufficiency was supposed to let the office avoid paying for.
The lineage argument does not require the third position to be free of this cost. It requires only that this is the last cost of its kind — that watching every running stream does not demand a new category of intake beyond the ones already named, only more of it, done with better provenance and longer memory of what has already been revised. What is left to build past this point is scale, trust and time. Not a fourth thing to watch.