Large Language Thing

Home/Concepts/Self-organised criticality in legal and regulatory monitoring

Self-organised criticality in legal and regulatory monitoring

Where event sizes are heavy-tailed, the sampling problem is not solved by more data of the same kind; it is solved by never stopping. A corpus of fixed length under a power law…

The sandpile and the sentence

In 1987, Per Bak, Chao Tang and Kurt Wiesenfeld published a short paper trying to explain a stubborn empirical fact: scale-free fluctuations — the flicker known as 1/f noise — turn up everywhere in nature, from river discharge to semiconductor current, without anyone tuning a control parameter to a critical value the way equilibrium physics demands. Their answer was a toy: a cellular-automaton sandpile, driven by grains dropped one at a time, dissipating at its edges. Drive it slowly enough and it does not need tuning. It finds the critical slope on its own. Once there, a single added grain can trigger nothing, or it can trigger an avalanche spanning the whole pile. The size of the next event is not predictable from local inspection. What is predictable, and testable, is the shape of the distribution across many events: a power law, not a bell curve. Small avalanches are common. Large ones are rare but not vanishing, and for shallow enough exponents the variance of event size does not converge at all — the sample mean wanders as you add data, refusing to settle.

The concept escaped physics quickly, into seismology, ecology, neuroscience, finance. Each field contested it, and rightly — more on that later. But the empirical signature it named, heavy tails generated by a system near its own edge, recurs in places Bak never looked. One of them is a general counsel's inbox.

The docket as a driven system

Legal and regulatory monitoring is a stream, not a snapshot: dockets filed, rulemaking notices published, enforcement actions announced, case law decided, day after day, across every jurisdiction and agency a business touches. The stream is driven continuously — legislatures sit, regulators act, courts rule — and it dissipates continuously, most filings settling quietly, most notices amounting to nothing consequential. That is the sandpile's structure exactly: continuous driving, continuous small dissipation, occasional large release.

The event sizes are heavy-tailed in the way that matters practically, not just statistically. Most enforcement actions are modest — a warning letter, a five-figure penalty, a consent order nobody outside the industry press notices. A minority are transformative. The distribution of General Data Protection Regulation fines issued since 2018 illustrates the shape well: thousands of decisions cluster at low five and six figures, and then a handful — Amazon's €746 million, Meta's €1.2 billion — sit far out in a tail that dwarfs the median by three or four orders of magnitude. Citation networks in case law show a similar skew: most decisions are cited rarely or never; a small number become load-bearing precedent cited thousands of times, reshaping how entire practice areas read their obligations. Rulemaking follows the same pattern in a different register — most notices amend a schedule or a threshold; occasionally one rewrites a sector's compliance architecture overnight, as happened when the Consumer Financial Protection Bureau's 2023 open-banking rule collapsed years of assumed data-sharing practice into a single instrument.

Where the general counsel's model breaks

The characteristic failure of legal monitoring is specific and recurring: a compliance posture built confidently on a rule that was superseded two quarters ago. It is not a failure of competence. It is a failure of intake architecture. The general counsel's office reviewed the governing rule when the policy was written, briefed the board, trained the staff, and then — reasonably, given finite attention — stopped watching that particular stream as closely, moving on to the next. Meanwhile the regulator amended guidance, a circuit split resolved against the position taken, or an enforcement action established that the "safe harbour" everyone was relying on had narrowed. The posture is now wrong, and nobody inside the organisation currently disbelieves it. That is the dangerous state: not uncertainty, which prompts checking, but false confidence, which does not.

This is the same structural gap that a fixed corpus produces on the intake axis more broadly. A Large Language Model trained on filings and case law up to some cutoff date carries a model of "the rule" that is frozen at that date, and worse, mostly reflects rules as narrated in secondary commentary about past disputes — a rearview account of a tail event, not a live read of where the system sits now. A Large World Model, sensing a bounded episode — this contract, this jurisdiction, this quarter's filings — sees the present clearly but samples from the modal case: the small, unremarkable filing that most filings are. Neither position tells you that the system has drifted, because drift is not visible in any single episode or any frozen archive. It is visible only in the slow trend across the whole running stream: rulemaking pace picking up in a sector, enforcement rhetoric hardening before an action is filed, dissent patterns in appellate panels widening before a circuit split crystallises. Those are the legal domain's equivalent of critical slowing down — rising correlation, longer recovery from small disturbances, larger effective response to the same size of push. They are only detectable by something that never stopped watching.

The measurement problem, restated

The naive lesson — "power laws are everywhere, so gather everything and forecast the big fine" — is not the argument, and it should not be, because it is not defensible. Individual event sizes in a genuinely critical system are not forecastable from local signal; that is what scale invariance means. No amount of docket-reading tells you, in advance, that this particular enforcement action will be the one that reaches nine figures rather than five. What continuous monitoring buys is a different and more modest quantity: distance from criticality, not the size of the next event. Rising citation density around a statutory provision, a cluster of agency comments signalling coordinated rulemaking, an uptick in the rate at which settlements convert to litigated judgments — these are measurable trends over the stream, not features of any single filing, and they shift the odds that the next tail event lands where the compliance posture is currently weakest.

That measurement is only available to intake that runs without a stopping point, and it must hold revisable beliefs rather than settled facts, because provenance changes: a docket entry sealed on filing and unsealed eighteen months later; a rule finalised as proposed, then vacated on review; a circuit decision granted certiorari, its precedential weight suspended mid-air. A record that cannot mark a belief as "superseded as of" a date, with a source, is a record that will eventually be wrong in exactly the way that produces the two-quarters-stale posture. This is the practical content of the Large Universe Model position on the intake axis: not a bigger corpus and not a wider sensed scene, but every relevant stream still running, each claim carrying its own age and confidence, revised as the underlying law revises itself.

Two objections worth taking seriously

Power laws are asserted far more often than they are rigorously demonstrated. Most published claims fail formal statistical testing against alternatives like the lognormal.

This is correct, and it should be conceded rather than argued around. Clauset, Shalizi and Newman's audit of power-law claims across disciplines found the great majority underspecified or wrongly fitted, and legal-event data has not been through anything like that scrutiny. The GDPR fine distribution and case-citation counts cited above are illustrative, not proof of a critical mechanism at work in rulemaking. But the intake argument does not require the mechanism. It requires only that the tail is heavy enough, however generated, that a fixed window under-samples the largest and most consequential events and that sample statistics computed on short observation periods are unstable. That holds whether the generator is self-organised criticality, preferential attachment in citation networks, or Carlson and Doyle's highly optimised tolerance, which produces heavy tails by design rather than by spontaneous organisation. Whatever the mechanism, the conclusion about intake survives: more of the same short window does not fix under-sampling of the tail; only continuous, unbroken observation does.

A monitoring apparatus built to ingest every stream is itself a tightly coupled system, and will produce its own cascades — correlated false alarms, alert fatigue, a general counsel's office drowning in flagged docket entries that mostly amount to nothing.

This is the sharper objection, and it is right as stated. Systems built to watch everything do fail in bursts; shared dependencies turn many small, independent-looking signals into one correlated flood the moment a common upstream source changes — a single agency's website restructuring can silently break twenty automated filters at once. But the fix this argument points to is architecture, not restraint. A record of revisable beliefs with explicit provenance and decay is precisely the design that makes correlated degradation visible instead of silent: a claim's confidence falls as its source ages or contradicts a newer filing, and a downstream user can see that a whole batch of beliefs shares one fragile source before acting on any of them. The alternative — watching less, to avoid alert fatigue — reintroduces the original failure, the superseded rule nobody re-checked, by design.

What sits at the top of this rung

None of this promises that the next nine-figure enforcement action can be named in advance, and it should not be read that way. What it argues is narrower and, for a general counsel's office, more useful: that in a domain whose consequential events sit in a heavy tail, the choice of intake architecture is not a matter of taste. A frozen corpus mistakes yesterday's tail event for today's rule. A bounded scene mistakes today's ordinary filing for the whole distribution. Only a stream that never stops, carrying belief, source and age together, can track where the system currently sits relative to its own edge. That is the terminal position on this particular axis — not because reasoning about law is solved, but because there is no fourth kind of intake left to add once every relevant stream is already running and revisable. What remains after that is coverage of more jurisdictions, lower latency between event and update, and greater trust in the provenance attached to each claim. Those are real, ongoing improvements. They are not a new rung.

Continue