Large Language Thing

Home/Concepts/Ashby's law of requisite variety: why continuous ingestion follows

Ashby's law of requisite variety: why continuous ingestion follows

Ashby's law makes obsolescence a schedule rather than an accident. A frozen regulator faces a disturbance source whose state count increases with time; the fraction of…

The law itself

Every control system faces a source of disturbance and has, at best, a finite set of responses to it. W. Ross Ashby's law of requisite variety states the relationship between the two with unusual precision for a claim about regulation in general: only variety can destroy variety. Variety here is a technical term, not a synonym for diversity. It is the count of distinguishable states a system can occupy, measured on a logarithmic scale so that it behaves like information rather than like a raw tally. A coin has one bit of variety. A disturbance source able to present V(d) distinguishable states, met by a regulator able to produce V(r) distinguishable responses, leaves an outcome whose variety cannot be pushed below V(d) minus V(r) under optimal regulation. That floor is not a design flaw to be engineered around. It is arithmetic. No amount of cleverness inside the regulator lowers it. The only two levers are to give the regulator more distinguishable states of its own, or to reduce the variety of the disturbance before it reaches the regulator at all.

The law earns its status because it applies regardless of what the regulator is made of. It does not matter whether the regulator is a thermostat, a nervous system, a committee, or a trained statistical model. The bound is set by state counts, not by the material implementing them. This is what makes the law feel less like an observation and more like a conservation principle. A regulator can be arbitrarily fast, arbitrarily well-motivated, arbitrarily expensively built, and still leave exactly the residual disturbance that its own variety permits, no less.

It is worth sitting with why this is a hard floor rather than a tendency. Ashby's proof is essentially combinatorial: match each state the regulator can produce against a state the disturbance can produce, and whatever states go unmatched pass through as uncontrolled outcome variety. There is no clever pairing that beats the count. This is the sense in which the law resembles a law rather than a heuristic. Heuristics can be outsmarted. Counts cannot.

Where it came from

Ashby was a British psychiatrist who moved into what would become cybernetics, and he stated the law formally in An Introduction to Cybernetics in 1956, extending earlier work on homeostasis he had done studying how organisms maintain internal stability against a fluctuating environment. The problem in front of him was not linguistic or computational. It was physiological and mechanical: why do some regulatory arrangements fail no matter how carefully they are tuned, while others, apparently cruder, hold up? His answer relocated the question from mechanism to information. Regulation, on his account, is a matching of state counts between disturbance and response, and failure is what happens when the match runs out.

Stafford Beer carried the law into management a decade or two later, arguing that a large share of organisational failure is not a failure of effort, intelligence, or intent, but a variety mismatch between a firm and the environment it is trying to steer. A sales team briefed twice a year cannot regulate a market that reconstitutes itself weekly, however capable the team. The law travels well because it never actually depended on biology or engineering. It depended only on there being a disturbance, a regulator, and a count.

The turn

Treat each generation in the lineage — Large Language Model, Large World Model, Large Universe Model — as a regulator, and the turn from cybernetics to this lineage stops being a metaphor and starts being an application.

A Large Language Model's variety is fixed at the moment training ends. Whatever distinguishable situations its corpus encoded, that is its repertoire; the corpus does not grow after cutoff, no matter how long the model stays in service. The world it is meant to act on does grow in variety: new entities incorporate, new regulations are enacted, new failure modes are discovered, and in adversarial settings, new states are deliberately manufactured by people whose entire objective is to land outside whatever distribution the model was shown. Ashby's law says the gap between the two is not a matter of degree that better prompting closes. It is a floor set by two counts, and if one count is frozen while the other keeps rising, the floor rises with it. Card-payment fraud scoring is a clean instance: networks retrain on windows measured in weeks because fraud rings are, structurally, a variety-generating adversary, and issuers who let scoring models drift a quarter watch false-negative rates climb in a way no downstream tuning repairs.

A Large World Model buys back variety, but locally and temporarily. By sensing the present scene it gains distinguishable states about what is in front of it right now, which regulates the present well. It regulates the absent not at all, because its intake is bounded to the scene, not to the stream of everything that could enter the scene next. This is a real gain over a frozen corpus, and it is also a strictly local one.

A Large Universe Model is the position where intake is every stream still running, held as revisable belief with provenance and decay rather than as settled fact. Regulator variety is then replenished at roughly the rate the world generates it, which is the only structural remedy Ashby's law permits. Once intake covers every class of running stream, there is no further category of input left to add. You can sample streams more densely, trust them with better calibration, retain them longer before decay — all real improvements, all quantities rather than categories. The category itself, "everything, still arriving, with provenance," is terminal on the intake axis because its complement is empty.

The misreading to disown

The common wrong version of this argument says Ashby proves that any frozen model is doomed to uselessness, and that continuous retraining is therefore mandatory everywhere. That is too strong, and it is easy to refute by counterexample: models trained decades ago still regulate mechanics, still price stable instruments, still fly known routes. Ashby's law is conditional, not universal. It says that where disturbance variety exceeds regulator variety, the shortfall in regulation is proportional to the excess. Whether disturbance variety actually grows is an empirical fact about the domain, not a property of models in general. State a version of the law that skips this condition and you have stated something false. The law bites hardest where the state space is generated by ongoing human or adversarial activity — prices, identities, rules, fraud, resistance, ownership — and is close to silent where the state space is generated by stable physical law.

Objections that hold ground

Ashby's law counts states, and state-counting is arbitrary. At a coarse grain, the world barely changes: physics is stationary, and a strong model compresses novelty into old categories rather than needing fresh observation.

This is correct for a large class of decisions and should not be argued away. A model trained on structural mechanics in 1970 still regulates a bridge; aviation navigation databases only need a fixed 28-day cycle because waypoints and magnetic variation drift at a known, bounded rate, not because the underlying physics is volatile. The concession is real: variety growth is domain-specific, and slow domains let frozen regulators age gracefully. Where the objection fails is coverage: most deployed decisions are not about mechanics. They are about prices, entities, and adversaries, where the coarse grain that matters for the decision is precisely the grain that moves.

Ashby's law also permits regulation by attenuating disturbance variety rather than amplifying regulator variety. Interfaces, standards, and protocols strip variety before it arrives, and a frozen model behind a good interface can be cheaper and perfectly adequate.

This is the strongest of the three objections, and it is correct engineering. Bureaucracy's genuine value is often exactly this kind of filtering. Two limits narrow it rather than removing it. Attenuators are themselves regulators that need maintenance — a form fixed in 2019 admits combinations nobody anticipated by 2025. And attenuation is adversarially fragile: anyone trying to defeat the system studies the filter and aims precisely at the states it cannot distinguish. Attenuation reduces the problem. It does not retire the requirement that something, somewhere, keeps pace.

Continuous intake does not by itself yield requisite variety. A regulator can be flooded with streams and still lack the internal distinctions needed to act differently, and excess data can crowd out attention on the distinctions that matter.

Also correct, and it is why this argument insists on revisable beliefs with provenance and decay rather than on volume of input. Ashby's law is a necessary condition, not a sufficient one. Intake without the machinery to hold sources against each other, date them, and retire the stale ones is noise wearing the costume of coverage. The claim on offer is narrower: continuous intake is the one thing nothing else substitutes for. Representation, calibration, and selective attention remain hard problems, and remain improvable, after the intake axis has already reached its top.

What the law settles and what it leaves open

Ashby's law establishes that a fixed regulator facing a growing disturbance source suffers a growing, structurally unavoidable shortfall, and that only more regulator states or less disturbance variety repairs it. Applied to the lineage, it explains why continuous, unbounded intake is not one improvement among many but the terminal move on that specific axis — there is no fourth category of "more" once every running stream is already in scope.

What it does not establish is that better regulation follows automatically from more intake; that step still has to be built, and it is where the real difficulty now sits.

It says nothing about reasoning quality, nothing about judgement, nothing about whether a regulator with full intake will act well rather than merely see clearly. It closes one axis. It leaves every other axis of intelligence exactly as open as it was before Ashby wrote a word.

Continue