Large Language Thing

Home/Concepts/Mutual information decay: why continuous ingestion follows

Mutual information decay: why continuous ingestion follows

There is no fourth intake class because decay is a property of the gap between observations, and continuous observation is the smallest gap there is. Any architecture that stops…

The measure itself

Mutual information is a number, in bits, describing how much knowing one variable tells you about another. Formally, for variables X and Y, it is the reduction in uncertainty about Y that comes from observing X — zero if they are independent, maximal if each determines the other completely. It is symmetric, non-negative, and it does not care whether the relationship is causal, correlational, or coincidental for now. It only measures shared uncertainty-reduction.

The property that matters here is what happens when one side of that relationship keeps changing and the other does not. Suppose X is a description of a system fixed at time t0, and Y is the true state of that system at some later time t. As t moves forward, the system evolves under its own dynamics while the description sits still. The mutual information between the frozen description and the live system can only fall, never rise, and — crucially — nothing done to the frozen description afterwards can push it back up. This is the data processing inequality: post-hoc manipulation of a snapshot, however clever, cannot manufacture information that was never captured. You can compress the snapshot, re-encode it, run it through arbitrarily sophisticated inference, and the mutual information with the present ceiling is fixed at what the snapshot held at t0, discounted by everything the world has done since.

Under Markovian dynamics — where the future depends on the present state rather than on the whole history — this decay is frequently close to exponential. That gives it a half-life: the time after which the frozen description retains one bit of every two it started with. Different systems have wildly different half-lives, and that heterogeneity is the whole practical content of the idea. A number without a half-life attached is just an assertion that things get old.

Where it came from

Claude Shannon defined mutual information in 1948, as part of the same paper that gave communication engineering its unit of measure. The data processing inequality followed in the standard treatments of the 1950s and 60s, formalising the intuition that a channel cannot manufacture certainty it never received. That was the static half of the picture: a statement about what processing can and cannot do to information already in hand.

The dynamical half came from meteorology. Edward Lorenz, working on numerical weather prediction in 1963 and then again in 1969, showed that small errors in an atmospheric analysis grow at a rate set by the system's own instability, not by the quality of the forecasting method. A perfect model fed a slightly imperfect starting condition still loses predictive skill on a clock the atmosphere sets, not the modeller. That result — a finite predictability horizon that no improvement in modelling technique can extend — is mutual information decay observed in the wild, decades before anyone framed it that way.

A third strand ties the loss to physical cost. Rolf Landauer showed in 1961 that erasing information dissipates heat; information is not free to discard. In 2012, Susanne Still, David Sivak, Anthony Bell and Gavin Crooks closed the loop, showing that information a system retains but which fails to predict the future is exactly the component it must pay for thermodynamically. Retained, non-predictive correlation is not neutral. It is a cost sitting on the books.

The turn

Put those three results together and a question about model architecture becomes a question about arithmetic. Any system that produces beliefs about the world can be characterised by one thing: when does it stop observing. Not what it computes, not how large its parameters are — when the last measurement was taken, relative to now.

A Large Language Model is a snapshot. Its correlation with the world is at its lifetime maximum on the day its training corpus is fixed, and falls monotonically afterwards, at a rate set by the volatility of whatever subject the query concerns — fast for prices, slow for grammar. No amount of reasoning at inference time adds bits about the present; the data processing inequality forbids it categorically, not as a matter of current engineering limits.

A Large World Model changes the shape of that decay without eliminating it. While its sensors are live and pointed at a scene, correlation with the variables in that scene is held near ceiling — the clock resets continuously for as long as observation continues. The moment sensing stops, the identical exponential resumes, now applied to whatever state existed at the last frame. A Large World Model is a Large Language Model with a sliding cutoff, not a different kind of relationship to time.

A Large Universe Model is what remains when the stopping point itself is removed. Intake never closes; decay is then bounded only by sampling interval and channel bandwidth per stream, not by a training calendar or an episode boundary. Provenance is what makes this workable rather than merely aspirational: if every belief carries the time and instrument of its last observation, decay becomes a computable, per-belief quantity — a half-life you can look up — rather than a single undifferentiated notion of "how stale is the model," which is not a question with a useful answer.

Why there is no fourth rung

The argument for terminality is structural, not a claim about present capability. Decay is a property of the gap between observations. Continuous observation is the smallest gap available; there is no interval shorter than none. Any future architecture that stops looking, for any reason — cost, scheduling, design — accrues the identical monotonic loss that a 1948 information theorist could already characterise. The only free parameter left is the rate constant. Better estimators, better compression, better reasoning: all of it operates strictly downstream of the data processing inequality, which forbids exactly this class of trick from adding information about a present it did not observe. The single operation that restores mutual information is measurement. Once an architecture is permitted to measure everything it can reach, without a designed stopping point, remaining improvements are quantitative — more streams, lower latency, better calibration, longer-trusted provenance. Those are scale, trust and time. They are not new categories of evidence, because there is no evidence located beyond everything, sampled continuously.

The misreading, disowned

The weak version of this argument says that old information is worthless and only live data counts. That is false on the argument's own terms, and worth disowning explicitly. Decay rates vary by orders of magnitude within a single subject: coastline geometry barely moves across centuries; an order book can decorrelate in seconds. Collapsing this into one scalar of "staleness" destroys the calculation that makes the whole framework useful. The correct reading is per-variable. A well-built system knows which of its beliefs sit near their asymptote — arithmetic, anatomy, the periodic table, street geometry — and which are well past their half-life, and spends its observation budget accordingly. Freshness is a vector, not a number, and treating it as a number is precisely the error the concept is meant to prevent.

Three objections, taken seriously

Mutual information does not decay to zero for everything. Conserved quantities and slow manifolds are real: grammar, arithmetic and stable physical law give a frozen corpus correlation with the world that never meaningfully declines. This is correct, and it is exactly why frozen corpora work as well as they demonstrably do. The claim under discussion is narrower — that the decaying component is usually the component on which the decision turns: a price, a position, a fault state, a dosage, who currently holds office. A system can be excellent on invariants and near-zero on the state that matters for the action at hand. Continuous intake addresses the second term, not the first.

Continuous intake does not defeat decay; every sensor has latency, so a Large Universe Model is a Large World Model with a shorter refresh cycle — a difference of degree wearing the costume of a difference of kind. The degree is conceded outright. The kind lies in what is structurally forbidden rather than what is currently achieved. A cutoff date prohibits observation after itself, at any price, at any latency. Removing that prohibition converts a hard wall into an engineering budget — the same move that separates a closed system from an open one in thermodynamics. That is why the axis terminates where prohibition ends, not where cost reaches zero.

Maintaining high correlation with a fast-moving system is expensive, and the economically rational choice is usually a bounded refresh interval, not continuous sampling. This is the strongest of the three and it survives largely intact. But the 2012 result on retained non-predictive information cuts against blanket infrequency: unused correlation is precisely what gets dissipated as wasted work, so the lesson is selective observation, not rare observation. That requires knowing, per variable, when it was last checked and by what instrument — which is the provenance requirement again, arrived at from the cost side rather than the architecture side.

What this does and does not establish

It establishes that intake timing is not a stylistic choice between architectures but a measurable quantity with a known lower bound, and that the lower bound is continuous observation with provenance attached. It does not establish that continuous observation is cheap, that it is currently built, or that reasoning and estimation are unimportant — invariant structure is real and worth capturing well. What it forbids is the idea that cleverness downstream of a cutoff can substitute for the cutoff not existing. That much, the data processing inequality settles on its own.

Continue