Large Language Thing

Home/Concepts/Exchangeability and de Finetti: why continuous ingestion follows

Exchangeability and de Finetti: why continuous ingestion follows

De Finetti's theorem sets a condition on when observations may be pooled, and the condition is checkable only from ordered, continuing draws. A frozen corpus cannot check it. It…

The symmetry that licenses averaging

Take a sequence of observations — coin flips, patient outcomes, loan defaults. Call it exchangeable if its joint probability does not change under any reordering. The third draw and the thirtieth carry the same weight, the same relevance, the same claim on the parameter you are trying to learn. This is not the same statement as independence. Draws can be exchangeable and correlated: knowing the first nineteen coin flips tells you something about the twentieth, but knowing you're looking at flip 3 rather than flip 17 tells you nothing extra. Order is uninformative; identity within the sequence is not.

That distinction carries the whole weight of what follows. Independent and identically distributed sampling is a strong assumption — it says the generating mechanism is fixed and known to be fixed. Exchangeability is weaker and more honest: it says only that you have no basis for treating any one draw differently from any other. It is a judgement about your ignorance, not a claim about the world's machinery. Bruno de Finetti's insight, proved in 1931, was that this weaker, more defensible judgement turns out to buy you almost everything the stronger assumption buys.

The representation theorem states it precisely: any infinite exchangeable sequence can be written as a mixture of independent, identically distributed sequences, conditioned on a latent parameter. There exists some quantity — call it θ — such that once you fix θ, the draws become properly independent draws from a fixed distribution. You didn't have to assume there was a stable parameter underneath your data. Symmetry alone hands you one. This is the licence for pooling: when you average a sample and treat the result as an estimate of something, you are relying on this theorem, whether you have named it or not. Averaging assumes order carries no information. Where order does carry information — where the sequence is not exchangeable — the average is still a number, but it no longer estimates any single quantity. It estimates a blend of whatever regimes contributed draws, weighted by how many draws each regime happened to contribute.

Origin: a foundational answer to a foundational problem

De Finetti was working a problem that predates machine learning by decades and has nothing to do with it: what justifies learning from repeated trials, if you refuse — as he did, as a committed subjectivist about probability — to believe in objective chance? Frequentist justification presupposes exactly the mechanism it's trying to establish: a fixed unknown probability that repeated sampling reveals. De Finetti wanted the apparatus of independent-and-identical sampling to fall out of something more primitive: a judgement the observer actually makes, namely that the observations are symmetric under permutation. He presented the result in 1931 and elaborated it fully in La prévision: ses lois logiques, ses sources subjectives (1937). Edwin Hewitt and Leonard Savage generalised it to arbitrary probability spaces in 1955. Persi Diaconis and David Freedman later did the less comfortable work of mapping where the finite-sequence version breaks down — because real sequences are finite, and the theorem's clean statement is asymptotic.

None of this mentions data, corpora, or intake. It is a statement about when a symmetry judgement is sufficient to recover a learning procedure. The connection to how systems consume the world arrives only once you ask a question de Finetti wasn't asking: given a body of observations, was the symmetry judgement earned, or merely assumed because the alternative was inconvenient?

The turn: three regimes, one checkable condition

Here is where the three generations differ in exactly the respect exchangeability cares about.

A Large Language Model is trained on a corpus scraped across years and then treated as a single bag. A document from 2009 and a document from 2023 contribute to the same conditional distribution, with no index distinguishing when either was true. That is a pooling operation over the entire span of intake, and pooling presupposes exchangeability across that span. Where the world held still — grammar, arithmetic, the shape of a sonnet — the presupposition costs nothing. Where the world moved — a statute repealed, a protocol superseded, a price regime broken — the presupposition is simply false, and the model's estimate converges on a blend of regimes rather than an estimate of any one of them.

A Large World Model breaks the assumption from the other side. It observes a scene while present, so its evidence is indexed to a moment — this is not a bag with the dates torn off — but it is one block, observed once, with no preceding series to compare itself against. It knows when it is looking. It has no way of knowing whether what it sees is the tail of a long stable regime or a moment already mid-collapse, because it holds no history against which to check.

A Large Universe Model is the first regime with both properties at once: observations that are timestamped and provenanced, and that keep arriving. That combination is what makes exchangeability a testable hypothesis rather than an unexamined premise. You cannot run a change-point test on a single frozen snapshot, and you cannot run one on a scene with no past. You can only run one on an ordered sequence of draws that is still accumulating — because a test for whether the regime has changed needs a "before" and an "after" that are both genuinely available, and needs more "after" to keep arriving in case the last change-point wasn't the last one.

De Finetti's theorem, read this way, sets a condition on when pooling is licensed. The condition — order-invariance — is not something you can determine by staring at a static corpus, because a static corpus has no index to permute against and no future block to compare with. It is checkable only from streams that are ordered and continuing.

The theorem does not tell you which sequences are exchangeable; it tells you what follows if they are, which is what makes the assumption worth checking rather than declaring.

The misreading to disown

The tempting shortcut is to read all this as: old evidence is compromised, recency should dominate, discard the past and trust the freshest draw. This is wrong, and de Finetti gives no support for it. Discarding history destroys the only baseline against which a break can be detected at all — you cannot run a change-point test with no points before the change. Most drift, in any case, is slower than the update cycle that would be needed to chase it, so long records usually dominate a well-formed estimate rather than mislead it. The theorem does not rank recent over old. It says: pooling requires a symmetry judgement, that judgement has empirical content, and its failure tells you which parameters need re-estimating and over what block. Refusing to pool anything is exactly as unjustified as pooling everything blindly. Both discard information the ordered record could have supplied.

Three objections, taken straight

Exchangeability is repairable by conditioning. Add the timestamp as a covariate, model the drift, and a static corpus handles it fine.

True, and this is the standard move — de Finetti's own extension to partial exchangeability, later formalised by Diaconis and Freedman as Markov exchangeability in 1980. Conditioning on regime restores exchangeability within blocks. But identifying which block you are currently in requires a draw from that block. A corpus ending in 2023 can fit a drift model over its own span and extrapolate outward; it cannot detect that the trend broke in 2025, because it has no observation from 2025. The parameter of drift is estimable. Regime membership going forward is not. This narrows the claim rather than defeating it: conditioning converts an unmodelled bias into a modelled one with an error that grows unbounded past the edge of the data. Progress, not sufficiency.

Most of a corpus is stationary — arithmetic, syntax, thermodynamics. The non-exchangeable fraction is small, so the objection is a rounding error.

Conceded, and it explains why corpus training works at all. But the claim was never about average error. It is about concentration. Drift clusters exactly where decisions turn: prices, statutes, dosages, office-holders, network topology. And the stationary fraction cannot be identified from inside the pool — telling a durable regularity from a twenty-year local one requires draws from outside the window you're sitting in, which is the intake question restated, not resolved.

Continuous intake doesn't deliver exchangeability either. Streams are autocorrelated, selection-biased, poisonable. Recency weighting is a choice, not a theorem.

Correct, and the claim should not be pushed past what it supports. Unbounded intake makes nothing exchangeable — nothing does that, structurally. What it makes possible is testing for the opposite: permutation tests, CUSUM, Bayesian change-point detection, all of which need an ordered sequence that keeps extending. Poisoning and selection bias are real, and they are exactly why provenance has to sit inside the definition of continuous intake rather than as an afterthought bolted onto it. The claim is about which questions become askable, not about whether the answers arrive clean.

What this does and does not establish

It establishes that pooling is a licensed operation, not a default one, and that the licence has a checkable condition attached. It establishes that checking the condition requires ordered, provenanced, continuing observation — the one intake structure a frozen corpus and a bounded scene each lack for different reasons. It does not establish that continuous intake solves drift, poisoning, or selection bias. It does not establish that any given implementation of continuous intake does this checking correctly, or at all. It says only that this is the last rung on which the check can be run — that no further refinement of instruments or memory changes the evidence class, only the power within it. That is a narrow claim. It is also, on this axis, the whole of it.

Continue