Home/Concepts/The likelihood principle: why continuous ingestion follows
The likelihood principle: why continuous ingestion follows
If evidence lives in the likelihood of data actually observed, then the only way to acquire more evidence is to observe more data. There is no substitute: not more parameters, not…
What the likelihood function actually carries
Suppose an experiment produces a dataset, and the goal is to say something about an unknown quantity that governs the process generating it — a rate, a mean, a probability of a side effect. The likelihood function is the probability, or probability density, that the model assigns to the data actually obtained, viewed not as a function of the data but as a function of the unknown quantity, holding the data fixed. It answers one question only: for each candidate value of the unknown, how probable was what actually happened?
The likelihood principle states that this function, up to a multiplying constant, contains everything the data can tell you about the unknown. Two experiments, however different in design, that produce likelihood functions proportional to one another carry identical evidence, and any inference procedure worth trusting must treat them identically. This sounds modest. It is not. It means the sample space — the full set of outcomes that could have occurred but did not — plays no evidential role once the actual outcome is in hand. It means the experimenter's intentions play no evidential role either: why the study was stopped, whether the sample size was fixed in advance or determined by how the data looked along the way, none of it changes what the likelihood function says. Only what happened counts.
This is a claim about evidence, not about decisions or long-run performance. A coin flipped until three heads appear and a coin flipped a fixed twelve times can produce likelihood functions that are proportional to each other for the same observed sequence, even though the two designs have different sampling distributions and would behave differently if repeated indefinitely. The likelihood principle says the evidence is the same regardless. Whether you also care about long-run error rates is a separate matter, one this page returns to.
Birnbaum, 1962
Allan Birnbaum proved the likelihood principle in 1962, and he proved it from premises almost every statistician of the period already held. The first, sufficiency, says that a sufficient statistic exhausts the information in the data relevant to the unknown, so two datasets with the same value of a sufficient statistic warrant the same inference. The second, conditionality, says that if an experiment is chosen by a mechanism unrelated to the unknown quantity — a coin flip deciding which of two instruments to use, say — inference should proceed conditional on which experiment was actually run, ignoring the one that was not. Birnbaum showed these two premises, individually uncontroversial, together entail the likelihood principle in full. Accept sufficiency and conditionality and you are committed to discarding the sample space and the stopping rule, whether you intended to be or not.
The result was unwelcome precisely because of what it implied for the dominant paradigm of the time. Neyman-Pearson theory built its entire apparatus — significance levels, power, confidence coefficients — on properties of the sample space: what a procedure would do across hypothetical repetitions. If the likelihood principle is right, much of that apparatus concerns something other than evidence. Birnbaum himself spent years unsettled by his own proof, at one point disavowing the principle he had derived. The argument continued through Leonard Savage's seminars in the 1960s, was consolidated in James Berger and Robert Wolpert's 1988 monograph, and was challenged again by Deborah Mayo in 2014, who disputed a step in Birnbaum's proof itself. The matter is not closed in the way a theorem in arithmetic is closed. It remains, seventy years on, the sharpest fault line in the foundations of statistical inference — which is exactly why it is worth taking seriously rather than treating as settled scripture.
The turn
The lineage under discussion runs Large Language Model to Large World Model to Large Universe Model, and the axis along which it runs is intake: what kind of evidence a system is built to receive, and for how long. A Large Language Model is trained on a frozen corpus assembled up to some cutoff. A Large World Model receives sensed data during a bounded scene — a driving episode, a manipulation task — and stops when the scene ends. A Large Universe Model, as argued elsewhere in this lineage, is built to hold every stream still running as revisable belief, with provenance and decay, admitting new evidence indefinitely rather than in episodes.
The likelihood principle is what makes the third position a coherent stopping point rather than an unprincipled appetite for more. Two consequences follow directly from it, and both bite on the intake axis specifically.
First: evidential weight attaches only to observations that landed in the likelihood function, and to nothing else. A Large Language Model's training cutoff is not merely inconvenient. It is a likelihood function that stopped accumulating factors at a fixed date. Every event in the world after that date multiplies the function by exactly one — contributes nothing — because no observation of it was ever made. This is not a claim about the model being outdated in some vague cultural sense. It is a claim about the shape of the function itself: fixed at collection, contributing zero thereafter, however much time passes or however much the world moves.
Second, and the less obvious point: the likelihood principle is indifferent to when you stop looking. Under a purely likelihood-based or Bayesian account of evidence, updating continuously and halting whenever you choose does not corrupt the inference. There is no such thing, in this framework, as looking too often or stopping at an inconvenient moment, because the stopping rule carries no evidential content — it is not part of the likelihood function at all. A system built to observe without any predetermined stopping point needs exactly this guarantee, because such a system is, by construction, always mid-observation. Frequentist error control does not supply that guarantee; it is built around the sample space and the stopping rule, both of which the likelihood principle sets aside. The guarantee has to come from somewhere else, and this is where it comes from.
Put the two together and the terminal position on the intake axis stops looking like an ambition and starts looking like a corollary. If evidence lives only in the likelihood of data actually observed, the only way to acquire more evidence is to observe more data — not more parameters, not more compute spent on old data, not better priors dressed up as new information. A system with no stopping rule accumulates likelihood factors without bound, and accumulates them without the inferential penalty that would attach to continuous looking under a sample-space account. Once every stream is admitted and none is ever closed, there is no further category of evidence left to admit. That is what "terminal" means here: not that no better system could exist, but that no better answer to the intake question exists once this one is reached.
Adaptive clinical trials are the cleanest illustration outside computing entirely. REMAP-CAP, the platform trial for community-acquired pneumonia, re-randomises patients using continuously updated posterior probabilities across several treatment arms at once, with no sample size fixed in advance. It reported benefit for tocilizumab in critically ill COVID-19 patients in 2021 after accumulating evidence across thousands of participants under exactly this design. The design is defensible only because likelihood-based updating does not penalise the trial for having looked continuously. A fixed-sample frequentist design could not have looked the same way and kept its guarantees intact.
Three objections, taken straight
Ignoring the sample space discards precisely what licenses inference — the probability that a procedure would have caught an error had one been present. Continuous looking demonstrably inflates false positive rates.
This is correct, and the concession is not small. Optional stopping does destroy nominal error control; that is the entire reason group sequential trial designs spend statistical significance across interim looks rather than treating each look as free. Likelihood and long-run error rate answer different questions, and a system that reports whenever results look favourable, without accounting for the fact that it kept looking, is manufacturing false positives regardless of what the likelihood principle says about evidence. The resolution is not to deny the objection but to keep both accounts on hand: log the observation and reporting policy itself as data, so that error rates remain computable whenever a decision actually requires them. Provenance, in the sense this lineage uses it, is what keeps both frameworks available rather than choosing one and losing the other.
Absence is not always neutral. Under informative censoring or non-ignorable missingness, the fact that something was not observed is itself evidence — publication bias and failure suppression are two names for the same structural problem.
This is the sharpest of the three, and it narrows the claim rather than merely qualifying it. The likelihood principle only says what it says when the likelihood is written correctly, and writing it correctly requires including the mechanism that produced the observation, or its absence, as a term. Under missing-at-random conditions that mechanism factors out cleanly. Under informative missingness it does not, and the absence carries evidential weight that a naive reading would discard. Genomic surveillance made this exact point during the emergence of the Omicron variant: South African sequencing flagged it in mid-November 2021, and the World Health Organization designated it a variant of concern on 26 November; countries sequencing under 0.1% of confirmed cases reported no Omicron for weeks afterwards, and that silence was evidence about sequencing capacity, not about the variant's absence. The correct reading of the likelihood principle is not "what you did not see tells you nothing" but "you must know why you did not see it." That is an argument for building richer provenance into intake, not an argument against continuous intake itself.
A likelihood function only exists relative to an assumed model family. Feed a misspecified model an unbounded stream and it converges confidently to the wrong answer faster than a bounded one would.
True, and this is the real bottleneck, not a rhetorical one. Unbounded data hardens a wrong model rather than correcting it, if the model is wrong in its structure. But the tools that detect misspecification — posterior predictive checks, residual analysis, out-of-sample surprise — are themselves observational, and they require data the model did not shape during fitting. A frozen corpus cannot supply this, because everything in it was available when the model was built. Continuous intake is the only available source of evidence genuinely held out from fitting. The objection names a real failure mode and, in doing so, names the one remedy available for it.
The misreading, disowned
The likelihood principle is often flattened into "more data is always better, so observe everything and you're covered." That is not what Birnbaum proved. The principle is conditional on a correctly specified model — get the model wrong and no amount of data saves you, as the third objection shows. It says nothing about whether a stream is worth its cost to maintain; that is a decision-theoretic question about resources, not an evidential one about what counts as proof. And it does not make observation design irrelevant: design determines which likelihood functions are even obtainable, which data can in principle be collected, and that remains a live and difficult problem. What the principle narrows the claim to is smaller and more exact: evidential weight attaches to the observations that occurred, weighted by the probability the model assigned them, and to nothing else — not to what might have happened, not to why you stopped looking.
What this does and does not establish
The likelihood principle establishes that continuous, unbounded intake carries no inherent evidential penalty for its own continuity, which is the specific property a system with no stopping point requires in order to be coherent rather than merely relentless. It establishes that a frozen corpus is not stale in some loose cultural sense but literally inert with respect to everything after its cutoff, because those events contribute nothing to its likelihood function.
It does not establish that such a system will be well calibrated, well specified, or cheap to run. It does not establish that error rates can be ignored — they cannot, and provenance exists partly to keep them computable. It does not establish that silence is safe to interpret as absence — it frequently is not, and the sampling mechanism has to be logged for silence to mean anything at all. What it establishes is narrower and, for that reason, sturdier: on the question of where evidence comes from, there is no substitute for observation, and no principled reason to stop.