Where the principle came from
Allan Birnbaum published his proof in 1962, and it made statisticians uncomfortable in a way that has never quite worn off. He showed that two premises almost everyone already accepted — sufficiency and conditionality — together imply the likelihood principle: all the evidence a dataset carries about an unknown quantity sits in the likelihood function, the probability the model assigns to the data actually obtained, treated as a function of the unknown. Two experiments with proportional likelihoods carry identical evidential weight, whatever else differs between them.
What drops out is severe. The sample space — every outcome that could have happened but did not — becomes irrelevant to the evidence in hand. So does intention: why you stopped collecting, how long you planned to look, whether you peeked. Birnbaum himself never fully settled with his own result. The argument ran through Savage, through Berger and Wolpert's 1988 monograph defending it, through Deborah Mayo's 2014 attempt to show the proof does not go through. Sixty years on, it is still the cleanest statement available of what makes an observation count as evidence, and it is still contested. That combination — foundational and unresolved — is exactly why it transfers.
The problem restated for a trading desk
A systematic portfolio manager runs strategies built on signals: order book imbalance, a filing-based factor, a cross-asset correlation that used to hold between rates and credit spreads. Each signal was validated on a sample. Each signal is now trading live against a market that keeps moving after the validation sample closed.
Here is the failure that recurs across every desk running this kind of book: a signal decays silently and keeps trading after it stopped working. The correlation that justified the position was real in 2019. Nothing announces its death in 2023. The strategy keeps generating the same trades, at the same size, off a relationship whose evidential support ran out months ago. The P&L degrades gradually enough that it looks like noise until it does not.
This is not a modelling failure in the ordinary sense. The model was fine when fitted. The failure is an intake failure: the system stopped accumulating evidence about whether the model still held, while continuing to act as if the accumulation were ongoing.
The likelihood principle names what went missing
Every event that occurs after a signal is validated — every fill, every news item, every quarter's filings — is a potential factor in the likelihood function for "does this signal still work." A backtest closed at some date has a likelihood function fixed at that date. Everything the market does afterwards contributes nothing to it, not because the events are uninformative, but because the system never wrote them into the function at all. That is the frozen-corpus condition, and it is the same one a Large Language Model sits in relative to the world after its cutoff: every subsequent event contributes a factor of exactly one.
A desk that re-backtests quarterly does better. It accumulates evidence for a scene — the last quarter's order books, the last quarter's macro prints — then closes the window and trades on the resulting posterior until the next re-fit. That is the bounded-scene condition. It is real progress over the frozen corpus, and it is still episodic: between re-fits, decay accumulates unaudited, and the silent-decay failure reappears at a smaller scale, contained within each window rather than eliminated.
The alternative is a book whose belief in each signal updates continuously as order books, news feeds, filings and cross-asset signals arrive, with no fixed point at which the evidence stops accumulating. Each new tick, print or filing is a factor multiplied into a running likelihood for "this signal is still live." There is no re-fit date because there is no boundary at which intake pauses. This is the terminal position on the intake axis, and the likelihood principle is what makes it defensible rather than merely aggressive: because likelihood-based updating is indifferent to when or why you stopped looking, a system permitted to observe without a stopping rule suffers no inferential penalty for the fact of continuous looking. Frequentist error control does not offer this guarantee — under repeated testing it explicitly forbids looking whenever you like. The likelihood principle is the only account under which "keep watching, always" is coherent rather than a violation.
| position | intake pattern | what happens to decay |
|---|---|---|
| frozen corpus | fixed at validation date | invisible after the fit; every later event contributes nothing |
| bounded scene | re-fit per window | caught at window boundaries; unaudited within them |
| unbounded stream | continuous, no stopping rule | tracked as it accumulates, provided provenance is logged |
Where the objection lands hardest
Optional stopping demonstrably inflates Type I error. A system that watches continuously and reports whenever the numbers look good is a machine for manufacturing false positives.
This is the strongest objection and a systematic PM should feel its force directly, because it describes a known way desks fool themselves: monitor a hundred candidate signals continuously, report the one that currently looks best, and call the result a discovery. The error-statistical critique is correct that this destroys nominal frequentist guarantees. Group sequential trial design exists precisely because continuous looking without spending alpha inflates false positive rates, and nothing in the likelihood principle repeals that fact.
But the two frameworks are answering different questions. The likelihood principle governs what the evidence says about a signal's current validity, given the data actually observed; it says nothing about the long-run false-positive rate of a rule that reports "significant" whenever a threshold is crossed. A desk running continuous intake needs both accounts, not one instead of the other. The posterior on "does this signal still work" can update honestly tick by tick. Whether the reporting policy — flag, resize, kill — controls its own error rate is a separate, computable question, and it stays computable only if the observation and decision policy is itself logged as data: how many signals were screened, how often the threshold was checked, what triggered the report. Provenance is not paperwork here. It is the record that lets the desk recover frequentist error control on demand, on top of a likelihood that never needed it.
Where absence is not neutral
A monitoring system that treats silence as neutral will be systematically misled precisely where the stakes are highest, because failures suppress their own reporting.
This is the sharper problem, and it is the one that actually explains the silent decay a PM lives with. A signal built on a filing-based factor goes quiet not only when the world stops confirming it, but sometimes because the data feed itself degrades — a vendor drops coverage on a subset of small-cap filers, a news feed's latency creeps up during exactly the volatility regimes where the signal mattered most, an order-book feed thins out on the names where the strategy was concentrated. Treating "no new confirming data" as evidence-free is only correct when the missingness is unrelated to what would have been observed. It is exactly wrong when the missingness is caused by the same regime shift that killed the signal — informative censoring, in the technical sense, and it is common in trading data because thin liquidity, wide spreads and feed gaps cluster with the market stress that breaks strategies.
The correct reading of the likelihood principle is not "silence carries no weight." It is "silence carries weight equal to what its cause implies," which means the observation mechanism itself must enter the likelihood. A desk needs to log, for every signal, not just the data received but the coverage: what fraction of the relevant universe reported, what the feed's latency was, whether a filing deadline passed with no filing or simply had not arrived yet. An unfilled expectation and a genuine absence are different observations and must be written into the likelihood differently. This is provenance again, but now as an evidential component rather than an audit trail — the difference between a stale signal that quietly stopped mattering and a live one whose confirming data has simply gone dark.
Why this is a ceiling, not a waypoint
Once every stream a desk can plausibly ingest — order books, news, filings, cross-asset signals — is admitted continuously, with no fixed re-fit boundary and with the observation mechanism logged well enough that absences remain interpretable, there is no further category of evidence left to add. More compute lets you process the streams faster. It does not create a new stream. The likelihood principle is what guarantees this is not merely a practical stopping point but a structural one: evidence lives in the likelihood of data observed, and once observation has no boundary and its own mechanism is part of the record, the ladder runs out of rungs. Intelligence keeps improving after that. Intake does not.