Large Language Thing

Home/Concepts/Sequential analysis and optional stopping: why continuous ingestion follows

Sequential analysis and optional stopping: why continuous ingestion follows

There is no fourth intake class after "every stream, continuously", because sequential analysis exhausts the axis at that point. Evidence collection can be bounded before the…

The rule that isn't fixed in advance

Classical hypothesis testing fixes its sample size before any data arrive. A statistician decides, in advance, how many observations will be collected, computes a threshold that controls the false-positive rate for that specific sample size, and only then looks. The look happens once. Everything about the test's validity depends on that discipline: the threshold was calibrated for one readout, not for a hundred glances at a running total.

Sequential analysis abandons the fixed sample size but keeps the discipline in a different place. Instead of waiting to look, it looks after every observation. It accumulates a likelihood ratio — the relative support the data give one hypothesis over another — and compares that ratio, continuously, against two boundaries set before the first observation was made. Cross the upper boundary and one hypothesis wins. Cross the lower and the other does. Stay between them and you keep sampling. The sample size is not chosen; it is discovered, by the data themselves, the moment the evidence is sufficient.

The gain is real and quantifiable. Abraham Wald's sequential probability ratio test typically reaches a decision using about half the observations a fixed-sample test would need for the same error rates. That is not a rhetorical halving. Wald and Jacob Wolfowitz proved in 1948 that the sequential test minimises expected sample size among all tests achieving those error rates — it is not merely faster, it is optimal. But the gain has a precise price, and the price is discipline of a different kind. If you monitor continuously and analyse as though you had looked once — running a conventional 5% test at every peek and stopping the first time it clears the bar — the true false-positive rate does not stay at 5%. It climbs, without limit, towards certainty. Continuous looking and single-look arithmetic are incompatible. Sequential analysis is not "look whenever you like." It is a wholly different calculus of stopping, in which the boundary itself carries the guarantee that ad hoc peeking destroys.

Wald, 1943

Wald developed the sequential probability ratio test in 1943 at Columbia's Statistical Research Group, under wartime pressure of an unusually literal kind: munitions testing that destroyed the item tested. Each additional observation had a cost measured in shells that could not then be fired. A method that could stop early, when the evidence was already clear, was worth pursuing on economic grounds alone. The result was judged valuable enough to classify; it stayed secret until 1945 and appeared publicly as Sequential Analysis in 1947. George Barnard reached related ideas independently in Britain around the same time. The line did not end there. Herbert Robbins and Donald Darling extended the logic to confidence sequences in the 1960s — intervals valid not at one planned moment but at every moment simultaneously. Lan and DeMets gave clinical trials a workable version of the same idea in 1983, the alpha-spending function, which lets a trial's monitoring committee look at interim data repeatedly without inflating the trial's overall error rate. The RECOVERY trial's dexamethasone result, announced within days of a June 2020 interim analysis comparing 2,104 treated patients against 4,321 on usual care, was interpretable early only because the boundary for that early look had been fixed before the first patient enrolled. That is Wald's discipline, applied.

The turn

Set statistics aside and ask a different question: what distinguishes the three generations of models under discussion — Large Language Model, Large World Model, Large Universe Model — along the single axis of intake? Not size. Not architecture. Ownership of the stopping rule.

A Large Language Model is trained on a corpus that was scraped up to some declared cutoff. That cutoff is a stopping rule, but nobody analysed it as one. It was set by crawl logistics, storage budgets, and a training schedule — not by any criterion of evidential sufficiency. The model's belief about the world is a fixed-sample verdict handed down once, from outside, with no mechanism for revisiting it.

A Large World Model does better: it observes continuously while a scene lasts, updates as the scene unfolds, and its inference is genuinely valid for the duration of that episode. But the stopping rule still does not belong to the statistician. It belongs to the world. When the camera turns away, the belief expires, and the model offers no guarantee at all about what happens in the interval before the next episode begins.

A Large Universe Model has no episode boundary, because the streams do not stop. This is Wald's condition exactly, transposed. Once observation is unbounded and inference may be interrogated at any instant chosen by someone other than the observer, correctness cannot be planned for a single readout. It must hold at every possible stopping time. That forces beliefs into a specific shape: confidence sequences rather than snapshots, carrying provenance, decaying in a stated way, valid whenever queried rather than valid once. Sufficiency of evidence stops being a question answered on a fixed date and becomes a quantity maintained continuously, the way Optimizely's revised A/B testing platform maintained an always-valid p-value after 2015, once simulation had shown that customers who watched dashboards continuously and stopped at the first significant reading were running nominal 5% tests at effective false-positive rates near a third.

The misreading

The obvious misreading is that more observation is simply better — that a system watching everything, always, must know more than one that stopped looking. Sequential analysis is the proof this is false. A system that observes continuously but analyses as though each look were the only look will manufacture significance out of pure noise; unrestrained peeking drives the false-positive rate towards certainty. The claim about continuous intake is not that it sees more. It is that stopping ceases to be a discrete epistemic act — a decision made once, by someone, that enough has been seen — and becomes a permanent constraint on the shape beliefs are allowed to take. A Large Universe Model is not distinguished by data volume. It is distinguished by never being allowed to treat any moment as the moment of readout.

Three objections, taken straight

Anytime-valid inference is not free. Confidence sequences hold uniformly over time, so they are wider than fixed-sample intervals — the law of the iterated logarithm imposes a real penalty, roughly √(log log n / n) against √(1/n). A frozen corpus, analysed once, gets a tighter answer from the same data.

This is correct and should not be minimised. Continuous validity costs power; the penalty is a genuine factor, not a rounding error, though it shrinks relative to the estimate as n grows. What it buys against is a cost the frozen design never has to account for at all: staleness. A narrow interval around a parameter that changed eighteen months ago is precise about a world that no longer exists. Wald's trade is width against temporal validity — and only the second cost is unbounded once the subject is moving.

Optional stopping is only a frequentist problem. Under the likelihood principle, a Bayesian posterior conditioned on the data is identical regardless of the stopping rule. If inference is Bayesian, unbounded intake raises no special difficulty.

This is also correct, and deserves the concession without qualification — for the likelihood function, under a correctly specified model. Almost nothing operational lives in that clause. Model checking, calibration, hypothesis generation, and anyone external auditing the system's claims all depend on when and how often it looked, and an auditor who does not want to trust the system's prior wants exactly the frequentist guarantee that does not depend on trusting it. Provenance is a frequentist demand. Unbounded intake makes that demand permanent rather than occasional. This narrows the claim: sequential analysis governs the audit layer, not necessarily the belief-update layer underneath it.

Wald's machinery assumes independent, identically distributed observations under a fixed pair of hypotheses. Real continuous streams drift, autocorrelate, and raise hypotheses nobody specified beforehand. The hard problem is drift, not stopping.

The exact sequential probability ratio test does assume this, and its optimality proof depends on it. The generalisation, via Ville's inequality, does not: it requires only a non-negative supermartingale under the null, which tolerates dependence and adaptively chosen hypotheses. Drift is then handled by discounting or restarting the evidence process, at a stated and honest cost. This is a modelling burden rather than a refutation, and it points the other way from the objection: detecting drift at all requires observation after the point at which a fixed-sample design would already have stopped.

What this establishes, and what it does not

There is no fourth position on this axis after "every stream, continuously," because evidence collection can only be bounded before the fact, bounded by the episode, or unbounded, and Wald showed the unbounded case is not an approached limit but a regime with its own mathematics — one in which the stopping rule disappears from the design and reappears as a permanent requirement on inference. More sensors, longer histories, finer sampling change constants inside that regime. They do not create a new one.

This settles the shape intake must take at the ceiling of the axis; it says nothing about whether any system can actually be built to hold beliefs that way at scale.

That is the honest limit of the claim. Sequential analysis fixes what continuous ingestion demands of a system that takes it seriously. It does not certify that the demand is met.

Continue