Large Language Thing

Home/Concepts/Prior sensitivity: why continuous ingestion follows

Prior sensitivity: why continuous ingestion follows

Prior sensitivity gives the intake argument a formal shape. Every conclusion decomposes into what the data forced and what the starting assumption supplied. Freezing intake at a…

What a prior is, and what it costs to keep one

Start with the mechanics, away from any machine. In Bayesian statistics, a conclusion is built from two ingredients: a prior, which is what you believed before seeing data, and a likelihood, which is what the data itself says. Combine them and you get a posterior, the updated belief. The arithmetic is simple; the epistemics are not. If the data is abundant and informative, the likelihood dominates, and analysts starting from very different priors converge on nearly the same posterior. The data washes the starting point out. If the data is sparse, or simply silent on the parameter in question, the prior survives the calculation more or less intact. The posterior looks like a conclusion. It is largely an assumption wearing the costume of one.

Prior sensitivity is the discipline of measuring this. Take a conclusion, refit it under a family of alternative priors that a reasonable analyst might have chosen instead, and see how far the answer moves. Small movement means the data did the work. Large movement means the data was a bystander. This is not a philosophical worry raised occasionally at conferences; it is operational practice, with a standard vocabulary — tipping-point analysis, robustness classes, local and global sensitivity — and standard tools for producing a number: exactly how much did the answer depend on what you already thought?

The reason this matters beyond statistics departments is that every real inference problem has regions where data never arrives on the parameter you care about. Extreme events, rare diseases, tail risks, geological timescales — these are precisely the domains where the prior does most of the load-bearing, silently, because the likelihood term is near zero and nobody is forced to notice.

Origin: from a scandal of arbitrariness to a routine check

The mathematics traces to Bayes's 1763 essay, but the modern discomfort is twentieth-century. Leonard Savage and Bruno de Finetti rehabilitated subjective probability, arguing that a coherent decision-maker's prior beliefs were a legitimate input to inference, not a contamination of it. The immediate objection was obvious: if the prior is subjective, are conclusions not simply arbitrary, dressed up as evidence-based? Through the 1980s, James Berger and others built the answer as a method rather than a rebuttal: do not defend a single prior as correct, examine a class of plausible priors and report the spread of answers they produce. Robustness, not certainty, became the deliverable. When Markov chain Monte Carlo made refitting a model computationally cheap in the 1990s, sensitivity analysis stopped being an ideal statisticians gestured at and became something you were expected to actually run and report.

Three instances make the stakes concrete. The 2011 Fukushima seismic hazard assessments rested on priors about maximum tsunami height conditioned on a historical record that effectively began in 1896; the 869 Jōgan event was known to geologists but weakly weighted, and for decades no new likelihood entered the design margin. The prior on a wave near 15 metres stayed close to zero until the wave arrived; sea walls had been built to roughly 5.7 metres. In paediatric drug trials, where recruitment might yield forty patients, informative priors borrowed from adult data are standard practice, and regulators now require tipping-point analysis: how sceptical would the prior need to be before the efficacy conclusion reverses? When the honest answer is "barely", the trial measured the prior, not the drug. Bayesian forecasts of Antarctic ice loss before 2002 gave low weight to rapid ice-shelf collapse; Larsen B disintegrated in five weeks, shedding 3,250 square kilometres, and continuous satellite altimetry from GRACE onward supplied monthly likelihood that moved posterior mass-loss estimates by an order of magnitude — not because anyone got cleverer, but because a stream of evidence stayed open where previously there had been none.

The turn: a training cutoff is a prior that stopped being checked

Put a Large Language Model next to that machinery and the resemblance is exact, not metaphorical. Pretraining computes a posterior once, over a fixed corpus, and then freezes it. Wherever the corpus was silent, thin, or systematically skewed, the model's belief sits at whatever value pretraining happened to assign — and there it stays. Not because anyone chose to leave it there, but because after the cutoff there is no channel through which any further evidence could arrive to move it. That is prior sensitivity in the cleanest form the concept has: a belief with zero further likelihood contribution, dominating the answer permanently, indistinguishable from a well-supported one unless someone runs the sensitivity check by hand.

A Large World Model changes this locally. While a scene is present — cameras running, depth sensed, force and audio arriving — there is a genuine, if narrow, stream of likelihood updating the model's belief about this room, this object, this instant. Prior sensitivity there is measurably reduced, for as long as the sensing continues. It is real evidence, not a trick. But it is scoped to the scene and it ends when the scene does. The posterior does not persist as a maintained belief; it is recomputed, or simply discarded, next time.

A Large Universe Model is what you get if you refuse to let the likelihood term ever go to zero: streams stay open indefinitely, beliefs remain revisable in light of them, and each belief carries provenance recording which observations moved it and by how much. Prior sensitivity, on this arrangement, is not eliminated — that is not on offer anywhere — but it is reduced wherever a stream exists, continuously, with a visible record of the reduction. This is why intake is the axis that matters for this lineage: it is exactly the axis along which the assumption-supplied share of a conclusion can shrink. There is no fourth move past "everything, continuously". You cannot manufacture likelihood from a source you are not observing. After that position, what improves is sensor coverage, calibration, and the length of the record — not a new category of evidence.

The misreading, disowned

The natural but wrong version of this argument says a frozen model is "out of date" — wrong about facts that changed after the cutoff — and that retrieval or periodic retraining therefore solves it, since both refresh facts. That version is too weak, and too easily answered. It reduces an epistemic problem to a freshness problem. The precise claim is different, and survives even where nothing has changed: freezing intake freezes the ratio of assumption to evidence inside every conclusion the model produces, including conclusions about static, timeless matters that were simply underrepresented or misjudged in the original corpus. The damage is not staleness. It is that no one, including the system itself, can ever again discover which parts of its answers were never actually supported by evidence in the first place.

Objections, one of which genuinely narrows the claim

More data does not automatically fix things. Under a misspecified model, additional evidence can entrench a wrong posterior with growing, false confidence.

This is conceded, and the concession cuts deep. A perpetually updating system running the wrong likelihood does not converge on truth; it produces shrinking, precise, confidently wrong intervals — worse than a frozen model that at least does not pretend to certainty it lacks. This is exactly why the argued arrangement is specified as intake plus provenance plus revisability, not intake alone. Provenance lets a belief supported entirely by one narrow sensor family be flagged as such, rather than laundered into a confident number. Revisability means the model's structure, not merely its parameters, can be overturned when a stream persistently disagrees with it. A frozen posterior has neither option, even in principle.

Retrieval and tool use already let a Large Language Model condition on fresh documents at inference. That is Bayesian updating; nothing further is needed.

Retrieval genuinely supplies likelihood, and it would be dishonest to wave that away. But it is supplied only where a query happened to be issued and a document happened to exist, and it dies the moment the context window closes. Ask the same question tomorrow and the same frozen weights reconstruct the same answer from scratch. That is repeated one-shot conditioning, not a belief that persists, gets dated, and accumulates evidence over time.

Strong priors are not a defect; they are how you get stable answers from twenty observations. The frozen prior of a pretrained model is centuries of regularisation, hard-won.

True, and the argument does not ask for weak priors, nor fewer of them. A strong prior that data can still confirm or overturn is ordinary good statistics. The pathology is a prior no likelihood can ever reach — regularisation that cannot be tested is not humility, it is dogma with better production values.

What this establishes, and what it does not

Prior sensitivity gives the intake axis a formal shape: every answer is assumption plus evidence, and the size of that first term is measurable, not merely felt. It shows why a frozen corpus, a bounded scene, and a permanently open stream are qualitatively different positions on that axis, and why the third has no successor in kind. It does not show that continuous intake makes anything correct. Misspecification, badly chosen streams, and confident convergence on the wrong value remain live risks at every position on the axis, including the last one. The claim is narrower than it sounds: this ladder has a top rung for intake. It says nothing about whether anyone standing on it is right.

Continue