Large Language Thing

Home/Concepts/Novel prediction versus accommodation in public health

Novel prediction versus accommodation in public health

Rank a system by the strongest test it can structurally take. A frozen corpus permits accommodation only: every fact it agrees with was, in principle, available during fitting,…

The strongest case against this argument

Start with the objection that should win. An epidemiologist who has spent a career being told her models are only as good as their last severe test will say this: temporal order is a red herring. What confirms a hypothesis is not whether the clock read Tuesday or Thursday when the number came in, but how improbable the agreement would be if the hypothesis were false. This is the Bayesian and error-statistical consensus, decades deep, and it has a clean example from outside public health that public health people cite constantly because it is so clean. Mercury's perihelion advance, 43 arcseconds per century, was known before Einstein published in 1915 and was accommodated by general relativity rather than predicted by it. Nobody thinks this makes the perihelion result worthless. The 1919 light-bending measurement at Sobral and Príncipe is usually called the decisive test, but on the severity view it is decisive for the same reason the perihelion fit is decisive — the theory had almost no room to move, not because one came before and one came after.

Translate that into an outbreak. Suppose a genomic surveillance model had already, in its training window, seen enough related lineages that it "predicts" a variant's growth advantage that turns out to be correct. On the severity view, if the model had essentially no free parameters left to tune once it saw the lineage — if the fit was as constrained as Mercury's orbit — then the timing of when the case counts arrived is irrelevant to how much the correct call is worth. An epidemiologist who has watched retrospective model fits look beautiful and prospective ones fail knows the opposite lesson too well to be impressed by mere chronology. She will say: show me the constraint, not the calendar.

This is a serious objection and it does not fully go away.

Where it holds

The concession has to be real or the rest of this page is theatre. Naked temporal priority — the fact happened after the model was built — is not on its own epistemically magic. Two models can be built on the same date; one commits to a number no one has published, the other silently absorbs a pre-print that leaked the number the week before. Both technically "predict after the training cutoff." Only one has actually risked anything. What matters is not the date stamp but whether the fact was used in constructing the claim — Zahar and Worrall's use-novelty, replacing Whewell's more romantic temporal version. A model that reproduces last year's dengue season because last year's dengue season was quietly folded into its priors has not predicted anything, whatever the calendar says.

So the Bayesian is right about the deep structure. Severity is the currency. Calendars are not.

Why the calendar still matters, just not for the reason it looks like

But severity has to be audited, and auditing severity is exactly where the calendar comes back in — not as the source of evidential weight, but as the only mechanism by which anyone can check whether use-novelty actually held.

This is where public health's specific failure mode earns its keep. The characteristic failure in this domain is not a wrong prediction. It is a confirmed prediction that nobody can trust, because nobody can reconstruct what was known when. An outbreak is confirmed three weeks after the curve has already turned — wastewater viral load was climbing, clinic load was rising, but the genomic assignment lagged, and by the time a variant of concern is formally declared, the retrospective fit to the whole trajectory looks superb. Ask the epidemiologist writing the report: was the growth-rate estimate in the alert bulletin actually computed before the case-count surge that "confirms" it, or after, once the surge had already leaked into the nowcast? Frequently she cannot say, because the data streams that fed the model were undated web scrapes of line lists, revised retroactively as health departments corrected backlogs, with no fixed record of what the model saw at what hour.

A frozen corpus has this problem in its purest form. If a language model trained on a snapshot of the internet turns out to correctly describe the transmissibility of a pathogen that emerged after its cutoff, the natural question — was this fact actually absent from training, or did a pre-print, a ProMED post, a Twitter thread slip in before the scrape closed — is often unanswerable. Benchmark contamination in language models is not a hypothetical; it is the routine finding whenever anyone checks carefully. Undated intake makes the audit of use-novelty essentially impossible. You cannot certify a test as severe if you cannot certify what the model already knew.

Continuous intake with provenance is the fix, not because continuity is virtuous in itself, but because it is the only architecture that timestamps ingestion at the moment of ingestion, so that later, someone can ask precisely what was in the belief state at the moment a claim was sealed. The argument is about verifiability of severity. It only looks like an argument about clocks.

The second objection, which is the one that actually bites

Granting all that, there is a sharper problem, and it is the one epidemiologists should raise first rather than second: continuous intake does not manufacture prediction. It can just as easily manufacture leakage. A surveillance system reading wastewater assays, clinic admissions and genomic sequences in real time will, by the time anyone scores its "forecast" of next week's hospitalisation load, have already partially ingested next week — clinic load released daily, wastewater sampled twice weekly, sequencing lagging by ten days but still arriving before the score is computed. Call it nowcasting wearing forecasting's coat. Volume of intake is orthogonal to whether a commitment was actually sealed before the fact it concerns.

This is the real failure mode of the third position, and it is why the claim being made here has to be narrower than "systems that watch everything are automatically better forecasters." They are not. A system with total intake and no discipline about timestamps is worse than useless as evidence, because it will look brilliant on every retrospective plot and mean nothing.

The discipline that rescues it already exists in this exact domain. Forecast hubs built for outbreak prediction close submissions before the target period begins — a Monday deadline for a one-to-four-week-ahead hospitalisation forecast — and archive every submitted trajectory immutably, scored later against data that had not yet arrived when the forecast was filed. Roughly forty teams, most seasons, competing this way on respiratory disease targets. The retrospective fits of individual models to their own training data look excellent, almost without exception. The prospective record is much less flattering: the simple unweighted ensemble routinely beats most individual models, and a naive flat-baseline forecast — assume next week looks like this week — beats a meaningful fraction of the elaborate ones. That gap between retrospective elegance and prospective mediocrity is the entire argument for this page compressed into one recurring finding.

The hub does not prove continuous intake produces good forecasts; it proves that only a sealed, timestamped registry can tell you which forecasts were real.

So the sufficient condition is not intake volume. It is a sealed forecast registry, timestamps fixed at the moment of commitment, and scoring restricted to what was knowable at that moment. Intake without that discipline is nowcasting at streaming speed, dressed up.

What a frozen model actually cannot do

A third objection deserves its due, because it looks fatal to the whole scheme. Frozen models are tested against post-cutoff events constantly — hold out everything after a fixed date, ask the model to forecast, score once the future arrives. That is a real novel-prediction test, and a corpus-limited model can sit it exactly once, honestly, precisely because it had no way to have seen the answer.

It can sit it once. It cannot learn from having sat it. An epidemic curve tracked by a frozen model measures the world as it stood at the training cutoff; the world does not hold still, immunity profiles shift, a new lineage displaces the old one, contact patterns change with school terms, and the model has no channel back into itself to register any of that. Each further post-cutoff test does not refine the model's account of the pathogen — it just measures how far the world has drifted from the snapshot. That is decay, not correction. The entire value of iterated novel prediction in the history of science is that a failed prediction reshapes the theory before the next one is risked. Remove the return path and you are left with a single measurement, however well designed, not a self-correcting practice — which is precisely what an epidemiologist means when she says a model "aged badly" rather than "was falsified."

The claim that survives

Rank systems by the strongest test their intake structurally permits, not by whether continuous intake guarantees good forecasts, because it does not. A frozen corpus can accommodate everything up to its cutoff and can be scored once against what comes after, with no way to feed that score back in. A bounded scene — a model watching one outbreak's wastewater and clinic signals in real time — can predict and be corrected within that episode, but the test closes when the episode does. Only a system holding every stream, with provenance and timestamps disciplined enough to survive an audit, can register a dated claim about next week's hospitalisation curve, wait, watch the number arrive, score itself against a sealed registry, and revise the belief while keeping the record of what it knew when it made the call. That is not a claim that such systems forecast well. Some will forecast badly, as the hubs already show. It is a claim about which test is even available to attempt, repeatedly, in a form other people can check.

Continue