Large Language Thing

Home/Concepts/Redundancy and predictability in natural language in public health

Redundancy and predictability in natural language in public health

Redundancy is borrowed stability. A predictor trained on a corpus can only recover what was already regular when the corpus was frozen, and the fraction of the world that is…

The objection that should win

Here is the strongest case against everything that follows, and it deserves to be stated in full before it is answered.

Redundancy in language is a fact about codes, not about the world. English is roughly 75 per cent redundant because of orthography, morphology, and grammatical agreement — constraints internal to the language itself. An epidemiologist reading a case report benefits from that redundancy for the same reason anyone does: 'q' predicts 'u', a plural noun predicts a plural verb, a familiar phrase completes itself before the eye finishes moving. None of this is knowledge about disease. It is knowledge about English. To claim that a language model's predictive success reflects a stable, recurring world is to mistake the structure of the channel for the structure of the referent. The whole argument for a Large Universe Model leans on that confusion, and public health, of all fields, should not be fooled by it — outbreaks are precisely the moments when the world stops being predictable, and no amount of grammatical redundancy will see one coming.

This is a serious objection and it should be allowed its full force before any answer is offered. Claude Shannon's 1951 guessing experiments, published as 'Prediction and Entropy of Printed English', measured the entropy of printed English at somewhere between 0.6 and 1.3 bits per character against an alphabet capable of carrying about 4.7 bits. That gap — roughly 75 per cent redundancy — is real, measured, and largely structural. A great deal of it has nothing to do with epidemiology, or with anything outside the language itself.

What the objection gets right

Concede the structural point in full. A great share of what makes English predictable is orthographic and syntactic. 'The patient presented with' predicts 'a' or 'an' with near certainty and predicts very little about the patient. An epidemiologist skimming a line list benefits from grammatical redundancy exactly as a novelist's reader does: fewer characters need to be parsed to recover the same sentence. None of that redundancy tracks a virus, a sewer main, or a hospital's bed occupancy. If the whole of a language model's predictive power came from this structural surplus, the objection would be fatal. A model that merely completed grammatical patterns would be a stylist, not an instrument.

It is also true that public health's characteristic failure looks like a failure of prediction rather than a failure of grammar. An epidemiologist confirms an outbreak roughly three weeks after the epidemic curve has already turned — not because sentences were parsed wrongly, but because the case reports, the genomic sequencing runs, and the wastewater assays that would have shown the turn earlier arrived late, arrived incomplete, or arrived in a form no model had been trained to expect. That lag has nothing obviously to do with entropy per character.

Where it breaks

The objection breaks on a measurable point: structural redundancy has a ceiling, and predictive performance in practice runs well past it. Complete the phrase 'the incubation period of' and a language model, or a trained clinician, produces a narrow, confident distribution of numbers — because incubation periods for known pathogens are stable facts about the world, encoded and re-encoded across thousands of documents. Complete 'the cluster in the dormitory has now reached' and the same model's confidence collapses, not because the grammar is harder, but because the referent — an active, unfolding case count — has not yet been written down anywhere in a form stable enough to compress. Cloze studies, descended from Wilson Taylor's 1953 readability procedure, make the separation explicit: once syntax is held constant, contextual predictability tracks world knowledge, not sentence shape. The gap between 'incubation period' and 'cluster size right now' is exactly the gap this page is about, and it is referential, not grammatical.

Public health phraseology makes the same distinction operationally, the way aviation radiotelephony does for pilots. A case report has fixed scaffolding — pathogen name, onset date, exposure history, laboratory method — that is almost fully predictable in form. The scaffolding is not the payload. The payload is the count, the sequence, the assay result: 14 confirmed cases, lineage BA.2.86, 220 copies per litre in the intake sample. That is the fragment no prior expectation reconstructs if it is lost or delayed, and it is precisely the fragment an outbreak turns on.

So the concession and the rebuttal sit side by side. Some of English's redundancy is a fact about the code. Some — a large, measurable share — is a fact about a recurring world, borrowed by the code because the people who wrote the sentences were describing things that hold still long enough to be described twice. A language model trained on public health literature learns both. It knows, with justified confidence, how tuberculosis transmits, how wastewater sampling works, how a genomic surveillance report is structured. It has no way of knowing whether the viral load in this week's intake sample is rising, because that fact did not exist stably enough, long enough, to be written into any corpus before the corpus was frozen.

The retrieval reply, and its limit

A natural answer follows immediately: give the model a live feed. Wire a frozen language model to the health department's case-reporting API, the wastewater dashboard, the sequencing pipeline's output queue, and the moving part is supplied at query time. This is not speculative; it is an existing engineering pattern, and it is the right instinct.

Its limit is inheritance. A retrieval layer is only as fresh and as complete as whatever it queries, and what it typically returns is text describing a measurement, not the measurement itself, taken with its own latency baked in. A wastewater dashboard updated every Tuesday answers a query on Wednesday with data that is, at best, six days old — inside the three-week lag that already defines the failure, but not eliminating it. Retrieval narrows the gap between corpus and event. It does not close the gap between event and current belief, because it still routes through documents that someone had to write, publish, and index before they became queryable.

The three-week outbreak lag is not a reporting delay to be fixed once; it is the visible edge of every stream's individual latency, stacked.

What continuous intake actually buys, and what it costs

The Large World Model's contribution is to stop reading about the scene and start sensing it — a camera on a production line, a sensor on a reactor, an occupancy count from a building. Applied loosely to public health, this looks like real-time wastewater assay data or hospital admission telemetry rather than a monthly bulletin describing the same. It closes real ground: the moving part is measured, not inferred from last month's report. But the recovery lasts exactly as long as the sensor is switched on and the analyst is looking. A wastewater plant not on the sampling schedule, a clinic whose load data feeds a system nobody reconciles overnight, a genomic run that finishes on a Friday and sits unread until Monday — each is a scene nobody is currently looking at, and the moment nobody looks, the world is again unmeasured, exactly as it was for the frozen corpus.

The step beyond that is holding every relevant stream — case counts, wastewater assays, genomic surveillance, clinic load — as a standing belief with a timestamp and a source, continuously, rather than as a document to be fetched or a scene to be glanced at. An outbreak belief under this regime does not say 'transmission is occurring'; it says 'transmission is estimated ongoing, wastewater signal last confirmed 6 hours ago, genomic lineage last confirmed 2 days ago, clinic load last confirmed 40 minutes ago, confidence weighted accordingly'. This is the Large Universe Model's proposal for public health: not a smarter predictor of the next case report, but a system that never stops checking, and that says plainly how stale each belief currently is.

LLMLWMLUM
what it holdscorpus of past reports, frozencurrent scene, while observedevery stream, continuously, with timestamps
outbreak signalinferred from precedentmeasured, but only in viewmeasured, revised, decayed if stale
characteristic failureconfirms three weeks lateblind to the plant not being watchedreconciliation load across contradicting sensors

The honest cost is the third objection, and it is real: more streams means more sensor drift, more clock skew, more duplicate case reports, more contradiction between a wastewater signal and a clinic count that has not yet caught up. Continuous intake does not lower entropy; it converts a prediction problem into an estimation and provenance problem, and that problem is where such a system will mostly fail in practice — which source, measured when, trusted how much, reconciled against what. The consolation is that disagreement between independent streams is itself informative in the way a single confident report is not: three sensors that disagree bound the truth better than one that agrees with itself. The epidemiologist's job does not disappear into the machinery. It moves upstream, into deciding which disagreements matter and how much delay in reconciliation the situation can tolerate — which is a narrower, harder-edged version of the job already being done, not a different one.

Continue