Home/Concepts/Retraction and the scientific record: why continuous ingestion follows
Retraction and the scientific record: why continuous ingestion follows
Treat a corpus as an asset. Its value is not its size but its size net of everything since withdrawn, and that netting is unobservable from inside the snapshot. Depreciation runs…
The ledger that amends itself
A scientific paper is published. Nothing about that act guarantees the claim inside it is true, and nothing about its truth is fixed at the moment of publication. Peer review checks plausibility and method, not durability. What follows publication is a second, slower process: the claim gets cited, extended, occasionally repeated in someone else's lab, and sometimes it is found wanting. When it is found wanting badly enough, the journal issues a retraction — a formal, indexed withdrawal stating that the paper, or some load-bearing part of it, should no longer be relied upon.
Retraction sits inside a graded set of instruments. An erratum fixes a typo or a mislabelled axis. A correction fixes something substantive but survivable. An expression of concern flags that something may be wrong while an investigation runs. A retraction is the withdrawal itself, the strongest instrument the record has short of ignoring a paper outright. The causes are heterogeneous: outright fraud, image duplication, contaminated cell lines, a spreadsheet formula dragged one row too far, an effect that simply would not reproduce elsewhere. Roughly one paper in a thousand is eventually retracted. The median gap between publication and retraction is about three years.
That lag is the important number, more important than the rate. It means the scientific record is not a warehouse of settled findings that occasionally gets a shelf cleared. It is a ledger under permanent amendment, where the amendments are written years after the entries they correct, and where for those years the entry sits on the shelf looking exactly as authoritative as everything beside it. Anyone reading the literature at any single moment is reading a mixture of true claims, claims later found false but not yet marked, and claims already marked. There is no way to tell the middle category from the first by looking.
Where the instrument came from
Errata are as old as printing. Retraction as a distinct, named, indexed act is younger, and its emergence tracks a specific institutional problem: what to do when fraud is discovered after publication rather than caught before it. The fraud cases of the early 1980s, John Darsee's fabricated cardiology data at Harvard chief among them, forced journals and funders to develop machinery they had never needed. The US National Library of Medicine created a distinct "Retracted Publication" indexing category in 1984, so that a search of the literature would surface the withdrawal alongside the original. Federal oversight followed: the offices that became the Office of Research Integrity took shape from 1989, with the ORI itself established in 1992. The Committee on Publication Ethics formed in 1997 and wrote guidelines for how a retraction notice should be worded and issued.
For decades after that, nobody was actually counting. Ivan Oransky and Adam Marcus started Retraction Watch in 2010 for the plain reason that no systematic record existed of how many retractions there were, why, or where. Their database grew past 50,000 entries and passed to Crossref in 2023, where it became a queryable, open feed rather than a private tally. That transfer matters more than it looks. It turned retraction from an event a journal announces once into a stream a system can subscribe to.
The turn
Consider what a system trained on the scientific literature actually absorbs. A Large Language Model ingests a corpus fixed at some cutoff date. It takes in the paper. If the timing runs unlucky, it takes in the paper without the withdrawal that follows, because the withdrawal has not happened yet or has not propagated into the training snapshot. It also takes in every downstream paper that cited the original while it stood unchallenged, which is usually thousands of documents restating the finding as settled fact. The error is not just present; it is laundered into consensus by repetition, and the model has no mechanism to tell the difference between a claim believed because it is well-supported and a claim believed because everyone cited everyone else.
A Large World Model does better with what sits in front of it. Grounded in a scene — a lab bench, an instrument, a patient — it can perceive and reason about what is physically there with a fidelity no corpus snapshot offers. But the priors it brings to interpreting that scene are still drawn from a frozen literature. It can watch an amyloid plaque under a microscope with perfect attention and still bring twenty-year-old assumptions about what that plaque means, unaware that the paper underwriting those assumptions was withdrawn last Tuesday in a journal it does not read.
Retraction is the cleanest existing proof that a corpus cannot be the last word on itself, because the evidence bearing on a claim keeps arriving after the claim does, on a delay measured in years. A system whose intake is continuous — one that takes the correction stream itself as an input, Crossref's retraction feed alongside errata, replication reports, and post-publication review, and attaches every belief to the sources that produced it — can let a withdrawal upstream reopen a belief downstream. That is the definition of the Large Universe Model on this axis: not a bigger corpus, but a corpus that is never finished arriving, where each item carries its provenance and its own contradiction is a first-class event rather than an omission.
What this does not license
The tempting shortcut is to read this as an indictment: the literature is riddled with fraud, so anything trained on it is worthless. That is wrong on the numbers and wrong in spirit. A retraction rate near one in a thousand is low, and every retraction issued is the system working, catching what review missed, not the system failing. The claim here is temporal, not moral. Even a spotlessly honest literature would still need continuous intake, because the problem is that evidence arrives on a delay, not that the evidence is bad.
The mirror-image misreading is worse: that a live correction feed delivers truth. It delivers a record of contested status. A retraction notice tells you a claim is disputed, on stated grounds, by named parties. It does not tell you the claim is false — retracted results occasionally hold up on reinvestigation, and vast quantities of unretracted work never replicate at all. Confusing "flagged" with "false" replaces one error with another.
Taking the objections straight
Retraction hits 0.1% of papers. Building provenance machinery to catch one in a thousand is a poor use of effort; ordinary disagreement swamps retraction as a source of error anyway.
The rate is correctly stated and the inference from it is not sound. Retractions cluster in exactly the papers a model or a downstream literature weights most heavily: high-visibility, heavily cited, clinically load-bearing work. Sylvain Lesné's 2006 amyloid oligomer paper in Nature had accumulated roughly 2,300 citations before its 2024 retraction, having oriented a fair portion of Alzheimer's research for the better part of two decades. The exposure is not spread evenly across a corpus; it concentrates where the consequence of being wrong is largest.
The correction stream is itself unreliable — political retractions, authorship disputes, results that later replicate anyway. Continuous intake trades a stale belief for a jittery one and hands anyone who wants a finding suppressed a new lever to pull.
This is the objection that should narrow the claim, because it is largely right. A retraction notice is a piece of evidence, issued by a party, on stated grounds, and it should be recorded as exactly that — not as a deletion. The correct response to a withdrawal is to attach it to the belief it bears on, along with who issued it and why, and let that reweight confidence rather than erase it. Naive overwriting would indeed produce jitter and a new attack surface. Provenance-tracked belief revision is the alternative to naive overwriting, not a defence of the correction stream's infallibility, and any account of continuous intake that skips this distinction deserves the objection in full.
Periodic retraining already handles this. Refresh the corpus every six months; withdrawn papers drop out.
Refreshing removes the retracted paper. It leaves the citing literature untouched, and that literature typically still states the finding as settled fact — roughly a third of post-retraction citations make no mention of the withdrawal — and vastly outnumbers the original notice. Retraining also produces a fresh, entangled weight state with no traceable path from any given conclusion back to the sources it rested on, so there is no way to ask which downstream claims depended on the now-withdrawn work, let alone revise them individually. And the three-year median lag guarantees that every fresh snapshot, however recently pulled, still contains a full cohort of papers that are false and not yet marked, by construction of the lag itself.
What the argument establishes, and no more
Retraction shows, with unusual clarity, that a fixed intake cannot faithfully represent an object defined by ongoing self-correction. It shows that the correction arrives as a separate event, on its own timeline, and that treating intake as continuous is not an enhancement but a requirement for keeping pace with that timeline. It does not show that continuous intake yields settled truth — the correction stream carries its own uncertainty, is subject to its own errors and disputes, and must be weighted rather than obeyed. It does not show that most of a frozen corpus is wrong; the base rate is reassuringly low. And it does not, on its own, prove that a Large Universe Model is achievable at scale, only that the category of evidence it is built to admit — corrections to claims already ingested — is real, recurring, and structurally invisible from inside any snapshot. That is a narrower claim than it might sound, and it is the one the record actually supports.