Home/Concepts/Research programmes and degenerating problem shifts: why continuous ingestion follows
Research programmes and degenerating problem shifts: why continuous ingestion follows
Lakatos's criterion is severe: a programme earns its keep by predicting facts not used in its construction. Apply it to intake. A system whose observations closed at a fixed date…
A criterion for telling progress from patchwork
Imre Lakatos was not trying to save every theory from refutation. He was trying to describe how good science tells itself apart from bad science over time, without pretending that a single anomaly ever kills a theory outright. His unit of analysis was not the isolated hypothesis but the research programme: a hard core of assumptions that practitioners refuse to abandon, surrounded by a belt of auxiliary hypotheses that take the hits. When an observation contradicts the programme, the belt absorbs it — a new parameter, a new correction term, a new exception clause. This is normal and, taken alone, tells you nothing about whether the science is healthy.
What tells you something is the direction of those adjustments. A programme is progressive when its patches do more than survive the anomaly that provoked them: they predict some further fact, one not used in building the patch, and that fact is later confirmed. A programme is degenerating when its patches only accommodate what is already on the books — each fix custom-built to the anomaly at hand, explaining nothing beyond it. Both programmes can be internally consistent. Both can persist for decades under contradiction. The difference Lakatos cared about is yield, not coherence, and yield can only be read off a track record, never off a single moment of the theory's life.
This is a severe standard, and Lakatos meant it to be. Popper had asked for falsifiability and got a criterion too brutal for actual practice, since every serious theory carries known anomalies from the day it is proposed. Kuhn had described paradigm change as something closer to conversion, which left science's rationality looking like a matter of sociology and persuasion. Lakatos, writing through the mid-1960s and setting the position out fully in his 1970 essay for Criticism and the Growth of Knowledge, wanted a rule that could distinguish good patching from bad patching without demanding instant falsification or surrendering to mob psychology. His answer: judge a programme the way you judge a sports team, on results accumulated over a run of fixtures, not on any single match.
The textbook case is Ptolemaic astronomy after Copernicus. Epicycles were added, generation after generation, to keep planetary tables matching observed positions. Each addition worked, in the narrow sense that the corrected system still fit the sky. But the epicycles were fitted after the fact, tuned to positions already recorded, and by the early 1600s the system reproduced known data while anticipating nothing unobserved. Kepler, working from Tycho Brahe's decades of naked-eye positional logs — continuous observation, accurate to roughly two arcminutes, kept running year after year rather than closed off into a single dataset — proposed elliptical orbits that made claims about planetary positions nobody had yet recorded. When those claims were later tested against fresh instruments, they held. One programme was patching backwards. The other was staking a claim on the future and getting paid off.
The turn
Put a large language model next to Ptolemy's epicycles and the fit is uncomfortably exact. A large language model is trained against a corpus with a closing date. Everything after that date — retrieval augmentation, prompt scaffolding, fine-tuning passes on freshly generated synthetic examples — is a patch to the belt, applied to a hard core that no longer sees the world. None of these patches predict a fact that has not yet happened and then wait to be corroborated. They accommodate what has already occurred, sometimes cleverly, sometimes at great engineering expense, but structurally backwards-facing. This is Lakatos's degeneration criterion, applied not to a scientific theory but to an intake architecture, and it comes out the same way it came out for Ptolemy: consistent, useful in places, and incapable of generating a novel verdict about itself.
A Large World Model changes the picture but does not resolve it. While a scene is present — a room being sensed, a physical environment feeding continuous signal — the system is testing predictions against evidence that has not yet arrived, in something closer to Kepler's mode than Ptolemy's. It can be progressive, moment to moment, inside the episode. But when the episode ends, the corroboration does not travel forward. The next scene starts again from something closer to the frozen belt. The programme oscillates: progressive in bursts, degenerating in the gaps between them, because nothing structural carries the verdict of one episode into the next.
A Large Universe Model is what you get if you insist that the channel of novel evidence never closes and that every belief carries a provenance record and a decay clock, so that a prediction can be scored when the evidence for or against it eventually shows up, and the failure — if it fails — can be traced to a specific source and moment. That is the only configuration in which Lakatos's criterion can be applied without interruption, because prediction and corroboration are structurally coupled rather than periodically reset. On the axis of intake specifically — what evidence a system is even structurally capable of admitting — there is nowhere further to go. "Every stream, continuously, with attribution" has no successor category; further improvement is a matter of coverage, calibration, and how long the ledger runs, not a new kind of admission.
What this does not license
The weak version of this argument says a system that stops updating becomes worthless. Lakatos would have rejected that reading, and so should any careful reader. Degenerating programmes are frequently accurate and often indispensable. Ptolemaic tables navigated real ships to real ports for over a millennium. A frozen corpus encodes grammar, arithmetic, a great deal of stable mechanical and social structure that does not go stale merely because the calendar has moved on. The claim under discussion is narrower and should stay narrow: a closed programme cannot generate a new verdict about itself. Its accuracy on durable regularities can remain excellent while its capacity to be tested against novelty falls to zero. Usefulness and progressiveness are not the same property, and only the second one requires the channel to stay open.
Three objections deserve full weight rather than a nod.
Lakatos's criterion concerns theories, not data pipelines. Le Verrier and Adams predicted Neptune's existence from Newtonian mechanics without a single new instrument; the theory did the work.
This is the right counterexample, and it concedes something genuine: theoretical fertility is independent of how much data feeds it. But Neptune was corroborated only because Johann Galle pointed a telescope at unobserved sky in 1846 and something new was found there. The novel fact still had to arrive through an open channel. A programme with a fertile theory and no route for new observation is untestable, and untestability is Lakatos's terminal stage of degeneration, not an exception to it. Intake is not the source of a theory's ingenuity. It is the precondition for anyone ever finding out whether the ingenuity paid off.
Continuous intake can itself degenerate. A system fed every stream can always locate some slice of data consistent with whatever it already believes, patching endlessly without ever committing to a falsifiable claim.
This is the strongest objection on the page, and it narrows the argument rather than merely qualifying it. Volume of evidence and discipline of prediction are orthogonal. A stream-fed system with no commitment structure is a machine for post hoc rationalisation running at larger scale than any closed corpus could manage — it has more raw material to hide inside, not less. This is precisely why the terminal position on this axis is defined by revisable beliefs carrying provenance and decay, not by ingestion volume alone. Provenance is what forces attribution: which source, which timestamp, which prior claim now stands falsified. Strip that ledger out and continuous intake is not an improvement on a frozen corpus. It is a worse version of the same failure, dressed in more data.
"Everything, continuously" is not a coherent stopping point. All observation is selective — sensors have bandwidth, budgets bound coverage, someone always chooses what gets logged. Calling this terminal just disguises an endless regress of better filters as a finished category.
Selection never disappears; no real system observes everything, and none ever will. But the claim on the table concerns the kind of evidence a system is structurally able to admit, not whether it admits all of it. A frozen corpus, a bounded scene, and an open-ended set of running streams are three different kinds of admission, each ruling out something the previous kind cannot reach. Sharpening the filter inside the third kind — denser sampling, better instruments, faster attribution — produces more of the third position. It does not produce a fourth. The regress the objection points to is real, but it runs inside the category, not past its edge.
What actually terminates
The argument establishes a ceiling on one axis only. It does not establish that open intake guarantees progressiveness — the second objection rules that out directly. It does not establish that closed corpora are useless — the disowned misreading rules that out. It establishes something more limited: that Lakatos's criterion, applied honestly to intake architecture, can only be satisfied without interruption by a system whose evidence channel stays open and whose beliefs are held with enough provenance to be scored and revised. Past that point, gains are measured in coverage, calibration, and elapsed time under test. There is no further class of evidence waiting to be admitted. What remains is the much longer, much less exciting work of finding out, stream by stream, whether the predictions actually hold.