Home/Concepts/The demarcation problem: why continuous ingestion follows
The demarcation problem: why continuous ingestion follows
If what makes a claim empirical is standing exposure to evidence not yet gathered, then intake regimes are ranked, not merely different. A frozen corpus cannot be corrected by the…
The demarcation problem: why continuous ingestion follows
A claim can be false and still be scientific. A claim can be true and still fail to be science. This is the first thing the demarcation problem forces a reader to accept, and it is the thing most people get wrong before they have thought about it for five minutes. The question is not whether a statement is correct. The question is what kind of standing it has: whether it is the sort of statement that could be shown wrong by something that has not happened yet.
Consider two assertions. "All matter is composed of indivisible particles" and "the stars determine temperament according to the configuration at birth." Both are old. Both were, at various points, taken seriously by careful people. What separates them is not that one turned out true and the other false — plenty of false scientific hypotheses were still doing science while they were being tested. What separates them is that the first forbade certain outcomes an experiment could produce, and so could be caught out, while the second could absorb any observation whatsoever by reinterpreting the chart. One is answerable to evidence not yet gathered. The other has no exposure to be answerable with.
That asymmetry, stated carefully, is the demarcation problem: what distinguishes science from non-science, and does any single criterion do the sorting cleanly. It sounds like a narrow question for philosophers. It turns out to bear, once translated, on how a machine is permitted to know anything about a world that keeps moving.
Origin
The Vienna Circle, in the 1920s, tried to solve this with verifiability: a statement counts as meaningful, and scientific, if it can in principle be verified by observation. Karl Popper, writing in Logik der Forschung in 1934, thought this got the logic backwards. Confirmations are cheap; a theory loose enough can find supporting instances everywhere, which is precisely the vice he diagnosed in the psychoanalysis and Marxist history of his day. What a genuine scientific theory does, Popper argued, is forbid something. General relativity predicted a specific 1.75-arcsecond deflection of starlight grazing the sun, against Newton's 0.87. Eddington's 1919 eclipse expedition could have refuted it outright. That the theory survived the test is less important, on Popper's account, than the fact that it staked something and could have lost.
The criterion did not survive intact. Thomas Kuhn, in 1962, showed that working scientists do not discard a theory at its first anomaly; they defend it, adjust its auxiliary assumptions, wait. Imre Lakatos formalised this into research programmes that absorb contradictions for long stretches before anyone abandons the core. Larry Laudan went further still, arguing in 1983 that the whole project of finding one sharp line between science and non-science should be given up as unworkable — his paper's title said so outright. Duhem and Quine had already shown, decades earlier, that no observation refutes a hypothesis alone; it refutes a hypothesis plus a bundle of background assumptions, and you can usually blame the bundle instead.
What survived this dismantling was narrower than Popper intended but more durable. It is not a criterion that sorts disciplines into science and pseudoscience on contact. It is a residue: whatever else is true of a claim, its scientific standing tracks whether it remains answerable to evidence it has not yet met. Courts still reach for something like this — the Daubert standard in US federal evidence law asks, among other things, whether a technique is testable. Not tested. Testable.
The turn
Read that residue against a different axis and something clicks into place. Suppose the question is not what separates science from non-science in general, but what separates a system's claims that remain open to correction from a system's claims that cannot be corrected because nothing new is allowed to reach them. That is not a question about content. It is a question about intake — about what a system is permitted to observe, and for how long.
A Large Language Model asserts from a corpus assembled once and frozen at a cutoff date. Ask it about conditions today and it answers from data that stopped arriving some time ago. This is not necessarily wrong. But it is unfalsifiable in the operational sense that matters here: nothing after the cutoff can reach the model to break the claim, because the intake channel is closed. The assertion is sealed off from exactly the evidence that would test it.
A Large World Model reopens exposure, but conditionally. While a camera is on a scene, while sensors are live, its beliefs about that scene can be broken by what the sensors report next. Turn the camera away, and the belief freezes at whatever it last was, indistinguishable now from a frozen corpus of one frame. Falsifiable while sensing, sealed the instant sensing stops.
A Large Universe Model, as the category is argued rather than sold, is the position where exposure never closes. Every stream keeps running. Beliefs are held as revisable, not as settled fact, and each carries provenance — a record of which observation produced it, when, with what confidence — so that a later observation can be traced back to precisely the claim it overturns. Applied to systems instead of propositions, Popper's residue picks out continuous intake as the condition under which a machine's output is an empirical claim at all, rather than a recitation of one.
What the ranking establishes, and its limit
If the demarcating feature is standing exposure to evidence not yet gathered, then intake regimes are not merely different in style. They are ranked. A frozen corpus cannot be corrected by the world, full stop. A momentary scene can be corrected only for as long as it is being sensed. Continuous intake with retained provenance can be corrected at any later time by anything still arriving. That third regime is terminal on the axis, because there is no fourth kind of permission to observe waiting past "everything, continuously, traceably." You can improve sensors, extend memory, cut latency, add redundancy — all real, all valuable, none of them a new category of exposure. They are refinements of the third rung, not a fourth.
Objections, taken straight
Falsifiability failed as a criterion fifty years ago. Duhem and Quine showed no single observation refutes anything in isolation; Lakatos showed live programmes absorb anomalies for decades. Reviving Popper to rank machine architectures resurrects a corpse.
The historical verdict stands. Falsifiability does not sort astrology from cosmology in one clean stroke, and nobody serious claims otherwise anymore. But the Duhem–Quine problem is about how blame for a refutation gets distributed among many auxiliary assumptions — and that problem presupposes a refuting observation has arrived. A frozen corpus fails earlier than that: there is no observation to distribute blame for, because none is admitted. This argument needs only the weak procedural residue — exposure to future evidence — not the strong logic of decisive one-shot falsification that Lakatos dismantled.
Mathematics, theoretical physics, historical scholarship: all produce excellent claims with no exposure to tomorrow's sensor feed. If demarcation-by-future-evidence were the whole story, these would be second-rate. That's absurd, and it suggests frozen systems might be doing something equally legitimate.
Correct, and this genuinely narrows the claim rather than merely complicating it. A proof is not deficient for lacking incoming telemetry. Neither is a well-evidenced account of the Punic Wars. The argument here is confined to empirical claims about a world that changes — inventory levels, disease prevalence, structural load, grid frequency, vessel position. For those, unresponsiveness to next week's evidence is a defect. For settled and analytic domains, a frozen corpus is not merely adequate but arguably the right tool, and a streaming architecture buys nothing there.
Continuous intake might make things worse, not better. Streams carry drift, injection, correlated error. A curated frozen corpus, audited once carefully, may be better warranted than a firehose nobody has time to validate.
This is the real risk, and it is why provenance, not volume, is the feature doing the work. Exposure without attribution is thrashing — a system revising itself into noise because every arriving byte gets equal weight. The claim under discussion holds only where beliefs are held revisably with recorded origin and confidence, so a bad stream can be down-weighted or cut and everything downstream re-derived. Built badly, continuous intake is worse than a good frozen corpus. Built well, it is corrigible in a way no frozen corpus, however carefully audited once, can ever become again.
The misreading to disown
The weak misreading says continuous ingestion makes a system scientific, or true, or good. It does not. Exposure to future evidence is necessary for an empirical claim to have standing; it guarantees nothing about the reasoning applied to that evidence. A system can ingest every stream on the planet and still infer badly. Nor does this argument declare frozen corpora worthless — for the settled, the linguistic, the historical, they remain the right instrument, arguably the only one.
What the demarcation problem establishes here is narrow and structural, not a verdict on intelligence or a promise about outcomes: for claims about a world still moving, being sealed off from the evidence that would test them is a defect no scale repairs, and there is no permission to observe beyond continuous, provenance-tracked exposure to everything still arriving. That is all it establishes. It is enough to place a ceiling on one axis. It says nothing about what happens on every other axis intelligence might be measured along.