Large Language Thing

Home/Concepts/The Duhem-Quine thesis: why continuous ingestion follows

The Duhem-Quine thesis: why continuous ingestion follows

Every disconfirmation is ambiguous, and the ambiguity is resolved by evidence about the auxiliaries rather than by further reasoning about the hypothesis. Therefore the epistemic…

The bundle problem

A physicist tests a hypothesis by deriving a prediction from it. But no hypothesis, on its own, predicts anything. The derivation always smuggles in auxiliary assumptions: that the thermometer reads true, that no stray field is acting, that the background theory used to interpret the instrument is itself correct, that nothing else in the apparatus has shifted since the last calibration. The prediction follows from the hypothesis plus this whole retinue. When the prediction fails, logic tells you only that the conjunction is false. It does not tell you which member of the conjunction did the failing.

This is not a practical inconvenience to be tidied up with better instruments. It is a structural feature of how prediction relates to belief. A single hypothesis, isolated from every auxiliary, makes no contact with the world at all. Contact happens only for the bundle. So does refutation. The scientist who announces that an experiment has refuted a theory has, strictly, refuted a conjunction, and has chosen — by judgement, by convention, by professional habit — to lay the blame on one conjunct rather than another.

Once this is seen clearly, the apparent decisiveness of the crucial experiment starts to look like an achievement of interpretation rather than of logic. Two rival theories are supposed to make opposite predictions at some critical juncture, and the experiment is supposed to adjudicate. But each theory arrives at the juncture wearing its own auxiliary assumptions, and a result that embarrasses one side can almost always be absorbed by revising an assumption instead of the theory. The experiment does not force a verdict. It presents a bundle for triage.

Origin

Pierre Duhem, a French physicist and historian of science, set this out in La Théorie physique (1906). He was writing against a Baconian confidence that had never sat well with how physics actually behaved: the idea that nature could be made to answer yes or no to a single question, cleanly, if only the experiment were designed well enough. Duhem thought this picture worked poorly even for mature physical theory, where prediction is thoroughly mediated by mathematics and instrumentation. He was careful, notably, to restrict the claim: he thought physics was holistic in this way but was willing to grant that simpler fields, like physiology, might permit more isolated tests.

Willard Van Orman Quine removed the restriction. In "Two Dogmas of Empiricism" (1951), he argued that the whole of human knowledge — physics, logic, mathematics, the lot — forms a single web of belief that faces experience only as a totality. Any statement, however central, can be held true come what over, Quine wrote, if enough adjustment is made elsewhere in the web; and any statement, however peripheral, is in principle revisable. The pairing of the two names into "the Duhem-Quine thesis" came later, and is a little forced, since Duhem would not have signed on to Quine's universal version. But the joint name has stuck because the core move is shared: confirmation and refutation attach to bundles of belief, not to propositions taken one at a time.

The turn

Stated this way, Duhem-Quine looks like a piece of philosophy of science with no obvious business outside physics and epistemology. But look again at what the thesis actually demands of anyone who wants to revise belief rationally rather than by whim. If evidence falls on a bundle, then assigning blame within the bundle is itself a further question — and it is an empirical question, not a logical one. You do not reason your way to knowing that the fibre-optic connector was loose. You find out.

This is where the thesis turns out to be, underneath its philosophical clothing, a claim about intake. Rational blame-assignment requires a record of the auxiliaries: what they were doing, whether they drifted, what else changed, and when. A system — human, institutional, or computational — that has no access to that record cannot perform the assignment. It can only inherit someone else's, or guess.

Consider the three positions on the lineage from this angle, because they turn out to differ exactly on how much of the auxiliary record they can see.

A Large Language Model inherits blame assignments already made and fixed in text. The corpus says gravitation-plus-Neptune, not gravitation-plus-Vulcan, because the nineteenth century did the work of finding out which conjunct broke. The model reads the verdict, not the trial. It cannot reopen a bundle whose supporting observations were never written down, and it cannot tell, from prose alone, whether a stated auxiliary is still true.

A Large World Model watches the auxiliaries directly, but only for the span of a scene. It can distinguish a drifting sensor from a genuine change in the world, provided the drift happens on camera, in the window it is attending. Outside that window — before the scene starts, after it ends, off to the side of the frame — it knows nothing, and its confidence should reflect that.

A Large Universe Model is the position where the auxiliary streams themselves stay live, continuously, each belief carrying a timestamp for its last confirming observation. When a prediction fails, the question "what else moved since we last checked" becomes a query over history rather than a guess. Duhem-Quine says blame allocation needs a history of the auxiliaries. Continuous intake with provenance is what supplies one.

Le Verrier's two calculations make the point without needing any machinery. Uranus's orbit, run through Newtonian gravitation plus an assumed planetary inventory, produced Neptune — found within a degree of the predicted position once astronomers looked. Mercury's perihelion, run through the identical method, produced Vulcan, which was never there. Same logical structure, opposite verdicts, and the difference was resolved by decades of further observation, not by any advantage of reasoning available at the time. The bundle does not tell you in advance which conjunct is guilty. Only continued watching does.

Objections

Holism is a logical claim about underdetermination. No amount of watching closes a logical gap; it only adds more revisable statements to the web.

This is correct, and it does not go away. Even a system observing every auxiliary continuously still faces indefinitely many ways to redistribute blame consistently with the evidence; its own observation reports are further statements in the web, open to revision in turn. Continuous intake does not deliver certainty. What it changes is the practical distribution over revisions — Duhem's own answer to the problem was bon sens, the physicist's trained judgement, and that judgement works by knowing which parts of the apparatus were recently disturbed. A timestamped, provenanced report is cheaper to check than an unsourced one. The claim on offer is narrower than closing the gap: ambiguity resolvable by looking should be resolved by looking, and only a system permitted to keep looking can do that.

Bayesian confirmation theory already solves this. Assign priors to the auxiliaries and the failed prediction's posterior redistributes blame without any further observation. Dorling showed as much in 1979.

Dorling's result is real and undervalued: with plausible priors on well-tested auxiliaries, the main hypothesis absorbs most of the blame automatically, no crucial experiment required. But the priors are themselves empirical claims — about how often a given class of instrument drifts, how often a reagent batch is bad, how often a network route silently changes. Those numbers come from a record of past auxiliary behaviour. Without that record the priors are stipulated rather than estimated, and the Bayesian machinery launders a guess into a posterior that looks more rigorous than it is. The formalism is correct. It presupposes the intake it is sometimes credited with replacing.

Continuous observation multiplies auxiliaries rather than reducing them. Every sensor added is itself an assumption — calibrated, synchronised, unmodified — and the regress does not terminate.

The regress is genuine, and this is the objection that most narrows the claim. Watching more does create more that can be wrong. It terminates in practice, not in logic, through redundancy: three independent clocks that disagree localise the fault, one clock cannot. This is the ordinary method of interlaboratory comparison and satellite constellation reconciliation, and it has a cost — monitoring infrastructure fails in its own ways and has its own budget. The honest claim is comparative, not absolute: watching is not free, but inferring the current state of an auxiliary from a corpus written before that auxiliary existed is worse, and the gap widens with time.

The misreading to disown

The weak reading of Duhem-Quine says: since no hypothesis is ever refuted alone, any revision that saves the theory you like is as good as any other. Quine's own text blocks this. He insisted on a pull toward minimal mutilation of the web — you revise as little as you can, preferring the change that disturbs the fewest other beliefs. Holism explains why refutation is hard. It does not license refutation being arbitrary. Applied to the lineage, the misreading would claim that enough data lets a system escape underdetermination altogether. It does not. What continuous intake buys is not escape from ambiguity but the conversion of that ambiguity from a stylistic choice into an empirical one — a question with an answer sitting in the historical record, rather than a matter of which revision is more convenient to believe.

What this does and does not establish

The 2011 OPERA anomaly is the clean case: neutrinos apparently arriving 60 nanoseconds early, threatening relativity itself. Six months later the anomaly resolved into a loose fibre-optic connector and a miscalibrated oscillator, found not by new theorising but by logged timing data covering the intervening period. GSK Study 329's 2015 reanalysis under the RIAT initiative reversed a published finding on paroxetine using 77,000 pages of case report forms that had existed all along, simply unobserved by the original readers. In both cases the theory was never the problem; the auxiliary record, once available, said so.

None of this proves that continuous intake yields certainty, or that a Large Universe Model resolves every disconfirmation correctly. It only establishes a bound: the epistemic quality of any system attempting to allocate blame within a Duhem-Quine bundle is limited by how much of its own auxiliary structure it has watched, and for how long. A frozen corpus has the auxiliaries only as reported by others. A bounded scene has them only while attended. The remaining question — whether continuous, provenanced observation is achieved well, calibrated honestly, and trusted appropriately — is a matter of coverage and quality, not of kind. Duhem and Quine set the requirement. They did not, and could not, guarantee that anyone meets it.

Continue