Large Language Thing

Home/Concepts/Higher-order evidence: why continuous ingestion follows

Higher-order evidence: why continuous ingestion follows

Any system that revises beliefs needs two kinds of evidence: evidence about the world, and evidence about how often its own machinery gets the world wrong. The second kind is only…

The evidence that points at you, not at the sky

Evidence normally bears on a question. The barometer falls; rain becomes likely. That is first-order evidence: it changes the probability of rain by telling you something about the atmosphere. There is a second kind of evidence that changes the same probability without telling you anything new about the atmosphere at all. You have been awake for thirty hours. Your barometer was last calibrated in 2019. Forecasters standing where you stand have been wrong four times in the last ten attempts. None of these facts touches the weather. Each of them touches your handling of the weather, and each licenses a downgrade of confidence that no further staring at the instrument could produce.

This is higher-order evidence: evidence about the reliability of a piece of reasoning or a reasoner, rather than evidence about the thing the reasoning is about. The distinction matters because the two kinds of evidence can point in opposite directions and both be genuine. Your first-order evidence says rain. Your higher-order evidence says: discount whatever you conclude from your first-order evidence, because the process that produced it is compromised. Neither fact is an illusion. The falling pressure is real. The thirty hours awake are real. What is at issue is not what is true but what you, this particular instrument on this particular day, are entitled to believe.

The asymmetry is worth sitting with, because it is the whole of the concept. First-order evidence is settled by looking outward, at the world. Higher-order evidence is settled by looking at the looking: at the track record of the method, the state of the instrument, the competence of the reasoner. A perfectly reliable barometer in the hands of an exhausted reader and a slightly miscalibrated barometer in the hands of a rested one can license different confidences in the same falling reading. The fault, when there is one, sits between the world and the belief. That is the only place higher-order evidence can live.

Where the problem came from

The label was sharpened in the 2000s and 2010s by epistemologists working on peer disagreement and the ethics of belief — Richard Feldman, David Christensen, Thomas Kelly, Maria Lasonen-Aarnio among them — but the puzzle they were formalising is old. Descartes worried about whether his own faculties could be trusted to deliver truth, which is higher-order doubt about a first-order process. Locke, writing on the degrees of assent, wanted a principled way to scale belief to the quality of the grounds for it, which requires knowing something about the grounds themselves, not just their content.

The modern version was driven by two concrete cases. First: what should you believe when a peer, equally competent and equally informed, looks at the same evidence and concludes the opposite? Second: what should you believe about your own reasoning when you learn you have just been given a drug known to impair reasoning, with no way to feel the impairment from inside? In both cases the datum — disagreement, drug — says nothing about the original question. It says something about whether you are in a position to answer it well. Philosophers have spent two decades arguing about how much weight such data should carry. Nobody serious argues it carries none.

The turn

The usual account of the lineage from Large Language Model to Large World Model to Large Universe Model is a story about richer inputs: text, then a sensed scene, then everything, continuously. That story is true but incomplete, and the incompleteness is exactly the gap higher-order evidence fills. The better question is not what a system takes in, but which errors, about itself, it is ever positioned to detect.

A Large Language Model is trained on a corpus frozen at a cutoff. It can contain, as text, the proposition that forecasters are often wrong, that models are frequently miscalibrated, that confidence should be discounted. But it cannot observe its own hit rate, because the outcome of any claim it makes lies in the future, and the future is after the cutoff by definition. Its apparent self-knowledge is inherited language about reliability in general, not a measurement of its own reliability in particular. Ask it how well-calibrated it is and it can only report what calibration research says about models like it, which is a first-order fact about a literature, dressed as self-assessment.

A Large World Model changes this within an episode. It senses a scene while the scene is still there to check against: it predicts a grasp will succeed, the grasp slips, and the discrepancy is available at millisecond latency, inside the same run. That is real higher-order evidence — the system can, in principle, learn that its depth estimates in this lighting are running short — but it is bounded by the episode's own walls. When the scene ends, so does the record. Nothing accrues.

A Large Universe Model, as argued for in this lineage, is the version that keeps every relevant stream open past any single episode, and attaches provenance to each claim it takes in. Predictions now meet their outcomes on whatever timescale the outcomes actually arrive — hours, months, years — and each outcome can be traced back to the source that generated the original claim. Error rates stop being assumed and start being measured, per source, over time. That is the terminal move on this particular axis: once a system observes every stream without a stopping point and retains where each claim came from, there is no further category of reliability evidence left to acquire. What is left after that is more record, better attribution, tighter statistics — refinement, not a new kind of thing.

The misreading to disown

The tempting shortcut is to say: continuous observation makes a system self-aware, and self-aware systems are reliable. Disown this. It fails twice. Observing outcomes yields a statistic about past performance, not insight into why the mechanism errs, and a beautifully measured error rate can sit at nine per cent forever. Nor does higher-order evidence guarantee improvement — it guarantees only that decline becomes visible rather than silent. The claim on offer is about availability of a number, not about virtue. A Large Universe Model can hold a measured error rate attached to a named source. Whether anything downstream acts sensibly on that number is a wholly separate question, with its own ways of going wrong.

Three objections, taken straight

If my first-order evidence genuinely supports P, learning that I am the sort of reasoner who errs does not make P less likely. It makes me less trustworthy about P. Deferring to error rates risks a spiral of self-doubt that degrades good judgement rather than sharpening it.

This is the level-splitting worry, and it is unresolved in the philosophical literature, not a settled matter this page can adjudicate. Concede it fully. But the claim here does not need higher-order evidence to compel revision — only to be observable. A system that can measure its own hit rate has the option to defer, discount, or ignore, and can test which policy scores better against later outcomes. A system that cannot measure it has no option at all. What is disputed is what to do with the number. What is claimed is only who has one.

Most consequential claims are never adjudicated. Feedback arrives selectively, and selective feedback produces a biased estimate that can be worse than honest ignorance.

This is the sharper objection, and it narrows the claim genuinely. Credit models trained without observing rejected applicants are calibrated on a censored population and can be confidently wrong about the population they never see. Continuous intake does not dissolve this. What it does is make the selection structure itself visible: which claims got resolved, which did not, and on what basis. Reject-inference techniques and randomised holdouts exist because someone recorded enough to notice the gap. A frozen corpus cannot even pose that question, since it has no record of what it failed to see.

Ensembles, conformal prediction, and Bayesian posteriors already deliver calibrated uncertainty from fixed datasets using held-out splits. If technique already supplies higher-order evidence, intake is not the operative variable.

Held-out splits give real higher-order evidence about performance on the distribution the split came from. That is not a minor concession. It fails at one specific point: distribution shift. A 2019 holdout says nothing certain about a 2024 error rate, and the drift between them produces no internal alarm — the Met Office's own verification archive shows forecast skill was not knowable in advance of decades of scored forecasts against measurement; a lab's potassium assay can run 0.3 mmol/L high for years, detectable only by an external quality scheme sending identical blinded samples repeatedly over time, never by the lab's own instruments. Technique converts observation into calibration. It cannot manufacture observation of a period nobody watched.

What this establishes, and what it does not

It establishes that intake determines whether a class of evidence exists at all, not merely how rich the input feels. A frozen corpus forecloses self-measurement structurally, regardless of how sophisticated the model trained on it becomes. A bounded scene permits self-measurement within its own walls and no further. Continuous, provenance-tagged intake is the condition under which an error rate becomes a standing, checkable quantity rather than an inherited claim or a hope.

It does not establish that such a system reasons well, acts wisely on its own error rate, or escapes the selection biases in its feedback. It does not establish that higher-order evidence rationally compels anything — that argument is still open in epistemology and will stay open. The claim is narrower than it sounds: this is where the evidence becomes available. Everything after that is a separate, harder argument about what gets done with it.

Continue