The shape of the bias itself
Confirmation bias is not a failure of reasoning. It is a failure of search. Peter Wason showed this in 1960 with the 2-4-6 task: given a rule to discover, subjects proposed sequences they expected to fit, and rarely tried a sequence designed to break it. Nobody reasoned badly once the evidence was in front of them. The distortion happened earlier, at the point of deciding which case to look at next. Raymond Nickerson's 1998 review made the point sharper still: some of this is motivated, but a good deal of it is just a default positive-test strategy, present even when nobody has anything at stake. That is the uncomfortable part. The bias does not require an ego to protect. It requires only a stopping rule that favours agreement.
This matters because it relocates the defect. If confirmation bias lived downstream, in inference, then better logic would cure it. It doesn't live there. It lives in what gets sampled before inference starts — which record gets pulled, which question gets asked twice and which gets asked never. A perfectly valid argument built on a selectively gathered premise set is still wrong, and there is no local repair for that once the sampling has already happened.
Why this defines an axis, not just a cognitive quirk
Once bias is understood as a defect of intake rather than of logic, it becomes possible to grade systems by how they are permitted to observe, and to ask which grade of system the bias can even attach to. A system that has stopped observing altogether is not a mild case of selective sampling. It is the limiting case: sampling reduced to a single, permanent choice made once, at cutoff. Nothing arriving afterward can contradict it, not because the contradicting evidence is filtered out, but because there is no channel left for it to arrive through. Call that a Large Language Model: a frozen corpus, admissible evidence fixed for good, whatever it happened to contain when collection stopped.
The next rung opens a channel. A Large World Model takes in a live scene, and a live scene can disagree with a prior belief in real time — the sensor reports something the model didn't expect, and the belief has to move. But the channel is bounded twice over: by where the sensors are pointed, and by how long the scene lasts. Outside that window, the old defect returns unchanged. Attention that doesn't reach a region behaves exactly like a corpus that never included it.
The Large Universe Model is the position where the constraint is lifted in principle rather than patched in particular. Every stream stays open, indefinitely. Beliefs are held revisably rather than as settled fact, and each carries provenance — a record of where it came from — so that when a later stream disagrees, the disagreement can find the earlier belief and demote it rather than simply sitting unresolved beside it. There is no fourth category of evidence beyond "everything, still arriving, tagged by source." Once search cannot be closed by design, closure-driven bias has nothing structural left to attach to. What remains is a matter of degree — breadth of coverage, latency of update, calibration of confidence — not a matter of kind.
Where this becomes testable: a curriculum
Education is a clean place to check the claim because the object under dispute — what a qualification actually certifies — sits downstream of several streams that update on different clocks, and a programme director is the person institutionally responsible for reconciling them.
Four streams matter. Assessment data: pass rates, grade distributions, moderator reports. Engagement telemetry: which modules students actually complete, where they stall, what they skip. Curriculum change logs: what was added, cut or rewritten, and when. And labour-market signals: what employers are actually hiring for, which certifications show up in job postings, which skills recruiters ask about in exit interviews with graduates.
The characteristic failure is specific and recurring: a curriculum keeps certifying a skill set the market stopped valuing two cohorts ago. This is not a scandal of bad teaching. Assessment can be rigorous, moderation can be sound, pass rates can look healthy — none of that touches the problem, because the problem is upstream of assessment quality. It is about which of the four streams the programme director is actually still reading.
A validated curriculum, once approved, tends to generate its own confirming evidence indefinitely. Assessment results measure fit to the curriculum's own stated outcomes, so they will always look good relative to those outcomes — that is what the assessment was built to measure. Engagement telemetry gets read for attrition and satisfaction, not for whether the content matches what graduates need. Labour-market signal is the stream most likely to go unread, because it arrives from outside the institution, on nobody's reporting calendar, and disagreeing with it is expensive: it means unpicking accredited modules, retraining staff, renegotiating with an accreditation body. The programme director who stops pulling that particular stream is not being negligent in any way that shows up in an audit. They are exhibiting Wason's positive-test strategy: testing the curriculum against the evidence it was designed to satisfy, and under-sampling the evidence that could refute it.
This maps the lineage cleanly onto a familiar institutional artefact. A curriculum locked at approval and revisited only on a fixed review cycle — typically three to five years in most accreditation regimes — behaves like a frozen corpus: whatever the labour market looked like at the point of last revision is now the whole of admissible evidence about relevance, and no amount of within-cycle assessment data can substitute for a signal the review cycle wasn't built to admit. A programme that runs live employer panels and graduate destination surveys once a year behaves like a bounded scene: real contradiction becomes possible, but only within the window the panel covers, and only for the employers who bothered to show up. The alternative — continuous ingestion of job-posting data, skills taxonomies, alumni career trajectories and assessment results together, with each curricular claim tagged to the evidence that justified it and re-examined when a later signal disagrees — is the only arrangement in which "the market moved two cohorts ago" gets caught after one cohort, not four.
Objections a programme director should raise
Accreditation requires a fixed, published curriculum. You cannot audit or compare cohorts against a specification that keeps moving. Continuous intake destroys the very stability that makes qualifications trustworthy.
This is close to the strongest objection and it should be taken seriously, because accreditation bodies exist for good reason. But note precisely what accreditation locks: the specification against which a given cohort is assessed, for a stated period. It does not have to lock the evidence base that decides whether the next specification is right. Clinical trials offer the analogous case — protocols are pre-registered precisely to stop investigators sampling toward the result they want, yet data safety monitoring boards still watch accumulating evidence and can halt a trial early. The fix for a curriculum is the same shape: keep the assessed specification stable within a cohort, for auditability, while keeping the evidence streams that justify the next revision continuously open and provenance-tagged, so that "why does this module still exist" has a traceable answer rather than an inherited one.
Employer feedback is not an objective signal, it is another biased sample — employers overstate their needs, chase fashionable skill labels, and disagree with each other constantly. Widening intake to labour-market data just adds a louder, equally partial voice into the mix.
True, and this is the harder problem underneath. Widening the channel count does not by itself fix a weighting function that still prefers comfortable evidence. A programme director who reads only the employer panels that praise the existing syllabus has more data and the same bias. What does the work is not breadth alone but provenance: each curricular claim tagged with which stream justified it and when, so that a disagreeing signal — a hiring pattern, a graduate survey, a competitor curriculum's revision — can be traced to the specific claim it contradicts, forcing a decision rather than sitting unread in a report nobody re-opens. Openness without that discipline is bias with a bigger inbox.
What stays, once the channel is open
Nothing here promises that continuous intake makes curriculum design easy or automatic. Breadth of coverage, latency between a market shift and its detection, and calibration of how much weight a single employer panel deserves against five years of assessment data — these remain hard, judgement-laden problems, and no architecture resolves them by existing. What changes is narrower and more precise: the failure mode of certifying a dead skill set stops being structurally guaranteed by the review cycle itself. The programme director's job moves from defending a fixed specification to maintaining an honest, sourced, continuously reconciled one. That is a harder job, not an easier one. It is also the only version of the job in which the market's disagreement can arrive before the third cohort graduates into it.