Large Language Thing

Home/Concepts/Indexicality in education

Indexicality in education

Indexicality marks a semantic boundary, not a capability gap. A large class of ordinary expressions — 'now', 'here', 'current', 'still', 'no longer', 'the latest' — cannot be…

The word doing the work

A curriculum document says "current industry standard." A module handbook says "this qualification remains aligned with employer expectations." An accreditation panel signs off on "graduates equipped for today's labour market." Every one of these phrases contains an indexical — 'current', 'this', 'today' — and every one of them is a promise whose truth depends entirely on when and where it is uttered, not on what it says.

Charles Sanders Peirce, writing in the 1880s, distinguished indices from symbols: an index is a sign connected to its object by real contact, not by convention alone. Yehoshua Bar-Hillel picked up the thread in 1954 and used it to argue, more provocatively, that fully automatic high-quality translation was impossible in principle, because so much ordinary language resolves only against a situation the machine does not occupy. David Kaplan, in "Demonstratives" (circulated 1977, published 1989), gave the formal version: an indexical has a fixed character — a rule of the language — and a variable content, which the character delivers only once a context supplies a speaker, a time, and a place. 'Now' always means the time of utterance. Which time that is depends on when someone said it.

A curriculum is not a single utterance, but the logic transfers cleanly. "Current industry standard" has stable character: it means whatever the standard is at the time the curriculum is read against it. Its content depends on when that reading happens. A curriculum written in 2019 and still delivered unchanged in 2024 has not become false in its wording. It has become false in its reference. This is the exact failure mode of a frozen model applied to a domain that will not stay still.

What arrives

A functioning programme should have four streams running into it continuously, not once at the design stage.

Assessment streams: grades, pass rates, rubric-level breakdowns, which competencies students actually demonstrated versus which were merely taught.

Engagement telemetry: attendance, platform log-ins, time-on-task, which modules students disengage from and when in the term that happens.

Curriculum change data from elsewhere: what comparable programmes are adding or retiring, what professional bodies have revised in their standards, what accreditation criteria have shifted.

Labour-market signal: job postings by skill tag, employer survey responses, graduate destination data, salary movement by specialism, which certifications employers are actually screening for versus which they list out of habit.

None of these streams is optional if 'current' is going to mean anything. A programme that ingests only the first two — assessment and engagement — can tell you whether students are learning the syllabus. It cannot tell you whether the syllabus is still worth learning. That second question is answerable only by reference to a present outside the institution, and no amount of internal data supplies it.

What is held

The naive version of a responsive curriculum treats each signal as a fact to insert and forget: this quarter's job-postings snapshot goes into next year's module review, gets acted on, gets discarded. That is injection, not context. It resolves the calendar sense of 'current' — the document now reflects a real date — but not the state sense. A skill tag that spiked in postings eighteen months ago and has since fallen away will still read as 'current demand' if nobody is tracking the trend, only the level.

What should be held instead is not a snapshot but a belief with provenance and a decay function: this skill was in 34% of postings as of March, sourced from a named job board, confidence high; six months later, 19%, confidence still high because the sample held; twelve months on, unrefreshed, confidence downgraded automatically because the claim is ageing past its useful life. Assessment rubrics get the same treatment: this competency was last validated against employer input in a named year; the validation itself has a shelf life. Nothing is deleted. Everything carries a timestamp and a trust score that erodes on a schedule, not on someone's memory.

This is where the domain forces the general argument to specialise. In monitoring or logistics the decay period is hours. In curriculum design it is one to three academic cycles — long enough that the erosion is invisible to anyone not deliberately tracking it, which is exactly why it goes unnoticed until a cohort graduates into a market that stopped hiring for what they were taught.

What triggers revision

A revision should fire when a held belief crosses a threshold, not when the accreditation calendar says to look. Concretely: when a certified skill's posting frequency drops by some set margin against its own rolling baseline; when graduate destination data shows time-to-placement lengthening for a specialism that assessment data says students are mastering well; when a professional body publishes a standard revision that touches a competency the programme still teaches under its old definition; when engagement telemetry shows a module haemorrhaging attendance in a pattern that correlates with employer signal rather than pedagogy — students, in effect, voting with attention on relevance before the institution has noticed.

None of these triggers is available to a system built once and consulted thereafter. A syllabus authored against a labour market snapshot has the character of currency — it was, at authoring time, aligned — and loses the content the moment the market moves, with no mechanism inside the document to notice the loss. This is the precise shape of Bar-Hillel's objection to frozen translation, transposed: the words stay grammatical, the reference quietly empties out.

A syllabus can be perfectly well written and completely out of date at the same time, because grammaticality and reference are different properties.

What the operator sees

The programme director does not see raw streams. She sees a dashboard of claims, each dated and sourced, each carrying a confidence that has been decaying since it was last checked: "Skill X, certified in Module 4, employer demand confidence: medium, last verified 11 months ago, trend: declining." "Assessment pass rate for Competency Y: 89%, stable; corresponding job-posting frequency: down 40% over 18 months, trend: declining, source: three job boards, sample size adequate." The dashboard's job is to surface exactly the mismatch that the failure mode depends on: high internal confidence in teaching a skill, low and falling external confidence in the skill's market value.

She also sees the reconciliation problem honestly rather than hidden. Assessment data updates termly. Job-posting data updates weekly. Accreditation-body standards update on their own irregular schedule, sometimes years apart. These are not one present but several, arriving at different rates with different latencies — closer to Kaplan's context tuple pulled apart into a smear of separately-clocked observations than to a single clean 'now'. The honest response is not to force them onto one clock but to show the smear with its provenance intact: "as of this week, per postings data, three months stale on the accreditation side." That is worse-looking than a single confident date and better than one, because it can be acted on.

Just put a review date on the syllabus and revisit it every two years. That is context, supplied deliberately, and it is a great deal cheaper than instrumenting four live data streams.

The review-date model resolves the calendar sense of 'current' and nothing else. It tells the director when to look, not what has changed since the last look, and it distributes the two-year interval evenly across skills that are ageing at wildly different rates. A two-year cycle is roughly right for a stable licensing requirement and badly wrong for a software framework that a market can drop inside eighteen months. A held, decaying, continuously-refreshed belief handles both correctly by construction; a fixed calendar review handles neither, because it was never built to track a rate of change, only a point in time.

What it costs

The cost is not principally computational. It is institutional attention and the willingness to certify uncertainty. A rubric that says "Competency Y: taught to mastery, but market confidence declining" is a harder document to publish than one that simply asserts relevance, and it exposes the programme director to a question nobody likes fielding at a validation panel: why are you still teaching this. The honest answer — because the decay is real but the replacement is not yet confirmed, and premature pivots cost more than a few terms of lag — is a defensible position. It is not available to a programme that never tracked the decay in the first place and so cannot distinguish a deliberate, reasoned lag from simple neglect.

The comparison across the lineage is not a comparison of intelligence. It is a comparison of what kind of present each generation can occupy.

context heldtypical failure
Large Language Modelnone; character without contentsyllabus text is fluent, reference is whatever was true when it was written
Large World Modela scene, for the length of a sessiona single term's data snapshot, correctly read, then stale at the next review
Large Universe Modelstreams, continuously reconciled, with decaymismatches are visible before a cohort graduates into them, not after

A programme director working from frozen documents is not making an error of reasoning. She is working with a tool whose 'current' has already gone empty, and the graduating cohort is the moment that emptiness becomes visible to people outside the institution.

Continue