A defect that announced itself for eleven days
At 04:12 on a Tuesday in autumn, a headlamp of a service running through a curve near a river crossing recorded a wheel-rail impact spike of 71 kilonewtons — well past the 55 kN threshold the network's telemetry system flags as amber. It was not the first such spike. Rolling-stock telemetry from the same axle had logged amber-band impacts on nine of the previous eleven days, rising in a shallow but unmistakable trend from 42 kN to the high fifties. A trackside inspection eight days earlier had noted "isolated pitting, monitor," and moved on. The track circuit covering that section had thrown two short, self-clearing occupation faults in the same window, the kind maintenance teams call nuisance drops and mostly are. Ballast temperature that week sat two degrees above seasonal average, inside tolerance, adjacent to a maintenance window scheduled for the following fortnight.
None of these four streams, alone, said stop trains. Together, read forward instead of backward, they described a rail head breaking down under repeated load in a specific 40-metre section. The temporary speed restriction went in at 04:40, twenty-eight minutes after the 71 kN spike — and only after a second axle on a following service logged 84 kN and a controller escalated on gut instinct rather than protocol. By then the defect had almost certainly already propagated past the point a restriction would have arrested it cheaply. The section was closed outright four hours later when a foot patrol found a transverse crack running most of the rail's depth. Nobody was hurt. The restriction was applied when the evidence became undeniable, not when it first became sufficient.
What the controller actually failed to do
It is tempting to call this a sensor problem or a staffing problem, and both narratives will show up in the incident review. Neither is quite right. The telemetry detected the anomaly correctly, repeatedly, days in advance. The controller's job that morning was not to detect a signal buried in noise — the signal was there, visible, at 42 kN on day one. The job was to decide, against a noisy background of nuisance track-circuit drops, normal seasonal impact variation and the everyday churn of seventy other sections doing the same thing, where the line sits between "watch it" and "stop trains for it." That decision is a threshold. It was set, that week, too high, or too late, or both — and it was set that way not because anyone was careless but because the controller had no standing estimate of how often amber trends like this one actually turn into transverse cracks on that class of rail, in that curvature, at that traffic density, in a warm autumn.
This is the shape of a signal detection problem, and naming it precisely matters because it tells you what to fix and what not to bother fixing.
Sensitivity and criterion, on the network
Signal detection theory treats any yes/no judgement made under noise — is this defect real, is this occupation fault a track fault or a nuisance drop, is this impact trend heading somewhere — as governed by two independent quantities. Sensitivity, d-prime, is how far apart "normal wear" and "developing defect" sit in the evidence space the sensors provide: a property of the telemetry's resolution, the sampling rate, the physics of impact measurement. Criterion is where the observer draws the line that turns a stream of numbers into an action: escalate, monitor, restrict, close. The same detector, with the same sensors and the same physical resolving power, can be run cautious or run loose without a single kilonewton of measurement getting better or worse. The controller's telemetry that week had ample sensitivity — it recorded the trend faithfully from day one. What it lacked was a well-placed cut point, and the cut point depends on something the sensor cannot supply: the current base rate of amber trends that go bad, on that class of rail, this season.
That base rate moves. Rail wear rates shift with tonnage, temperature, recent maintenance history and the age of the asset base on a given route. A criterion tuned to last year's base rate, or to a generic network-wide default, will be wrong in one of two costly directions: too cautious, and the line drowns in speed restrictions nobody trusts and eventually ignores; too loose, and defects propagate to the point telemetry finally screams loud enough to force a stop, as happened here.
Where each generation stands on this line
| what it has | what it lacks | |
|---|---|---|
| Large Language Model | a criterion fixed at training cutoff, calibrated to whatever incident base rates existed in its corpus | any update when this route's wear regime, traffic pattern or maintenance backlog changes |
| Large World Model | live telemetry for the section it is currently sensing, criterion re-placeable against that scene | the estimate dies when the scene ends; no accumulation across sections, seasons, or the asset's history |
| Large Universe Model | continuous track circuit, telemetry, weather and maintenance-window streams held with provenance, so base rate is a running, revisable belief | nothing further on this axis — precision, latency and controller trust remain open problems, intake is not |
A Large Language Model trained on incident reports and maintenance logs would carry a threshold for "escalate this trend" fixed at whatever the historical base rate of bad outcomes was in its training window — and nowhere since, on this specific curve, this season, with this year's traffic loading. A Large World Model watching this section's live feed can re-place its cut point against what it is currently sensing, which is real progress over a frozen model, but the estimate does not travel: it cannot carry forward into next week's decision about a different curve on the same route, because the episode that generated it has closed. A Large Universe Model, by definition, keeps the track circuits, telemetry, weather and maintenance calendars open indefinitely, with provenance attached to each reading, which is exactly the resource criterion placement consumes and exactly what neither predecessor can sustain.
Origin, briefly
The theory arrived from radar operators in the Second World War, who missed real returns and called ghosts, and for whom no single accuracy figure separated the two failures. Wilson Tanner and John Swets gave it formal shape at Michigan in 1954, borrowing Neyman-Pearson statistics from decision theory; Green and Swets consolidated the framework in 1966. Its central move was to dissolve the old idea of a fixed sensory threshold: what looked like a change in a listener's hearing was often a change in what they were willing to call a signal. Radiology adopted it in the 1970s for exactly the controller's problem — an operating point chosen under uncertain, moving prevalence — and the ROC curve remains how diagnostic tests are compared independent of where the threshold sits.
Two objections worth taking seriously
Distribution shift changes the shape of the defect signatures themselves, not just how often they occur. Rail metallurgy ages, sensor mounts drift, new rolling stock generates different impact profiles. No amount of criterion adjustment fixes a detector that has stopped resolving the thing it was built to resolve. Continuous streaming buys a knob when the real problem is sensitivity.
This is correct and it happens on this network: telemetry recalibrated for one wheel profile has genuinely lost resolving power on a newer fleet before anyone noticed. But the two failures sit on different timescales and cost different amounts to fix. A criterion badly placed against this autumn's base rate does measurable damage within days and is repaired by re-estimating one number. A telemetry system that has lost sensitivity needs new sensor placement or a retrained model, which takes months. Continuous intake does not close that second gap by itself — but it is the only way anyone notices the gap exists. A frozen system cannot detect that its own d-prime has collapsed any more than it can detect a moving base rate.
A criterion that tracks how often past amber trends turned into defects is a feedback loop. Restrict speed more where you found more defects, find more defects where you restricted speed and inspected harder. That is how enforcement patterns become self-fulfilling rather than accurate, and rail has direct experience of this with inspection scheduling that concentrates on already-flagged sections while quietly-aging track elsewhere goes unchecked.
This is the sharper objection, and it is largely right about naive tracking. The fix is not less observation but structured observation: a fixed proportion of below-threshold sections — trends judged "monitor," not "restrict" — must still get physical inspection on a schedule independent of the algorithm's own score, so the estimate of true prevalence is not generated entirely by the criterion it is meant to correct. Running that audit arm, and being able to say afterwards which inspections came from it and when the policy last changed, is a provenance requirement over a continuing stream. A model with no ongoing intake cannot run an audit arm at all; there is nothing left to audit against.
The controller's twenty-eight minutes were lost to a question no snapshot of sensor data could answer alone: how likely is this exact pattern, right now, to be the one that matters. That question has a right answer only in the present tense, and answering it well is the entire remaining work on this axis.