Home/Concepts/The block universe and the problem of change in rail operations
The block universe and the problem of change in rail operations
Grade intake by how much of the block is admitted, and the ladder has three rungs and then stops. One frozen slab of the past. One moving sliver of the present. Every stream,…
Three ways to know a railway
A Large Language Model reading a railway would hold a timetable frozen at publication, an engineering standards manual as it stood on the day it was scanned, and no way of knowing that either has since been revised. A Large World Model would hold what one signal box camera or one train's onboard sensor suite can see for as long as the shift lasts, and nothing before or after. A Large Universe Model holds the network as it is now streaming: track circuits, axle telemetry, weather, possessions, all still running, all subject to correction. Rail operations is a useful test case precisely because the third position is not aspirational there. Something like it already has to exist, informally, in the head of a network controller, because the alternative is a railway that only reacts to damage already done.
| generation | what it holds | temporal shape |
|---|---|---|
| Large Language Model | manuals, rulebooks, incident reports as archived | a slab, sealed at a cutoff |
| Large World Model | one signal box's current view, one train's current sensors | a thin present, closes when the shift ends |
| Large Universe Model | track circuits, telemetry, weather, possessions, continuously | an open interval, nothing closes it |
What arrives
A busy mixed-traffic route generates several distinct streams, none of which waits for the others. Track circuits report section occupancy by interruption, typically refreshed every two to five seconds per section, and there are hundreds of sections on a control area of any size. Rolling stock telemetry — axle bearing temperature, brake pipe pressure, wheel load — arrives from onboard systems at roughly 1 hertz when the train is in a coverage zone, and from fixed trackside detectors such as Wheel Impact Load Detectors and Hot Axle Box Detectors at fixed points, commonly spaced 30 to 50 kilometres apart on a heavy freight corridor. Weather arrives as hourly feeds from met stations plus continuous rail-temperature sensors embedded in the ballast on routes prone to buckling. Maintenance and possession data arrive as planned windows set days ahead and revised, sometimes hourly, as work overruns or track access is withdrawn.
None of these streams is a snapshot of the network. Each is a partial, asynchronous trickle, and each has its own latency and its own failure mode. A track circuit tells you a section is occupied; it does not tell you by what, or how heavily loaded, or how fast. A HABD tells you one axle at one location and one moment ran hot; it says nothing about whether that reading is rising, falling, or a sensor fault. Any one of them, read alone and discarded once read, is close to useless. That discarding is exactly what a scene-bound system does — it is the Large World Model's characteristic loss, applied here to a railway rather than a room.
What is held
What a Large Universe Model architecture requires, and what the informal practice of an experienced controller already approximates, is that each reading becomes a belief with a timestamp, a source, and a decay rate, rather than a fact that is simply true from now on. A rail temperature reading of 32°C taken at 14:02 on a stretch of jointless track is not "the temperature" — it is a belief about buckling risk that decays as new readings supersede it and as the diurnal cycle moves on. A HABD reading of 68°C on axle three of a freight service is not "that axle is fine" or "that axle is failing" — it is a data point that must be held alongside the previous HABD reading from the site 40 kilometres back, because the trend, not the value, is what indicates a developing defect.
This is where provenance stops being an abstraction and becomes an operational necessity. A speed restriction issued on the strength of a single detector reading is issued on weaker grounds than one issued on the strength of a confirmed trend across two sites, and the controller needs to know which is which. A belief carried with provenance says: this figure came from WILD site 14, at 14:02, and the previous reading from site 9 at 13:31 was 51°C. A belief without provenance is just a number on a screen, indistinguishable from noise, indistinguishable from a fault in the sensor itself.
What triggers revision
Revision happens when a new observation contradicts, confirms or exceeds a threshold set against a held belief, and the interesting cases are the ones where no single reading crosses any threshold on its own. A wheel with a developing flat spot produces impact loads that rise gradually across several WILD sites over hundreds of kilometres; each individual reading may sit comfortably under the trigger value that would generate an automatic alarm. The defect is visible only across the interval, as a slope, not at any instant, as a value. This is the block universe's central claim about change arriving on the shop floor: an instantaneous reading contains no trend, no rate, no direction, and a system that discards each reading after acting on it can never see the slope at all.
The same holds for rail-temperature stress. A single sensor reading tells you today's peak. A belief that has been carried across the preceding fortnight, with decay weighting recent days more heavily, tells you whether today's peak sits on a rising run that raises buckling probability well beyond what the instantaneous figure alone would suggest.
What the controller sees
>The controller doesn't need philosophy. She needs the screen to tell her which alarm to trust and which one is yesterday's sensor drift dressed up as today's emergency.
That is the right complaint, and the answer to it is that provenance and decay are exactly what let a screen make that distinction rather than presenting every alarm as equally urgent. Without them, the controller sees an undifferentiated stream: a HABD flag here, a track circuit anomaly there, a possession running long somewhere else, each demanding attention with no indication of which is a single stale outlier and which is the third confirmation of a worsening trend. With provenance and decay in place, what she sees instead is a ranked picture — this reading is fresh and corroborated, that one is fresh but isolated, this one is three hours old and has already been superseded twice. The difference is not cosmetic. It is the difference between a screen that reports and a screen that has already done the work of deciding what still deserves belief.
What it costs when the loop fails
The characteristic failure on this axis is a restriction applied after the defect has already propagated rather than when it first appeared. Concretely: a wheel flat is first detectable, marginally, at WILD site A. It does not trigger an automatic threshold there. It shows up more strongly at site B, 40 kilometres on, still under threshold. By site C it crosses the line, an alarm fires, and a speed restriction is applied — but the wheel has now run at speed and under load for over 100 kilometres longer than it needed to, the track has taken additional impact loading at every unrestricted crossing and switch in between, and the margin between "restrict now" and "the axle fails in service" has narrowed considerably. The system did not fail to detect anything. It failed to carry the first two readings forward as a belief that could accumulate into a trend. Each reading was true and each was, on its own, insufficient, and insufficiency compounded across three sites because nothing held the interval together.
The two objections worth taking seriously
The first: track circuits sample discretely, telemetry arrives in packets, nothing is truly continuous, so calling this "every stream, continuously" is overstated. That is correct as a description of the hardware and irrelevant to the claim. The distinguishing feature is not sample rate but the absence of a closing bracket. A WILD reading that is discarded once acted on behaves like a Large World Model's scene, however fast the sample rate; a WILD reading carried forward with a timestamp, ready to be reconciled against the next one from the same axle, is doing something a fast snapshot never does. Continuity of commitment, not continuity of signal, is the property in play.
The second, stronger, objection: modern interlocking and automatic train protection systems already act from a state vector at an instant, and where that state is a sufficient statistic of history, nothing earlier needs to be retained — this is the whole point of a Markov formulation, and it works. It works for the problem it was built for: safe separation, given that the state variables were chosen correctly. The trend-detection problem is different in kind. Choosing which variables suffice, and noticing when a defect's signature lies in a slope no instantaneous state vector was built to represent, is precisely the work that requires an interval to have been watched. A compressed state also drops provenance: it cannot say which detector reported what, or that an earlier reading has since been found unreliable and must be withdrawn from the trend. Markov efficiency and revisable belief are solving different problems, and rail safety needs both — one for the split-second protection layer, the other for the slower, cumulative layer where wheels, bearings and rails actually fail.