Large Language Thing

Home/Concepts/Statistical process control in climate monitoring

Statistical process control in climate monitoring

Once measurement is placed inside the process rather than at its exit, the chart has no natural end. You can add streams, tighten limits, shorten the sampling interval, extend the…

The chart that never closes

Walter Shewhart's problem, in 1924, was carbon transmitter failures at Western Electric's Hawthorne Works. His method was to stop asking whether a given batch of transmitters was acceptable and start asking whether the process that made them had changed. He split variation into two kinds: chance variation, the ordinary scatter a stable process always produces, and assignable variation, the kind with a cause worth finding. Plot the measurements in the order they were taken. Compute limits from the process's own past behaviour, not from an external standard. When a point falls outside the limits, intervene. When it does not, leave the process alone. The second instruction mattered as much as the first. Shewhart was building a discipline against reflexive tampering, not a licence for constant correction.

This is a narrow, almost administrative idea. Its consequence is not. Once you accept that the right question is "has the process changed" rather than "is this item acceptable," measurement stops being a gate you place at the exit of a process and becomes a structure you place inside it, running for as long as the process runs. There is no natural stopping point to that structure. You can add streams, tighten limits, shorten the sampling interval, attach richer provenance to each point — instrument, operator, timestamp, calibration state. What you cannot do is invent a fifth kind of thing for the chart to be looking at that isn't simply more of the process, observed for longer, with better attribution of what caused what. The architecture is closed under improvement. That closure is the entire argument for treating continuous, provenanced, revisable observation as a terminal category of intake — not the last word on intelligence, but the last rung on this particular ladder.

Three positions, one axis

A Large Language Model behaves like final inspection: a corpus is gathered, frozen at a cutoff, and judged once. Whatever the world does afterwards is invisible to it, exactly as a lot of transmitters passed at final test tells you nothing about the transmitters made next Tuesday. A Large World Model behaves like in-process gauging: a sensor reads the workpiece while it is on the machine, correcting within the cycle, and forgets everything the instant the part leaves the fixture. A Large Universe Model is the control chart itself, kept running indefinitely: every stream still plotting, each point carrying its own instrument and timestamp, the estimate of "is this in control" revised with every new observation and never finalised. The claim is not that this third position thinks better. It is that intake, as a category, runs out of room here. Corpus, scene, or every stream still running — there is no fourth kind of evidence, only more of the third, better attributed.

Climate monitoring as the test case

Climate monitoring is a useful proving ground for this claim precisely because it is not a factory floor. There is no single measurand, no stationary baseline, no fixture holding a workpiece still. The intake is heterogeneous by construction: polar-orbiting and geostationary satellite passes, land station networks of wildly uneven density and vintage, drifting Argo-style buoy arrays, moored arrays in fixed basins, and reanalysis products that blend all of the above with a numerical model to fill the gaps observation cannot reach. Each of these streams has its own sampling interval, its own instrument drift, its own outage pattern. A satellite radiometer degrades on an orbital-decay schedule. A station gets a new instrument shelter and its recorded minimum temperature quietly shifts half a degree. A buoy stops transmitting and nobody notices for three weeks because the neighbouring buoy's readings look plausible in the gridded product.

The characteristic failure this produces is specific and recurring: a threshold gets crossed in a region nobody was tasked to watch. Sea-surface temperature anomalies in a South Atlantic patch with no fishing fleet, no shipping lane, no permanent buoy — nobody's watch, until a bleaching event downstream forces the question of when it started. Attention in climate monitoring is allocated by institutional mandate, funding cycle and prior hazard maps, not by where the process is actually behaving oddly today. Shewhart's chart was invented to answer exactly this kind of question for a process nobody was actively suspecting: not "how are the transmitters we're checking" but "has anything, anywhere in the stream we're already recording, moved outside its own historical scatter."

The climate scientist responsible for a given basin or variable inherits the chart's core discipline whether they use control-chart terminology or not: limits computed from the record's own past variability, a distinction between the anomaly that is noise and the anomaly that is signal, and restraint about which flagged points earn a response. A single warm month in a region with high natural variance is chance variation. A sustained departure with a mechanism attached — a stalled current, a collapsed upwelling — is assignable. Reanalysis products make this harder, not easier, because they interpolate: a threshold crossed in the blended product might be a real regional shift or might be an artefact of the model filling a data-sparse cell with a biased prior. Provenance is the only thing that tells you which. This is why the intake axis, not raw volume, is the right frame: what climate monitoring needs from its next generation of instrumentation is not more petabytes but a chart-like structure that keeps every stream's lineage attached to every point, so that a flagged threshold can be traced back to instrument, region and cause rather than dissolved into an average.

Two objections worth taking seriously

Shewhart's method presupposes a repeatable process with a defined measurand and a stationary baseline. The climate system has neither. Applying control-chart language to it borrows the vocabulary and discards the conditions that made the vocabulary valid.

This is largely right, and the climate case makes the limitation sharper than the manufacturing case does. A pressed washer has a diameter with a fixed engineering target; global mean sea-surface temperature does not have a "target," and its baseline is itself moving under long-term forcing, which is precisely the non-stationarity that breaks the classical ±3-sigma arithmetic. Naïve control limits computed against a twentieth-century baseline will flag ordinary present-day behaviour as an emergency, because the baseline itself has shifted. What survives the transfer is not the arithmetic but the architecture: continuously estimated state, uncertainty derived from the record's own history rather than an external standard, and provenance tight enough to trace a signal to a cause. Sequential Bayesian updating and Kalman-style filtering, both descendants of the same lineage, already handle non-stationary measurands by letting the baseline itself be part of what is estimated. The climate system is the hard case for this architecture, not evidence against it.

More streams produce more false alarms, not more insight. Run enough regional charts simultaneously and something will cross a threshold by chance somewhere on the globe every week. The binding constraint is attention, not access to data.
A chart with a thousand regional cells watched at conventional significance will manufacture a "signal" almost every week by chance alone.

Also right, and it is the more dangerous failure mode in climate monitoring specifically, because a false regional alarm competes for scarce expert attention with a real one, and the two are cheap to confuse in a press cycle. But this is an argument about response policy, not about intake. The fix is not to observe less — it is to filter at the point of alarm rather than at the point of observation, exactly as multirule schemes in other fields raise the bar for declaring a signal without discarding any of the underlying measurements. The intake axis terminates at continuous, provenanced observation. What happens after a threshold is crossed — who is notified, what counts as corroboration from an independent stream, how long a flagged basin stays flagged before being stood down — remains a live, unsolved problem of institutional design. Conceding that is not a retreat from the claim. It is the claim: observation has a ceiling, judgement does not, and the two should not be confused with each other.

Continue