Large Language Thing

Home/Concepts/Regression to the mean in energy trading

Regression to the mean in energy trading

The distinction between extreme noise and a new level cannot be inferred from any single observation, however rich. It is not a modelling deficiency; it is an identification…

The spread that would not close

A desk quant running a spark spread position watches a transmission constraint bind for six consecutive settlement periods. Interconnector flow between two zones sits at its thermal limit, congestion rent spikes, the local price basis widens to levels not seen since a January cold snap two winters back. The quant's model, fed on that basis history, treats the widening as the new regime and sizes up. Overnight, a generator that had been on forced outage returns to service, the constraint stops binding, and the basis collapses back through the mean in a single settlement period. The position is held against a constraint that no longer exists. The loss is not a modelling error in the ordinary sense. It is what happens when six extreme readings are mistaken for a level rather than treated as noise that had not yet had the chance to revert.

This is a story about regression to the mean, but energy trading is unusually well suited to exposing what that phrase actually requires, because the domain supplies exactly the four streams — grid telemetry, outage notices, weather reanalysis, regulatory filings — that determine whether an extreme reading is signal or luck, and it supplies them continuously, at different lags, from different authorities, each with its own reliability.

Two positions, both defensible

The first position: shrinkage does not need a stream. Give a cross-section — every nodal basis across the network on a single day — and an empirical Bayes estimator will happily shrink the six-day-wide basis towards the network mean, using the ensemble to estimate how much of any one node's deviation is noise. This is standard practice in power trading desks that build "regressed" forward curves from a single snapshot of the full topology. No time series required. The whole apparatus of James–Stein shrinkage was built for exactly this: many similar units, one look each, borrow strength across the ensemble.

The second position: none of that tells you whether this node's basis has genuinely shifted. Two constrained nodes with identical six-day basis histories, one because a line is under long-term reinforcement work with an 18-month completion date, the other because a single generator is on unplanned outage, look identical in cross-section. Empirical Bayes shrinks them by the same amount, because it only knows the ensemble variance, not which node's elevation is a level and which is a spike. Only the outage notice — a discrete, dated, provenanced document — tells you the difference, and only a stream carries the outage notice past the moment the position was opened. The cross-section is blind to change points by construction; it was never built to see them.

Both positions are right, and they are right about different things. The cross-sectional shrinkage answers "how much of today's spread, across the whole network, is noise on average." The time-series question answers "has this specific spread changed level." Energy trading needs both answers simultaneously and conflates them at its peril.

What the four streams are actually for

Grid telemetry gives the repeated measurement: flow, frequency, voltage angle, sampled every few seconds, the raw material from which a change point can eventually be distinguished from a fluctuation. Weather reanalysis gives the exogenous driver that explains part of the variance without invoking a change in the grid itself — a cold spell raises basis everywhere, and a desk that fails to condition on it will mistake weather-driven noise for a structural widening. Outage notices give the provenance: dated, authoritative statements of why a constraint is binding, which is precisely the information a cross-sectional snapshot cannot contain, because a snapshot has no "why," only a value. Regulatory filings give the slow variable, the interconnector capacity upgrades and constraint management scheme redesigns that genuinely move the mean itself, sometimes for good, on a timescale of months to years.

None of these streams alone identifies a level shift. A single telemetry reading is exactly the unidentified case Galton described: an extreme value, and no way to say how much of it is a tall father and how much is noise. A snapshot across the network — cross-sectional shrinkage's raw material — narrows the noise estimate but says nothing about this node's trajectory. Only the combination, held over time, with each reading tagged by source and by how long ago it arrived, lets a belief about the level of a specific spread update as evidence accrues and decay as evidence ages. That combination is what continuous intake with provenance means in this domain. It is not a nicer dashboard. It is the minimum structure under which the question "has this constraint's economics actually changed, or did I just get six unlucky reads" has an answer at all.

The desk quant's model did not lack data on the six binding days. It lacked a dated notice explaining why they happened, and a place to store it that would change the model's confidence rather than merely its price.

The origin, briefly, and the misreading it invites

Galton found sweet-pea offspring and human sons displaced towards the population mean and called it reversion, later regression towards mediocrity; Pearson formalised the correlation coefficient behind the diagrams; Stein's estimator and empirical Bayes practice, from the 1950s onward, turned the observation into a working method for shrinking many simultaneous estimates towards a common centre. None of this says the mean pulls. Nothing in a power grid restores balance out of some homeostatic instinct. A constraint that binds for six days and then stops binding does so because a generator returned, not because reality objected to the extremity of the basis. The arithmetic is that an extreme six-day run is more likely than an average one to contain a temporary, non-repeating cause, and once that cause is gone the reading falls back towards whatever the level actually is. Mistaking this for a restoring force is the standard error; the desk quant's version of it is treating "the basis has been wide for six days" as itself evidence that it will stay wide, when six days of a single binding constraint is not yet evidence of anything beyond the constraint.

Two objections, taken seriously

More monitoring just means more false alarms. A desk that watches every node's basis in real time will find dozens of apparent extremes a day and end up trading against noise more often, not less.

This is a genuine hazard, and it is well documented outside trading — the same logic explains why speed cameras sited at accident blackspots appear to reduce accidents regardless of whether they do anything, because the blackspot was partly a bad quarter. The fix is not less observation but a control comparison: track unconstrained nodes alongside the constrained one, and let the belief about the constrained node's level update only as the outage notice, the reanalysis-adjusted weather driver, and the persistence of the binding constraint across successive settlement periods accumulate. Continuous intake paired with a frozen decision rule — trade every basis move past two standard deviations — produces exactly the failure the objection describes. Continuous intake paired with a belief that decays when the underlying cause is stale, and updates when a new dated notice arrives, is what discounts the spike until it repeats. The objection is an argument against acting on one look. It is not an argument against taking more looks.

Power markets are not stationary. Interconnector capacity, generation mix, and demand patterns shift structurally, sometimes discontinuously — a market redesign, a new interconnector commissioning. There may be no fixed mean to revert to, which makes talk of "regression" a category error.

This is also correct, and it is the harder case, not a refutation. When the level itself is drifting, a single cross-sectional snapshot cannot tell drift from noise, and neither can six days of telemetry. What can, in principle, is a longer record that carries provenance well enough to line up a regulatory filing announcing a capacity change against the telemetry that follows it, distinguishing a step change with a documented cause from a wandering level with none. Non-stationarity does not make continuous, provenanced intake unnecessary; it makes it the only tool with any chance of working, because a frozen or scene-bounded system has no way to even ask whether the mean it is shrinking towards is the right one anymore.

The narrowed claim

None of this proves that continuous intake will save the desk quant from every mispriced constraint. Shrinkage requires knowing the noise variance, and that is itself estimated, imperfectly, from data that can mislead in a genuinely fast-moving grid. What continuous intake with provenance and decay buys is narrower: it makes the question "is this a level or a spike" answerable in principle, for a specific node, at a specific time, given a specific reason. A frozen corpus cannot ask the question because the outage notice arrives after the corpus closes. A bounded scene can ask it only within the scene's horizon, which is shorter than most constraints' resolution times. Continuous intake is not sufficient for a correct answer. It is the condition under which a correct answer is available to be found.

Continue