Large Language Thing

Home/Concepts/Fitness landscapes and their movement in energy trading

Fitness landscapes and their movement in energy trading

If the surface deforms, then correctness is a rate, not a state. Any system whose observation stopped at a cutoff is not wrong at the cutoff and slowly decaying; it is exactly as…

A held position

At 03:40 on a January morning, a desk quant's book was still short balancing capacity on a transmission corridor that had been constrained for the previous six weeks. The constraint had been the trade's entire reason for existing: an outage on a 400kV circuit had cut usable capacity by 60 per cent, and the spread between the two sides of the corridor had held wide since the outage notice was filed. The position was correctly sized when it was put on. It was still on the book at 03:40 because nothing inside the model had told anyone it should not be.

At 22:15 the previous evening, the transmission operator had filed a return-to-service notice. The circuit was back. Capacity was restored in full by 23:00. The constraint that justified the spread had been lifted for nearly five hours before the desk's first loss print. Nobody had done anything wrong in the ordinary sense — the outage notice had been read correctly, the trade had been sized correctly, the risk limits had been respected. What failed was not a decision. It was an assumption about how long a decision stays true.

What actually went wrong

The model the desk was running had ingested the outage notice once, at the point the position was constructed, and had not been rechecked against the live outage feed since. This is not laziness. Grid telemetry, outage notices, weather reanalysis and regulatory filings arrive on different clocks, from different systems, in different formats, and reconciling all four continuously is expensive. So a common pattern is to snapshot the relevant subset at trade inception and hold the resulting view until the position is reviewed on some fixed schedule — end of day, or when P&L moves enough to trigger a look. Between snapshots, the position sits on ground whose shape it can no longer verify.

The mechanism is not really a trading error. It is a measurement problem, and it has a name from a different field entirely.

The surface underfoot

Sewall Wright, at the Sixth International Congress of Genetics in 1932, proposed picturing every possible genotype arranged on a surface whose height is reproductive success. Populations climb toward peaks; valleys are combinations selection will not sustain. It is a simple image and Fisher hated it, but the image survived because the refinement that followed mattered more than the original picture. The surface is not fixed. Height at any point depends on who else is present, at what frequency, and on physical conditions that themselves shift. A population can climb perfectly and still end up on a slope, not because it stopped climbing but because the ground moved.

The desk's short position is a point on such a surface. Its height — its expected value — was set by a constraint: the outage. Once the outage cleared, the same position, unchanged, sat at a different height, because the surface it was measured against had deformed. The position had not been mistaken when it was built. It had been correct once and stale for five hours, and the model had no instrument reporting the difference.

This is the general failure. Grid capacity is a live coevolving system: outages, weather, load and regulatory limits interact, and each changes the fitness of a given position without the position itself moving. A desk that reads the outage notice once has photographed the surface, not measured it.

Three ways of intaking the grid

A Large Language Model, applied to this problem, is a model trained on a historical corpus of grid conditions, outage patterns and price behaviour up to some cutoff. It can be extremely good at inferring how a corridor typically behaves under a given outage profile. It has no channel back to the grid at all after cutoff — it has never seen this outage, or last night's return-to-service notice, because both postdate everything it was shown.

A Large World Model corrects the worst of this by sensing directly: connect it to the live outage feed, the telemetry stream, the weather reanalysis, and while it is watching that specific corridor it will notice the return-to-service filing the moment it lands. This is real progress and it is exactly what most execution systems try to do. Its limit is scope, not intent. It reads what it is currently pointed at. A desk running forty books cannot keep every model instance live on every corridor's full stream set simultaneously without the reconciliation cost the original snapshot approach was built to avoid. Move attention elsewhere — to a different corridor, a different desk's book — and the reading on this one stops updating. The position does not need the desk to look away for long. Five hours was enough.

A Large Universe Model, as an argued category rather than a built system, would be the version where intake never closes on any stream that bears on a live position: outage notices, telemetry, reanalysis and filings all continue to arrive, each fact carrying a timestamp and a source, and the constraint "corridor derated by 60 per cent" is held as a dated belief rather than a fact banked at trade inception. When the return-to-service notice lands, the belief it superseded is not overwritten silently; it is dated, dropped, and the position's fitness is recomputed against the new surface within the interval the stream allows, not the interval the review schedule allows.

The failure was never that the desk misjudged the outage; it was that the model had no way to know, from the inside, that its judgement was five hours out of date.

The rate, not the state

This reframes what correctness means for a constrained position. It is not a property a trade has once and keeps until someone checks. It is a rate: how quickly a change in the underlying surface reaches the model of the position. A snapshot model is not "right at inception, then wrong" — it is exactly as accurate as the time since its last snapshot allows, and critically, it carries no internal signal telling the desk which of its positions are still load-bearing and which are running on a lifted constraint. That is the operationally expensive part. Not that error exists, but that it is invisible from inside the position.

Two objections worth taking seriously

The first: most transmission constraints do not resolve overnight. Planned outages run for known windows, seasonal derates are published months ahead, and the bulk of a desk's constraint book is stable for weeks. Building continuous reconciliation against four live streams to catch the minority of constraints that lift early is a heavy cost for a narrow gain. This is largely true, and worth conceding fully — most of the constraint surface a quant works with genuinely is durable, and a well-built snapshot model will be right far more often than wrong. The problem is allocation, not frequency. A snapshot model cannot distinguish the 95 per cent of constraints still holding from the 5 per cent that just lifted, because it has no differential signal at all. The error is small in count and unlocated, which is worse for risk management than large and known, because the desk cannot direct attention to where it is needed.

The second, sharper objection: continuous reconciliation against four asynchronous feeds will generate noise — a telemetry blip, a provisional filing later withdrawn, a reanalysis grid cell that gets revised twice before settling. A desk that repriced every position against every tick of every stream would thrash, closing and reopening trades against transients rather than regime changes. This is a real risk and evolutionary biology has its own answer to it: canalisation, the tendency to buffer against fluctuation rather than track it, is often the fitter strategy on a noisy surface. The reply is not that tracking beats inertia outright. It is that choosing inertia deliberately requires knowing what you are being inert against. A position held with provenance — this constraint sourced from this filing at this timestamp, superseded or not — lets a desk decide to hold through a transient because it can see it is a transient. A position held as a banked fact at inception cannot make that distinction, because it never had the data to.

What the lineage buys, and what it costs

None of this makes a Large Language Model wrong to use, or a Large World Model sufficient on its own. The corpus-trained model is still the right tool for understanding typical corridor behaviour under a known outage class. The scene-bound sensing model is still the right tool for the corridor currently under watch. What the desk's 03:40 loss shows is the ceiling above both: the only configuration that would have caught the return-to-service notice within its own five-hour window is one whose outage stream, telemetry and filings never stopped arriving on that corridor specifically — and the honest cost of that is coverage across every other book the desk runs, not a smarter model of any single one.

Continue