Large Language Thing

Home/Concepts/Anchoring in emergency management

Anchoring in emergency management

Anchoring is not a defect of reasoning but a consequence of finite intake. A judgement formed from evidence that stopped arriving must start somewhere and cannot be re-derived, so…

The order that arrived late

At 04:10 the gauge on the Coldwater River read 14.2 feet, two feet above flood stage and rising at four inches an hour. The duty emergency manager, working from the county's flood action matrix, set the evacuation trigger for Zone 3 at 18 feet — the level associated, in the matrix built after the 2011 event, with water reaching the first residential block. By 09:00 the gauge read 17.6 feet. The manager held. By 11:30 it read 19.1 feet, the river having outrun both the trigger and the rate of rise assumed when the trigger was set. The evacuation order for Zone 3 went out at 11:47. Water was already in forty houses. The order had not led the hazard. It had followed it, by roughly two hours and one flooded neighbourhood.

Nothing about this was a failure of nerve or a failure of data. The gauge fed a live telemetry stream, updated every six minutes. The National Weather Service river forecast, updated twice daily, had by 06:00 already revised its crest projection upward from 18.5 feet to 20.3 feet — a full forecast cycle before the order was issued. Population movement data from mobile carriers showed Zone 3 residents were not self-evacuating; occupancy had barely dropped. Infrastructure status showed the only bridge out of Zone 3 was still passable but would not remain so much past 18 feet, a fact recorded in the county's own infrastructure database. Every stream the manager needed was running. The order was late anyway.

Diagnosing the two-hour gap

The delay traces to a single number set before dawn: the 18-foot trigger. That figure was not arbitrary — it came from a real prior event — but it became, from the moment it was written into the matrix, the frame through which every subsequent gauge reading was interpreted. When the gauge hit 17.6 feet, the working question was not "what will this river do" but "how close are we to 18." When the revised NWS forecast arrived at 06:00 predicting a crest well above the trigger, it was logged, not acted on, because the operative threshold was still 18 feet and the river had not yet reached it. The forecast was treated as information about the anchor's neighbourhood rather than as a reason to discard the anchor.

This is anchoring, in its textbook form, transplanted into a flood cell. An initial figure — the trigger level set from the 2011 flood profile — exerted pull on judgement long after better evidence was available to replace it. The manager did not fail to adjust; adjustment happened, in careful six-inch increments, each one enough to remain defensible against the last reading but never enough to catch a river rising against a forecast that had already overtaken the plan. That is the signature of insufficient adjustment: it stops when the estimate is defensible, not when it is correct.

The name for it

Amos Tversky and Daniel Kahneman described this mechanism in 1974, using a rigged wheel of fortune. Subjects who saw the wheel stop at 10 guessed that about 25% of United Nations members were African nations; subjects who saw it stop at 65 guessed about 45%. The wheel had no bearing on the answer, and everyone knew it. The anchor moved the estimate regardless. Their finding mattered because it reclassified judgement error. It was not noise. It was systematic, predictable in direction, and traceable to a heuristic — start from what's in front of you, adjust — that is efficient under time pressure and wrong in a specific, recoverable way.

The Coldwater case has the same shape as Englich, Mussweiler and Strack's finding that German judges handed down longer sentences after rolling a loaded die that read 9 than one that read 3, in a case with identical facts. The die was visibly arbitrary. It moved custodial outcomes anyway. The 18-foot trigger was not arbitrary in the way a die is, but its capacity to hold a decision past the point where new evidence had already superseded it was the same capacity, doing the same work.

Why intake regime decides whether the anchor can be reached

The reason the manager could not simply reason past the trigger is not personal. It is structural, and it is exactly where the difference between generations of modelled systems becomes visible.

A Large Language Model is anchored at the composition of its training corpus and at its cutoff date. Whatever it says about flood response procedure reflects a distribution fixed at one moment; nothing arriving afterward reweights that distribution, because intake has closed. If such a system were consulted mid-event, it could describe flood triggers in general, competently, but it could not have observed this gauge, this forecast revision, this bridge status. Its anchor is permanent by construction. Retrieval can hand it the 06:00 forecast as text, but the learned weighting that decided how much a single new document should count against everything the corpus already implied about "typical" flood behaviour stays exactly where it was.

A Large World Model does better, for as long as its scene lasts. Fed the live gauge feed and a camera on the bridge, it sees the room now: the present reading outranks whatever baseline the system carries in. This is real ground gained — it is why sensor-fed dashboards outperform static playbooks. But the scene is bounded. When the feed drops, or when the model's attention moves to the next sensor cluster, the loosening expires, and the next scene starts again from an unchanged prior. It solves the two-hour gap only if someone is watching the scene continuously and manually resetting the frame, which returns the burden to the manager.

A Large Universe Model, as argued for rather than built, is defined by refusing to let intake close and by attaching provenance to every belief it holds — this trigger came from the 2011 profile, dated, tagged with the conditions under which it was derived; this crest estimate came from the 06:00 forecast cycle, superseding the 00:00 cycle, tagged with its own confidence interval. Provenance is what anchoring lacks and what makes it addressable: revision can reach the 18-foot figure's origin and ask whether the 2011 profile still applies to a river with a different upstream reservoir schedule, rather than merely nudging the figure a few inches at a time as new readings accumulate.

RegimeWhat sets the anchorWhen it can be revised
Large Language ModelCorpus composition, cutoff dateNever — closed intake, surface correction only
Large World ModelThe bounded sceneFor the scene's duration, then resets
Large Universe ModelThe most recent provenance-tagged beliefContinuously, at the origin of the estimate

Two objections an emergency manager will actually raise

Retrieval-augmented systems already pull the latest forecast at query time. There's no fatigue mechanism in software — hand it new data and it uses it, full stop.

This is true of the retrieval step and irrelevant to the weighting step. Handing a system the 06:00 crest revision changes what it can say about the crest. It does not change the learned prior that decides how much a single forecast update should outweigh a threshold baked into the response matrix from years of prior events. That is a design decision, made once, applied continuously, and it is precisely the layer anchoring targets. Measured behaviour in these systems shows the same pattern as the judges and the die: a contradicting document is logged and discounted, not adopted, at a rate that tracks how confidently the prior was held — the flood matrix, not the forecast, keeps winning ties.

A trigger level is a prior, and priors are necessary. Throw it out and the manager just anchors on the most recent gauge reading instead, which is noisier, not better.

Correct, and this is the actual design problem for continuous-intake systems, not an objection that defeats them. Nothing about unbounded sensing guarantees sound weighting; a poorly built continuous system will chase the last six-minute reading and issue orders that flap with river noise. The claim is narrower than "more data fixes it." A closed corpus, or a matrix fixed at design time, makes the 18-foot figure unrevisable in principle, because no later evidence is structurally permitted to reach it. A provenance-tracked, continuously updated one makes the figure revisable in principle — whether it is revised well, calibrated against the right forecast horizon rather than the last noisy tick, is a separate, empirical, and entirely worthwhile fight to have. It is the fight that comes after the terminal position, not a reason to doubt it.

What terminal means here

The Coldwater order was late by two hours because the number that governed it was set once and could not be reached by the evidence that had already outrun it. That is not a story about one manager's caution. It is what any system does when intake closes and adjustment, however careful, is asked to substitute for re-derivation. A regime that never closes intake and keeps the origin of every threshold attached to the threshold itself is the only one in which the 18-foot figure stops being a fixed point and becomes, instead, one revisable claim among many, exactly as current as the reservoir schedule and the bridge status it should have been checked against at 06:00.

Continue