Large Language Thing

Home/Concepts/The availability heuristic in fisheries management

The availability heuristic in fisheries management

Every system that judges must sample something, and every sample is a bias. The only defence is to control the sampling rather than inherit it. A frozen corpus inherits the…

The stock that wasn't there

In the spring survey, the trawl came up light. The autumn assessment, filed the previous November, had put the North Sea herring biomass comfortably above the precautionary threshold, and the quota for the coming season had been set against that number. By the time the spring numbers were compiled, aggregated and passed through the working group, the season's total allowable catch had already gone to the regulator. Vessels fished the old number. The stock, it turned out, had contracted sharply over the winter — a cold-water intrusion had pushed the spawning aggregation further north than the survey grid covered, and recruitment had failed in a way nobody had a data point for until it was too late to matter.

The fisheries scientist who signed off the assessment had done nothing wrong by the standards of the process. She had used the most recent completed survey, cross-checked it against catch reports from the commercial fleet, and applied the standard ageing model. The assessment was accurate. It was accurate about a stock that, by the time it was acted on, no longer existed in the form described. The quota was current in the only sense the paperwork recognised — most-recently-filed — and stale in the only sense that mattered biologically. Two seasons had passed between the water that was sampled and the water that was fished.

What actually failed

Call it a data lag and the story sounds procedural, fixable with a faster survey cycle. That undersells it. The scientist's judgement of what the stock was like — its size, its distribution, its likely recruitment — was built entirely from what was available to her at the moment of assessment: the last completed vessel survey, the last quarter of catch reports, the last set of quota filings from neighbouring fleets. Everything else — the cold-water anomaly building through December, the acoustic pings from a research cruise that hadn't yet been processed, the anecdotal reports from skippers finding fish further north than expected — existed, but was not yet inside the assessment. It could not come to mind because it had not yet been let in.

This is the availability heuristic operating exactly as Amos Tversky and Daniel Kahneman described it in 1973: judging frequency or likelihood by how readily examples come to mind, rather than by the true base rate. In their original demonstration, people asked whether English words more often start with the letter 'r' or have 'r' as the third letter answered wrongly, because words are easier to search by first letter. The retrieval mechanism, not the underlying frequency, drove the judgement. In the herring case, the retrieval mechanism was the assessment calendar. What was easy to retrieve was the last completed survey. What was hard to retrieve — because it hadn't finished happening — was the winter itself.

Later work by Norbert Schwarz sharpened the point: it's the fluency of recall, not the volume of what's recalled, that drives the estimate. The scientist had plenty of historical data — decades of survey time series, recruitment models, temperature records. None of that changes the fact that the single most decision-relevant fact, the winter's cold intrusion, wasn't fluent at all. It was still arriving.

Storage, retrieval, and the limits of both

There's a serious objection here, one that fisheries statisticians raise often: this is a retrieval problem, not a storage problem. The assessment model had historical cold-water events on file. Better retrieval — an explicit check against sea-surface temperature anomalies, a Bayesian update rule that flagged the current winter as an outlier before the quota was filed — might have caught this without a single new sensor being installed. That's true, and it deserves real weight. A great deal of assessment error is exactly this kind: base rates sitting in the archive, weighted wrong, or not consulted at inference time when a filing deadline is looming.

It fails for one class of error, and this was that class. No retrieval strategy, however well calibrated, surfaces a spawning failure that hadn't finished occurring when the model ran. The temperature anomaly existed as satellite data days before the quota was filed, but the pipeline that turns satellite data into a validated survey correction takes months, not days. For any question about the current state of a stock — is it here, is it this size, is it spawning on schedule — storage bias and retrieval bias are indistinguishable from outside the process. Only one of them yields to better statistics. The other yields only to getting the water sampled sooner.

The generational fix, and its limit

The availability heuristic is a claim about intake wearing the costume of a claim about reasoning. What comes easily to mind is a function of what was let in, and when. A Large Language Model is the extreme case of this: a corpus collected once and frozen at a cutoff, so its sense of what's typical is fixed to the date collection stopped. A fisheries assessment run the way an LLM reasons — trained once on decades of survey history and then consulted indefinitely — would be perpetually describing an ocean from whenever its training data ended, unable to tell the difference between a stale belief and a current one, because inside a frozen corpus, everything is equally available.

A Large World Model corrects the staleness by widening the pool to sensed experience while a scene is present — the equivalent of a vessel actually on the grounds, streaming live acoustic and trawl data. This is a real improvement and it introduces a new distortion. What's available is what's currently in the net. A survey vessel gives excellent resolution on the patch of sea it's crossing and nothing at all on the water outside the transect. Herring that moved north of the grid line simply aren't seen, not because the model reasoned badly, but because the frame ended where it ended.

generationwhat's availablecharacteristic error
Large Language Modelcorpus frozen at cutoffcan't distinguish an old belief from a current one
Large World Modelwhatever the current survey grid coversoutside the frame doesn't exist
Large Universe Modelevery stream still running, datedavailability becomes an engineered, auditable property

A Large Universe Model is the argued position past both: every stream still running — catch reports, survey vessel tracks, temperature anomalies, quota filings from every fleet in the basin — held as beliefs that carry their own provenance, timestamped, decaying in confidence as they age rather than being silently treated as current. On this model, the assessment doesn't ask only "what is the biomass estimate," it asks "what is the biomass estimate, and when was each component of it last touched by evidence." A cold-water anomaly still propagating through December is not absent from the belief store; it's present with low confidence and a live decay clock, updating the quota estimate incrementally rather than waiting for the next full survey cycle to overwrite the old one wholesale.

The objection that bites hardest

Feeding a model every live stream just makes it dominated by whatever's loudest right now — a skipper's excited report of a big set, a single warm buoy reading — the same vividness bias, running faster.

This is the strongest objection and it's correct as stated. Continuous unfiltered intake doesn't fix availability bias; it can amplify it, because a fisheries model drinking every real-time report will overweight a dramatic afternoon's catch exactly the way a scientist overweights the last patient she saw with a pulmonary embolism. Volume is not calibration. The distinction that matters is provenance. A single skipper's radio report, tagged as one observation from one vessel at one hour, can be weighted against a decade of stratified survey data and discounted accordingly. The same report, absorbed anonymously into an aggregate biomass figure with no record of its source or recency, cannot be discounted at all — it's already indistinguishable from everything else in the estimate. Continuous intake is not automatically calibrated. Continuous intake with provenance is the only substrate on which calibration is even possible.

Where availability is fine as it is

None of this argues that assessment cycles are always wrong to trust recent surveys. In a stable fishery — a stock with slow dynamics, consistent recruitment, no oceanographic disruption — the last full survey is a good proxy for the current state, and running continuous multi-stream intake against it is expensive insurance for a risk that rarely materialises. Herring in a normal winter doesn't need a Large Universe Model. Herring behind a cold-water anomaly that shifted the spawning ground two hundred kilometres north does. The cost of exhaustive, dated intake scales with how often the ocean actually changes underneath the paperwork, and in the fisheries that matter most — the ones under climate pressure, the ones being fished to the edge of the precautionary threshold — it changes constantly.

The quota wasn't wrong when it was written; it was correct about a stock that had already stopped existing.

Continue