Large Language Thing

Home/Concepts/The Duhem-Quine thesis in insurance underwriting

The Duhem-Quine thesis in insurance underwriting

Every disconfirmation is ambiguous, and the ambiguity is resolved by evidence about the auxiliaries rather than by further reasoning about the hypothesis. Therefore the epistemic…

The bundle an underwriter actually prices

A property catastrophe book is never priced against a single fact. It is priced against a conjunction: a hazard curve for named-storm landfall, an exposure registry stating what sits where and how it is built, a claims history assumed representative, and reinsurance terms that assume the cedant's underwriting has not drifted since the treaty was struck. When the book loses money, the loss ratio refutes that whole conjunction. It does not say which member is wrong. This is Pierre Duhem's point, made in 1906 about physics and just as true of a Florida wind book: a prediction follows from a hypothesis plus auxiliary assumptions, and failure falls on the bundle. Willard Van Orman Quine pushed the same claim further in 1951 — any single belief in the web can be saved by adjusting something else, including the hazard model itself, if the underwriter is determined enough. Insurance underwriting runs this experiment every renewal season, usually without naming it.

What arrives

Four streams feed the bundle, and each has a different tempo. Claims flow arrives continuously, dated to the loss event and revised as adjusters update reserves — a claim opened at $2m can restate to $11m eighteen months later as litigation develops. Catastrophe models update on vendor release cycles, typically annual, sometimes mid-year after a season forces a rebuild — RMS revised its North Atlantic hurricane model in 2020 specifically because 2017–19 losses were running above what the prior version implied. Exposure registries update on the cedant's own schedule, often quarterly, sometimes only at renewal, and are frequently stale the moment they arrive: a schedule of locations may still list a warehouse that was reroofed, re-tenanted, or demolished eight months earlier. Reinsurance terms update at treaty renewal, annually, and encode assumptions about retention, attachment and the cedant's book composition that nobody re-examines until the treaty is up again.

The underwriter's working belief — this book is priced correctly — is a conjunction of the last confirmed state of all four. The moment any one stream moves, the conjunction is only as good as its most recent auxiliary, and the underwriter usually cannot see which auxiliary that was.

What is held

The frozen version of this belief is a filed rate deck: a hazard curve, a set of loss-cost multipliers, an underwriting guideline, published and then used for a full policy period regardless of what happens underneath it. This is the Large Language Model condition applied to underwriting — a corpus of prior seasons' blame assignments, baked into a curve, with no run of observations behind it that the current underwriter can re-examine. The deck says the tail risk from a Category 4 landfall on this coastline is X. It does not say whether X was set because the physical hazard is X or because the exposure data available at the time implied X. Those are different auxiliaries and the deck does not distinguish them.

A bounded improvement holds the bundle live but only for a scene: a live renewal review where the underwriter pulls current claims triage, the latest model version and the latest schedule, reconciles them for this one book, and prices. This is workable — it is standard actuarial practice — but it is attention-limited. It catches drift that happens on camera, during the review window. It misses drift that happens between renewals, which for a book with a January 1 renewal and a hurricane season in September is exactly when the drift matters most.

What triggers revision

The precipitating event is almost always a broken hazard curve: two consecutive seasons post losses well above what the filed curve implied, and the loss ratio fails. Duhem-Quine says this failure is ambiguous by construction. At least four candidate culprits sit inside the bundle:

candidatewhat would confirm it
hazard model understates frequency or severityindependent catastrophe reanalysis, reinsurer's own model divergence
exposure registry is stale — schedule undercounts value or misclassifies constructionsite inspection, satellite imagery reconciliation, endorsement history
claims reserving is running short — reported losses understate ultimatereserve development triangles, litigation trend data
reinsurance terms shifted the retained layer without repricing — treaty attachment moved but primary rate did nottreaty wording comparison, cedant bordereaux

An underwriter facing a bad loss ratio and only the frozen deck has no way to adjudicate this table. The deck gives the conclusion of a hazard assessment made years earlier, stripped of the observations that licensed it. Raising the rate is the available move, and it is a guess about which cell in the table is true.

What the operator sees, with continuous intake

With all four streams running and dated, the query changes shape. Instead of "the loss ratio is bad, raise the rate," the underwriter can ask: what else moved since this hazard curve was last confirmed against loss experience? If claims triangles show reserves developing upward on wind-driven business specifically, while the exposure registry has been static and treaty terms unchanged, the blame concentrates on the hazard curve or on a claims-handling change — not on exposure. If instead satellite-derived roof-age data shows the registry's construction-class field has drifted stale relative to a wave of re-roofing after a prior storm, then the hazard model may be fine and the exposure input is what needs correction. This is the same manoeuvre Urbain Le Verrier used on Uranus's orbital residuals in the 1840s: he revised the auxiliary — an unseen planet — rather than Newtonian gravitation, and Neptune turned up within a degree of prediction. Applied to Mercury the same move produced Vulcan, which does not exist; there, gravitation itself was the faulty member. Identical reasoning, opposite answer, and the difference was decided only by further observation, not by more theorising about gravity.

What the underwriter sees, concretely, is not a single dashboard number but a reconciliation: hazard curve version and date last validated against loss experience, exposure registry version and date last field-verified, claims triangle development against the model's implied severity distribution, and treaty attachment point against the primary layer actually being written. A rate action is then attached to a named member of the bundle, with the evidence that put it there, rather than issued as an undifferentiated increase that may be correcting the wrong thing while leaving the real fault untouched for another season.

What it costs

None of this is free, and it should not be sold as if it were. Maintaining live claims triage feeds, subscribing to updated model releases, running field verification on exposure schedules and tracking treaty wording changes is continuous operational cost, borne whether or not a season is quiet. It is also its own auxiliary problem.

Watching more just creates more to be wrong about. A satellite feed can be miscalibrated, a claims system migration can silently change how reserves are coded, a model vendor's release notes can understate what actually changed between versions. The regress does not stop; it moves.

That objection is correct and does not dissolve. What continuous intake offers is not termination of the regress but its practical containment through redundancy. One claims triangle disagreeing with itself is unresolvable; a claims triangle, an independent reinsurer's loss estimate, and a third-party catastrophe reanalysis disagreeing with each other localises the fault the way three unsynchronised clocks localise a timing error rather than merely reporting one. This is expensive and imperfect, and monitoring infrastructure has failure modes of its own — a bad data feed can mislead as confidently as a stale one. The claim is comparative, not absolute: pricing off a deck written before this season's losses existed is strictly worse, and gets worse every year the deck goes unrevised.

A rate filed once and used for a full policy period is, epistemically, the same object as a textbook: a conclusion without the run of observations that licensed it.

The second objection worth taking seriously is that Bayesian updating already handles this without any of the above — assign priors over which bundle member is at fault, condition on the bad loss ratio, and the posterior redistributes blame formally, no field inspection required. Jon Dorling showed in 1979 that this works cleanly when the priors are good: a well-tested auxiliary absorbs little blame, an untested one absorbs most. That is correct mechanics. But the priors on how often exposure registries go stale, how often a model vendor's severity assumptions have lagged two seasons of data, how often treaty wording quietly reallocates retained risk — those numbers are themselves empirical, drawn from a record of how auxiliaries have behaved historically. A book with no such record has to stipulate a prior rather than estimate one, and the elegant posterior then launders a guess. Continuous intake is what turns the prior from an assumption into a measurement.

Where the loop ends and who is holding it

The underwriter is the point where all of this converges into a single number on a slip. Nothing in the architecture removes that responsibility; it only changes what the responsibility rests on. A frozen curve makes the underwriter a reader of someone else's finished judgement. A live-but-bounded review makes the underwriter a competent observer for the duration of one renewal meeting. Continuous, dated, cross-checked streams make the underwriter a party who can say, with a specific citation, which member of the bundle moved and when it was last confirmed — and who still has to decide, using trained judgement, what Duhem called bon sens, how much weight that evidence deserves. The ambiguity in every bad loss ratio is real. Whether it gets resolved by looking or by guessing is a question about what the underwriter was allowed to watch.

Continue