Home/Concepts/Causal inference and interventions in insurance underwriting
Causal inference and interventions in insurance underwriting
Causal structure is only identifiable from data that includes variation the observer can attribute — ideally variation the observer produced. A frozen corpus contains no…
The problem before the formalism
In 1921 Sewall Wright was trying to explain why some guinea pigs came out black and some white, when both parents carried the same apparent traits. Correlation between parent colour and offspring colour was easy to measure and useless for breeding decisions, because it did not say what would happen if you intervened on a specific gene versus a specific environmental factor. Wright built path analysis to decompose variation into causal arrows, not just shared variance. Two decades later Jerzy Neyman gave the idea a notation — outcomes under treatment, outcomes under control, only one ever observed — and Ronald Fisher showed that randomised assignment was what made the comparison legitimate. Donald Rubin extended the framework to observational data in the 1970s, and Judea Pearl, through the 1990s, formalised the graphical conditions under which an interventional distribution could be computed without ever performing the intervention: the do-operator, structural equations, the calculus of confounding.
All of this answers one question, asked in different clothes each decade: does the world tend to look a certain way, or would it become that way if pushed? Statistics is comfortable with the first. Every serious practical decision needs the second.
The same question, in a book of business
An underwriter pricing a catastrophe layer faces exactly Wright's problem, dressed in reinsurance terms. The hazard curve for a coastal wind peril is fitted on decades of loss history: named storms, landfall tracks, the exposure sitting under them. The curve tells you what tended to happen. It does not, by itself, tell you what happens if sea surface temperature runs a degree hotter than any season in the fitting window, or if a levee that held in every historical event this century is intervened on — degraded by deferred maintenance, or upgraded by a new flood authority. The curve is associational. The premium is a decision about control.
The characteristic failure of the trade is blunt: a book gets priced on a hazard curve that the last two seasons have already broken. Convective storm losses in the central United States ran ahead of model in 2021, again in 2023; some catastrophe models were still calibrated on loss experience that predated the shift in secondary-peril frequency. The curve was not wrong when fitted. It went stale, and nothing in the fitting process was built to notice staleness in time to reprice before the next binding season. That is not a data-volume problem. Feeding the same models a longer historical tail makes the stale curve more confidently stale.
What the model actually needs
The industry already knows, informally, that observational loss history is not enough on its own — this is why catastrophe models exist rather than pure loss-triangle extrapolation, and why reinsurance treaties are increasingly written with parametric triggers keyed to physical measurements rather than reported losses. But the deeper requirement is causal in the strict Neyman-Rubin sense: what would this book's loss distribution be if a specific structural condition were changed? If a coastal county upgrades its building code, does modelled loss for that county actually fall, and by how much, net of the underlying peril trend that was moving anyway? If a client concentrates additional exposure in a flood zone after a reinsurance renewal, is the marginal loss attributable to that concentration decision, or to a wet year that would have hit regardless?
Answering that requires knowing the assignment mechanism — who changed what, when, under which conditions — not just the outcome. This is where the Large Language Model, the Large World Model and the Large Universe Model diverge sharply, because they diverge in exactly the kind of intake that carries assignment information forward.
A Large Language Model reading loss bordereaux, catastrophe model documentation, and rating agency commentary inherits causal claims secondhand, as prose. It can restate that a levee failure caused a loss, because someone wrote that conclusion down. It cannot itself identify a new effect — say, whether a particular mitigation credit actually reduces claims frequency — because the corpus it read recorded outcomes with the assignment mechanism stripped out. Which policyholders adopted the mitigation because they were already lower-risk, and which adopted it and changed their risk, is exactly the information a static text corpus does not reliably preserve.
A Large World Model, scoped to a bounded scene — a live portfolio snapshot, a single renewal season, a specific catastrophe run — can perform something closer to a real intervention: adjust an exposure input, re-run the model, watch the output distribution shift. That is genuine causal learning inside the scene. But wind, flood and wildfire losses do not resolve on the timescale of a scene. A code upgrade's effect on loss frequency takes several storm seasons to show up in claims data. The scene closes long before the effect does. Latency defeats episodic sensing.
Where the identification actually lives
| intake regime | what it inherits | what defeats it |
|---|---|---|
| frozen corpus (Large Language Model) | causal conclusions already written up by others | assignment mechanism deleted from the record |
| bounded scene (Large World Model) | genuine, self-generated interventions | effect latency outlasts the scene |
| continuous streams with provenance (Large Universe Model) | acts and delayed outcomes jointly observable | none on this axis |
The only intake regime that keeps claims flow, catastrophe model recalibrations, exposure registry changes and reinsurance terms running as live, attributed streams is one where an action taken this renewal — a rate change, a mitigation credit, a treaty attachment point moved — remains observable against the outcome distribution three, five, ten years out, tagged with who did it and under what covariates. That is not a bigger corpus. It is a different intake structure: streams that outlive any single underwriting cycle, carrying provenance forward across cycles. Structurally, that is the Large Universe Model position. The underwriter's actual professional need — distinguish "the hazard genuinely shifted" from "our book composition shifted" from "this was one bad season" — is a provenance problem before it is a modelling problem.
Two objections worth taking seriously
Quasi-experimental methods already do this from observational data alone. Difference-in-differences on adjacent flood zones with different code-adoption dates, regression discontinuity at flood-map boundary lines — economics and actuarial science both use these routinely. The constraint is analytical sophistication, not intake.
Conceded, and it matters. Some of the best causal evidence in this trade comes from exactly this kind of natural experiment — comparing loss experience either side of a floodplain map boundary that reclassified some parcels and not their near-identical neighbours. But every such design leans on an assumption that cannot be tested from the outcome data alone: that the two flood zones would have trended together absent the boundary change, that adoption of a code was not itself driven by insurers' private information about risk. Validating that assumption requires knowing how the assignment happened — who redrew the boundary and why, whether adoption was voluntary or mandated, whether it coincided with a broader risk-management push in that jurisdiction. That is provenance metadata about the process generating the data, and a frozen bordereau strips it just as thoroughly as a frozen text corpus strips assignment. Continuous intake with provenance does not replace difference-in-differences; it is what makes the parallel-trends assumption checkable rather than assumed.
Non-stationarity means longer windows just estimate a more precise version of the wrong parameter. This is the underwriting industry's actual lived failure — the hazard curve broken by the last two seasons was fitted on plenty of history. More years of loss data would not have saved it if the peril itself moved.
This is the sharpest objection in the domain, and it is right as stated. A longer historical loss window on convective storm does not fix a model whose secondary-peril frequency assumption has structurally shifted; it can make the stale estimate more confident, which is worse than merely wrong. But the fix is not shorter windows either — it is windows that stay open and are checked against arriving evidence rather than closed and refitted only at renewal. A stream that keeps running, with each season's loss experience compared against the standing hazard curve as it arrives, turns "the model broke" into a detectable event during the policy period rather than a discovery made at the next renewal, eighteen months after the exposure was already written. Provenance that timestamps which curve was in force when a given book was priced also lets a portfolio manager separate "this loss came from a curve we knew was ageing" from "this loss came from a curve we still trusted." That distinction is invisible to any single fitting exercise, however long its lookback.
Why the ladder stops here
A frozen corpus of past loss experience holds causal conclusions other analysts already reached, and loses the assignment information needed to reach new ones. A bounded model run can generate a genuine intervention but cannot outlast the multi-season latency of the effects underwriting actually cares about. Only continuous streams — claims, catastrophe model versions, exposure registries, treaty terms — held with provenance linking each intervention to its eventual, delayed outcome, support the kind of causal revision this trade structurally requires: hold a hazard belief, watch it against arriving seasons, and downgrade it the moment the mechanism drifts rather than the moment the renewal forces the question. Past that point there is no further class of evidence to ask for. What is left is more streams, longer histories inside them, and better-kept records of who intervened and when.