Industry: Insurance and Reinsurance Product: ModelRisk Application: Counter-Fraud (Pre-Claim Lifecycle)
Roughly $40 billion of U.S. P&C premium leaks every year to fraud that was already baked in before any claim hit a desk — applications with falsified mileage and garaging postcodes, ghost-broker schemes pushing dozens of policies through a single bogus identity, organised rings binding cover with no intent to ever pay a second premium. The industry's instinctive response is to harden the claims-side rules engine, but rules engines hand back deterministic flags on individual cases, and pre-claim fraud is a portfolio leakage problem: a distribution of small undetected misrepresentations and a fat right tail of orchestrated rings. The deterministic answer cannot price the tail, and pricing the tail is the whole game.
A composite carrier writing $4B of gross written premium on a book of roughly 850,000 new policies and 2.3M renewals per year built a counter-fraud framework in ModelRisk that explicitly modelled leakage before the first claim. The remit was the policy lifecycle from quote to first renewal — application fraud, premium evasion, synthetic identity, and organised-ring exposure — and was deliberately separated from the carrier's claims-stage anti-fraud program, which is run by a different unit.
The distribution above is what a rules engine can never produce. The simulated annual pre-claim leakage on this book has a mean near $695M, a 95th percentile of $1.86B, and a 99th percentile of $3.72B — a right tail so heavy that the expected shortfall beyond the 99th percentile reaches $5.84B. There is a 17% chance leakage exceeds $1,000M in any given year. The deterministic answer cannot price that tail, and pricing the tail is the whole game.
The framework decomposes pre-claim fraud into four channels, each with its own rate and severity:
Application fraud. Misrepresented mileage, garaging postcode, occupation, claim history. Base rate fitted as Beta(3, 350), giving a mean of about 0.85% of new business — a probability distribution rather than a point because the realised rate varies by channel and time. Loss per incident is LogNormal with median around $11K (foregone premium adequacy across the policy term).
Premium evasion. Fully deceptive declarations to attract a lower premium. Rate Beta(8, 220) ≈ 3.5%, severity LogNormal with median ~$1.8K — high frequency, low severity, in aggregate the largest mean leakage channel.
Synthetic and stolen identity. Rate Beta(2, 800) ≈ 0.25%, severity LogNormal with median ~$46K and σ = 0.95 — lower frequency, much heavier severity because synthetic-ID exposures often coincide with imminent claim activity.
Organised-ring policies. Rate Beta(1.5, 1200) ≈ 0.125%, severity LogNormal with median $180K and σ = 1.10 — the smallest channel by count, the dominant channel by tail. Beta is the right family for each rate because rates are bounded on [0, 1] and the data is informative about both the mean and the dispersion; modelling rates as Normal lets the simulator draw negative percentages, which historically has been the first error in-house leakage models make.
The carrier's existing rules-only baseline reported "annual fraud leakage" as a single number — about $240M, derived from suspicious-activity samples grossed up to the book. The Monte Carlo run on the same book says something rather different:
The simulation mean is $695M, the 95th percentile is $1.86B, and the probability that annual leakage exceeds $1,000M is around 17%. The deterministic baseline saw only a point estimate — and saw the wrong one, because it ignored the long-tail contribution of organised rings. A single bad year, in which two or three large rings successfully bind cover, can easily push leakage past $3.7B at the 99th percentile, an event the rules-only view simply did not represent.
Decomposing expected annual leakage by archetype turns the rules-engine intuition on its head:
The first two archetypes — application misrepresentation and premium evasion — account for the bulk of expected leakage by sheer volume of cases. But organised-ring policies, despite occurring in roughly one new policy per thousand, are the dominant contributor to the tail. That is the central counter-fraud insight: a single screening framework cannot serve both problems. High-volume archetypes need cheap, fast probabilistic scoring at quote stage; tail archetypes need expensive network-analytic and identity-graph investigations on a small fraction of binds.
The counter-fraud program — probabilistic scoring at quote stage, third-party identity verification on flagged binds, and organised-ring network detection on the broker channel — was modelled with channel-specific detection reductions and re-simulated:
Mean annual leakage drops from $695M to about $292M — a 58% reduction — the P90 contracts from $1.33B to $549M, and the probability of exceeding the $400M budget tolerance falls from 61% to 19%. The shape of the histogram tightens as well — exactly the outcome the carrier needs to underwrite the program's ROI to the executive committee.
Ring-policy severity tail dominates, which validates the network-analytic investment as the single highest-leverage counter-fraud project. Synthetic-ID detection threshold sits second — a calibration choice the carrier can tune in real time as third-party identity providers improve. Application-fraud base rate and premium-evasion severity follow. False-positive review cost ranks last, which directly counters the most common executive objection to probabilistic screening ("won't this drown my underwriters in false flags?"). It will not, and the sensitivity analysis says so quantitatively.
Pre-claim counter-fraud is not a rules-engine problem; it is a distribution-shaping problem. ModelRisk is what makes the distribution explicit — and once explicit, every decision about investigation capacity, third-party screening, and channel-level appetite can be made against the shape of leakage rather than against a single number that was wrong on both the mean and the tail.