Industry: Insurance and Reinsurance Product: ModelRisk Application: Claims Processing Optimization
At 92% average adjuster utilization, an auto-claims back office looks healthy on paper. Mean cycle time sits around 18 days, complaints are stable, and the headcount line in the operating plan looks defensible. Then a Tuesday morning brings 1.4× the usual FNOL volume, and by Friday the queue has tripled, the 30-day SLA is in breach on roughly 20% of files, and senior leadership wants a staffing answer by Monday. The deterministic capacity plan — which divides expected daily volume by expected per-file handle time — cannot answer this question, because the answer lives in the non-linear queueing penalty that kicks in above ~85% utilization, and in the distribution of per-file processing time, not its mean.
A personal-lines auto carrier processing roughly 420 FNOLs a day across its U.S. operation rebuilt the cycle-time model for its claims back office in ModelRisk. The whole staffing case comes down to one before-and-after picture: the same total headcount, rebalanced to pull average utilization down from 92% to 78%, collapses the heavy right tail of the cycle-time distribution.
The objective was not "faster claims"; it was the right shape of staffing plan to balance cycle-time targets, surge resilience, and operating cost — a problem with no defensible deterministic answer.
End-to-end cycle time is the sum of stage times across six steps: FNOL intake, triage, assignment, investigation, adjustment, and settlement. Each stage was fitted to its own distribution rather than represented by an average.
Triangular distributions handled intake (0.1, 0.3, 0.8 days), triage (0.2, 0.6, 1.4), assignment (0.1, 0.4, 1.0), adjustment (1.0, 2.5, 6.0), and settlement (0.5, 1.5, 4.0) — expert-shaped families with a defensible minimum, mode, and maximum.
Investigation time was fitted as LogNormal with ln-mean of 4.0 days and σ = 0.55, because investigation outcomes have a long right tail that triangular cannot capture: 80% of investigations finish in under 6 days, but a meaningful tail runs past 12 days when third-party reports, body-shop estimates, or recorded statements drag.
Queue wait under capacity pressure was the load-bearing piece. When adjuster utilization rose above 85%, expected queue time was modelled as Gamma(shape 2.2, scale (utilization − 0.85) × 40) — a heavy-tailed wait distribution that captured the M/M/c-like blow-up real teams experience when load creeps near saturation. Modelling queue time as a fixed adder, as the spreadsheet baseline did, hides the very non-linearity that turns a small mis-staffing into a major SLA breach.
The deterministic plan said the back office could handle steady-state volume at 92% utilization with a comfortable mean cycle time. The Monte Carlo run on the same staffing said: yes, the mean is around 18 days — and the P90 is 24 days, the P95 is 27 days, and roughly 2% of files breach the 30-day SLA every month even in steady state. The mean was not wrong; the planning question was wrong. The right question is not "what is average cycle time?" but "what fraction of files breaches SLA, and how does that fraction respond to staffing changes?"
The rebalanced plan — same total headcount, redistributed to bring average utilization down to 78% by adding capacity at the investigation stage and trimming over-staffed intake — produces the noticeably tighter distribution shown in the opening chart: mean cycle time falls from 17.6 to 12.3 days, P90 contracts from 24 to about 16 days, and the share of files breaching the 30-day SLA drops from roughly 2% to under 1% — without adding a single FTE. That is the queueing non-linearity working in reverse: a 14-point reduction in utilization shaves the heavy tail far more than it shaves the mean.
Pure cycle-time optimisation misses the question that matters most to the COO: what happens when a hurricane lands and FNOL volume spikes 60% in a week? The same simulation, with utilization driven up by the surge multiplier, answers it:
Under surge with the current staffing plan, cycle time's P95 runs past 40 days and more than one file in five breaches the 30-day SLA — an outright service collapse. Under the rebalanced plan plus a pre-arranged reserve adjuster pool that can be activated within 48 hours of a Cat-3 landfall, the surge-stressed P95 holds around 35 days and the breach rate roughly halves. The two-line difference between those CDFs is precisely the operational case for the reserve pool — a contract the deterministic capacity plan could neither price nor justify.
Adjuster utilization dominates by a wide margin — confirming that staffing rebalance, not process redesign, is the highest-leverage action. Investigation-time variance ranks second: tightening variance (e.g., better third-party report SLAs with body shops) shifts the P90 noticeably even when the mean investigation time is unchanged. FNOL arrival rate ranks third, which is mostly out of the back office's control but feeds the surge planning case directly.
The lesson the cycle-time team kept repeating after deployment: at high utilization, the difference between the mean and the P95 is not a rounding error — it is the gap between an SLA you meet and an SLA you breach. Monte Carlo is the only way to plan for the second number while still hitting the first.