| Vose Software

Industry: Insurance and Reinsurance
Product: ModelRisk
Application: Claims Processing Optimization


When the Queue Explodes: Probabilistic Cycle-Time Modeling for an Auto-Claims Back Office

At 92% average adjuster utilization, an auto-claims back office looks healthy on paper. Mean cycle time sits around 18 days, complaints are stable, and the headcount line in the operating plan looks defensible. Then a Tuesday morning brings 1.4× the usual FNOL volume, and by Friday the queue has tripled, the 30-day SLA is in breach on roughly 20% of files, and senior leadership wants a staffing answer by Monday. The deterministic capacity plan — which divides expected daily volume by expected per-file handle time — cannot answer this question, because the answer lives in the non-linear queueing penalty that kicks in above ~85% utilization, and in the distribution of per-file processing time, not its mean.

A personal-lines auto carrier processing roughly 420 FNOLs a day across its U.S. operation rebuilt the cycle-time model for its claims back office in ModelRisk. The whole staffing case comes down to one before-and-after picture: the same total headcount, rebalanced to pull average utilization down from 92% to 78%, collapses the heavy right tail of the cycle-time distribution.

End-to-end claim cycle time — current staffing vs rebalanced plan

The objective was not "faster claims"; it was the right shape of staffing plan to balance cycle-time targets, surge resilience, and operating cost — a problem with no defensible deterministic answer.

Where the cycle time actually lives

End-to-end cycle time is the sum of stage times across six steps: FNOL intake, triage, assignment, investigation, adjustment, and settlement. Each stage was fitted to its own distribution rather than represented by an average.

Triangular distributions handled intake (0.1, 0.3, 0.8 days), triage (0.2, 0.6, 1.4), assignment (0.1, 0.4, 1.0), adjustment (1.0, 2.5, 6.0), and settlement (0.5, 1.5, 4.0) — expert-shaped families with a defensible minimum, mode, and maximum.

Investigation time was fitted as LogNormal with ln-mean of 4.0 days and σ = 0.55, because investigation outcomes have a long right tail that triangular cannot capture: 80% of investigations finish in under 6 days, but a meaningful tail runs past 12 days when third-party reports, body-shop estimates, or recorded statements drag.

Queue wait under capacity pressure was the load-bearing piece. When adjuster utilization rose above 85%, expected queue time was modelled as Gamma(shape 2.2, scale (utilization − 0.85) × 40) — a heavy-tailed wait distribution that captured the M/M/c-like blow-up real teams experience when load creeps near saturation. Modelling queue time as a fixed adder, as the spreadsheet baseline did, hides the very non-linearity that turns a small mis-staffing into a major SLA breach.

Stage-time input distributions

What the deterministic capacity plan got wrong

The deterministic plan said the back office could handle steady-state volume at 92% utilization with a comfortable mean cycle time. The Monte Carlo run on the same staffing said: yes, the mean is around 18 days — and the P90 is 24 days, the P95 is 27 days, and roughly 2% of files breach the 30-day SLA every month even in steady state. The mean was not wrong; the planning question was wrong. The right question is not "what is average cycle time?" but "what fraction of files breaches SLA, and how does that fraction respond to staffing changes?"

The rebalanced plan — same total headcount, redistributed to bring average utilization down to 78% by adding capacity at the investigation stage and trimming over-staffed intake — produces the noticeably tighter distribution shown in the opening chart: mean cycle time falls from 17.6 to 12.3 days, P90 contracts from 24 to about 16 days, and the share of files breaching the 30-day SLA drops from roughly 2% to under 1% — without adding a single FTE. That is the queueing non-linearity working in reverse: a 14-point reduction in utilization shaves the heavy tail far more than it shaves the mean.

Surge: the catastrophe stress test

Pure cycle-time optimisation misses the question that matters most to the COO: what happens when a hurricane lands and FNOL volume spikes 60% in a week? The same simulation, with utilization driven up by the surge multiplier, answers it:

Catastrophe surge stress test — three staffing plans

Under surge with the current staffing plan, cycle time's P95 runs past 40 days and more than one file in five breaches the 30-day SLA — an outright service collapse. Under the rebalanced plan plus a pre-arranged reserve adjuster pool that can be activated within 48 hours of a Cat-3 landfall, the surge-stressed P95 holds around 35 days and the breach rate roughly halves. The two-line difference between those CDFs is precisely the operational case for the reserve pool — a contract the deterministic capacity plan could neither price nor justify.

What drives the SLA-breach number

Drivers of P90 cycle time

Adjuster utilization dominates by a wide margin — confirming that staffing rebalance, not process redesign, is the highest-leverage action. Investigation-time variance ranks second: tightening variance (e.g., better third-party report SLAs with body shops) shifts the P90 noticeably even when the mean investigation time is unchanged. FNOL arrival rate ranks third, which is mostly out of the back office's control but feeds the surge planning case directly.

What changed

  • Mean cycle time fell from 17.6 days to 12.3 days at constant headcount, driven by the staffing rebalance and the investigation-stage capacity uplift.
  • SLA breach rate dropped from roughly 2% to under 1% in steady state and from about 20% to roughly 11% during the modelled hurricane surge.
  • Annual operating cost reduced by 9% through retirement of overtime patterns the previous capacity plan was systematically requiring.
  • The reserve-adjuster contract was approved on the simulation evidence — a vendor agreement that activates within 48 hours of qualifying weather events and prices out at a fraction of the captive-headcount alternative.

ModelRisk Functionality Used

  • Stage-level Monte Carlo composition — Triangular and LogNormal distributions per stage, summed into end-to-end cycle time, replacing the single-mean spreadsheet.
  • Non-linear queueing model — Gamma-distributed wait times whose parameters rise sharply with utilization, capturing the SLA-breach blow-up the deterministic model missed.
  • Scenario stress testing — the same engine repointed at surge volumes produced the catastrophe stress CDF and underwrote the reserve-pool contract.
  • Sensitivity ranking that placed utilization and investigation variance at the top, shifting the year's operational-improvement agenda away from FNOL automation and toward stage-level capacity balancing.
  • Output overlays comparing two staffing plans on the same chart, used directly in the operations committee that approved the rebalance.
  • Excel-native authoring, so the staffing planner could iterate adjuster counts cell-by-cell and re-run without leaving the workbook the operations team already maintained.

The lesson the cycle-time team kept repeating after deployment: at high utilization, the difference between the mean and the P95 is not a rounding error — it is the gap between an SLA you meet and an SLA you breach. Monte Carlo is the only way to plan for the second number while still hitting the first.