Industry: Defense Product: ModelRisk Application: Threat Detection
A defense security operations centre screens roughly 9,000 benign events a day against an average of 2.2 real intrusions. Set the detection threshold low and you catch more true threats — but the false-alarm flood buries the analyst queue and real alerts age past the response window. Set it high and the queue is calm but threats slip through undetected. The legacy ISR-style model treated detection probability Pd as a fixed sensor spec. Rebuilt in ModelRisk over 120,000 simulated SOC-days, the truth is operational: at the chosen operating point Pd is 78%, yet the per-threat miss rate is 30% — because on 20% of days the alert queue overflows triage capacity and detected threats die in the backlog. The threshold is not a detector setting; it is a staffing decision.
This is an operations problem — detection rates, queue load, time-to-respond — not a loss-dollar or kill-chain problem. The deliverable is the operating point and the staffing level that keep true threats from aging out.
"Our detector has Pd = 78%" is true and useless. It ignores that Pd is one end of a trade governed by the ROC operating point: you cannot raise Pd without raising the false-alarm rate, and false alarms are not free — they consume the same finite analyst-hours that true positives need. A point Pd also assumes every day looks average. Real threat days cluster: the same elevated-tempo day that brings extra intrusions also brings extra benign anomalies (patch storms, exercises, ops surges), so the queue overflows precisely when it matters most. Only a simulation that draws daily volume and serves it through a finite queue reveals the gap between detector Pd (78%) and delivered detection (70% — a 30% miss rate).
Daily arrivals. True intrusions are Poisson (mean 2.2/day) and benign events Poisson (mean ~9,000/day) — but both rates are scaled by a single shared daily threat-tempo factor (Gamma, mean 1). High-tempo days inflate true and benign volume together. The model verifies the coupling: the correlation between daily true-threat and benign-event counts is 0.684 with the shared factor, versus 0.001 in an independent model. That clustering is load-bearing — it drives queue-overflow probability to 20%, where an independent model shows 0%. Modelling busy days as uncorrelated would declare the SOC safely staffed when it is not.
Detector. The operating point is one point on a power-law ROC (Pd = FAR^(1/k), k = 18 for a strong detector). At the chosen threshold the false-alarm rate is 1.2% of benign events → Pd = 78%. Each true threat is detected with probability Pd; each benign event raises a false alarm with probability FAR.
Pd = FAR^(1/k)
Queue. True-positive and false-positive alerts join one triage queue served at 40 alerts/analyst/shift × 4 analysts = 160/day. When inflow exceeds capacity, the backlog buries a proportional share of all queued alerts — including detected true positives, which then miss the response window. A true threat is "missed within the MTTD window" if it was never alerted (1 − Pd) or alerted but stuck in backlog.
At the operating point: mean 109 alerts/day, of which 98% are false positives; P90 203/day, P99 333/day.
Sweeping the threshold reveals a trade that a fixed Pd hides completely:
That non-monotone miss curve — better detection, worse outcomes — is the entire reason a SOC cannot be tuned on Pd alone.
The daily alert distribution is right-skewed by the threat-tempo clustering: mean 109/day, P90 203/day, P99 333/day, against a fixed capacity of 160/day. The shaded tail beyond capacity is the 20% of days the queue overflows — and because tempo couples true and benign volume, those overflow days are disproportionately the days carrying real intrusions.
The threshold is set by the detection-engineering team; the shift manager controls headcount. Holding the threshold fixed, the model sweeps analysts on shift.
The marginal analyst from 4 → 5 removes 9 points of overload and 4 points of miss rate; beyond 6, you are paying to chase the detector floor, and the better spend is a higher-quality detector (a steeper ROC), not more headcount.
Threat detection is not a sensor specification — it is a queueing system where every threshold choice spends analyst-hours, and Monte Carlo is what turns "our detector catches 78%" into "we actually action 70%, and here is the staffing level that closes the gap."