| Vose Software

Industry: Defense
Product: ModelRisk
Application: Threat Detection


Tuned for 78% Detection, the SOC Still Misses 30% of Real Threats

A defense security operations centre screens roughly 9,000 benign events a day against an average of 2.2 real intrusions. Set the detection threshold low and you catch more true threats — but the false-alarm flood buries the analyst queue and real alerts age past the response window. Set it high and the queue is calm but threats slip through undetected. The legacy ISR-style model treated detection probability Pd as a fixed sensor spec. Rebuilt in ModelRisk over 120,000 simulated SOC-days, the truth is operational: at the chosen operating point Pd is 78%, yet the per-threat miss rate is 30% — because on 20% of days the alert queue overflows triage capacity and detected threats die in the backlog. The threshold is not a detector setting; it is a staffing decision.

This is an operations problem — detection rates, queue load, time-to-respond — not a loss-dollar or kill-chain problem. The deliverable is the operating point and the staffing level that keep true threats from aging out.

Detection ROC curve with the chosen operating point and its analyst load

Why a point estimate fails here

"Our detector has Pd = 78%" is true and useless. It ignores that Pd is one end of a trade governed by the ROC operating point: you cannot raise Pd without raising the false-alarm rate, and false alarms are not free — they consume the same finite analyst-hours that true positives need. A point Pd also assumes every day looks average. Real threat days cluster: the same elevated-tempo day that brings extra intrusions also brings extra benign anomalies (patch storms, exercises, ops surges), so the queue overflows precisely when it matters most. Only a simulation that draws daily volume and serves it through a finite queue reveals the gap between detector Pd (78%) and delivered detection (70% — a 30% miss rate).

Modelling the SOC as a detector feeding a finite queue

Daily arrivals. True intrusions are Poisson (mean 2.2/day) and benign events Poisson (mean ~9,000/day) — but both rates are scaled by a single shared daily threat-tempo factor (Gamma, mean 1). High-tempo days inflate true and benign volume together. The model verifies the coupling: the correlation between daily true-threat and benign-event counts is 0.684 with the shared factor, versus 0.001 in an independent model. That clustering is load-bearing — it drives queue-overflow probability to 20%, where an independent model shows 0%. Modelling busy days as uncorrelated would declare the SOC safely staffed when it is not.

Detector. The operating point is one point on a power-law ROC (Pd = FAR^(1/k), k = 18 for a strong detector). At the chosen threshold the false-alarm rate is 1.2% of benign events → Pd = 78%. Each true threat is detected with probability Pd; each benign event raises a false alarm with probability FAR.

Queue. True-positive and false-positive alerts join one triage queue served at 40 alerts/analyst/shift × 4 analysts = 160/day. When inflow exceeds capacity, the backlog buries a proportional share of all queued alerts — including detected true positives, which then miss the response window. A true threat is "missed within the MTTD window" if it was never alerted (1 − Pd) or alerted but stuck in backlog.

At the operating point: mean 109 alerts/day, of which 98% are false positives; P90 203/day, P99 333/day.

The threshold sweep — detection, miss rate, and load in one picture

Threshold sweep showing detection probability, miss rate, and analyst load

Sweeping the threshold reveals a trade that a fixed Pd hides completely:

  • At a low threshold (FAR 0.2%): detector Pd is 71%, load is a quiet 20 alerts/day, and the per-threat miss rate is 29% — almost all of it the detector simply not firing.
  • At the operating point (FAR 1.2%): Pd 78%, load 109/day, miss rate 30%.
  • At a high threshold (FAR 4%): detector Pd climbs to 84%, but load explodes to 363 alerts/day — far past the 160/day capacity — and the per-threat miss rate jumps to 65%. Raising the threshold made the detector better and the SOC far worse, because the false-alarm flood buried the true positives.

That non-monotone miss curve — better detection, worse outcomes — is the entire reason a SOC cannot be tuned on Pd alone.

The queue load against capacity

Daily alert queue load distribution against triage capacity

The daily alert distribution is right-skewed by the threat-tempo clustering: mean 109/day, P90 203/day, P99 333/day, against a fixed capacity of 160/day. The shaded tail beyond capacity is the 20% of days the queue overflows — and because tempo couples true and benign volume, those overflow days are disproportionately the days carrying real intrusions.

Staffing is the lever the SOC manager actually controls

The threshold is set by the detection-engineering team; the shift manager controls headcount. Holding the threshold fixed, the model sweeps analysts on shift.

Staffing sweep of overload probability and per-threat miss rate versus analysts

  • 2 analysts: 60% of days overload, 52% per-threat miss rate.
  • 4 analysts (baseline): 20% overload, 30% miss rate.
  • 5 analysts: 11% overload, 26% miss rate.
  • 7 analysts: 3% overload, 23% miss rate — asymptotically approaching the detector's own 22% floor (1 − Pd), the irreducible miss the queue cannot fix.

The marginal analyst from 4 → 5 removes 9 points of overload and 4 points of miss rate; beyond 6, you are paying to chase the detector floor, and the better spend is a higher-quality detector (a steeper ROC), not more headcount.

What the model changed

  • The SOC stopped reporting detector Pd (78%) as its detection metric and adopted delivered detection / per-threat miss rate (30%) — the number that accounts for the queue.
  • Staffing was set at 5 analysts rather than 4, the point where the marginal analyst still buys meaningful miss-rate reduction (30% → 26%) before diminishing returns.
  • The detection-engineering roadmap reprioritised toward ROC quality over simply lowering the threshold, after the sweep showed a higher threshold raises the net miss rate by flooding the queue.

ModelRisk Functionality Used

  • Poisson arrival processes for both true intrusions (2.2/day) and benign events (~9,000/day), feeding a frequency-driven alert model.
  • A shared daily threat-tempo factor (Gamma) coupling true and benign volume — verified to lift their correlation from 0.001 to 0.684 and queue-overflow probability from 0% to 20%.
  • A power-law ROC operating point (FAR 1.2% → Pd 78%) swept across thresholds to expose the non-monotone miss curve (miss rate 29% → 65% as the threshold rises).
  • A finite-capacity triage queue translating alert inflow into backlog and a delivered per-threat miss rate (30%) distinct from detector Pd.
  • A staffing sweep quantifying overload and miss rate from 2 to 7 analysts (60%/52% down to 3%/23%) to set headcount at the marginal-value point.

Threat detection is not a sensor specification — it is a queueing system where every threshold choice spends analyst-hours, and Monte Carlo is what turns "our detector catches 78%" into "we actually action 70%, and here is the staffing level that closes the gap."