Industry: Healthcare and Epidemiology Product: ModelRisk Application: Medical Device Testing — 510(k) Validation Campaign and FDA Gating
The validation plan for a Class II AI-assisted cardiovascular imaging device pre-specified sensitivity ≥ 88% and specificity ≥ 90% as the dual primary endpoints for the FDA 510(k) submission. The deterministic V&V dashboard reported point estimates of sensitivity = 0.887 and specificity = 0.929 — both above the threshold, the device "passes." The Monte Carlo simulation on the same cohort data said the joint probability of meeting both endpoints in the as-submitted readout is roughly 58% — barely better than a coin flip on a $14M validation programme. The point estimate hid that the sensitivity figure was sitting fractions of a percentage point above the cliff edge, with a 95% credible interval that crossed it.
Each dot is one simulated as-submitted readout under the posterior uncertainty in the two operating characteristics. The two endpoints are weakly correlated, and the joint probability that both pass — the green region — is about 58%, not the implied 100% of a deterministic plan that just says "0.887 ≥ 0.88." The cliff edge is on the sensitivity axis: 93% of trials beat the specificity bar but only 63% beat sensitivity, and the joint is 58%. Specificity is comfortable; sensitivity is the gate.
The device manufacturer rebuilt the validation-readout and submission-gating analysis in ModelRisk. The result was a re-scoped validation campaign that added 320 cases to the reader study before lock — and the lock-day readout showed both endpoints met with P(both) > 95%, with no further FDA-cycle delay.
The validation cohort contained 1,200 patients (480 disease-positive by ground truth, 720 disease-negative). On the curated dataset the device flagged 426 true positives, 54 false negatives, 669 true negatives, 51 false positives. The deterministic readout divides these and reports two numbers. The probabilistic readout treats each operating characteristic as a Beta posterior on the underlying Bernoulli rate:
Beta is the natural conjugate prior for a probability — bounded on [0, 1], parameterized by event counts. The deterministic point estimate ignores the posterior width, which on a 1,200-case study is non-trivial.
Pre-submission bench testing ran 12 devices for 6 months = 72 device-months. Five failure modes were tracked, each as an independent Bernoulli trial per device-month. The submission plan triggers a 30–120 day bench rework (Triangular(30, 60, 120)) whenever total failures cross the tail threshold of 7 or more. The simulation gives the probability of a rework trigger as ~16% and an expected rework delay (conditional on trigger) of 70 days.
The Pareto ranks contribution to expected rework days. Sensor drift > 3σ ranks first — at 0.022/device-month and a 22-day rework if it occurs, it accounts for ~40% of expected rework time. Calibration shift after temperature cycling ranks second (~22%): rarer at 0.009/device-month but a costly 30-day rework. That ranking redirected the engineering team's pre-submission attention to the sensor-drift and calibration sub-systems, where a firmware fix dropped the per-device-month calibration-drift probability from 0.009 to 0.003 in a follow-up run.
The single highest-leverage lever for the joint endpoint probability is true sensitivity — moving the underlying rate from 0.86 to 0.92 changes the joint pass probability by ±28 percentage points. Cohort size ranks third: doubling the reader study from 800 to 1,600 cases moves the joint probability by ±12 pp by tightening the Beta posterior. That comparison is what told the regulatory team that adding 320 cases (cheaper, faster, deterministic schedule impact) outperformed any further pre-submission algorithm tuning (slower, uncertain, may regress the specificity figure).
Submission readiness was day 240 from project start. FDA 510(k) review for AI/ML-enabled software-as-a-medical-device runs LogNormal centered around 4.5 months (135 days, σ_log = 0.30) — review queue plus information-request cycles. Combined with the bench-rework triggered at p ≈ 16%, the simulation produces a stochastic finish date:
The deterministic plan landed in December 2026; the simulation says P50 finish is January 2027, P90 is April 2027 — three to four months of true contingency the deterministic schedule did not show. That contingency was real money: the device launch had been promised to two distribution partners on the deterministic date, and at the P50 the company would have been negotiating delay penalties.
The deterministic readout said the device passed; the probabilistic readout said the device passed with a coin-flip's worth of confidence, which on a $14M programme is a different sentence entirely. Monte Carlo simulation in ModelRisk is what turns a single "above the threshold" point estimate into a number the submission committee can actually buy down to a defensible level before the lock.