| Vose Software

Industry: Healthcare and Epidemiology
Product: ModelRisk
Application: Medical Device Testing — 510(k) Validation Campaign and FDA Gating


A Class II AI-Imaging Device on the Razor's Edge of Its Sensitivity Endpoint

The validation plan for a Class II AI-assisted cardiovascular imaging device pre-specified sensitivity ≥ 88% and specificity ≥ 90% as the dual primary endpoints for the FDA 510(k) submission. The deterministic V&V dashboard reported point estimates of sensitivity = 0.887 and specificity = 0.929 — both above the threshold, the device "passes." The Monte Carlo simulation on the same cohort data said the joint probability of meeting both endpoints in the as-submitted readout is roughly 58% — barely better than a coin flip on a $14M validation programme. The point estimate hid that the sensitivity figure was sitting fractions of a percentage point above the cliff edge, with a 95% credible interval that crossed it.

Joint endpoint outcomes — green = both endpoints met

Each dot is one simulated as-submitted readout under the posterior uncertainty in the two operating characteristics. The two endpoints are weakly correlated, and the joint probability that both pass — the green region — is about 58%, not the implied 100% of a deterministic plan that just says "0.887 ≥ 0.88." The cliff edge is on the sensitivity axis: 93% of trials beat the specificity bar but only 63% beat sensitivity, and the joint is 58%. Specificity is comfortable; sensitivity is the gate.

The device manufacturer rebuilt the validation-readout and submission-gating analysis in ModelRisk. The result was a re-scoped validation campaign that added 320 cases to the reader study before lock — and the lock-day readout showed both endpoints met with P(both) > 95%, with no further FDA-cycle delay.

The endpoint distributions, treated honestly

The validation cohort contained 1,200 patients (480 disease-positive by ground truth, 720 disease-negative). On the curated dataset the device flagged 426 true positives, 54 false negatives, 669 true negatives, 51 false positives. The deterministic readout divides these and reports two numbers. The probabilistic readout treats each operating characteristic as a Beta posterior on the underlying Bernoulli rate:

  • Sensitivity: Beta(110, 14) — posterior mean 0.887, 95% CI [0.83, 0.94]
  • Specificity: Beta(170, 13) — posterior mean 0.929, 95% CI [0.89, 0.96]

Beta is the natural conjugate prior for a probability — bounded on [0, 1], parameterized by event counts. The deterministic point estimate ignores the posterior width, which on a 1,200-case study is non-trivial.

Bench failures: the second gate

Pre-submission bench testing ran 12 devices for 6 months = 72 device-months. Five failure modes were tracked, each as an independent Bernoulli trial per device-month. The submission plan triggers a 30–120 day bench rework (Triangular(30, 60, 120)) whenever total failures cross the tail threshold of 7 or more. The simulation gives the probability of a rework trigger as ~16% and an expected rework delay (conditional on trigger) of 70 days.

Pareto of bench-rework drivers — failure mode prob × rework days

The Pareto ranks contribution to expected rework days. Sensor drift > 3σ ranks first — at 0.022/device-month and a 22-day rework if it occurs, it accounts for ~40% of expected rework time. Calibration shift after temperature cycling ranks second (~22%): rarer at 0.009/device-month but a costly 30-day rework. That ranking redirected the engineering team's pre-submission attention to the sensor-drift and calibration sub-systems, where a firmware fix dropped the per-device-month calibration-drift probability from 0.009 to 0.003 in a follow-up run.

Drivers of the joint endpoint probability

Tornado — drivers of P(both 510(k) endpoints met)

The single highest-leverage lever for the joint endpoint probability is true sensitivity — moving the underlying rate from 0.86 to 0.92 changes the joint pass probability by ±28 percentage points. Cohort size ranks third: doubling the reader study from 800 to 1,600 cases moves the joint probability by ±12 pp by tightening the Beta posterior. That comparison is what told the regulatory team that adding 320 cases (cheaper, faster, deterministic schedule impact) outperformed any further pre-submission algorithm tuning (slower, uncertain, may regress the specificity figure).

From submission to clearance

Submission readiness was day 240 from project start. FDA 510(k) review for AI/ML-enabled software-as-a-medical-device runs LogNormal centered around 4.5 months (135 days, σ_log = 0.30) — review queue plus information-request cycles. Combined with the bench-rework triggered at p ≈ 16%, the simulation produces a stochastic finish date:

510(k) clearance date — submission to FDA decision

The deterministic plan landed in December 2026; the simulation says P50 finish is January 2027, P90 is April 2027 — three to four months of true contingency the deterministic schedule did not show. That contingency was real money: the device launch had been promised to two distribution partners on the deterministic date, and at the P50 the company would have been negotiating delay penalties.

What the model changed

  • Reader-study extension funded before submission lock: 320 additional cases at $260,000 total, moving simulated P(both endpoints met) from 58% to over 95%. The alternative — submit on the original cohort and accept a coin-flip on a deficiency letter — would have cost an estimated 6 months of cycle time and $1.4M of opportunity cost.
  • Sensor-drift and calibration sub-system firmware fix identified by the Pareto as the highest-ROI engineering change pre-submission; per-device-month calibration-drift probability dropped from 0.009 to 0.003 in the rerun.
  • Distribution-partner contracts re-negotiated to a clearance-conditioned launch window of Jan 2027 – Apr 2027 rather than a fixed Dec 2026 commit, eliminating the penalty exposure.
  • FDA pre-submission meeting restructured: rather than presenting point-estimate operating characteristics, the team brought the posterior distributions and the simulated joint probability, which the agency reviewers described as the clearest readout they had received on a similar device that quarter.
  • Q-Sub strategy template propagated to the company's other two in-flight 510(k) programs.

ModelRisk Functionality Used

  • Beta posteriors on cohort counts (Beta(110, 14) for sensitivity, Beta(170, 13) for specificity) replacing the deterministic point estimates and exposing the 58% joint-pass probability the dashboard had hidden.
  • Joint endpoint scatter under the dependence structure of the two operating characteristics, producing the green/red separation chart that the submission committee used as the gating visual.
  • Per-failure-mode Binomial draws across 72 device-months, aggregated to a bench-rework trigger with a Triangular(30, 60, 120) rework-delay tail.
  • Tornado on the joint pass probability that ranked cohort size above any further algorithm tuning, redirecting $260K of pre-submission spend to the highest-leverage lever.
  • Composite finish-date simulation combining submission-prep duration, conditional bench rework, and LogNormal FDA review, producing the P10/P50/P90 clearance-date distribution.
  • Pareto on failure modes with probability × rework-day impact, ranking sensor-drift and calibration-shift firmware as the engineering team's first call.

The deterministic readout said the device passed; the probabilistic readout said the device passed with a coin-flip's worth of confidence, which on a $14M programme is a different sentence entirely. Monte Carlo simulation in ModelRisk is what turns a single "above the threshold" point estimate into a number the submission committee can actually buy down to a defensible level before the lock.