Industry: Biotech Product: ModelRisk Application: Biomarker Discovery
A discovery team screens 2,000 candidate markers across a small case–control cohort, looking for the ~25 that are genuinely associated with disease. The planning spreadsheet multiplies expected power by the marker count and predicts roughly 14 true positives against about 20 false positives — a hit list that looks two-thirds real. Run the same screen 100,000 times with the biology and the batch noise allowed to vary, and the picture inverts: the list carries 14 true positives drowning in 59 false positives, a realised false-discovery rate of 80%. The point estimate was not merely optimistic; it was on the wrong side of the 50% line.
That gap is not rounding error. It comes from a source the deterministic calculation cannot see: cohort-to-cohort quality. A dirty extraction or a noisy batch simultaneously weakens every true marker's signal and inflates the rate at which null markers cross the threshold. The two effects move together, so bad cohorts are bad on both counts at once — and the screen-level outcome stays genuinely uncertain instead of averaging away over 2,000 tests.
Multiple testing is the whole problem, and it is multiplicative. With 2,000 candidates and an uncorrected per-test threshold of α = 0.01, the expected number of false positives is already 20 even when everything goes to plan. But "to plan" assumes a fixed effect size and a fixed noise floor. In reality a shared, per-cohort quality factor scales the effect size of every true marker (modelled as a quality-dependent Cohen's d) and inflates the per-null false-positive rate by up to five-fold on the worst batches. A single mean-of-means calculation collapses that common factor to a constant and reports a tidy 14-to-20 split. The simulation keeps it, and the list balloons.
Across 100,000 simulated screens the true positives recovered average 14 (P10–P90: 7–21) while the false positives average 59 (P10–P90: 40–79). The deterministic plan marks 14 true positives — coincidentally close to the simulated mean — but its companion estimate of 20 false positives understates the real figure threefold, because it ignores the noise inflation on poor cohorts. The realised false-discovery rate averages 80% and reaches 91% at the P90. A team validating the top of this list with no FDR control would spend most of its follow-up budget chasing markers that do not exist.
Carrying the top 12 hits into a validated diagnostic panel, the probability that the panel clears both an 85% sensitivity and a 90% specificity target rises steeply with cohort size. At 40 cases per group the panel hits the joint target only 47% of the time; at the planned 60 per group it reaches 71%; by 130 per group it is 99%. The 80% reliability line is crossed at roughly 70 cases per group — a concrete, defensible enrolment target rather than a guess. The leverage is real because a larger cohort raises per-marker power, which raises the genuine fraction of the carried panel, which is what the panel's sensitivity and specificity ultimately depend on.
Against a baseline realised FDR of 80%, the per-test threshold dominates: tightening α from 0.05 to 0.001 cuts FDR from 93% to 48%, while loosening it does the reverse. Cohort size matters next — 40 versus 150 per group moves FDR from 85% to 71% — followed by true effect size (d 0.30 vs 0.65 → 89% vs 73%). The batch-noise inflation factor and the number of genuinely associated markers round out the ranking. The message for the protocol is unambiguous: a stringent threshold and an adequately powered cohort are the two highest-value design choices, and they are choices the team controls.
Plotting each screen's true-positive yield against its realised FDR shows the shared-quality factor at work: low-yield screens are also the dirtiest screens. The probability of landing in the worst quadrant — fewer than 15 real hits and a majority-false list — is 57%. Across all screens the median yield is 13 true positives at a median FDR of 81%, and 100% of screens exceed a 50% FDR at this threshold. The bad outcomes are correlated, not independent, which is exactly why a deterministic "expected" hit list is misleading.
The team stopped treating the screen as an exercise in counting expected hits and started treating it as a multiple-testing problem with correlated, cohort-level uncertainty. Three decisions followed directly from the simulation: adopt FDR control rather than an uncorrected per-test threshold (the baseline 80% FDR is otherwise indefensible); set enrolment at roughly 70 cases per group, the point where the validated panel clears its joint target with 80% probability; and budget validation around the realistic 14-true-positive yield rather than the optically larger raw hit list. The probabilistic view replaced a single, comforting count with an honest range — and changed where the discovery budget went.