Industry: Engineering Product: ModelRisk Application: Quantifying the Statistical Power of a Proof-Test Qualification Program
A pressure-vessel supplier qualified each production lot the conventional way: proof-test 20 units at 1.25× design load, accept the lot if none fail. The contract required a demonstrated reliability of R ≥ 0.999 at design load. Passing the test felt like proving the target had been met. The reliability engineer asked the question the pass/fail rule cannot answer: of the lots that pass this test, how many are actually below 0.999? The Monte Carlo answer was 4.9% — about one passing lot in twenty is a sub-target lot that slipped through. Worse, the same test was rejecting 38% of genuinely good lots. The qualification program was simultaneously too weak to catch bad lots and too harsh on good ones.
This is not a statement about any single vessel's stress margin — it is a statement about the test program's statistical power. The team modelled the whole qualification scheme in ModelRisk: an uncertain population of lots, a proof test applied to each, and the gap between what the test demonstrates and what is actually true.
The key realism is that the vendor ships a population of lots, each with its own true reliability — median strength wanders around 1.55× design load and the coefficient of variation around 12%, lot to lot. The histogram splits those lots by test outcome. Lots that pass have a mean true reliability of 0.99980; lots that fail average 0.99708. But the two groups overlap, and the overlap to the left of the 0.999 target is the consumer's risk: sub-target lots that passed anyway. With a zero-failure test on 20 units, that overlap is 4.9% of all passing lots.
A unit survives the proof test if its true strength exceeds the proof load — a Bernoulli outcome whose probability p = P(strength > proof load) is governed by the lot's (uncertain) strength distribution. The number of failures in a lot of N units is then Binomial(N, 1−p), and the lot is accepted under a zero-failure plan if that count is zero. The classic reliability-demonstration identity follows directly: for a zero-failure plan, the probability a marginal lot passes is p^N.
Each lot's true design-load reliability and its proof-pass probability are both derived from a LogNormal strength model whose median and CV are themselves drawn per lot. Modelling the lot quality as uncertain — rather than assuming every lot is the nominal "good" lot — is what makes consumer's and producer's risk meaningful quantities rather than zero by assumption.
A single pass/fail outcome is one Bernoulli draw on top of an uncertain lot. It carries no statement of how discriminating the test is. A test can pass for two completely different reasons — the lot is genuinely strong, or the lot is marginal and 20 units happened not to find the weak one. The deterministic mindset ("it passed, therefore R ≥ 0.999") cannot separate them. Only by simulating across the population of lots, and tracking each lot's true reliability alongside its test outcome, do the two risks become visible:
The 38% producer's risk is the finding that surprised the team: a 1.25× proof factor with a zero-failure rule is so demanding that more than a third of compliant lots get scrapped or re-tested. The test was expensive in exactly the way nobody had measured.
For a lot sitting just below target, consumer's risk falls as p^N. Sweeping the zero-failure sample size shows how slowly a marginal lot is caught:
A lot at R = 0.998 — only a hair below the 0.999 target — needs 14 units just to push its pass probability under 10%, and far more to discriminate it sharply. The current N = 20 (dashed line) sits in the region where a marginally-bad lot still passes uncomfortably often, because the closer a bad lot is to the target, the harder it is to tell apart from a good one.
The OC curve plots P(lot passes) against the lot's true reliability — the test's ability to tell good lots from bad. Comparing the current N = 20 plan with a proposed N = 60 plan:
At N = 60 the curve drops far more steeply through the target: a lot at exactly 0.999 passes only 0.1% of the time instead of 9.6%, and a slightly-better 0.9995 lot passes 0.8% instead of 20%. Tripling the sample size turns a soft, ambiguous accept/reject boundary into a sharp one — at the cost of testing 40 more units per lot, a trade the simulation prices directly.
Ranking the program levers against the baseline 5% consumer's risk:
A proof test that a lot passes does not prove the reliability target was met — it proves only that this many units survived this load this time. Monte Carlo is what converts "it passed" into "it passed, and here is the probability that a sub-target lot would have passed too" — which is the number a qualification program actually exists to control.