| Vose Software

Industry: Aerospace
Product: ModelRisk
Application: Certification Test Campaign Duration and Cost Forecasting


The deterministic test plan booked 28 months for certification — the simulation puts the chance of beating it at 67%, the P90 finish at 32 months, and a $456M cost tail

A clean-sheet narrow-body certification campaign comprises roughly 2,400 individual flight-test points, 36 system-level qualification tests, and 11 major coupon and full-scale structural test articles before EASA/FAA Type Certificate sign-off. The program plan that the chief test engineer signed off on booked a deterministic 28-month duration from first flight to Type Certificate — and treated that 28 months as a safe ceiling the campaign would comfortably beat.

The Monte Carlo schedule simulation, parameterised against four prior certification campaigns and the actual catalogue of failed-test-rerun frequencies, told a sharper story. It put the P50 finish at 26.1 months, the P90 at 32.1 months, and the mean at 26.4 months. The probability of certifying on or before the 28-month plan date was 66.9% — a one-in-three chance of overrunning, not the near-certainty a padded plan implies. The 28-month line is not a ceiling; it is roughly the two-thirds point of the distribution, with a long right tail that runs four months past it.

Certification campaign duration S-curve from first flight to Type Certificate

The S-curve spans a P10 of roughly 21 months to a P90 of 32.1 months, with the deterministic 28-month plan line sitting just past the P50 — about a third of the probability mass lands beyond it. The distribution carries a long right tail because a program-wide difficulty factor stretches every critical-path activity together and the heavy-tail risks — stall-departure reruns, regulator queueing — cascade on the hardest programs rather than averaging away.

Modeling test campaign duration as a stochastic network

The campaign was decomposed into 47 milestone activities — flight envelope expansion, flutter clearance, hot-and-cold weather trials, MEL clearance, system-integration tests, water-ingestion certification, etc. — with explicit precedence relationships.

Each activity's duration was modeled with a PERT-Beta distribution parameterized from prior-campaign data: minimum, most likely, and maximum. The flutter clearance activity, for example, was PERT(4 weeks, 8 weeks, 22 weeks) — the long upper tail reflects the documented possibility of having to redesign and retest a wing-mounted antenna fairing that fails the flutter boundary.

Test-rerun frequency was the dominant uncertainty. For each major test point, a Bernoulli flag triggered a rerun with a calibrated probability — for instance, the high-AoA stall-departure characteristic test reruns with probability 0.34 in the prior-campaign dataset, and conditional on a rerun, requires a mean 6.5 weeks of rework before the next attempt (LogNormal, σ_log 0.42). Reruns are not independent: a failure in stall-departure raises the probability of a related failure in roll-control characteristics by a factor of 1.8, captured with a Gaussian copula across the seven test-point pairs where this dependency was documented.

Regulator response time was modeled as LogNormal with median 11 weeks per documentation submission and σ_log 0.48 — a long right tail driven by the documented variability of FAA and EASA review queues, with the queue length itself modeled as a correlated process across both regulators.

Weather-driven test-window loss for envelope expansion at high-altitude / hot-and-cold sites was modeled as a daily Bernoulli with site-specific weather-availability probabilities, integrated over the activity duration.

The schedule simulation

25,000 schedule runs over the 47-activity network produced the certification-date distribution shown above. Deterministic plan: 28 months. P10 (best 10% of runs): 21.1 months. P50: 26.1 months. P90: 32.1 months. Mean: 26.4 months. The probability of certifying on or before the 28-month plan date was 66.9% — meaning a one-in-three chance of an overrun, with the P90 running to 32.1 months, four months past the plan. The decision-relevant question is "how much schedule and cost reserve does the program need to hold against a tail that materialises a third of the time?" — not whether the date is comfortable.

Cost follows duration — and follows it non-linearly

Each test-month carries a roughly $14M run-rate (test fleet, instrumentation, flight crew, ground support, regulator interaction). Beyond month 32, the launch-customer delivery penalty clauses begin to bite at $2.8M per aircraft per month of delay across the first 18 deliveries. Total campaign cost is therefore a piecewise-linear function of duration with a kink at 32 months.

Total certification campaign cost distribution

Mean cost: $384M. P50: $365M. P90: $456M. The deterministic plan estimated $392M (28 months × $14M run-rate). The mean sits just below the plan, but the P90 runs $64M above it — because in the upper tail the finish date crosses month 32, where the delivery-penalty clauses activate and the cost curve steepens sharply. The plan priced the median campaign and ignored the kink; the simulation shows the cost risk lives entirely in the right tail, where schedule overrun and delivery penalties compound.

What drives the schedule spread

Tornado: drivers of certification duration (P10–P90)

The single largest driver of the P10–P90 spread is the program-difficulty common factor — the systematic risk that test-article quality, instrumentation maturity, and team experience stretch every critical-path activity together — contributing ±4.6 months around the 26.4-month mean. It is the reason the 47 activities do not average away into a tight bell. Stall-departure test-rerun probability is second at ±3.2 months, regulator response time σ third at ±2.7 months, and flutter-clearance variability fourth at ±2.1 months. The Gaussian copula on related test-point reruns adds ±1.4 months on its own — invisible in the prior independent-test schedule model.

What the model changed

  • The program contingency reserve was sized against the simulated P90 finish date (32.1 months) rather than the deterministic 28-month plan + a generic 15% buffer — because the P90 lands four months past the plan, the reserve was held (and the cost reserve raised to the $456M P90), not released, and explicitly funded against the one-in-three overrun the deterministic plan had hidden.
  • The two parallel test articles the program had weighed at a $42M up-front cost to break the hot-and-cold weather-trial bottleneck now have a real schedule case — the simulation shows a third of programs cross the 28-month plan and the upper tail reaches month 32, where delivery penalties begin, so de-bottlenecking buys schedule the program is genuinely at risk of needing.
  • The launch-customer delivery commitment was set against the P90 (32 months) with the penalty exposure priced explicitly, rather than against the median — both sides now hold a date and a penalty bond sized to the tail the simulation exposed, not to the comfortable mean.
  • The stall-departure test plan was rewritten to front-load the most likely failure mode (low-altitude high-AoA), because together with the program-difficulty factor it drives most of the schedule spread — running it earlier shortens the right tail where the cost penalties live.

ModelRisk Functionality Used

  • PERT-Beta for activity durations, parameterised from four prior certification campaigns' as-flown data — replaced the program's prior fixed-duration schedule that had no expression of variability.
  • Bernoulli rerun flags with calibrated per-test probabilities (0.34 for stall-departure, 0.18 for flutter, etc.), with conditional rerun probabilities captured by Gaussian copulas across the seven documented related-test-pair dependencies.
  • LogNormal regulator response time with explicit FAA / EASA queue-length correlation — the right tail of the schedule was being driven primarily by regulator queueing, not test-engineering risk.
  • S-curve analysis producing P10 / P50 / P90 certification dates and the explicit probability of meeting the published plan date (66.9%, i.e. a one-in-three overrun — not the near-certainty the padded plan implied).
  • Piecewise-linear cost model with the kink at month 32 where the delivery-penalty clauses activate — modelling the kink explicitly is what showed the upper-tail finishes cross it, putting the P90 cost $64M above the plan and locating the cost risk in the right tail.
  • Shared program-difficulty common factor scaling every critical-path activity together — the systematic risk that prevents the 47 activities from averaging into a deceptively tight bell, and the largest single driver of the schedule spread.

A 28-month deterministic certification plan is not "the schedule" — here it was roughly the two-thirds point of a right-skewed distribution whose P90 lands at 32.1 months and whose P90 cost runs $64M above the program plan. The deterministic number was not a safe ceiling; it was a coin-flip-plus, and the simulation is what showed the program where the schedule-and-cost tail actually lives — and how much reserve to hold against it rather than release.