| Vose Software

Industry: Aerospace
Product: ModelRisk
Application: Satellite Design Under Uncertainty


A Single-String Bus Misses Its 0.90 Reliability Floor in Every Trial — the $14M Redundancy Step That Clears It

A 12-year LEO Earth-observation mission carries a non-recurring engineering cost north of $280M and a launch bill that scales at roughly $7,500 per kilogram to a sun-synchronous orbit. Every additional kilogram of redundant hardware buys reliability — and costs hard cash twice, once to build and once to lift. The fundamental design question is therefore not "is the spacecraft reliable?" but "what is the probability the spacecraft completes its 12-year mission, and what is the cheapest redundancy architecture that hits the contractual 0.90 reliability floor?"

A commercial Earth-observation operator with three insurance-rated buses on contract rebuilt its bus-level reliability allocation in ModelRisk. The team's frustration with the traditional parts-count MIL-HDBK-217 approach was simple: it produced a single MTBF number per subsystem and a multiplicative "system reliability" that was either implausibly optimistic or implausibly pessimistic depending on the calibration assumptions, with no visibility into where the actual mission-loss probability lived. The Monte Carlo answer is the chart below: across 800 parameter-uncertainty trials of the full six-subsystem stack, the baseline single-string bus carries a mean 12-year completion probability of 0.844 (P10 0.828, P90 0.859) — and in not one trial does it clear the contractual 0.90 floor (P(meet) = 0%). The legacy parts-count workbook had signed off the same bus at 0.93.

Mission-reliability distribution for the baseline single-string bus

Subsystem failure as a Weibull process, not a constant rate

A satellite subsystem in orbit does not fail at a constant hazard rate. Infant mortality dominates the first 90 days (manufacturing escapes, launch-induced damage), wear-out dominates the last 2–3 years (radiation dose to CMOS, propellant depletion, battery cycling), and the middle years carry a relatively flat residual rate. The team replaced the constant-λ exponential model with a per-subsystem Weibull with a shape parameter β fitted to on-orbit failure data from the operator's prior 14 buses:

  • Reaction wheels: Weibull(β = 1.8, η = 83 yr) — clear wear-out, β > 1; ~0.97 over 12 years.
  • Star tracker: Weibull(β = 0.7, η = 3,200 yr) — radiation-driven infant mortality, β < 1; ~0.98.
  • Battery string: Weibull(β = 2.5, η = 43 yr) — strong wear-out from cycling; ~0.96, the weakest link.
  • Transponder: Weibull(β = 1.0, η = 594 yr) — effectively constant rate, β ≈ 1; ~0.98.
  • Other electronics: Weibull(β = 1.2, η = 310 yr) — mild wear-out on the residual avionics; ~0.98.
  • Propulsion: depletion based on stochastic delta-v consumption (see below).

For the propulsion subsystem, the "failure" is fuel exhaustion. Drag and station-keeping delta-v draws were modeled as LogNormal(μ = ln 28 m/s/yr, σ = 0.18), with the tank sized to 375 m/s total — which gives a fuel-depletion-age distribution (~0.98 survival to 12 years), not a single nominal lifetime.

Subsystem redundancy is the design lever. Cold-redundant pairs are switched in on failure; the pair survives if at least one element survives at age t. The unreliability of a 1-of-2 cold-redundant pair is not (1 − R)² (that's hot redundancy) — it is the convolution of two Weibull lifetimes, and ModelRisk computes it directly by sampling.

What the multiplicative parts-count model hides

The legacy parts-count workbook reported a 12-year mission reliability of 0.93. The Weibull simulation reports something lower and, more importantly, distributed. Mean 12-year mission completion probability across modeled uncertainty in the Weibull shape and scale parameters themselves is 0.844, with a P10 of 0.828 and a P90 of 0.859. The probability of clearing the contractual 0.90 floor on the baseline single-string architecture is 0% — the entire distribution sits below the line, not at the comfortable 0.93 the parts-count number implied. The ~9-point gap is the compounding penalty the deterministic constant-rate model never shows: six series subsystems, each individually ~0.96–0.98 over the mission, but whose wear-out-dominated tails (battery β = 2.5, reaction wheel β = 1.8) and propellant depletion multiply together into a system reliability that lands well short of the optimistic parts-count figure.

Where the unreliability budget lives

Sensitivity analysis on the subsystem contributions to mission-failure probability:

Tornado: subsystem contributions to 12-year mission-failure probability

Against a baseline mission-failure probability of 15.3%, the battery string is the single largest contributor — a 3.6-percentage-point swing — followed by propulsion (propellant depletion) at 3.1 pp and the reaction wheel assembly at 2.6 pp. "Other electronics" and the transponder tie at 1.7 pp each, with the star tracker at 1.6 pp. This ranking is the redundancy-investment work plan: the battery and reaction wheel are the two wear-out items worth cold-sparing, while propulsion — second on the list but not a candidate for a redundant tank — is bought down by the 375 m/s propellant margin instead.

Trading mass for reliability

The team priced three candidate redundancy architectures against the same launch-mass and unit-cost budget:

  • Baseline (N): no redundancy beyond the legacy single-string design. Bus dry mass 1,420 kg. NRE $260M, recurring $42M.
  • N+1 critical: redundant battery string and redundant reaction wheel (cold). +38 kg. NRE $268M, recurring $48M.
  • N+1 full: above plus redundant star tracker and transponder (cold). +52 kg. NRE $276M, recurring $54M.

Each architecture was simulated for 12-year mission completion, then total program cost computed as NRE + recurring + launch mass × $7,500/kg.

Reliability vs total program cost across redundancy architectures

The deterministic recommendation was N+1 full because "more redundancy is always better." The probabilistic answer prices the decision exactly: cold-redundant pairs lift mission reliability from 0.844 (baseline, $312.6M) to 0.910 (N+1 critical, $326.9M) to 0.946 (N+1 full, $341.0M) — each step roughly a $14M program-cost increment. The crucial finding is that N+1 critical — redundant battery and reaction wheel only — is the cheapest architecture that clears the 0.90 floor, at $326.9M (a $14.3M premium over the non-compliant baseline). N+1 full buys a further 3.6 points of reliability (0.946) for another $14.1M; whether that margin is worth it depends on the insurance-premium reduction it unlocks, but it is no longer required to meet the contract. The baseline, despite the parts-count sign-off at 0.93, is simply not deliverable.

What changed

  • The baseline single-string bus was shown to be non-compliant — its entire reliability distribution (mean 0.844, P90 0.859) sits below the 0.90 contractual floor, overturning the parts-count sign-off of 0.93 and making redundancy mandatory rather than optional.
  • N+1 critical was identified as the cheapest compliant architecture — redundant battery and reaction wheel lift reliability to 0.910 at $326.9M, a $14.3M premium that clears the floor; N+1 full (0.946, $341.0M) became a margin-vs-insurance trade rather than a requirement.
  • The redundancy-investment ranking was set by the tornado — battery string (3.6 pp) and reaction wheel (2.6 pp) are the two cold-spare targets; propulsion (3.1 pp) is addressed by the 375 m/s propellant margin, not a redundant tank.
  • A reliability allocation the systems team owns. When a subsystem vendor proposes a parts substitution, the team re-runs the affected Weibull and sees the mission-reliability impact in 20 minutes, not the four weeks the prior parts-count rebuild required.

ModelRisk Functionality Used

  • Weibull lifetime fits per subsystem from on-orbit failure histories, with β values from 0.7 (radiation-driven infant mortality) to 2.5 (battery wear-out) — the correct distribution family for time-to-failure, not the constant-rate exponential the legacy workbook used.
  • Cold-redundant pair convolution evaluated by direct sampling rather than the (1 − R)² shortcut that applies only to hot redundancy.
  • LogNormal delta-v consumption model for the propulsion subsystem, fed through a fixed tank size to produce a fuel-depletion-age distribution rather than a single nominal lifetime.
  • Tornado ranking on mission-failure probability that produced the prioritized redundancy-investment list, identifying the battery and reaction-wheel subsystems as the only two worth N+1 hardening on this mission.
  • Program-cost comparison of three architectures — total of NRE, recurring, and launch mass at $7,500/kg — plotted against simulated reliability, so the redundancy decision reflected cost and reliability jointly rather than "more redundancy is always better."

Reliability is not the product of point estimates. It is the joint distribution of subsystem lifetimes, the shape parameters of each Weibull, and the convolution arithmetic of every cold-redundant pair. Monte Carlo simulation in ModelRisk makes that joint distribution visible — and once it is visible, "how much reliability does each redundancy dollar actually buy, and is the contractual floor even reachable?" becomes a calculation, not an argument.