Industry: Telecommunications and IT Product: ModelRisk Application: IT System Resilience
A payments platform with a four-nines availability target gets a downtime budget of 52.6 minutes per year. The architecture diagram — six tiers, active-passive everywhere, hot-standby database, multi-AZ — said the budget was comfortable. A chaos-engineering campaign that injected real failure modes into a staging clone said something different. When the team rebuilt the same composition probabilistically in Vose Software's ModelRisk, the 4-nines budget was breached in 30% of simulated operational years, and a P99 year saw 6.2 hours of accumulated downtime — more than seven times the budget.
The gap is not poor engineering. The gap is what happens when an architecture diagram is multiplied through in averages: redundancy, correlation, and the heavy-tailed repair-time distribution disappear, and what is left is a single optimistic number. The chart below shows it directly — the mean year sits just inside the 52.6-minute budget, but the distribution's right tail spills far past it.
Each of the six tiers — load balancers, auth service, payments core, database tier, message bus, CDN/edge — has its own failure intensity (events per year, Poisson) and its own repair-time distribution (LogNormal in minutes, with geometric standard deviation between 1.5 and 2.6). Redundancy is captured as a multiplicative factor on the per-tier failure rate: an active-passive pair with healthy failover cuts the effective rate to 9–11% of the bare-metal number; the CDN's anycast mesh cuts it to 7%.
Two design decisions matter at the modeling layer:
Annual downtime is the sum across all tiers and the regional shock. Availability is 1 − downtime / 525,600.
1 − downtime / 525,600
The deterministic resilience worksheet, built by summing each tier's mean-time-between-incidents and mean-time-to-repair, reported an expected availability of 99.992% — comfortably inside the 4-nines target with 10 minutes of headroom. The Monte Carlo run produced a mean of 51 minutes/year, a P90 of 130 minutes and a P99 of 374 minutes. The 4-nines budget just held on the mean, but breached 30% of operational years — close to one in three. The averaged-out diagram had no language for "most years fine, one year in thirty catastrophic" — yet that is exactly the operational pattern the platform was about to inherit.
Ranking each source by its mean contribution to annual downtime tells the team which tier is buying them the least resilience per dollar invested.
The rare regional event is the single largest contributor — roughly 22 minutes/year on average — more than any individual tier. Among the tiers, the payments core (≈9 min/yr) and auth service (≈8 min/yr) lead, with the database tier close behind (≈7 min/yr); the payments core's MTTR is wide (GSD 2.4) because rolling back a poisoned write is sometimes a multi-hour exercise. Crucially the load-balancer (≈2 min/yr) and CDN (≈1 min/yr) tiers, where the team had been planning to invest, contribute very little. The chart redirected the next $3M of resilience capex from edge to the regional-failover and database tiers.
Two architectural changes were modeled head-to-head against the as-is baseline:
The full package drops the mean annual downtime from 51 to 27 minutes/year, the P99 from 374 to 182 minutes, and the probability of breaching the 4-nines budget from 30% to 17%.
Downtime cost was modeled at $48,000 per minute — the industry midpoint for a mid-size payments platform once direct revenue, SLA credits, and per-incident remediation are summed. The cost distribution is more revealing than the downtime distribution, because it inherits the heavy tail and converts it into a budgeting number a CFO can act on.
Expected annual outage cost is $2.4M as-is, $1.8M with cross-region failover, and $1.3M with the full active-active package. The P95 annual cost — the figure used in the CIO's risk register — falls from $9.1M to $4.8M. The full $3.0M five-year programme pays back in expected terms inside three years and removes roughly $4M of P95 tail-year exposure.
VosePoisson
VoseLognormal2
Resilience is not the diagram; it is the joint distribution of how every tier fails and how long it takes to come back, plus the rare event that takes several tiers down at once. Monte Carlo simulation in ModelRisk is what turns "we are nominally four-nines" into "we breach four-nines roughly one year in three — and here is the $3M that roughly halves that exposure."