| Vose Software

Industry: Construction and Infrastructure
Product: ModelRisk
Application: Risk Register Quantification on a $2B Metro Project


38 discrete risks, a $70M expected exposure, and a P99 of $171M the contingency never funds

A $2B underground metro project carried a risk register with 38 named events — ranging from "TBM cutter-head obstruction" (40% likelihood, $4–18M impact) to "archaeological discovery requiring rerouting" (5% likelihood, $30–95M impact). The deterministic risk roll-up summed the expected values to a $70M risk-weighted exposure. The Monte Carlo aggregation said the probability of total risk impact exceeding the $100M contingency was 17%, and the 99% VaR was $171M — 1.7× the booked contingency and nearly 2.5× the deterministic exposure the project had reserved against.

The difference is not because the deterministic sum was arithmetically wrong. It is because expected value is the wrong number to fund against.

Risk matrix — expected exposure by likelihood × impact band

The heat map is the familiar likelihood-versus-impact register view — but here each cell is the expected dollar exposure (probability × mean impact) contributed by the risks that fall in it, not a subjective red/amber/green rating. The dense band of moderate-probability, moderate-impact risks is where most expected value lives; the sparse high-impact / low-probability cells in the top-left are where the tail lives. The two require different funding logic, and the rest of this analysis quantifies both.

A risk register is a frequency-severity model in disguise

Every line in a project risk register is implicitly a compound distribution: a probability of occurrence multiplied by an impact distribution. Adding them up correctly requires Monte Carlo aggregation, not arithmetic on the means. The 38 risks were rebuilt as:

  • Bernoulli occurrence per risk, with probability calibrated either from frequency on the firm's database (16 risks) or from expert elicitation using the modified-Delphi protocol (22 risks).
  • PERT(min, mode, max) impact in $M, conditional on occurrence.
  • Discrete event multiplicity for three risks where the event could happen more than once (TBM obstruction, weather stoppage, change-order cascade) — modelled as Poisson conditional on at least one event.
  • Dependency structure — five risk-pair correlations were calibrated as a copula: TBM obstruction and ground-condition surprise (ρ = 0.55, same root cause), labour dispute and productivity loss (ρ = 0.40), regulatory delay and permit-cost overrun (ρ = 0.35), weather stoppage and structural-rework demand (ρ = 0.30), supply-chain disruption and price escalation (ρ = 0.45).

Treating the risks as independent under-states the aggregate P99 only modestly here — about 2% ($168M vs $171M) — because the register is large and the correlated pairs are a minority of it. But the dependency bites locally: where two risks share a root cause, the joint tail is far heavier than independence predicts, and that joint tail is exactly what overruns a tunnelling contingency. The aggregate number hides it; the pairwise view (below) does not.

Pareto of expected impact by risk event

The Pareto view showed that 7 risks accounted for 78% of expected impact — but the tail of the aggregate distribution was dominated by a different set: the four risks with low probability and very high impact (ground collapse, archaeological discovery, contractor insolvency, force-majeure event). These never made it to the top of the expected-value ranking, but they owned the 99th percentile.

The aggregate exposure distribution

50,000 trials of the 38-risk register, with the calibrated copula on the five correlated pairs, produced a sharply right-skewed distribution.

Aggregate risk exposure with VaR and Expected Shortfall

  • Mean $70M (matches the deterministic sum, as it should).
  • Median $64M.
  • P90 $115M.
  • P95 $133M.
  • P99 (VaR 99%) $171M.
  • Expected Shortfall ES = E[loss | loss > VaR99] = $192M.

The $100M contingency reserve sat at the P83 — meaning a 17% chance that the aggregate impact would breach the contingency. Pushing the contingency to P95 required $133M; pushing to P99 required $171M. The board's risk appetite — a 5% probability of breaching — demanded $133M, not $100M. The gap to the booked contingency was $33M at the board's own stated appetite, and it was visible only after Monte Carlo aggregation.

Where correlation actually bites

A risk register added up line-by-line implicitly assumes the risks are independent. The most consequential dependency in this project is between the TBM cutter-head obstruction and the ground-condition surprise — they share a root cause (the same uncertain geology), so they tend to fire together, and when they do, both fire large. The scatter below plots the two risks' per-trial impacts against each other.

Joint impact of the two correlated geotechnical risks

The mass in the upper-right quadrant — both geotechnical risks severe in the same simulated project — is the joint tail. A sum-of-independent-risks model lands there with probability 0.56%; the copula-aware model lands there 0.72% of the time — roughly 30% more often — because the geology that triggers one triggers the other. The gap looks small in percentage points, but it is concentrated entirely in the most expensive corner of the distribution: that upper-right quadrant is exactly the scenario that overruns a tunnelling contingency, and it is invisible to a deterministic register.

Which assumptions move the answer

Tornado: drivers of P99 aggregate exposure

The biggest driver of P99 was the TBM obstruction risk (40% probability, $4–18M impact, Poisson multiplicity averaging 1.5 events). The single-event tail of the archaeological-discovery risk was second. Third was the calibrated correlation between TBM and ground-condition risks — though, as the aggregate showed, the register-wide effect of correlation is small; its leverage on P99 comes almost entirely through this one geotechnical pair.

The bottom of the tornado was illuminating: risks that the project team had spent hours debating in workshops — supplier insolvency, IT-system failure, key-personnel turnover — moved the answer by less than $4M. The Monte Carlo redirected workshop time away from these and toward the four risks that owned the P99: TBM obstruction, ground condition, archaeological discovery, and labour productivity loss.

Risk-response framing — fund the distribution, not the mean

The project team had been arguing about whether to insure or self-insure each risk based on its expected value. Monte Carlo reframed the question:

  • Transfer the discrete high-severity, low-probability risks (archaeology, ground collapse, contractor insolvency, force-majeure). Their tail dominates the P99, and the insurance market prices them efficiently — moving them off the retained register is the single biggest lever on aggregate tail capital.
  • Self-insure / accept the frequent, lower-severity risks (weather stoppage, minor change orders). The actuarially fair premium exceeds the expected loss because the variance is low.
  • Mitigate the calibrated-correlated pairs specifically (TBM/ground, labour/productivity) — this barely moves the aggregate number, but it thins the geotechnical joint tail that owns the worst-case tunnelling scenario.

After applying these decisions — transferring the four tail risks and halving the modelled correlation on the retained pairs — the re-simulated P99 dropped from $171M to $131M and the P95 from $133M to $107M. The contingency was right-sized at $110M (the retained P95), and roughly $40M of P99 capital was released for use elsewhere on the programme.

ModelRisk Functionality Used

  • Compound Bernoulli × PERT aggregation across 38 risk-register lines — preserving the per-risk shape that an expected-value sum collapses to a point.
  • Gaussian copula on 5 calibrated risk-pair dependencies — modest at the aggregate (~2% on P99) but the dominant shaper of the geotechnical joint tail, which fires ~30% more often than independence predicts.
  • Poisson event multiplicity for the 3 risks (TBM, weather, change-order) that can fire more than once — capturing the cluster tail.
  • Pareto chart showed 7 risks owned ~78% of expected value but a different 4 risks owned the P99 — the workshop time had been spent on the wrong list.
  • Sensitivity ranking on P99 distinguished the high-leverage assumptions (TBM probability, archaeological tail) from the low-leverage ones (supplier insolvency, IT failure).
  • Insurance / mitigation framing by tail position rather than expected value cut P99 from $171M to $131M and released ~$40M of contingency.

A risk register is not a list of numbers to add up. It is a portfolio of distributions to convolve — and the convolution lives in the tail, which the average never sees.