Industry: Telecommunications and IT Product: ModelRisk Application: Data center optimization
A modern hyperscale-style data center runs at PUE between 1.2 and 1.6 and consumes electricity at industrial-tariff rates of $0.06–$0.18/kWh depending on region — meaning a 40 MW IT-load campus burns through $25M–$70M in electricity alone per year before any allocation of capital, staff, or maintenance. Multiply by a five-campus estate over a 10-year horizon and the cumulative power bill alone clears $1.5 billion, with a spread driven by three things the deterministic model treats as constants: demand growth, PUE drift, and tariff volatility.
A telecom-and-IT services operator rebuilt its 10-year TCO and capacity-expansion model in ModelRisk to make the spread visible. The deterministic plan said "TCO ≈ $2.5B". The simulation said "mean $3.56B, P10 = $2.88B, P90 = $4.52B — and with the current incremental build plan there is a 63% chance you breach the capacity envelope before year 6."
The deterministic $2.5B sat below even the 1st percentile of the simulated distribution: it was not a central estimate at all, but an optimistic floor. The whole purpose of the rebuild was to replace that single number with the curve — its mean, its tail, and the capacity-breach probability the point estimate could never expose.
The previous plan used a single point-estimate of 8%/yr compound growth in IT load. The rebuilt model treats annual growth as a truncated Normal(mean 8%, σ = 4%, clipped to [0%, 20%]) — bounded because runaway demand growth saturates at physical and budgetary limits, and a plain Normal would draw negative growth. Compounded over 10 years against a 200 MW year-0 estate load (five campuses of ~40 MW IT load each), this produces a year-10 estate-load distribution where:
That ~130 MW spread between P10 and P90 is the entire capacity-planning problem in one number.
PUE is not a constant — it drifts as hardware ages, free-cooling availability changes with climate, and workload mix shifts. The model treats per-site PUE as LogNormal(median 1.35, σ_log = 0.06), bounded below at 1.10 (Carnot/thermodynamic floor for the cooling architecture). A LogNormal is the right shape: PUE can spike upward when a chiller fails or ambient conditions degrade, but cannot fall below the physical floor.
Energy per year per site is then:
E_year = IT_load_MW × 8760 h × PUE × (1 - free_cooling_hours_fraction)
producing energy distributions that combine demand uncertainty with PUE drift.
The first iteration of the model used geometric Brownian motion for electricity prices. That is the standard textbook choice and it is wrong for industrial-tariff power: prices in the 2021–2024 European and North American records show regime switches (the gas-price shock of 2022 doubled industrial-tariff power overnight in many countries), not the smooth log-Normal drift GBM produces. The rebuilt model uses a two-regime mean-reverting process with Poisson regime-switch arrivals at λ = 0.12/yr (about one shock per decade per region). In the "calm" regime, price reverts to the long-run mean of $0.10/kWh with annual vol 12%; in the "shock" regime, price reverts to $0.18/kWh with annual vol 22%. Without the regime-switching layer the model under-estimates 10-year energy-cost variance by a factor of nearly 2.
Per-year cash flow per site:
Cost_year = Energy_cost + Maintenance + CapEx_amortised + Labor TCO_10yr = Σ_t Cost_t / (1 + r)^t with r = 8%
across five regional sites. The simulation runs 50,000 trials over the 10-year horizon.
The deterministic plan reported TCO of ~$2.5B. The simulation produces a mean of $3.56B with a P90 of $4.52B — the deterministic figure understates the mean by more than $1B and the P90 by roughly $2B. The right tail comes from the joint event "demand at P75 and a sustained price shock in years 4–8" — neither extreme on its own, but together they swing TCO by hundreds of millions.
A pure-TCO view hides the operationally critical question: when do we run out of capacity? With the current build plan (modular adds at years 2, 5, and 8), the simulation says:
These are not symmetric problems: a capacity breach forces emergency colocation procurement at a premium and risks customer-facing outages; stranded capacity is a sunk-cost annoyance. The simulation makes plain that the current incremental plan errs heavily toward the breach side — capacity runs hot for most of the decade — which is exactly why the build-strategy comparison below matters.
Sensitivity on P90 TCO ranks the inputs by tail leverage:
Demand-growth volatility dominates — a wider σ on the growth rate moves P90 TCO by ~$140M. The electricity-price regime-shift probability is second; PUE drift third. The standard sensitivity treatment that locks PUE at 1.35 and varies it ±5% gets the ranking wrong because it ignores the joint event of high demand × shock-regime electricity.
Three CapEx strategies were simulated against the same demand/price distributions:
The result overturns the usual intuition: on this model the front-loaded build is cheapest on both mean and tail, and the flexibility of the hybrid is paid for in a materially higher TCO. The decision is therefore not "which is cheapest" but how much the operator will pay in expected TCO to avoid committing CapEx against demand that might not arrive — exactly the trade-off a single deterministic number cannot frame.
The deterministic plan's $2.5B TCO is one number sitting below the bottom of a curve that runs past $2.9B at the P10 and $4.5B at the P90. Monte Carlo turns the curve into the actual planning artefact — and the artefact, not the single number, is what gets the capacity-strategy and hedge decisions right.