Industry: Utilities Product: ModelRisk Application: Energy savings program evaluation under uncertainty
A utility's demand-side efficiency program — a portfolio of LED, heat-pump, controls, and insulation measures — was filed with the regulator at a gross first-year saving of 240 GWh, and booked over a 10-year measure life as a flat 2,400 GWh. Two forces quietly erode that headline. Free-ridership: some participants would have installed the measures anyway, so those savings are not the program's to claim. Persistence: installed measures fail, get removed, or degrade, so each year's savings decay. Neither shows up in a number that multiplies a first-year gross figure by ten.
Rebuilt in ModelRisk over the full 10-year horizon, the program's cumulative net savings have a mean of 1,683 GWh — 70% of the 2,400 GWh gross claim — with a P10 of 1,256 GWh and a P90 of 2,099 GWh. The deterministic claim is not a forecast of net savings; it is the ceiling none of the simulated futures reach. The fan chart below shows the gap opening year by year.
The flat 2,400 GWh assumes net-to-gross equals 1.0 and savings persist undiminished for a decade — two assumptions the evaluation literature contradicts on every program ever metered. Worse, the two erosions compound: free-ridership shrinks the base in year one, and persistence shrinks what is left every year after. A point estimate cannot represent a quantity that drifts away from itself over time; it can only report one of the curve's endpoints and hope. The simulation shows the mean cumulative net saving running 30% below the gross claim, with a P10–P90 band of 843 GWh — nearly the program's entire annual output of uncertainty that the single number renders as zero.
Each simulated program path starts from a gross first-year saving drawn LogNormal around 240 GWh (±7% on install volume) and applies two erosions:
A shared program-execution factor, drawn Beta(5,2), drives both free-ridership and persistence: a poorly-run program recruits more free-riders and commissions measures that fail sooner. This common factor is what keeps the 10-year cumulative from collapsing to its mean — without it, the year-by-year draws would diversify and the cumulative band would be implausibly tight. With it, the bad-execution scenarios are coherent and the P10 (1,256 GWh) reflects a program that goes wrong on several fronts at once.
Decomposing the average path into gross claim versus net delivered shows the erosion is not a one-time haircut — it widens every year.
In year one, net delivered averages 196 GWh — 81% of the gross 240 GWh — the free-ridership haircut alone. By year ten, net has decayed to 144 GWh, just 60% of gross, as persistence compounds on top. The deterministic flat-240 line sits above the green net bars in every single year, by a margin that grows from 19% to 40%. That widening wedge is the entire case for probabilistic evaluation.
Ranking the drivers by their effect on the P10 cumulative net saving (1,256 GWh) tells the program manager where the answer is decided.
Free-ridership / net-to-gross is the largest driver at ±235 GWh of P10 cumulative saving, with persistence second at ±175 GWh and the shared execution factor third at ±140 GWh. Gross install volume — the input the program tracks most closely in its dashboards — moves the P10 by only ±90 GWh. The implication is that a better net-to-gross study and a measure-retention program buy more verified savings than simply installing more units does.
The program lives or dies on the cost-effectiveness test — its levelised cost per net kWh saved against the regulator's avoided-cost benchmark of $0.072/kWh.
The program passes the cost test in 96% of simulated futures, with a mean levelised cost of $0.0484/kWh against the $0.072 benchmark. But the deterministic calculation reports $0.0325/kWh — a third cheaper than the real mean — because it divides the $78M program cost by the un-eroded 2,400 GWh. The point estimate makes a comfortable program look like a slam-dunk and hides the 4% of futures in which net benefit turns negative (mean net benefit +$43M, but a P10 of just +$12M). The program is sound; its margin of safety is thinner than the single number suggests.
For an efficiency program, the gross saving is a marketing number and the net saving is the real one — Monte Carlo is what measures the decade-long gap between them, and turns "we saved 2,400 GWh" into "we will deliver 1,683 GWh, with a 1-in-10 chance below 1,256."