Industry: Construction and Infrastructure Product: ModelRisk Application: Equipment Maintenance Optimization
The contractor's fleet of 18 bulldozers has been on a 500-hour preventive-maintenance cadence for a decade because that is what the manufacturer recommends. The probabilistic sweep in ModelRisk across six different PM intervals showed that extending PM from 500 hours to 750 hours cut the expected annual cost per machine by 12% — about $9,500 each — while still keeping the catastrophic-failure probability below the contract's downtime cap. Across 18 machines that is a $170k annual saving, paid for by nothing more than a re-read of the failure data.
Preventive-maintenance interval optimisation is one of the cleanest demonstrations of where the deterministic answer is wrong. A point estimate of expected failures per year, run against expected PM cost, gives a single number. The real distribution is dominated by the small probability of a catastrophic failure that costs 5×–10× a normal repair — and the right PM interval is the one that minimises expected total cost including the catastrophic tail, not the one that the manufacturer's brochure prints.
Time-to-failure for a bulldozer transmission, hydraulics or final-drive is Weibull with shape β = 2.2 and scale (characteristic life) η = 1,900 hours, calibrated from five years of maintenance-management system data. The shape parameter above 1 is the signature of wear-out failure — the hazard rate rises with cumulative use, the textbook reason PM cadence exists at all.
An exponential time-to-failure — the constant-hazard alternative many in-house tools use — would have implied that the failure probability in the next hour is the same whether the machine is freshly serviced or has 4,000 hours on the clock. The Weibull captures the wear-out reality.
Per-failure repair cost is LogNormal (mean $18k, σ = 0.55 in log-space) but with a 5% probability of escalation to a catastrophic event in the $80k–$150k range. This is the single most important modelling choice: industry experience says one in 20 failures becomes a major rebuild or component replacement, and the catastrophic tail is what makes "wait until it breaks" so expensive. A plain LogNormal without the catastrophic mixture would under-state expected cost per failure by roughly 35%.
Per-PM event cost is also LogNormal but much tighter (mean $3.2k, σ = 0.35). Downtime hours per failure are LogNormal (mean 24 hours, σ = 0.7); downtime per PM is a fixed 4 hours. Both translate to $/hour at $380 of idle labour + lost productivity.
The Weibull(β=2.2, η=1,900h) hazard turns directly into the probability of surviving any candidate PM interval without failure. The threshold-sweep chart traces both "no failure in the interval" and "no catastrophic failure across the year" as functions of PM cadence:
At PM=300h the in-interval failure probability is ~3%; at PM=500h it is ~9%; at PM=750h it is ~19% (still acceptable because the catastrophic-failure conditional probability is ~5%); at PM=1,000h it climbs to 33% and the catastrophic-failure annual probability begins to dominate. The point of cost optimisation — the famous U-curve for total expected cost per machine — falls between PM 600h and PM 800h, with PM=750h giving mean annual cost per machine of ~$69k versus $78k at PM=500h and $96k for run-to-failure. The P90-per-machine on run-to-failure is roughly $145k — more than 2× the PM=750h optimum — driven by the catastrophic tail the threshold sweep makes visible. Across 18 machines, the choice between PM 500h and PM 750h is worth roughly $160k expected per year, pure parameter optimisation, zero capex.
The PM-750h fleet-level total is roughly $1.24M mean per year. The month-by-month cumulative fan shows the envelope:
The deterministic burn line trends to $1.24M by month 12. The P10–P90 envelope widens steadily — by mid-year the difference between a "lucky" and an "unlucky" fleet-year is already ~$0.3M, and by year-end the P90 trajectory sits around $1.5M (the budget cap line). The P50 trajectory undershoots the cap; the P90 trajectory hits it almost exactly. The right-tail trajectory (above P90) is where the catastrophic failures live — the months in which a single $115k transmission rebuild punches a step into the cumulative curve.
The largest mover of expected cost at the 750-hour optimum is the Weibull shape parameter — i.e., the engineering judgement about whether the fleet is in wear-out mode or still in random-failure mode. The second is the catastrophic-failure probability (the 5% baseline assumption). A re-read of the failure data every six months — the model is parameterised, the data is in the MMS — is the cheapest single way to keep the 750-hour optimum correct as the fleet ages.
The interval-optimisation result is "use PM=750h on average". The per-machine, per-checkpoint decision is sometimes different: at the 750h checkpoint the supervisor is occasionally tempted to defer one more cycle. The decision tree quantifies that temptation:
At age 750h, the conditional probability of failure during the next 750h is about 37% (the Weibull hazard rises with age, and the survival fraction shrinks rapidly past one characteristic life). Doing the PM now costs ~$4.7k with certainty. Deferring carries a 63% chance of $4.7k (no failure, just PM 750h later) and a 37% chance of ~$37k (failure repair plus the PM you ultimately still have to do). The expected cost of "defer" is ~$16.9k — more than 3× the expected cost of "PM now", which is why the footer says choose PM now. The decision-tree visual makes the asymmetry obvious in a way the cost-curve does not: the defer branch is cheaper two times in three, but the one-in-three failure carries an ~8× cost spike. The same framework re-priced for a harsh-duty project (10% catastrophic rate) shifts the optimum inward to ~600h; for a mature fleet (3% catastrophic) the optimum pushes outward to ~900h. The deterministic "PM every 500 hrs" rule is right for none of these.
PM interval optimisation is a problem deterministic models cannot solve correctly because the answer lives in the tail of the failure-cost distribution. Monte Carlo simulation in ModelRisk makes that tail visible — and once it is visible, the right PM interval is no longer the manufacturer's recommendation but the parameter that minimises the full expected cost including the catastrophic failures that defined the problem in the first place.