| Vose Software

Industry: Construction and Infrastructure
Product: ModelRisk
Application: Equipment Maintenance Optimization


18 Bulldozers, Six PM Strategies, and the 750-Hour Sweet Spot

The contractor's fleet of 18 bulldozers has been on a 500-hour preventive-maintenance cadence for a decade because that is what the manufacturer recommends. The probabilistic sweep in ModelRisk across six different PM intervals showed that extending PM from 500 hours to 750 hours cut the expected annual cost per machine by 12% — about $9,500 each — while still keeping the catastrophic-failure probability below the contract's downtime cap. Across 18 machines that is a $170k annual saving, paid for by nothing more than a re-read of the failure data.

Preventive-maintenance interval optimisation is one of the cleanest demonstrations of where the deterministic answer is wrong. A point estimate of expected failures per year, run against expected PM cost, gives a single number. The real distribution is dominated by the small probability of a catastrophic failure that costs 5×–10× a normal repair — and the right PM interval is the one that minimises expected total cost including the catastrophic tail, not the one that the manufacturer's brochure prints.

Why Weibull, and why a catastrophic-failure mixture

Time-to-failure for a bulldozer transmission, hydraulics or final-drive is Weibull with shape β = 2.2 and scale (characteristic life) η = 1,900 hours, calibrated from five years of maintenance-management system data. The shape parameter above 1 is the signature of wear-out failure — the hazard rate rises with cumulative use, the textbook reason PM cadence exists at all.

An exponential time-to-failure — the constant-hazard alternative many in-house tools use — would have implied that the failure probability in the next hour is the same whether the machine is freshly serviced or has 4,000 hours on the clock. The Weibull captures the wear-out reality.

Per-failure repair cost is LogNormal (mean $18k, σ = 0.55 in log-space) but with a 5% probability of escalation to a catastrophic event in the $80k–$150k range. This is the single most important modelling choice: industry experience says one in 20 failures becomes a major rebuild or component replacement, and the catastrophic tail is what makes "wait until it breaks" so expensive. A plain LogNormal without the catastrophic mixture would under-state expected cost per failure by roughly 35%.

Per-PM event cost is also LogNormal but much tighter (mean $3.2k, σ = 0.35). Downtime hours per failure are LogNormal (mean 24 hours, σ = 0.7); downtime per PM is a fixed 4 hours. Both translate to $/hour at $380 of idle labour + lost productivity.

The PM-interval choice in one curve

The Weibull(β=2.2, η=1,900h) hazard turns directly into the probability of surviving any candidate PM interval without failure. The threshold-sweep chart traces both "no failure in the interval" and "no catastrophic failure across the year" as functions of PM cadence:

P(failure-free) vs preventive-maintenance interval

At PM=300h the in-interval failure probability is ~3%; at PM=500h it is ~9%; at PM=750h it is ~19% (still acceptable because the catastrophic-failure conditional probability is ~5%); at PM=1,000h it climbs to 33% and the catastrophic-failure annual probability begins to dominate. The point of cost optimisation — the famous U-curve for total expected cost per machine — falls between PM 600h and PM 800h, with PM=750h giving mean annual cost per machine of ~$69k versus $78k at PM=500h and $96k for run-to-failure. The P90-per-machine on run-to-failure is roughly $145k — more than 2× the PM=750h optimum — driven by the catastrophic tail the threshold sweep makes visible. Across 18 machines, the choice between PM 500h and PM 750h is worth roughly $160k expected per year, pure parameter optimisation, zero capex.

Cost trajectory through the year

The PM-750h fleet-level total is roughly $1.24M mean per year. The month-by-month cumulative fan shows the envelope:

Cumulative annual maintenance cost trajectory — PM=750h fleet of 18 bulldozers

The deterministic burn line trends to $1.24M by month 12. The P10–P90 envelope widens steadily — by mid-year the difference between a "lucky" and an "unlucky" fleet-year is already ~$0.3M, and by year-end the P90 trajectory sits around $1.5M (the budget cap line). The P50 trajectory undershoots the cap; the P90 trajectory hits it almost exactly. The right-tail trajectory (above P90) is where the catastrophic failures live — the months in which a single $115k transmission rebuild punches a step into the cumulative curve.

Drivers of the optimum-interval cost

Tornado: drivers of expected cost at PM=750 hrs

The largest mover of expected cost at the 750-hour optimum is the Weibull shape parameter — i.e., the engineering judgement about whether the fleet is in wear-out mode or still in random-failure mode. The second is the catastrophic-failure probability (the 5% baseline assumption). A re-read of the failure data every six months — the model is parameterised, the data is in the MMS — is the cheapest single way to keep the 750-hour optimum correct as the fleet ages.

The defer-or-do decision at each checkpoint

The interval-optimisation result is "use PM=750h on average". The per-machine, per-checkpoint decision is sometimes different: at the 750h checkpoint the supervisor is occasionally tempted to defer one more cycle. The decision tree quantifies that temptation:

Defer-or-PM decision at the 750h checkpoint

At age 750h, the conditional probability of failure during the next 750h is about 37% (the Weibull hazard rises with age, and the survival fraction shrinks rapidly past one characteristic life). Doing the PM now costs ~$4.7k with certainty. Deferring carries a 63% chance of $4.7k (no failure, just PM 750h later) and a 37% chance of ~$37k (failure repair plus the PM you ultimately still have to do). The expected cost of "defer" is ~$16.9k — more than 3× the expected cost of "PM now", which is why the footer says choose PM now. The decision-tree visual makes the asymmetry obvious in a way the cost-curve does not: the defer branch is cheaper two times in three, but the one-in-three failure carries an ~8× cost spike. The same framework re-priced for a harsh-duty project (10% catastrophic rate) shifts the optimum inward to ~600h; for a mature fleet (3% catastrophic) the optimum pushes outward to ~900h. The deterministic "PM every 500 hrs" rule is right for none of these.

What the model changed

  • Bulldozer PM cadence extended from 500 to 750 hours, saving roughly $170k per year across 18 machines — pure parameter change, zero capex.
  • Re-parameterisation cycle introduced: the maintenance team re-fits the Weibull shape and catastrophic probability quarterly from the MMS data, and the optimum interval is recomputed automatically.
  • Catastrophic spares pre-positioned on the three most catastrophic-prone components (final drive, hydraulic pump, transmission) — a $0.2M inventory investment that the model showed shaved another 1.5 percentage points off the P90.
  • CBM investment justified for cranes and high-value excavators on a separate sweep — but not for haul trucks, where the catastrophic-failure cost is too low for CBM sensors to pay back.

ModelRisk Functionality Used

  • Weibull time-to-failure with shape β = 2.2 — fitted to MMS data; the textbook wear-out model where the constant-hazard exponential would have mis-stated PM benefit.
  • LogNormal repair cost plus a 5% catastrophic-event mixture — the single most important parameter for PM-interval optimisation; without it, every interval looks too long.
  • Parameter sweep across six PM intervals with full distribution output (mean, P10, P90) for each — the U-shape that point estimates cannot draw.
  • Scenario CDFs by duty cycle — showing why the optimum interval is site-specific, not a fleet-wide constant.
  • Tornado on the optimum's expected cost — directing periodic re-parameterisation effort to the Weibull shape, the single biggest mover.

PM interval optimisation is a problem deterministic models cannot solve correctly because the answer lives in the tail of the failure-cost distribution. Monte Carlo simulation in ModelRisk makes that tail visible — and once it is visible, the right PM interval is no longer the manufacturer's recommendation but the parameter that minimises the full expected cost including the catastrophic failures that defined the problem in the first place.