| Vose Software

Industry: Construction and Infrastructure
Product: ModelRisk
Application: Managing Equipment Downtime Under Uncertainty


Why a $3.8M Maintenance Budget Runs Out One Year in Three

A heavy-civil contractor running a $260M dam-and-tunnel project carries 24 critical machines — six tower cranes, four TBM components, eight haul trucks, four excavators, two batching plants. The deterministic downtime budget is $3.8M annually. The probabilistic rebuild in ModelRisk puts the mean at $5.0M, the P90 at $10.3M, and a 35% probability of spending above $5M — driven less by failures themselves than by spare-parts shortages and by the rare catastrophic failure of a single TBM component, which together create a tail far heavier than the mean suggests.

The diagnostic question is not "how often do machines break?" — fleet maintenance teams know that. It is "what is the joint cost of failure events plus the lead time on imported spare parts in the worst quarter of the year?" The mean answer is comfortable. The 90th-percentile answer pays for redundancy.

Why Weibull, why NegBin, why LogNormal

Time-to-failure for heavy construction equipment is Weibull with shape β > 1 — wear-out, not random. β between 1.8 (haul trucks, intensive duty) and 2.6 (TBM components, infrequent but punishing) reflects the failure pattern the maintenance team actually sees. An exponential time-to-failure — the standard textbook starting point — assumes constant hazard rate and would systematically under-state failure risk for an aging fleet.

The annual failure count per machine is Negative Binomial rather than plain Poisson. The reason is over-dispersion: in months of high site activity (concrete pour cycles, tunnel-drive pushes) failures cluster, because the machines are working harder and operators are pushed. Dispersion parameter 1.6 is calibrated to the maintenance-management system's last three years of data.

Repair durations are LogNormal — right-skewed because the long tail is dominated not by the technical complexity of the repair but by parts lead time. An imported gearbox can sit in customs for two weeks. The model lays this on top: each failure carries an 18% probability of a 2.5×–4.5× repair-time multiplier representing the spare-parts shortage that the project's offshore supply chain delivers more than the field crew would like.

When in the year, and which equipment

The annual aggregate hides a seasonal pattern that the maintenance planner can act on. The heatmap of expected monthly downtime hours by fleet class makes both visible:

Expected monthly downtime hours by fleet class

Haul trucks peak in mid-summer (high cycles per shift, heat-related component stress). Batching plants and excavators peak in the winter months (cold-start hydraulics, cement-line freezing risk). TBM components run a long flat tail across the year but pop higher in November–February when scheduled cutter-head changeouts coincide with parts-shipment cycles. The Pareto by group totals these to annual fleet downtime of about 1,500 machine-hours mean, P90 about 2,700, P99 over 3,900, with annual downtime cost mean $5.0M, P50 $3.6M, P90 $10.3M, P95 $13.7M. Note the gap between the median ($3.6M) and the mean ($5.0M): the distribution is sharply right-skewed, so the typical year looks affordable while the average year is dragged up by the spares-shortage tail. The deterministic budget of $3.8M sits just above the median — so in a typical year it holds, which is exactly why the deterministic view felt safe — but it is breached 45% of the time and the overruns, when they come, are large.

Where the dollars are concentrated

Downtime cost by fleet group — Pareto (annual)

The Pareto by fleet group makes the priority obvious: TBM components account for roughly 55% of expected downtime cost despite being only 4 of 24 machines, because their idle cost is $4,800/hr — an order of magnitude above any other group. Haul trucks have many failures but cheap downtime; tower cranes are expensive per hour but fail less often. The classic 80/20 picture is more like 70/20 here: two fleet groups carry the cost.

How much spares inventory for what confidence?

Spares-pool size is a continuous lever. The threshold-sweep chart traces P(annual cost stays within budget) against spares-pool capex, for three candidate budget thresholds:

P(annual downtime cost <= budget) vs on-site spares inventory

The three curves are three candidate budgets: a tight deterministic plan ($3M), a mean-plus-buffer figure ($6M), and a P90 contingency ($10M). At zero spares capex (relying on offshore air-freight), even the mean-plus-buffer $6M budget holds only ~65% of the time, and the tight $3M plan barely a third. Investing ~$1.4M in an on-site spares pool — which cuts spare-shortage probability from 18% per failure toward 2% — lifts the $6M budget to ~84% (just touching the 85% target) and keeps the $10M P90-contingency comfortably above 85% across the whole range. The deterministic $3M plan never clears 55%, no matter how many spares are bought, because the irreducible failure-and-repair variability alone exceeds it. The lesson: spares inventory buys real tail protection, but only against a budget that was set to the distribution, not to the deterministic point estimate.

What changed at the steering committee

  • Downtime budget reset to $6M (mean + buffer) replacing the deterministic $3.8M — eliminating the structural under-budgeting that the mean-based forecast had hidden.
  • ~$1.4M of capex authorised for an on-site spares pool plus TBM-component redundancy, justified on P90 protection rather than expected value.
  • Spare-parts strategy reframed. The model showed that lead time, not failure rate, was the dominant downtime driver; the maintenance contract was renegotiated to bring critical spare inventory on-site rather than relying on offshore air-freight.
  • Tower-crane PM cadence held. The model also showed crane PM at 750 hrs was already near optimum — a finding that prevented a planned over-investment in additional crane preventive maintenance.

ModelRisk Functionality Used

  • Weibull time-to-failure with group-specific shape (β 1.8–2.6) — the right wear-out model where exponential would have under-stated aging risk.
  • Negative Binomial annual failure counts with dispersion 1.6 — capturing the failure clustering during high-duty months that a plain Poisson smooths away.
  • LogNormal repair durations with a Bernoulli spare-shortage multiplier — separating the technical repair distribution from the parts-supply-chain tail that actually drives the worst weeks.
  • Pareto-by-group output that re-prioritised maintenance investment from haul-truck (high count, low cost) to TBM (low count, very high cost per hour of downtime).
  • Before/after CDF on the $5M budget breach probability — the explicit metric the steering committee used to authorise the redundancy capex.

Downtime risk lives in the joint distribution of failure events and parts lead times — not in either separately. Monte Carlo simulation in ModelRisk is what makes that joint distribution visible, and once it is visible, redundancy investment can be priced against the tail it actually prevents.