Three Strategies, Three Different Bets on Failure
Every maintenance strategy is really a bet about when and how an asset will fail, and each of the three dominant strategies — reactive, preventive, and predictive — makes that bet differently. Reactive maintenance (also called run-to-failure or corrective maintenance) makes no bet at all: the asset runs until it breaks, and repair happens after the fact. Preventive maintenance (PM) bets that failure probability is reasonably correlated with time or usage, so intervention on a fixed schedule (every 500 operating hours, every 90 days, every 10,000 cycles) will catch most failures before they happen. Predictive maintenance (PdM) bets that failure is preceded by measurable physical degradation — rising vibration amplitude, increasing bearing temperature, degrading oil chemistry — and that monitoring those signals lets you intervene only when the asset actually needs it, not on an arbitrary calendar.
None of these strategies is universally correct. The right choice is a function of failure mode, consequence of failure, and the cost of monitoring relative to the cost of unplanned downtime — and most real plants run all three simultaneously across different asset classes, which is the point this comparison is built around.
Reactive Maintenance: When Run-to-Failure Is Actually Correct
Reactive maintenance gets a bad reputation in most improvement literature, but it is the economically correct strategy for a specific, well-defined category of assets: low-cost, non-critical, easily replaceable components where the cost of monitoring or scheduled replacement exceeds the cost of simply replacing the part when it fails. A $4 conveyor idler roller with no safety implication and a five-minute swap is a textbook run-to-failure asset — building a PM schedule or condition-monitoring program around it destroys value rather than creating it.
The failure mode of reactive maintenance is in its misapplication: run-to-failure on a single point of failure asset (a main gearbox, a critical PLC, a compressor feeding an entire line) converts a preventable stoppage into an unplanned one, and unplanned downtime typically costs 3 to 5 times more than the equivalent planned maintenance action once you account for expedited parts freight, overtime labor, secondary damage from a cascading failure, and lost production. Industry surveys (Plant Engineering, ARC Advisory Group) consistently put unplanned downtime cost in discrete manufacturing at $10,000–$50,000+ per hour depending on the line, and in continuous-process industries (refining, chemicals) into six figures per hour — numbers that make reactive maintenance on critical assets one of the most expensive mistakes a plant can make, even though it appears cheapest on paper because no maintenance budget line item shows up until the failure occurs.
Preventive Maintenance: The Time-Based Middle Ground
Preventive maintenance schedules intervention — inspection, lubrication, part replacement, calibration — at fixed intervals defined by time (monthly), usage (every 2,000 hours), or cycles (every 50,000 actuations), independent of the asset's actual condition at that moment. Interval determination typically comes from OEM recommendations, historical failure data, or a Weibull analysis of time-to-failure distributions for that component class. PM's core statistical assumption is that failure probability rises predictably enough with age or usage (a wear-out failure pattern) that a fixed interval set safely ahead of the average failure point will catch most failures.
The well-documented weakness is the bathtub curve mismatch: the classic reliability bathtub curve shows most industrial components spend the bulk of their service life in a low, roughly constant "useful life" failure-rate region, not a rising wear-out region — the 1978 United Airlines/Nowlan-Heap study behind reliability-centered maintenance found only about 11% of aircraft component failure modes actually followed the age-related wear-out pattern that time-based PM assumes. For components with random (not age-correlated) failure modes, a fixed-interval PM program does two costly things at once: it fails to prevent the random failures it wasn't designed to catch, and it wastes labor and parts replacing components that still had significant remaining useful life — a phenomenon maintenance engineers call "infant mortality induced by maintenance," where the act of intervention itself (reassembly, reseating, recalibration) introduces new failure risk into an otherwise healthy component.
PM remains the right strategy where the underlying physics genuinely produces an age-correlated failure mode — belt and hose replacement, filter changes, lubricant replenishment, calibration drift — and it requires no sensor investment or specialized analysis skill, which makes it the default, lowest-barrier-to-entry strategy for most plants building out a maintenance program from a pure reactive baseline.
Predictive Maintenance: Condition-Based Intervention
Predictive maintenance replaces the fixed calendar with continuous or periodic condition monitoring, intervening only when measured data indicates the asset is actually degrading. The most common PdM techniques, each suited to different failure modes:
- Vibration analysis — accelerometers detect bearing wear, misalignment, imbalance, and gear tooth defects, typically weeks to months before failure, by tracking overall vibration amplitude (ISO 10816/20816) and frequency-domain signatures (FFT) tied to specific defect frequencies (bearing pass frequency, gear mesh frequency).
- Infrared thermography — thermal imaging catches electrical connection hot spots, bearing friction heat, and insulation breakdown; widely used on electrical switchgear and motor control centers where a 1°C-per-year baseline drift is a leading indicator of connection degradation.
- Oil analysis — tracks wear metal concentration (iron, copper, chromium via spectrometry), viscosity shift, and contamination (water, particulates) in lubricated systems, revealing internal wear before it produces detectable vibration or heat.
- Ultrasonic analysis — detects high-frequency sound from early-stage bearing friction, compressed air/gas leaks, and electrical arcing/tracking, often the earliest-available signal in the P-F interval for certain failure modes.
- Motor current signature analysis (MCSA) — detects broken rotor bars, bearing defects, and eccentricity from electrical signature alone, useful where physical sensor access is difficult.
The theoretical foundation for PdM's timing is the P-F curve: the interval between the point a potential failure becomes detectable (P) and the point it becomes a functional failure (F). Effective PdM requires a detection technique whose lead time is shorter than the P-F interval for that specific failure mode — vibration analysis on a bearing might give weeks of P-F interval, while ultrasonic detection of the same bearing's earliest friction signature might extend that warning window further, but a catastrophic shaft fracture may have a P-F interval measured in seconds, making no PdM technique fast enough and pointing instead toward redesign or protective interlock rather than monitoring.
PdM's costs are real and often underestimated in ROI pitches: sensor hardware, wiring or wireless infrastructure, analysis software, and — the most commonly underfunded line item — a trained vibration analyst or reliability engineer (Category II/III certification per ISO 18436) capable of correctly interpreting spectral data rather than just collecting it. A plant that buys sensors without building analysis capability ends up with expensive data no one acts on, which is a documented and common failure mode of PdM rollouts.
Cost Structure Compared
The three strategies trade fixed monitoring/labor cost against unplanned-failure risk along a predictable curve:
- Reactive — lowest planned cost, highest unplanned cost; total cost of ownership is dominated by downtime, secondary damage, and expedited-parts premiums. Appropriate only for low-consequence, low-cost, non-cascading failure modes.
- Preventive — moderate, predictable planned cost (parts and labor on a fixed schedule); reduces but does not eliminate unplanned failures, and adds waste from unnecessarily replacing healthy components. Best for genuinely age-correlated wear-out failure modes.
- Predictive — higher upfront capital cost (sensors, software, training) and ongoing analysis labor, but typically the lowest total cost of ownership on critical, capital-intensive assets, because intervention happens only when actually needed and with enough lead time to plan parts, labor, and a production window without expediting anything. Industry benchmarks commonly cited (DOE, SMRP) suggest PdM programs deliver 8–12% cost savings over PM-only programs and can reduce breakdowns by 70–75% versus a purely reactive baseline, though these figures depend heavily on asset criticality mix and program maturity.
Reliability-centered maintenance (RCM) is the formal framework (per SAE JA1011/JA1012) for making this selection rigorously rather than by habit: for each significant failure mode of each asset, RCM asks what the failure's consequence is (safety, environmental, operational, economic), whether a suitable condition-monitoring technique exists for that specific failure mode, and only then assigns reactive, preventive, predictive, or redesign as the appropriate response. This is why mature reliability programs never describe themselves as "we do predictive maintenance" or "we do PM" as a blanket plant-wide policy — the strategy is assigned failure-mode by failure-mode, asset by asset.
Asset Criticality: The Actual Decision Driver
The single most important input to strategy selection is asset criticality, typically scored through a matrix combining probability of failure with consequence of failure (safety risk, environmental exposure, production impact, repair cost, and redundancy). A criticality-ranked asset register — the output of an FMEA or a simpler risk-priority-number exercise — should drive strategy allocation directly:
- High criticality, no redundancy, high failure consequence (main line compressors, kiln drive motors, primary PLCs) — predictive maintenance, often layered with online continuous monitoring rather than periodic route-based collection, because the cost of an undetected failure is severe.
- Moderate criticality, age-correlated wear mode (belts, filters, seals, lubricants) — preventive maintenance on OEM- or history-derived intervals.
- Low criticality, low replacement cost, redundant or non-cascading (individual sensors on non-critical loops, low-cost rollers, indicator lamps) — reactive maintenance is the economically correct choice.
A common and costly mistake is applying the same maintenance philosophy uniformly across an entire plant regardless of this criticality gradient — either over-investing PdM infrastructure on assets where run-to-failure would have been cheaper, or leaving genuinely critical single points of failure on a reactive or under-scoped PM schedule because "that's just how we've always done it here."
CMMS and Data Infrastructure Requirements
Each strategy places different demands on a computerized maintenance management system (CMMS). Reactive maintenance needs only a work-order and parts-inventory system to track repairs after the fact. Preventive maintenance requires the CMMS to trigger work orders automatically on time or usage thresholds, which in turn requires accurate runtime or cycle-count data feeding the system — a PM program built on a CMMS with stale usage data silently degrades into either wasted early replacements or missed intervals. Predictive maintenance requires the CMMS to integrate with a condition-monitoring platform (vibration data historian, IIoT sensor gateway, or a standalone reliability software package) so that alarm thresholds and trend data can automatically generate work orders when a monitored parameter crosses a defined limit — without that integration, PdM data lives in a separate silo that maintenance planners have to check manually, which in practice means it often doesn't get checked consistently enough to realize the strategy's lead-time advantage.
Building a Blended Program in Practice
Almost no real plant runs a single strategy plant-wide, and pursuing "100% predictive" as a goal is usually a misallocation of capital. A practical rollout sequence that most reliability engineering teams follow: first build accurate asset criticality rankings and a clean CMMS foundation; apply preventive maintenance broadly to known age-correlated wear items as a baseline improvement over pure reactive; then layer predictive techniques selectively onto the highest-criticality, highest-consequence assets where the P-F interval and failure-mode physics genuinely support condition monitoring; and deliberately leave low-criticality, low-cost components on run-to-failure rather than over-engineering a monitoring program around them. Measured against baseline KPIs — mean time between failures (MTBF), mean time to repair (MTTR), planned maintenance percentage (PMP), and overall equipment effectiveness — this blended, criticality-driven allocation consistently outperforms any single strategy applied uniformly.