PONOPT FIELD NOTES · Данные, GIS и AI

Predictive Maintenance Without the Hype: Which Failures Can Be Forecast?

Which equipment failures can predictive maintenance actually forecast? A practical guide to age-related vs. random failures, P-F intervals, method limits, and decision tools.

Predictive maintenance genuinely works only for failures that develop gradually and leave measurable clues before the machine stops: wear, corrosion, fatigue, contamination and insulation or capacity loss you can catch through vibration, temperature, oil or current. It cannot forecast sudden random events such as operator error, debris ingestion or external shocks; those need protection and root-cause removal, not prediction. Classify each failure mode before spending on sensors and models.

Key takeaways

  • Roughly 80% of failures are not predictable from age or running hours: only wear and degradation that leave a measurable trace can be forecast on a calendar.
  • The P-F interval — time between a defect becoming detectable and functional failure — is the core concept; prediction is only possible where such a window exists.
  • Confidently forecastable modes include bearing and bushing wear, misalignment and unbalance, gear fatigue, impeller wear, filter clogging, and motor insulation degradation.
  • Sudden failures — operator error, ingested debris, impacts, power surges — have no warning window and are handled with protection, redundancy and root-cause removal.
  • Forecast reliability falls as the horizon grows: a model accurate for tomorrow degrades sharply for next week, and quoted accuracy percentages are context-specific.
  • The biggest implementation killer is false alarms and the inability to act: a signal must reach a work order and root-cause review, or trust in the tool collapses.

The failure-pattern reality: most breakdowns aren't age-related

The single most useful correction to predictive-maintenance hype is understanding how equipment actually fails over time. Reliability research that began with United Airlines engineers Stan Nowlan and Howard Heap in their 1978 Reliability-Centered Maintenance report — and was later repeated in naval and industrial studies — found that most components do not wear out predictably with age. Only a small share of failures follow the classic bathtub curve; in the civil-aviation study roughly four percent of components did, and only about 11 to 20 percent of failures across the main studies proved strongly age-related.

In practice, most reliability engineers work with a simpler summary: only about 15–20% of failures can be forecast from age or running hours, while the other 80–85% appear random with respect to time. That does not mean they have no cause — misalignment, contamination, moisture, operator error and external shocks are all real causes — it means their timing cannot be predicted from a calendar. The consequence is that time-based overhauls alone will never prevent most breakdowns, and an analytics model that only counts operating hours will miss most of the value.

  • Infant mortality: a burst of failures right after installation or overhaul.
  • The chance or random period: a relatively constant failure rate for most of the service life.
  • The wear-out zone: a sharp rise for components that genuinely age or cycle.

What condition monitoring can actually forecast

What can be forecast is degradation — a physical process that begins to change measurable condition well before the machine stops working. This is the potential-to-functional-failure (P-F) interval concept used in reliability-centered maintenance. Vibration analysis, oil and debris analysis, thermography, ultrasound and current monitoring detect the moment a defect becomes detectable (point P), leaving a window before functional failure (point F) in which you can plan a repair rather than react to a breakdown.

Concretely, the classes that lend themselves to prediction include bearing and bushing wear, shaft misalignment and unbalance, gear fatigue, impeller and pump wear, filter and screen clogging, lubrication breakdown visible in oil samples, insulation and winding degradation in motors, and gradual capacity loss in batteries. Each leaves a signature — rising vibration, rising temperature, metal particles in oil, rising current — that a sensor or an inspector can catch in time.

The failures prediction cannot catch

Prediction fails for failures that reach functional failure without a usable advance window, or whose cause is external and instantaneous. Operator error, debris or oversized material ingested into a machine, lightning and power surges, shipping or installation damage, and sudden single-point electronic failures typically offer no detectable P-F interval at all. A model cannot sound an early warning about an event that has no warning.

For these, the honest strategy is not more analytics but protection, redundancy and cause removal: guarding and interlocks, surge protection, spares and redundancy for single points of failure, and root-cause analysis that eliminates the triggering condition. Predictive maintenance that only flags these after the fact — or that mislabels a sensor fault as a real machine fault — wastes money and erodes trust.

How forecasts are made — and why they degrade

Where a measurable degradation process exists, three families of methods estimate remaining useful life (RUL). Purely data-driven models learn patterns from sensor histories; physics-based models simulate wear, fatigue and stress from first principles; hybrid physics-informed models combine both. Data-driven models are easier to deploy but are only as good as their training data, extrapolate poorly to never-seen conditions, and can produce physically impossible answers. Physics-based models generalize better but demand detailed knowledge of each component.

A recurring practical limit is that forecasts get less reliable the further they reach: accuracy that is high for tomorrow degrades sharply for next week. On top of that, few plants have clean full-lifecycle failure data, because machines are repaired before they fail. Treat any quoted "percent accurate" figure as context-specific rather than a promise.

Making predictions you can act on: false alarms, data and trust

The biggest operational killer of predictive maintenance is not the algorithm but the way alarms are used. Route-based monitoring only sees a machine at the moment of inspection, so a machine can fail between rounds; continuous sensors close that gap but can drown teams in noise. When many warnings turn out to be false positives, or cannot be acted on because crews are busy with today's breakdowns, technicians stop trusting the tool — a pattern sometimes called the predictive-AI paradox.

Good practice includes sizing the inspection interval as a fraction of the P-F interval so a symptom is seen several times before failure, checking that sensors actually sit where the signal exists (poor placement can make a fifth of them useless), confirming a reading before acting, and ensuring every confirmed prediction feeds a work order and a root-cause review. Detection alone does not remove the cause: unless the failure mode is eliminated, the same random fault can recur.

  • Set measurement frequency to capture several points inside the P-F interval.
  • Confirm a warning with a repeat measurement before stopping the line.
  • Route every confirmed prediction to a work order and root-cause analysis.
  • Avoid the "data graveyard": a sensor that is not connected or analyzed delivers nothing.

Where prediction belongs in your maintenance mix

Prediction is one tool in a portfolio, not a replacement for strategy. For strongly age- and cycle-driven parts that are cheap to replace, scheduled time-based maintenance may still be the most cost-effective answer. For degradation that offers a usable P-F window and where downtime is expensive, condition-based and predictive maintenance earns its keep. For high-consequence failures with no warning, invest in design, protection and root-cause elimination.

The screener below is a practical way to classify your own failure modes before you buy sensors and models. Most sites find that only a minority of assets justify full continuous prediction; the rest are better served by inspections, scheduled replacement, protection, or well-managed run-to-failure.

Failure-mode predictability screener: which failure deserves which strategy

Run your recurring failures through this matrix, starting with work orders from the last two to three years and a Pareto chart of what fails most. Your answers decide whether to invest in sensors and models or stick with scheduled replacement, protection and cause removal.

  1. List your 10–15 most frequent and most expensive failures from work orders and long-serving maintainers.
  2. For each, ask: is there a measurable early sign before the stop — vibration, heat, metal in oil, a leak, rising current?
  3. Yes → candidate for condition-based prediction; estimate the shortest P-F interval you have seen in history.
  4. No, but the failure is strongly age/cycle-driven and the part is cheap → scheduled replacement before the wear-out zone.
  5. No warning window and low consequence → managed run-to-failure with spares and protection in place.
  6. No warning window and high consequence → add redundancy and protection and remove the root cause; prediction will not save you here.
  7. Match the technique to the symptom: vibration for bearings, oil analysis for wear and contamination, thermography for electrical, current for motors.
  8. Verify the sensor sits where the signal exists and that data actually reaches analysis rather than a "data graveyard."
  9. Assign owners and an action path: a confirmed prediction must become a work order and a root-cause review.
  10. Track implementation metrics: false-alarm rate, confirmed detections, and average lead time from warning to failure.

Questions people ask

Which types of failure can predictive maintenance realistically forecast?

Failures that develop gradually and leave a measurable trace before the stop: bearing and bushing wear, misalignment and unbalance, gear fatigue, impeller wear, filter clogging, lubricant and insulation degradation, and gradual battery capacity loss. Each gives a window — the P-F interval — between a defect becoming detectable and functional failure, inside which intervention is possible. Sudden events with no advance signs cannot be forecast.

Why are most failures described as "random," and what does that mean for scheduled maintenance?

Research dating to Nowlan and Heap's 1978 report found that only about 15–20% of failures depend on age or running hours; the rest are random in time. "Random" does not mean causeless — it means the moment cannot be predicted from a calendar. As a result, fixed-interval overhauls alone will not prevent most breakdowns and can even introduce installation defects. Condition monitoring plus root-cause removal addresses the majority of real-world failures.

What is the P-F interval and how does it set inspection frequency?

The P-F interval is the time between a defect becoming detectable (P, potential failure) and functional failure (F). A single failure mode can have several symptoms, each with its own interval — vibration may be visible months ahead, heat only weeks. Set inspection frequency as a fraction of the interval so you observe the symptom several times before failure, leaving room to confirm the reading and plan the repair. Shorter intervals demand more frequent rounds or continuous sensors.

Why do predictive-maintenance programs often disappoint, and how do you avoid that?

Common causes are false alarms, data graveyards from poorly placed sensors, route-based readings that miss failures between rounds, and a lack of authority to stop production when a signal appears. Forecasts also degrade as the horizon grows. To make it work, confirm warnings before acting, drive inspection frequency from the P-F interval, route every prediction to a work order and root-cause review, and treat claimed accuracy percentages as context-specific rather than guarantees.

What replaces prediction for sudden failures with no warning signs?

For events without a P-F interval — operator errors, ingested debris, lightning, power surges, sudden electronic failures — predictive models add little. Use guarding and interlocks, surge protection, redundancy and spares for single points of failure, and root-cause analysis that removes the triggering condition. This is usually cheaper and more reliable than trying to predict a physically unpredictable event.

Sources and further reading

Sources were checked when this page was generated. Confirm changing dates, rules and prices with the original publisher.

  1. Effective maintenance strategies begin with an understanding of why assets failSchneider Electric Blog
  2. What is Condition Based Maintenance Strategy?Accendo Reliability
  3. Use P-F Intervals to Map, Avert FailuresReliable Plant
  4. An Introduction to Equipment Failure PatternsLimble CMMS
  5. RCM Failure Charts: Age Related or RandomReliabilityweb.com
  6. Closing Status Gap Between Maintenance and AIReliable Plant