The short answer
Prioritize a maintenance backlog by scoring each open work order across three dimensions: the asset's criticality and probability of failure (risk), the business and repair cost if it fails or stays deferred (cost), and how long the job really takes including parts lead time (repair time). Gate safety and compliance work to the top, weight the three factors into one score, add an aging bump, and let a trained planner make the final call.
Key takeaways
- Backlog age is a weak ranking signal; rank by the risk, cost-of-failure and duration profile of each deferred job, not just by how long it has waited.
- Express risk as probability of failure multiplied by consequence of failure; a cheap single-point-of-failure asset can outrank an expensive one with redundancy.
- Include the cost of deferral (lost production, quality, energy, safety exposure), which often exceeds the cost of the repair itself.
- Add repair time and parts lead time so long-lead jobs are started early rather than scheduled only after a failure has occurred.
- Safety, compliance and legal triggers should always override the numeric score and jump to the top of the queue.
- Recalculate priorities periodically and refresh asset criticality at least every 12–24 months or after any major process change.
Why order age is a poor ranking tool
The most common shortcut is to work the oldest orders first. It is easy to understand but dangerous: a ten-day-old cosmetic paint request can outrank a fresh bearing replacement on the single compressor that feeds the whole line. Age only tells you how long a job has waited; it says nothing about what happens if it waits longer.
Deferred maintenance usually begins with minor problems pushed aside because resources are limited, and the longer tasks sit, the larger the repair, downtime, risk and energy cost become. Mature teams therefore separate strategic deferral, when a low-priority job is consciously moved forward so a critical one runs, from accidental deferral, when the queue simply grows from emergencies and staffing gaps.
Replacing “oldest first” with “most dangerous first” needs a model. At minimum a planner needs two inputs: the asset's criticality from a systems criticality analysis and the operational criticality of the specific job. Multiplying the two gives a more accurate priority. The sections below add cost of failure and repair time to that base.
Factor 1 — Risk: criticality and probability of failure
Risk is best expressed as probability of failure multiplied by consequence of failure. Consequences are scored across several impact categories — safety and environment, production stoppage, repair cost and quality — in the same spirit as FMEA/FMECA criticality reviews, where you take the highest consequence score and multiply it by the probability score.
Criticality is a property of the asset, not of one work order, and it is fairly stable. Before scoring, build a clean asset hierarchy (site, area, unit, equipment, component) and assess criticality at the equipment level, not the bearing level. Refresh the ranking every 12–24 months or when the process, regulations or demand change.
A common mistake is treating expensive equipment as automatically critical. If a plant runs three identical compressors but needs only two, each one's criticality drops because redundancy softens the consequence of failure. Conversely, a low-cost sensor that is the single point of failure for a multi-million-dollar line is a highly critical asset.
- Probability of failure: failure history, MTBF, age, operating environment, vibration and temperature trends.
- Consequence of failure: injuries and violations, line stoppage or available backup, scrap volume, replacement cost.
- Criticality is assigned to the asset once; priority is the dynamic attribute of each order placed on it.
Factor 2 — Cost of failure and cost of deferral
The second factor answers a business question: what does the company lose if the work is not done now but postponed for a week or a month? This is not the repair bill but the price of inaction: lost output and margin, overtime, excess energy, scrap, penalties and reputation. Comparing “do it now” versus “do it later” yields the return on investment of each job in the backlog — the larger the gap, the higher the priority.
In practice, estimate a weekly deferral cost in money for each order in the upper part of the queue. A job on a non-redundant packaging line costs an hourly stoppage rate if postponed; a lubrication job on a ventilation fan costs far less. The model automatically lifts the former above the latter.
The cost of failure also includes the full repair if neglect turns into a breakdown. A complete failure is almost always more expensive than prevention, so all else being equal you close a job on a critical asset with a high probability of failure before cosmetics on a secondary one. This is the heart of maintenance economics: directing scarce resources to where risk and loss are greatest.
Factor 3 — Repair time and logistics reality
The third factor is how long the job realistically takes: labor hours, availability of the right skill, parts lead time and the shutdown window. Mean time to repair and mean time between failures are core reliability metrics that inform both criticality and planning. A job with a long parts lead time must start early even if its “pure” risk is lower, or the failure will arrive before the part does.
This factor also captures windows: some jobs need a line stoppage and production sign-off. They cannot run tomorrow regardless of score; they must be planned into the next available window with a complete package — spares, tooling, permits. The planner assembles materials before the order reaches the crew, not after.
Logistics is why ranking by risk and money alone fails: a high score on a job that physically cannot be done this week without parts is only an illusion of control. Its right home is the next real window with a ready kit, while freed hours go to executable medium-priority work in the meantime.
Combining the three factors into one score
A practical scheme is two-stage. First run a gate: any safety, compliance, legal or breakdown work goes to the top regardless of score. Everything else is scored and sorted. The same logic appears in asset criticality programs, where a high ranking dictates a more sophisticated maintenance strategy.
For scoring, use an additive-multiplicative model. Assign the asset a criticality of 1–5. For the order, rate the probability of failure before the next window (1–5) and the consequence of failure (1–5). Multiply probability by consequence for a risk sub-score from 1 to 25. Then add a weighted cost-of-deferral term and a correction for repair time and parts lead time.
Example: a seal job on a non-redundant feed pump — criticality 5, probability 4, consequence 4 — gives a risk of 16; high weekly deferral cost and parts on hand for a two-hour job push it into the top ten. Painting a pipe rack — criticality 2, probability 1, consequence 1 — gives a risk of 2 and low deferral cost; it falls to the end of the queue but is not forgotten.
- Gate first: safety, compliance and legal work go to the top with no arithmetic.
- Risk = probability of failure × consequence of failure (1–25).
- Add weekly deferral cost in money and a repair-time/lead-time correction.
- Sort descending, then verify the top ~20% is executable with current skills and spares.
Refresh cadence, data quality and limits
The model is only as good as the data. First purge duplicates, obsolete and merged orders from the queue, and review older open orders alongside new ones so crews do not run an alignment this week and a motor repair with no alignment next week. Without decent failure history in the CMMS, probability estimates are guesses — acknowledge it and start capturing failure codes and running hours.
Recalculate priorities weekly or after a material event: an incident, a demand shift, an arrival of spares. Field evidence shows that cleaning up criticality and removing over-scheduled work can cut backlog substantially and free time for genuinely critical jobs.
Keep the limits in view. The score supports a decision; it does not replace the planner and supervisor who see site context, crew skill and agreements with production. Do not promise a fully empty queue — a small healthy backlog is normal; uncontrolled growth is the danger. The model orders work and makes the trade-off among risk, money and time explicit and defensible, but the final call stays with a person.
Put it into practice
Backlog prioritization scorecard (printable worksheet)
A reusable planner worksheet: five scoring steps per open order, an override gate, and a sorting rule. Score each order in the queue, then rank by the composite score and sanity-check the top tier.
- Pull all open work orders and remove duplicates, obsolete and merged jobs before scoring.
- Record the asset ID and its current criticality (1–5) from the latest criticality review.
- Score the probability of failure before the next maintenance window on a 1–5 scale (failure data, MTBF, trends).
- Score the consequence of failure on a 1–5 scale across production, safety, environment and quality.
- Multiply probability by consequence for the risk sub-score (1–25).
- Estimate weekly deferral cost in money: lost output, overtime, energy, penalties, scrap.
- Note repair effort in hours and parts lead time in days; flag anything beyond your planning horizon.
- Add the repair-time/lead-time correction and an aging bump (e.g., +1 per overdue week).
- Apply the override gate: safety, compliance and legal work go to the top regardless of score.
- Sort descending and confirm the top ~20% is executable with available people and spares in the next windows.
Questions people ask
What is the difference between asset criticality and work-order priority?
Criticality is a relatively stable property of the equipment and its role in the process — what happens if the asset fails. Priority is the dynamic attribute of a specific work order: a leaking seal on a critical pump is high priority, while painting that same pump is low priority. Criticality sets the base and the maintenance strategy for the asset; priority decides the order in which jobs are executed right now.
How should safety and compliance work be treated in a scoring model?
Put it outside the arithmetic behind an explicit gate: any job tied to injury risk, regulatory exposure or inspector findings goes to the top of the queue regardless of its computed score. Such work should not compete on weights with economic factors, because its cost cannot be fairly expressed in money per week of deferral.
How often should asset criticality scores be refreshed?
The common recommendation is every 12 to 24 months, plus whenever something material changes: new equipment, a modified process, added redundancy or a new single point of failure, changed regulations or shifts in demand. If the ranking is never refreshed it goes stale, and the model will systematically over- or under-prioritize certain assets.
Does an expensive repair raise or lower a job's priority?
The price of the repair alone does not set priority — the comparison between repair cost and the cost of failure or deferral does. If an expensive job prevents an even more expensive stoppage, its priority is high. If the asset is secondary, redundant and unlikely to fail, an expensive job can be deferred, but it must not be lost: budget for it and order long-lead parts early.
What is a realistic healthy backlog size?
A common rule of thumb is to hold the total backlog at roughly two to four weeks of the crew's planned capacity, with a “ready” backlog of one to four weeks per craft. A completely empty queue is abnormal and usually means work is not being detected. The danger is uncontrolled growth: the longer planned or corrective work goes undone, the higher the risk of an expensive failure.
Why does redundancy change an asset's criticality?
Criticality is defined by the consequence of failure, not by purchase price. If three identical compressors are installed but only two are needed, a failure of one will not stop production because the backup softens the consequence, so each unit's criticality drops. Conversely, a cheap valve or sensor that is a single point of failure for an entire expensive line ranks highly critical because its failure halts a costly process.
Sources and further reading
Sources were checked when this page was generated. Confirm changing dates, rules and prices with the original publisher.
- Calculating RIME: Ranking Index for Maintenance ExpenditureseMaint (Fluke Reliability)
- Reliability workflowBIC Magazine
- Maintenance KPIs for Every RoleeMaint (Fluke Reliability)
- How to Find Critical Assets: Asset Criticality Ranking GuideF7i / Factory AI
- AIM Strategy Reduces FPSO Maintenance Backlog 40%ABS Group
- What Is Deferred Maintenance?MaintainX