Overview
Maintenance Strategy Development is the process used to determine how an asset, system, or failure mode should be managed so that required function is preserved at an acceptable level of risk and lifecycle cost.
A maintenance strategy is not simply a PM list. It is the logic that connects required function, functional failure, failure mode, consequence, criticality, detectability, technical feasibility, task selection, interval or trigger, corrective response, and ownership.
The objective is not to maximize maintenance activity. The objective is to select the least intrusive, technically effective, and economically justified action for the failure mode and operating context.
What It Is
A maintenance strategy defines:
- what function must be protected;
- how that function can fail;
- what happens if it fails;
- how significant the consequence is;
- whether degradation can be detected;
- whether a scheduled task is technically effective;
- when the task should occur;
- when run-to-failure is acceptable;
- when redesign or procedural change is required.
Possible strategy outputs include:
- time-based preventive maintenance;
- usage-based preventive maintenance;
- condition-based or predictive maintenance;
- failure-finding;
- scheduled restoration;
- scheduled discard;
- operator care;
- corrective maintenance;
- run-to-failure;
- redesign;
- operating controls;
- spare-parts and contingency planning.
Why It Matters
A good maintenance strategy:
- reduces unnecessary maintenance;
- focuses effort on credible failure modes;
- improves PM and PdM quality;
- supports asset availability;
- improves planning and scheduling;
- strengthens spare-parts decisions;
- reduces maintenance-induced failures;
- makes run-to-failure decisions explicit;
- identifies when maintenance cannot solve the problem and redesign is required.
Without a clear strategy, organizations often inherit OEM checklists, duplicate PMs, over-maintain low-risk assets, under-maintain critical failure modes, and continue repeating ineffective tasks.
When to Use It
Use Maintenance Strategy Development when:
- commissioning new assets;
- creating or rebuilding a PM program;
- reviewing chronic failures;
- implementing RCM or FMEA outputs;
- deploying PdM;
- reviewing high-criticality assets;
- standardizing maintenance across similar equipment;
- evaluating run-to-failure;
- changing operating context;
- introducing new technology or controls.
Do not use it as a substitute for work prioritization, job planning, scheduling, RCA, or capital-project selection.
Core Principles
- Required function comes before maintenance task.
- Failure modes drive strategy selection.
- Consequence determines the rigor of the decision.
- Asset criticality informs where analysis effort should be concentrated.
- The least intrusive effective strategy is preferred.
- More maintenance is not automatically better.
- Run-to-failure can be a valid strategy.
- Hidden failures require failure-finding.
- Redesign is appropriate when maintenance cannot adequately control risk.
- OEM recommendations are inputs, not automatic final decisions.
- Operating context matters.
- Strategy effectiveness must be reviewed using failures and condition findings.
Process or Lifecycle
- Define Required Function leads to Identify Functional Failures.
- Identify Functional Failures leads to Identify Failure Modes.
- Identify Failure Modes leads to Evaluate Consequences.
- Evaluate Consequences leads to Apply Criticality and Operating Context.
- Apply Criticality and Operating Context leads to Select Strategy.
- Select Strategy leads to Build Tasks and Job Plans.
- Build Tasks and Job Plans leads to Configure CMMS.
- Configure CMMS leads to Execute and Capture Results.
- Execute and Capture Results leads to Review Failures, Findings, and Cost.
- Review Failures, Findings, and Cost leads to Select Strategy.
A strategy that is never reviewed becomes an assumption rather than a controlled maintenance decision.
Roles and Responsibilities
| Role | Responsibility |
|---|---|
| Reliability Engineer | Leads strategy development and optimization |
| Maintenance Planner | Converts approved strategy into executable work |
| Scheduler | Integrates strategy-generated work into schedules |
| Technician | Executes tasks and provides condition feedback |
| Maintenance Supervisor | Verifies execution and task quality |
| Operations | Defines operating context and supports access |
| Engineering | Supports failure analysis and redesign |
| EHS | Reviews safety and environmental consequence |
| Quality / Food Safety | Reviews product and compliance consequence |
| CMMS Administrator | Maintains plans, task lists, counters, and master data |
| Maintenance Manager | Owns performance and strategy governance |
| Site Leader | Resolves business-risk and resource conflicts |
Required Inputs
Typical inputs include:
- asset hierarchy;
- required function;
- performance standard;
- operating context;
- asset criticality;
- failure history;
- failure modes;
- OEM information;
- drawings;
- process data;
- condition-monitoring data;
- regulatory requirements;
- quality and food-safety requirements;
- safety and environmental consequence;
- spare-parts lead times;
- labor capability;
- shutdown opportunities;
- redundancy;
- contingency options;
- lifecycle cost information.
Required Outputs
A completed strategy should define:
- asset or asset class;
- required function;
- performance standard;
- functional failure;
- failure mode;
- failure effect;
- consequence;
- criticality;
- selected strategy;
- task type;
- interval or trigger;
- technical basis;
- acceptance criteria;
- labor and material requirements;
- owner;
- review date.
Step-by-Step Implementation
1. Define the required function
Describe what the asset must do and the required performance standard.
Example:
Function: Deliver 450 gallons per minute of process water at 60 psi. Functional failure: Unable to deliver at least 450 gallons per minute at 60 psi when demanded.
2. Identify functional failures
Functional failures may be complete, partial, intermittent, hidden, degraded, unsafe, inefficient, quality-related, or unavailable on demand.
3. Identify credible failure modes
Failure modes should be technically specific and actionable.
Examples include bearing lubrication contamination, seal wear, impeller erosion, winding insulation breakdown, sensor drift, valve sticking, filter blockage, and belt fatigue.
4. Evaluate consequences
Consider safety, environmental, food safety, quality, regulatory, production, customer, financial, asset damage, recovery time, and redundancy.
5. Select the strategy
| Strategy | Appropriate when |
|---|---|
| Time-based PM | Failure is meaningfully age-related |
| Usage-based PM | Failure relates to hours, cycles, mileage, or throughput |
| Condition-based maintenance | Degradation can be detected with sufficient warning |
| Failure-finding | Failure is hidden until demand |
| Scheduled restoration | Restoration renews resistance to failure |
| Scheduled discard | Replacement before a life limit is effective |
| Operator care | Simple safe checks can be performed by operators |
| Corrective maintenance | Defect is identified before functional failure |
| Run-to-failure | Consequence is acceptable and response is controlled |
| Redesign | Maintenance cannot adequately control risk |
| Procedural control | Operating method materially influences failure |
| Spare/contingency strategy | Fast recovery is more effective than prevention |
6. Set interval or trigger
Consider age, usage, P-F interval, failure consequence, environment, regulatory requirement, historical findings, lead time for corrective work, and shutdown constraints.
7. Build executable work
Translate strategy into job plans, task lists, material requirements, acceptance criteria, condition limits, work triggers, counters, and inspection routes.
8. Review effectiveness
Use work history, failures, findings, costs, and condition data to decide whether the strategy remains effective.
Decision Rules
- Failure Mode Identified leads to Consequence Acceptable?.
- Consequence Acceptable?, when Yes, leads to Effective Preventive or Predictive Task?.
- Consequence Acceptable?, when No, leads to Can Risk Be Controlled by Maintenance?.
- Effective Preventive or Predictive Task?, when No, leads to Run-to-Failure with Contingency.
- Effective Preventive or Predictive Task?, when Yes, leads to Select Least Intrusive Effective Task.
- Can Risk Be Controlled by Maintenance?, when Yes, leads to Select Least Intrusive Effective Task.
- Can Risk Be Controlled by Maintenance?, when No, leads to Redesign or Operating Change.
- Select Least Intrusive Effective Task leads to Set Interval or Trigger.
- Set Interval or Trigger leads to Build Job Plan and CMMS Controls.
Run-to-failure is appropriate only when
- consequence is low;
- failure is evident;
- secondary damage is limited;
- repair is straightforward;
- spares are available;
- downtime is acceptable;
- no better task exists;
- contingency is defined.
CMMS Considerations
A CMMS may need to support asset hierarchy, criticality, failure-mode coding, maintenance plans, task lists, counters, condition triggers, job plans, BOMs, strategy classification, regulatory flags, version control, work-order linkage, measurements, condition findings, and strategy review dates.
The CMMS should hold the executable controls. It should not become the place where strategy logic is improvised without technical basis.
SAP PM Considerations
The MKS identifies SAP PM concepts that may support strategy execution, including equipment, functional locations, maintenance plans, maintenance items, task lists, strategy plans, counters, measurement points, notifications, maintenance orders, catalog profiles, work centers, planner groups, BOMs, permits, revisions, and status management.
Exact configuration varies by SAP release and site design.
Metrics and KPIs
Strategy Coverage
Definition: Portion of in-scope assets or failure modes with an approved maintenance strategy.
Formula:
Strategy Coverage = Assets or Failure Modes with Approved Strategy ÷ Total In-Scope Assets or Failure Modes × 100
Units: Percent
Interpretation: Indicates how much of the intended population has documented strategy logic.
Limitations: High coverage does not prove strategy quality.
Failure-Mode Coverage
Definition: Portion of significant identified failure modes with an assigned strategy.
Units: Percent
Interpretation: Indicates technical completeness.
Limitations: Depends on the quality of failure-mode identification.
PM Effectiveness
Definition: Measures whether PM tasks are detecting or preventing the failure conditions they were designed to address.
Units: Organization-defined
Interpretation: Supports PM optimization.
Limitations: Requires reliable work-order and finding data.
PdM Conversion Rate
Definition: Portion of valid condition findings converted into controlled corrective work.
Units: Percent
Interpretation: Tests whether condition monitoring produces action.
Limitations: A high rate is not necessarily desirable if alarm quality is poor.
Repeat Failure Rate
Definition: Portion of failures that repeat within the defined analysis boundary.
Units: Percent
Interpretation: Indicates whether strategy and corrective actions are eliminating recurring problems.
Limitations: Requires consistent failure coding.
Common Failure Modes
- Starting with PM tasks instead of functions and failure modes.
- Copying OEM checklists without validating operating context.
- Applying identical PM programs to all identical equipment.
- Using criticality to select task type instead of failure behavior.
- Treating run-to-failure as neglect rather than a deliberate decision.
- Never considering redesign.
- Using calendar PM on random failure modes.
- Setting intervals without technical rationale.
- Creating PdM routes with no corrective-response process.
- Never reviewing the strategy after failures.
- Updating strategy decisions without updating CMMS controls.
Best Practices
- Start with required function.
- Analyze significant failure modes.
- Use criticality to decide analysis depth.
- Select the least intrusive technically effective action.
- Document run-to-failure decisions.
- Escalate failure modes that maintenance cannot control to redesign.
- Document interval logic.
- Translate strategy into executable work.
- Review failures against the current strategy.
- Standardize similar asset classes only when operating context is comparable.
- Use technician and operator feedback.
- Keep strategy connected to PM, PdM, RCA, FMEA, and RCM.
Maturity Levels
| Level | Characteristics |
|---|---|
| 1 — Reactive | Maintenance response begins after failure |
| 2 — Calendar-Based | PMs exist but often lack failure-mode basis |
| 3 — Controlled | Strategy templates, criticality, and change control exist |
| 4 — Reliability-Based | PM, PdM, RCM, FMEA, and RCA are integrated |
| 5 — Optimized | Enterprise strategy libraries and continuous feedback are used |
Real-World Example
A process-water pump is required to deliver 450 gpm at 60 psi.
A reliability review identifies several credible failure modes: bearing lubrication contamination, mechanical seal wear, impeller erosion, and winding insulation breakdown.
The team does not assign one generic monthly PM to all four failure modes. Instead, bearing condition is monitored using condition-based methods; seal leakage is inspected and trended; impeller performance is monitored through process performance; motor condition is monitored with applicable electrical and condition techniques; low-consequence consumable failures are evaluated for run-to-failure; and recurring installation-related failures trigger redesign or precision-maintenance action.
The result is a strategy package tied to failure behavior rather than a calendar checklist.
Audit Questions
- Are required asset functions defined?
- Are significant functional failures identified?
- Are failure modes technically credible?
- Are consequences evaluated?
- Is asset criticality used appropriately?
- Are time-based tasks limited to age-related failure behavior?
- Are condition-based tasks supported by sufficient warning time?
- Are hidden functions covered by failure-finding?
- Are run-to-failure decisions explicit?
- Are redesign triggers defined?
- Are intervals supported by evidence?
- Are strategy outputs translated into job plans and CMMS controls?
- Are failures reviewed against the strategy?
- Are strategy changes traceable?
- Are similar assets standardized only when operating context is comparable?
Related Reliability Method Topics
Confirmed related topics:
- Asset Criticality Analysis
- Preventive Maintenance
- Condition Monitoring & Predictive Maintenance
- Maintenance Planning
- Maintenance Scheduling
- Failure Analysis & Failure Mechanisms
- Failure Modes and Effects Analysis
- Reliability-Centered Maintenance
- Root Cause Analysis
Related Tools and Assets
Potential future assets:
- Maintenance Strategy Development Worksheet
- Strategy Selection Decision Tree
- Run-to-Failure Decision Form
- PM/PdM Selection Matrix
- Maintenance Strategy Audit
- Strategy Library Template
- Strategy Review Checklist
These assets are not created by this document.
References
References are limited to the authoritative source and confirmed related Reliability Method sources:
- RM-MKS-6009 — Maintenance Strategy Development
- RM-MKS-6008 — Asset Criticality Analysis
- RM-MKS-6002 — Maintenance Planning
- RM-MKS-6003 — Maintenance Scheduling
Public derivative source: Draft. Publication status: Not Authorized.
