Source and Scope Boundary
This page is the public derivative of `RM-MKS-7002 — Failure Modes`, v1.2, Approved Internal. It owns physical failure mechanisms and engineering examination: fatigue, wear, corrosion, erosion, cavitation, lubrication, thermal, electrical, and human-induced degradation. Root Cause Analysis (RCA) owns investigation and causal reasoning; FMEA owns structured risk/action analysis. This page does not define failure-mode coding taxonomy or CMMS data structure — those remain owned by the underlying source and Failure Coding & Asset Taxonomy.
Definition
Failure Analysis is the systematic process of investigating how and why equipment failed by identifying the physical failure mechanism, contributing factors, and underlying causes.
A failure mechanism is the physical, chemical, electrical, thermal, or human process that causes a component to degrade until it can no longer perform its intended function.
The objective is to understand how failures develop so they can be prevented in the future.
Executive Summary
Equipment rarely fails without warning.
Failures usually begin with a defect that develops through one or more failure mechanisms until functional failure occurs.
Organizations that understand failure mechanisms can:
- Prevent recurring failures
- Improve maintenance strategies
- Improve equipment design
- Select better materials
- Improve reliability
- Reduce lifecycle costs
- Improve safety
- Increase asset availability
Why Failure Analysis Matters
Repairing failed equipment restores production.
Understanding why it failed prevents the next failure.
Failure Analysis transforms maintenance history into engineering knowledge.
What Failure Analysis Is
A complete Failure Analysis program includes:
- Failure investigation
- Physical examination
- Failure mechanism identification
- Material evaluation
- Operating condition review
- Maintenance history review
- Root Cause Analysis
- Engineering recommendations
- Knowledge management
What Failure Analysis Is Not
Failure Analysis is not:
- Replacing failed parts without investigation
- Assigning blame
- Guessing the cause of failure
- Limited to catastrophic failures
- An activity performed only by engineers
It is a structured engineering discipline supported by maintenance, operations, and technical specialists.
Objectives
An effective Failure Analysis program should:
- Identify physical failure mechanisms
- Prevent repeat failures
- Improve maintenance strategies
- Improve equipment design
- Reduce maintenance costs
- Improve asset reliability
- Support engineering decisions
- Capture organizational knowledge
- Drive continuous improvement
Failure Analysis Philosophy
Every equipment failure contains valuable engineering information.
Organizations that learn from failures become more reliable over time.
Core Components
A complete Failure Analysis program includes:
- Evidence preservation
- Failure mechanism identification
- Operating condition review
- Maintenance review
- Material evaluation
- Engineering assessment
- Corrective action development
- Verification
- Knowledge sharing
Relationship to the Maintenance Process
Failure Analysis supports:
- Reliability Engineering
- Root Cause Analysis
- Failure Modes and Effects Analysis
- Reliability-Centered Maintenance
- Defect Elimination
- Precision Maintenance
- Condition Monitoring
Inputs
Successful Failure Analysis depends on:
- Failed components
- CMMS history
- Inspection records
- Condition monitoring data
- Operating data
- Maintenance records
- Photographs
- Witness observations
Outputs
An effective program produces:
- Identified failure mechanisms
- Engineering recommendations
- Improved maintenance strategies
- Updated standards
- Reduced repeat failures
- Reliability improvements
- Organizational learning
Fatigue Failure
Fatigue is progressive cracking caused by repeated cyclic loading below the material's ultimate strength.
Common causes include:
- Repeated stress cycles
- Misalignment
- Vibration
- Stress concentrations
- Improper design
Indicators include beach marks, crack propagation, and sudden final fracture.
Wear Mechanisms
Wear removes material from contacting surfaces over time.
Common wear types include:
- Abrasive wear
- Adhesive wear
- Erosive wear
- Fretting wear
- Surface fatigue
Proper lubrication, alignment, filtration, and material selection reduce wear.
Corrosion
Corrosion is the deterioration of materials through chemical or electrochemical reactions.
Common forms include:
- Uniform corrosion
- Galvanic corrosion
- Pitting
- Crevice corrosion
- Stress corrosion cracking
Control methods include coatings, material selection, cathodic protection, and environmental control.
Erosion
Erosion results from high-velocity fluids or particles removing material from a surface.
Typical applications include:
- Pumps
- Valves
- Piping
- Elbows
- Fans
Velocity control and improved materials reduce erosion.
Cavitation
Cavitation occurs when vapor bubbles form and collapse within a liquid.
Common symptoms include:
- Pitting
- Noise
- Vibration
- Reduced pump performance
- Impeller damage
Prevent cavitation through proper suction conditions and pump selection.
Lubrication Failures
Poor lubrication is a leading cause of bearing and gearbox failures.
Common causes include:
- Incorrect lubricant
- Over-lubrication
- Under-lubrication
- Contamination
- Improper intervals
Lubrication programs should be standardized and verified.
Thermal Damage
Excessive temperature accelerates material degradation.
Examples include:
- Insulation breakdown
- Bearing overheating
- Lubricant oxidation
- Seal degradation
- Thermal distortion
Trend temperature changes using condition monitoring.
Electrical Failure Mechanisms
Electrical failures may result from:
- Insulation deterioration
- Voltage imbalance
- Loose connections
- Harmonics
- Partial discharge
- Overheating
Routine testing can detect many electrical defects before failure.
Human-Induced Failures
Many failures originate from maintenance or operational practices.
Examples include:
- Improper installation
- Incorrect torque
- Misalignment
- Operating outside design limits
- Poor troubleshooting
Standard work and competency reduce human-induced failures.
Failure Progression
Failures typically progress through stages:
- Defect introduction
- Early degradation
- Detectable condition
- Functional failure
- Secondary damage
Understanding progression improves inspection timing and maintenance strategy.
Failure Analysis Governance
Failure Analysis should operate under documented governance with standardized investigation methods, defined responsibilities, and measurable objectives.
Governance should establish:
- Program ownership
- Investigation criteria
- Evidence handling requirements
- Documentation standards
- Review frequency
- Approval authority
- Continuous improvement expectations
Cross-Functional Investigation Teams
Effective Failure Analysis requires participation from:
- Reliability Engineering
- Maintenance
- Operations
- Engineering
- Production
- Quality
- Safety
- OEMs or subject matter experts, when needed
Collaboration improves the accuracy of engineering conclusions.
Evidence Preservation
Preserve evidence before repairs begin whenever practical.
Best practices include:
- Photograph components
- Record equipment position
- Tag failed parts
- Prevent contamination
- Document operating conditions
- Preserve failed components for examination
Lost evidence often results in incorrect conclusions.
Component Examination
Examine failed components for:
- Crack patterns
- Wear surfaces
- Corrosion products
- Heat discoloration
- Lubrication condition
- Fracture characteristics
- Manufacturing defects
Physical evidence should support every conclusion.
Engineering Assessment
Engineering assessment should determine:
- Primary failure mechanism
- Contributing factors
- Failure progression
- Preventability
- Design improvements
- Maintenance improvements
- Operational improvements
Knowledge Management
Capture lessons from every significant investigation.
Document:
- Failure mechanism
- Root causes
- Corrective actions
- Photographs
- Engineering recommendations
- Standards updated
- Lessons learned
Performance Measurement
Evaluate the program using:
- Repeat failure reduction
- Failure investigations completed
- Engineering recommendations implemented
- Corrective action effectiveness
- Investigation cycle time
- Lessons learned published
- Business value delivered
- Reliability improvement achieved
Case Study
The following is an illustrative composite drawn from common patterns across maintenance organizations, not a specific documented case.
A centrifugal pump experienced repeated shaft failures.
Failure analysis identified:
- Pipe strain
- Misalignment
- Cyclic fatigue
- Inadequate installation practices
Engineering corrected the piping, updated alignment procedures, and revised installation standards.
The team would define a controlled event, population, and exposure basis and track repeat-failure rate over a defined observation window before attributing any measured MTBF change to the correction. Any claimed result would require that controlled before-and-after evidence and disclosure of other changes affecting performance. No specific outcome is claimed here.
Continuous Improvement
Improve the program through:
- Reliability reviews
- RCA integration
- Standards updates
- Technician training
- AI-assisted diagnostics
- Failure mechanism library expansion
- Benchmarking
Knowledge Graph Updates
Future Knowledge Library topics introduced:
- Failure Mechanisms
- Material Failure Analysis
- Physical Evidence Examination
- Engineering Failure Assessment
- Failure Progression
- Failure Knowledge Base
- Engineering Recommendations
- Failure Classification
- Fatigue Analysis
- Wear Mechanisms
- Corrosion Mechanisms
- Cavitation Analysis
- Thermal Damage
- Electrical Failure Mechanisms
- Human-Induced Failures
- Failure Progression Models
- Failure Evidence Preservation
- Component Examination
- Failure Investigation Workflow
- Engineering Lessons Learned
- Failure Analysis KPIs
- Continuous Failure Improvement
Industry Applications
Food Manufacturing
Failure Analysis should prioritize:
- Food safety critical equipment
- Refrigeration systems
- Packaging equipment
- Utilities
- CIP systems
- High-speed rotating assets
Distribution and Warehousing
Priority applications include:
- Conveyor systems
- Gearboxes
- Motors
- Sortation equipment
- Dock equipment
- Automated storage systems
Municipal Utilities
Focus on:
- Pumps
- Motors
- Blowers
- Valves
- Electrical distribution
- Backup generators
Commercial Facilities
Typical applications include:
- HVAC systems
- Chillers
- Boilers
- Cooling towers
- Emergency generators
- Building automation systems
Small Manufacturing
Prioritize:
- Chronic equipment failures
- Rotating equipment
- Utility systems
- Air compressors
- Production bottlenecks
Failure Analysis for Small Business Owners
Small organizations should:
- Preserve failed parts.
- Document failures with photographs.
- Identify recurring failure patterns.
- Correct underlying failure mechanisms.
- Maintain a library of lessons learned.
Every investigated failure improves future reliability.
Failure Analysis Maturity Model
Level 1 — Reactive
- Replace failed parts
- Minimal investigation
- Repeat failures common
Level 2 — Developing
- Basic failure investigations
- Failure documentation
- Initial engineering recommendations
Level 3 — Managed
- Formal Failure Analysis program
- Standard investigation methods
- Failure mechanism library
- Cross-functional reviews
- Performance measurement
Level 4 — Optimized
- Enterprise Failure Analysis program
- AI-assisted diagnostics
- Digital failure knowledge base
- Predictive engineering insights
- Reliability-centered culture
Failure Analysis KPIs
| KPI | Formula or definition | Interpretation limit |
|---|---|---|
| Reviewed Classification Coverage | Applicable failure events with completed qualified classification review / applicable failure events due for review × 100 | Review threshold must be explicit; completion does not prove correctness. |
| Repeat-Mode Event Rate | Qualifying repeat-mode events / defined operating exposure for the same population | Requires stable population, event, restoration, and exposure definitions. |
| Classification Correction Rate | Reviewed events whose classification changed after evidence review / reviewed events × 100 | A change can reflect learning, not original negligence; segment by reason. |
| Evidence Completeness | Reviewed events meeting the applicable evidence rule / reviewed events sampled × 100 | Evidence rules vary by consequence and investigation method. |
Example: if 18 of 20 applicable failure events due for review have a completed qualified classification, reviewed classification coverage is `18 / 20 × 100 = 90.0%`. Completion does not prove the classification is correct — pair this measure with a classification correction rate over time. Numeric targets are local governance decisions, not universal benchmarks.
Common Mistakes
Organizations frequently:
- Discard failed components too early.
- Jump to conclusions without evidence.
- Ignore physical failure mechanisms.
- Fail to document findings.
- Skip engineering review.
- Never verify corrective actions.
- Repeat the same investigations.
- Separate failure analysis from reliability improvement.
Best Practices
- Preserve evidence.
- Base conclusions on facts.
- Identify physical failure mechanisms.
- Integrate RCA with Failure Analysis.
- Maintain a searchable knowledge base.
- Verify engineering improvements.
- Share lessons learned.
- Continuously refine engineering standards.
Reliability Method Standards Roadmap
Future Reliability Method standards proposed to support this topic (not yet published):
- RM-FA-001 — Failure Analysis Standard
- RM-FA-002 — Failure Mechanism Classification Standard
- RM-FA-003 — Evidence Preservation Standard
- RM-FA-004 — Component Examination Standard
- RM-FA-005 — Engineering Assessment Standard
- RM-FA-006 — Failure Analysis Governance Standard
- RM-FA-007 — Fatigue Failure Standard
- RM-FA-008 — Wear Mechanisms Standard
- RM-FA-009 — Corrosion Analysis Standard
- RM-FA-010 — Lubrication Failure Standard
- RM-FA-011 — Failure Progression Standard
- RM-FA-012 — Evidence Preservation Standard
- RM-FA-013 — Engineering Failure Assessment Standard
- RM-FA-014 — Failure Knowledge Management Standard
- RM-FA-015 — Failure Analysis Performance Standard
- RM-FA-016 — Failure Mechanism Library Standard
- RM-FA-017 — Enterprise Failure Analysis Program Standard
- RM-FA-018 — Failure Knowledge Management Standard
- RM-FA-019 — Failure Signature Atlas Standard
- RM-FA-020 — Failure Analysis Excellence Framework
Product Opportunities
The items below are potential future product ideas for roadmap and planning purposes. They are not existing Reliability Method products, features, or services.
Templates
- Failure Analysis Report
- Evidence Collection Checklist
- Component Examination Worksheet
- Failure Mechanism Library
- Engineering Recommendation Register
- Lessons Learned Report
Calculators
- Failure Cost Calculator
- MTBF Improvement Calculator
- Downtime Cost Calculator
- Reliability Improvement ROI Calculator
AI Tools
- Failure Analysis Assistant
- Failure Mechanism Identifier
- Component Inspection Advisor
- Engineering Recommendation Generator
- Failure Knowledge Search
Facility Manager Features
- Failure Knowledge Base
- Failure Mechanism Library
- Evidence Repository
- Engineering Recommendation Tracker
- Reliability Analytics Dashboard
- Reliability Method Standards Library
Related Knowledge Topics
- Reliability Engineering
- Root Cause Analysis
- Failure Modes and Effects Analysis
- Defect Elimination
- Precision Maintenance
- Condition Monitoring & Predictive Maintenance
References
- SMRP Body of Knowledge
- ISO 55000 — Asset Management
- ASM Handbook, Volume 11: Failure Analysis and Prevention
- Reliability Method Internal Standards
Revision History
| Version | Date | Change |
|---|---|---|
| 1.0 | 2026-07-27 | Initial Failure Analysis & Failure Mechanisms foundation created; failure mechanisms, governance, engineering assessment, industry applications, maturity model, KPIs, Reliability Method Standards roadmap, product opportunities, references, and revision history completed. |
| 1.1 | 2026-08-03 | Reconciled to approved `RM-MKS-7002`; added source boundary, controlled KPI formulas and worked example, hedged the illustrative case study to remove an unsupported MTBF claim, and populated the source reference register for publication review. |