RELIABILITYMETHOD

Reliability Engineering

Root Cause Analysis (RCA)

Root Cause Analysis (RCA) is a structured process used to identify the underlying causes of a problem so that the problem can be permanently prevented from recurring.

Status: PublishedDifficulty: BeginnerUpdated: 2026-07-24

Publication Status

This page is the owner-authorized public-release preparation record for `KN-4002`.

This article is published as part of the Reliability Method Knowledge Library at `https://reliabilitymethod.com/knowledge/root-cause-analysis-rca`.

The templates, calculators, AI tools, Facility Manager features, dashboards, reports, SOPs, training assets, and consulting offers named below are planned opportunities only. They are not currently available product assets.

Plain-English Definition

Root Cause Analysis (RCA) is a structured process used to identify the underlying causes of a problem so that the problem can be permanently prevented from recurring.

Rather than asking, "Who made the mistake?" RCA asks, "Why did the failure occur, and what changes will prevent it from happening again?"

The objective is permanent improvement—not simply restoring equipment to operation.


Executive Summary

Every maintenance organization experiences failures.

Reactive organizations repair the equipment and return it to service.

Reliable organizations investigate why the failure occurred and eliminate the conditions that allowed it to happen.

Root Cause Analysis is one of the most valuable reliability improvement tools because it transforms failures into learning opportunities.

RCA supports:

  • Improved reliability
  • Reduced repeat failures
  • Better maintenance strategies
  • Improved safety
  • Lower maintenance costs
  • Increased equipment availability
  • Continuous improvement

When consistently applied, RCA changes maintenance from reacting to failures into preventing them.


Why Root Cause Analysis Matters

Most failures are symptoms of deeper problems.

Replacing a failed bearing may restore production, but it does not answer:

  • Why did the bearing fail?
  • Why did it fail earlier than expected?
  • Why wasn't the defect detected?
  • Why has it happened before?
  • What system allowed the failure to occur?

Without answering these questions, organizations often repeat the same repairs while expecting different results.

RCA breaks that cycle.


What Root Cause Analysis Is

Root Cause Analysis is a systematic investigation used to identify and eliminate the underlying causes of undesirable events.

It can be applied to:

  • Equipment failures
  • Safety incidents
  • Quality defects
  • Environmental events
  • Repeat maintenance work
  • Production interruptions
  • Process failures
  • Human errors
  • Business process failures

Although commonly associated with maintenance, RCA is an organizational improvement methodology.


What Root Cause Analysis Is Not

Root Cause Analysis is not:

  • Finding someone to blame
  • Completing paperwork after a breakdown
  • Guessing at the cause
  • Replacing failed parts
  • Writing a work order
  • A warranty investigation
  • A maintenance report

An RCA is successful only when corrective actions reduce the likelihood of recurrence.


Objectives of Root Cause Analysis

An effective RCA should:

  • Identify the true underlying causes
  • Eliminate repeat failures
  • Reduce business risk
  • Improve maintenance strategies
  • Improve equipment reliability
  • Improve operating procedures
  • Improve training
  • Improve asset design
  • Capture organizational learning

The Root Cause Philosophy

Every significant failure has multiple contributing causes.

Most failures are not caused by a single event.

Instead they develop through a chain of technical, human, organizational, and process-related factors.

Effective RCA investigates the entire chain—not just the final event.


Types of Causes

Reliability organizations generally classify causes into three levels.

Physical Causes

The physical reason the failure occurred.

Examples:

  • Bearing fatigue
  • Broken shaft
  • Seal leakage
  • Loose electrical connection
  • Corroded piping

Human Causes

Actions or decisions that contributed.

Examples:

  • Improper installation
  • Incorrect lubrication
  • Missed inspection
  • Incorrect operating practice

Latent (System) Causes

Weaknesses in the management system that allowed the failure.

Examples:

  • Missing procedures
  • Inadequate training
  • Poor planning
  • Weak PM program
  • Inadequate spare parts
  • Poor design standards

Long-term improvement usually comes from correcting latent causes.


When to Perform RCA

Not every failure requires a formal investigation.

RCA should normally be performed when:

  • Safety was affected
  • Environmental impact occurred
  • Production losses were significant
  • A critical asset failed
  • The failure repeats
  • Repair costs were unusually high
  • Customer service was affected
  • Leadership requests investigation

Organizations should define clear RCA trigger criteria.


Relationship to Reliability Engineering

Root Cause Analysis is a core discipline within Reliability Engineering.

Reliability Engineering identifies opportunities for improvement.

RCA determines what improvements should be made.

The findings from RCA frequently result in:

  • PM optimization
  • New PdM inspections
  • Equipment redesign
  • Procedure revisions
  • Training improvements
  • Capital projects

Inputs

Typical RCA inputs include:

  • Work orders
  • Equipment history
  • Failure codes
  • Maintenance records
  • Operator interviews
  • Photographs
  • Failed components
  • OEM documentation
  • PdM reports
  • Process data
  • Inspection records

Outputs

A completed RCA should produce:

  • Problem statement
  • Timeline of events
  • Root causes
  • Contributing causes
  • Corrective actions
  • Preventive actions
  • Action owners
  • Due dates
  • Lessons learned
  • Updated standards

The Root Cause Analysis Process

A successful RCA follows a repeatable process rather than relying on intuition.

Recommended process:

  1. Define the problem.
  2. Secure evidence.
  3. Assemble the investigation team.
  4. Develop an event timeline.
  5. Identify failure modes.
  6. Identify contributing causes.
  7. Determine root causes.
  8. Develop corrective actions.
  9. Implement improvements.
  10. Verify effectiveness.
  11. Capture lessons learned.

The goal is not to complete a report. The goal is to prevent recurrence.


Defining the Problem

A clear problem statement establishes the scope of the investigation.

An effective problem statement answers:

  • What happened?
  • When did it happen?
  • Where did it happen?
  • Which asset or process was involved?
  • What business impact occurred?

Avoid assumptions or assigning blame.

Poor example:

"The mechanic caused the failure."

Better example:

"Packaging Line 3 stopped unexpectedly when the drive motor seized, resulting in four hours of production downtime."


Collecting Evidence

Root Cause Analysis should be based on evidence rather than opinions.

Sources include:

  • Failed components
  • Work orders
  • CMMS history
  • Operator interviews
  • Technician interviews
  • Photographs
  • PdM reports
  • Process data
  • Alarm history
  • Maintenance procedures
  • OEM documentation

Evidence should be collected as soon as possible before conditions change.


Building the Investigation Team

The investigation team should include people who understand both the equipment and the process.

Typical participants include:

  • Reliability Engineer
  • Maintenance Supervisor
  • Technician
  • Operator
  • Planner
  • Production Representative
  • Engineer
  • Safety Representative (when appropriate)

Cross-functional participation improves both technical accuracy and acceptance of corrective actions.


Developing an Event Timeline

A timeline reconstructs the sequence of events leading to the failure.

Typical timeline elements include:

  • Normal operation
  • First abnormal condition
  • Alarm or warning
  • Operator response
  • Equipment shutdown
  • Maintenance response
  • Repair activities
  • Restart

Timelines frequently reveal overlooked contributing factors.


The 5 Whys Method

The 5 Whys is one of the simplest RCA techniques.

It repeatedly asks "Why?" until the underlying cause is identified.

Example:

Problem: Motor failed.

Why? Bearing failed.

Why? Bearing received insufficient lubrication.

Why? Lubrication interval was missed.

Why? PM was not scheduled.

Why? Asset was omitted from the CMMS PM program.

The investigation identifies a system weakness rather than stopping at the failed bearing.


Fishbone (Cause-and-Effect) Diagram

A Fishbone Diagram organizes possible causes into logical categories.

Common categories include:

  • People
  • Methods
  • Machines
  • Materials
  • Measurement
  • Environment

This technique helps teams consider multiple contributing factors instead of focusing on one suspected cause.


Cause Mapping

Cause Mapping visually connects causes and effects.

Example:

Production stopped

Motor seized

Bearing failed

Lubrication absent

PM task omitted

Asset missing from PM library

Cause Mapping is especially useful for complex investigations involving multiple interacting events.


Barrier Analysis

Barrier Analysis identifies controls that should have prevented the failure.

Questions include:

  • What barriers existed?
  • Which barriers failed?
  • Which barriers were missing?
  • Which barriers should be added?

Examples of barriers:

  • PM inspections
  • PdM monitoring
  • Operator inspections
  • Safety devices
  • Procedures
  • Training
  • Engineering controls

Identifying Contributing Causes

Contributing causes increase the likelihood or severity of failure but may not be the primary root cause.

Examples include:

  • Delayed repairs
  • Poor communication
  • Missing documentation
  • Inadequate supervision
  • Incorrect spare parts
  • Weak planning

Multiple contributing causes often exist.


Determining Root Causes

Root causes are the underlying conditions that allowed the failure to occur.

A true root cause should be:

  • Supported by evidence
  • Correctable
  • Capable of preventing recurrence when addressed

If eliminating the cause would likely prevent similar failures, it is probably a true root cause.


Verifying Findings

Before finalizing an RCA, verify that conclusions are supported by facts.

Review:

  • Physical evidence
  • CMMS history
  • Interviews
  • Measurements
  • Inspection reports
  • Operating data

Avoid speculation whenever possible.


Developing Corrective Actions

The primary output of an RCA is not the report.

It is the corrective actions that prevent recurrence.

Effective corrective actions should:

  • Address the verified root cause
  • Be practical to implement
  • Reduce risk
  • Be assigned to an owner
  • Include a completion date
  • Be measurable

Replacing a failed component is usually a repair—not a corrective action.


Corrective vs. Preventive Actions

Corrective actions eliminate verified causes of an observed failure.

Examples:

  • Revise a PM procedure
  • Install contamination controls
  • Update alignment standards
  • Redesign a component

Preventive actions reduce the likelihood of similar failures elsewhere before they occur.

Examples:

  • Update maintenance standards across all facilities
  • Revise technician training
  • Modify design specifications
  • Improve spare parts standards

Both should be considered during every RCA.


Verifying Effectiveness

Corrective actions are only successful if they produce lasting improvement.

Verification methods include:

  • Monitoring repeat failures
  • Reviewing MTBF
  • Inspecting completed work
  • Auditing PM changes
  • Reviewing PdM findings
  • Tracking downtime trends

If failures continue, the investigation should be reopened.


RCA Integration with Preventive Maintenance

Many RCAs identify weaknesses in existing PM programs.

Typical improvements include:

  • Adding inspections
  • Removing ineffective PM tasks
  • Changing frequencies
  • Revising procedures
  • Adding acceptance criteria
  • Improving lubrication practices

PMs should evolve as new failure knowledge is gained.


RCA Integration with Predictive Maintenance

RCA frequently identifies failure modes that could have been detected earlier.

Possible improvements include:

  • Add vibration monitoring
  • Add infrared inspections
  • Expand oil analysis
  • Introduce ultrasound routes
  • Increase inspection frequency

PdM should target failure modes identified through RCA.


RCA Integration with Work Management

Corrective actions should flow directly into the maintenance work process.

Examples include:

  • New work orders
  • PM revisions
  • Job plan updates
  • Planner work requests
  • Capital project requests
  • Engineering change requests

Recommendations without execution create no value.


RCA Integration with CMMS

The CMMS should capture information that supports future investigations.

Recommended fields include:

  • Failure Mode
  • Problem Code
  • Cause Code
  • Remedy Code
  • Component
  • Downtime
  • Corrective Action Reference
  • RCA Number

Standardized data improves future reliability analysis.


Organizational Learning

Every completed RCA should improve organizational knowledge.

Lessons learned should be:

  • Documented
  • Shared
  • Incorporated into standards
  • Included in technician training
  • Used during future investigations

Knowledge that remains inside a report has little long-term value.


RCA Audits

Organizations should periodically audit completed RCAs.

Verify:

  • Trigger criteria were followed.
  • Evidence supports conclusions.
  • Root causes were validated.
  • Corrective actions were implemented.
  • Effectiveness was verified.
  • Lessons learned were documented.
  • Standards were updated where appropriate.

Case Study

The following is an illustrative composite drawn from common patterns across maintenance organizations, not a specific documented case.

A wastewater treatment facility experienced repeated failures of a sludge transfer pump.

Initial repairs focused on replacing bearings and seals.

An RCA identified:

  • Chronic dry running
  • Poor level control logic
  • Inadequate operator alarms
  • PM tasks that never inspected suction conditions

Corrective actions included:

  • Revising PLC logic
  • Installing additional level alarms
  • Updating PM inspections
  • Training operators

Results after one year:

  • Pump failures reduced dramatically.
  • Emergency work declined.
  • Equipment availability improved.
  • Maintenance costs decreased.

The greatest improvement came from correcting the operating process rather than changing hardware.


Continuous Improvement

Root Cause Analysis should become a continuous improvement process rather than an occasional investigation.

Organizations should routinely review:

  • Repeat failures
  • Significant downtime events
  • Safety incidents
  • High-cost repairs
  • Bad actor assets
  • Warranty claims

Every investigation should strengthen maintenance standards and improve organizational knowledge.


These related concepts may become separate Knowledge Library records or supporting resources later. They are listed as conceptual extensions only, not as claims that public pages or tools currently exist:

  • 5 Whys
  • Fishbone Diagram
  • Cause Mapping
  • Barrier Analysis
  • Event Timeline
  • Corrective Actions
  • Preventive Actions
  • Lessons Learned
  • Defect Elimination
  • RCA Facilitation
  • Event Timeline Analysis
  • Evidence Collection
  • Investigation Facilitation
  • Evidence Validation
  • Corrective Action Development
  • Corrective Action Management
  • Preventive Action Programs
  • RCA Auditing
  • Lessons Learned Programs
  • Engineering Change Management
  • Maintenance Standards
  • Organizational Learning
  • CAPA Management
  • Reliability Governance

Industry Applications

Food Manufacturing

Root Cause Analysis helps food manufacturers reduce unplanned downtime while improving food safety, quality, and regulatory compliance.

Common RCA investigations include:

  • Product contamination
  • Packaging failures
  • Refrigeration failures
  • Conveyor breakdowns
  • Utility interruptions
  • Sanitation-related failures

Corrective actions should improve both equipment reliability and manufacturing processes.


Distribution and Warehousing

Typical investigations include:

  • Conveyor failures
  • Sortation issues
  • Dock equipment failures
  • Forklift reliability
  • Battery charging problems

The objective is reducing shipping delays and improving operational continuity.


Municipal Utilities

Utilities commonly apply RCA to:

  • Pump failures
  • Lift station outages
  • Blower failures
  • Electrical faults
  • Chemical feed interruptions
  • Regulatory events

Reliable public service depends upon preventing recurrence rather than repeatedly repairing equipment.


Commercial Facilities

Commercial facilities use RCA to investigate:

  • HVAC failures
  • Boiler outages
  • Chiller problems
  • Elevator incidents
  • Fire protection impairments
  • Building automation failures

Small Manufacturing

Small manufacturers should prioritize RCA for failures that repeatedly affect production, customer deliveries, or maintenance costs.


Root Cause Analysis for Small Business Owners

Formal RCA software is rarely necessary for small organizations.

A structured discussion using questions such as:

  • What happened?
  • Why did it happen?
  • What allowed it to happen?
  • What should change?
  • How do we prevent it next time?

can dramatically improve long-term reliability.


RCA Maturity Model

Level 1 — Reactive

  • Equipment repaired after failure
  • No formal investigations

Level 2 — Developing

  • Occasional 5 Whys
  • Limited documentation

Level 3 — Managed

  • Standard RCA process
  • Cross-functional investigations
  • Corrective actions tracked
  • Lessons learned documented

Level 4 — Optimized

  • RCA integrated into Reliability Engineering
  • Enterprise standards
  • Trend analysis
  • Continuous organizational learning
  • Preventive improvements across multiple assets

Root Cause Analysis KPIs

Recommended metrics include:

  • Repeat Failure Rate
  • RCAs Completed
  • Corrective Action Completion Rate
  • Corrective Action Overdue Percentage
  • Repeat Failures After RCA
  • Mean Time Between Repeat Failures
  • Bad Actor Reduction
  • Downtime Avoided
  • Maintenance Cost Avoidance
  • Lessons Learned Implemented

Metrics should evaluate the effectiveness of improvements rather than the number of investigations performed.


Common Mistakes

Organizations frequently:

  • Stop at the first apparent cause.
  • Assign blame instead of identifying system weaknesses.
  • Skip evidence collection.
  • Ignore operator input.
  • Fail to verify corrective actions.
  • Never update maintenance standards.
  • Close RCAs before actions are complete.
  • Repeat the same investigations without organizational learning.

Best Practices

  • Define clear RCA trigger criteria.
  • Base conclusions on evidence.
  • Include cross-functional teams.
  • Address physical, human, and latent causes.
  • Assign accountable action owners.
  • Verify effectiveness after implementation.
  • Update PMs, PdM routes, and standards.
  • Share lessons learned across the organization.
  • Review recurring failures regularly.
  • Treat every significant failure as an opportunity to improve.

Fault Tree Analysis

Fault Tree Analysis works backward from the failure event using logical relationships between contributing events.

Applications include:

  • High-risk systems
  • Safety investigations
  • Reliability studies
  • Complex equipment failures with multiple interacting causes

Fault trees help quantify failure pathways and are most useful when the 5 Whys or Fishbone diagram alone cannot capture how multiple conditions combined to produce the failure.


RCA Governance

Root Cause Analysis should operate under documented governance with defined responsibilities and consistent investigation standards.

Governance should establish:

  • RCA ownership
  • Investigation criteria
  • Facilitation responsibilities
  • Approval authority
  • Documentation standards
  • Review cadence
  • Continuous improvement expectations

Governance connects RCA Prioritization (see When to Perform RCA above) to a consistent, repeatable investigation standard rather than an ad hoc response to each failure.


RCA Facilitation

An effective facilitator should:

  • Remain objective
  • Encourage participation
  • Challenge assumptions
  • Keep discussions evidence-based
  • Document findings
  • Drive consensus
  • Focus on system improvements

The facilitator manages the process, not the outcome. This role is distinct from the investigation team assembled in Building the Investigation Team above — the facilitator guides the method; the team supplies the subject-matter knowledge.


Potential Future Resource Concepts

The items below are potential future resource ideas for roadmap and planning purposes. They are not existing Reliability Method products, features, or services.

Templates

  • RCA Report Template
  • 5 Whys Worksheet
  • Fishbone Diagram Template
  • Event Timeline Worksheet
  • Corrective Action Tracker
  • Lessons Learned Register

Calculators

  • Downtime Cost Calculator
  • Cost Avoidance Calculator
  • Repeat Failure Calculator
  • Corrective Action Completion Dashboard

Potential Future AI Tool Concepts

  • RCA Assistant
  • 5 Whys Generator
  • Cause Mapping Assistant
  • Corrective Action Advisor
  • Lessons Learned Generator

Potential Future Facility Manager Concepts

  • RCA Module
  • Corrective Action Tracking
  • Failure Timeline
  • Cause Code Analytics
  • Lessons Learned Library
  • Reliability Dashboard

Training

  • Root Cause Analysis Fundamentals
  • RCA Facilitation
  • 5 Whys Workshop
  • Fishbone Analysis
  • Evidence Collection
  • Corrective Action Development

Consulting

  • RCA Facilitation
  • Reliability Assessments
  • Repeat Failure Elimination
  • Maintenance Strategy Improvement
  • Reliability Program Development

  • Reliability Engineering
  • Failure Modes
  • Preventive Maintenance
  • Predictive Maintenance
  • FMEA
  • Reliability-Centered Maintenance
  • Asset Criticality Analysis
  • Maintenance Planning
  • Work Order Management
  • CMMS Fundamentals

References

  • SMRP Body of Knowledge
  • ISO 55000 — Asset Management
  • ISO 14224 — Reliability and Maintenance Data
  • SAE JA1011
  • SAE JA1012
  • Apollo Root Cause Analysis Methodology
  • OEM Maintenance Documentation
  • Reliability Method Internal Standards

Revision History

Version 1.0 Initial Root Cause Analysis foundation created.

Version 1.1 Expanded investigation methods, evidence collection, and implementation guidance.

Version 1.2 Completed industry guidance, maturity model, KPIs, product alignment, references, and revision history.

Version 1.3 Merged unique content (Fault Tree Analysis, RCA Governance, RCA Facilitation) from the retired duplicate record `root-cause-analysis-rca-kn-8002` into the main article flow; removed redundant overlapping subsections.

Version 1.4 Removed internal merge notes from article flow and reframed knowledge-graph/product-roadmap language for public-release preparation.