Resources/Agentic AI for Root Cause Analysis: From 5-Why Whiteboards to Autonomous Investigation
Sustainability & Trends

Agentic AI for Root Cause Analysis: From 5-Why Whiteboards to Autonomous Investigation

Agentic AI can search CMMS history, sensor trends, procurement records, and operator notes to propose ranked causal hypotheses for engineers to test. Here is how the architecture works and how to pilot it safely.

12 min read
By Monitory team

A critical gearbox fails. A cross-functional team gathers, fills a whiteboard with fishbone branches, and agrees the cause was lubrication. The greasing interval changes. A few weeks later, the replacement fails the same way. A second investigation, with different people asking different questions, finds that a bearing supplier was substituted during a procurement consolidation months earlier, and the new part did not suit that duty.

Stories like this are familiar to reliability engineers because the first investigation was not careless. It was limited. The team could only review the data it could pull together in a few hours, and the first plausible explanation shaped every question after it. The supplier change sat in a purchasing system nobody thought to open.

Agentic AI does not replace that team. It changes what the team starts with: a broad, systematic search across maintenance, sensor, procurement, and operating records, returned as ranked hypotheses with the evidence behind each. This article explains how that works, where it helps most, and how to pilot it without handing judgment to a model.

Why whiteboard RCA hits a ceiling

Structured methods such as 5-Why, fishbone diagrams, and fault tree analysis are sound. Their limits come from how they are run:

  • Data breadth. In a single session, a team can review a limited number of work orders, a few trend plots, and a handful of interviews. The relevant record may be in a system nobody pulled.
  • Anchoring. Once the room settles on a likely cause, later questions tend to confirm it.
  • Facilitator dependence. Results vary with who leads the session and which branches they push.
  • Time. Gathering data from separate systems takes longer than analyzing it.

The cost of getting RCA wrong is a repeat failure, and repeat failures are expensive. NIST's study of U.S. manufacturers found that establishments relying most heavily on reactive maintenance were associated with 3.3 times more downtime and 16.0 times more defects than others [1]. PNNL's O&M guidance notes that run-to-failure events can also damage secondary devices, adding costs that can be significant [2].

MethodData usually reviewedMain bias riskBest use
5-WhyInterviews and recent work ordersAnchoring on the first branchSimple, single-cause events
FishboneInterviews, some trends, work ordersCategory-driven thinkingBrainstorming candidate causes
Fault tree analysisDrawings, work history, sensor dataEffort limits how far it is appliedKnown failure modes, safety-critical systems
Agentic RCAAll connected sources, searched systematicallyData quality and model errorsGenerating and ranking hypotheses for engineers to test

How multi-agent RCA works

The design splits a large investigation into smaller searches, each handled by an agent with access to one kind of data, coordinated by an orchestrator.

The orchestrator

The orchestrator receives the failure event (asset, time, symptom), defines the time window to search, assigns work to specialist agents, and merges what they return. It also enforces limits: which systems agents may read, what they may write, and when a human must review.

Specialist agents

AgentSearchesTypical finding
CMMS historianWork orders on the asset and its neighboursRepeat repairs, recent maintenance on related equipment
Condition dataVibration, temperature, current, oil analysisWhen degradation started and how fast it progressed
Process dataHistorian tags for load, speed, temperature, productOperating changes that preceded the failure
ProcurementPurchase orders, part substitutions, supplier changesA different part or supplier installed before the failure
Operator notesShift logs and free-text commentsObservations nobody coded into the CMMS

Hypothesis merging

Each agent returns candidate contributing factors with timestamps and evidence. The orchestrator lines them up on a timeline, looks for factors that several sources support, and ranks hypotheses by how well the evidence fits. Each hypothesis carries a confidence score and a suggested verification step, such as inspecting a removed part, checking a supplier specification, or pulling a specific trend.

From correlation spaghetti to ranked hypotheses

The output should be short and testable. A useful hypothesis report for an engineer looks like this:

text
Event:        Gearbox GB-201 seized, press section
Hypothesis 1: Bearing substitution (supplier change on prior PO), confidence high
  Evidence:   Different part number installed at last rebuild; bearing temperature
              trend rose after that rebuild; no matching change on sister gearbox
  Verify:     Compare installed bearing cage material with OEM specification
Hypothesis 2: Lubrication interval, confidence medium
  Evidence:   Two greasing PMs completed late in the prior quarter
  Verify:     Oil and grease sample from failed unit

Engineers then do what they have always done: test the hypotheses, decide the cause, and approve corrective actions. The difference is that the supplier change is on the list on day one instead of after the second failure.

Keep the human decision explicit

Agentic RCA should propose and evidence, not conclude. Require an engineer to accept, reject, or modify each hypothesis and record why. Those decisions are also the best training data you will ever collect for improving the system.

The failure chains humans tend to miss

Slow drift

A gradual decline across a bank of similar assets may never trip an alarm on any single one. Searching trends across the fleet over long windows surfaces drift that a single-asset review misses.

Procurement-induced failures

Part substitutions, supplier changes, and specification changes live in purchasing systems that RCA teams rarely open. An agent that checks what was actually installed against what was specified can catch them.

Maintenance-induced failures

Faults introduced by recent work, such as misalignment after a rebuild or the wrong lubricant, show up as changes that start right after a work order. Lining up condition data against work order completion times makes this pattern visible.

PNNL's KPI guidance suggests tracking rework below 3% of work orders, trending downward [3]. Maintenance-induced failures show up in that number, and agentic RCA can help explain it.

Integration architecture

The agents are only as good as their access:

SourceAccess patternNotes
CMMSRead-only API for work orders, assets, failure codesConsistent failure codes matter more than volume
HistorianRead-only queries for tags around the event windowInclude operating state so normal transients are not flagged
Condition monitoringAlerts and trends by assetAsset IDs must map to the CMMS; see why alerts point to the wrong asset
ERP and purchasingRead-only access to part numbers, suppliers, substitutionsOften the most valuable and least connected source
Shift logsText searchTreat as weak evidence unless confirmed

Keep agents read-only by default and keep plant control systems out of reach. NIST SP 800-82r3 provides guidance on securing operational technology while addressing its unique performance, reliability, and safety requirements [4]; use it when deciding how investigation tools reach plant data. Our CMMS integration guide covers the CMMS side.

Governance and measurement

The NIST AI Risk Management Framework is intended to help organizations incorporate trustworthiness considerations into the design, development, use, and evaluation of AI systems [5]. For RCA, that translates into:

  • Documented scope. Which systems the agents read and which decisions remain human.
  • Measured accuracy. How often the top-ranked hypothesis was confirmed, tracked over time.
  • Traceability. Every hypothesis links to the records it relied on.
  • Feedback. Engineers' accept and reject decisions are stored and reviewed.

The DOE O&M Best Practices Guide (Release 3.0) makes a point that applies to any new diagnostic technology: proper application and training are critical, and equipment should not be bought for in-house use without a serious commitment to implementation and training [6]. Budget for the people who will run and review the system, not just the software.

Implementation playbook

1. Pick one asset class and one recurring failure mode, such as bearing failures on a pump family, with enough history to test against. 2. Connect read-only sources for that class: CMMS, historian, condition data, and purchasing. 3. Back-test on closed investigations. Run the system on past failures with known causes and see whether it ranks the confirmed cause highly. 4. Run in parallel with your normal RCA on new events for a defined period. Compare the hypotheses with the team's conclusions. 5. Decide from the record. Expand to another failure mode only when the confirmed-hypothesis rate meets the threshold you agreed at the start.

Frequently asked questions

Does agentic RCA replace human engineers?

No. It handles data gathering and cross-referencing and proposes ranked hypotheses. Engineers verify them, decide the cause, and approve corrective actions.

How much historical data does it need?

Enough closed work orders and condition history on the target asset class to back-test against known causes. Consistent failure coding matters more than years of records.

What if our CMMS data quality is poor?

Start by standardizing failure codes for the asset class you pilot, and treat free-text notes as weak evidence. Poor data limits any RCA method, human or AI.

How do we handle wrong hypotheses?

Require a verification step for every hypothesis, record engineers' accept or reject decisions, and track the confirmed-hypothesis rate before widening use.

References

[1] NIST, "Research Suggests Significant Benefits to Investing in Advanced Machinery Maintenance," 2020. https://www.nist.gov/news-events/news/2020/06/research-suggests-significant-benefits-investing-advanced-machinery

[2] Pacific Northwest National Laboratory, "O&M Best Practice Issue Discussion: Maintenance Approaches." https://www.pnnl.gov/projects/om-best-practices/maintenance-approaches

[3] Pacific Northwest National Laboratory, "Applying Key Performance Indicators." https://www.pnnl.gov/projects/om-best-practices/applying-key-performance-indicators

[4] NIST, "SP 800-82 Rev. 3, Guide to Operational Technology (OT) Security," September 2023. https://csrc.nist.gov/pubs/sp/800/82/r3/final

[5] NIST, "AI Risk Management Framework." https://www.nist.gov/itl/ai-risk-management-framework

[6] U.S. Department of Energy, Federal Energy Management Program, "Operations and Maintenance Best Practices Guide, Release 3.0," 2010. https://www.energy.gov/sites/prod/files/2020/04/f74/omguide_complete_w-eo-disclaimer.pdf

Ready to put this into practice?

See how Monitory helps manufacturing teams implement these strategies.