LearnRoot cause analysis
Root Cause
Root Cause Analysis

Most investigations stop one question before the useful answer.

Root cause analysis is a structured investigation that works backwards from an event to the conditions that allowed it, and stops only at a cause the evidence supports and the organisation can actually control. The methods — five whys, cause-and-effect diagrams, fault trees, barrier analysis, causal factor charting — are containers for the same discipline: build a timeline from evidence, generate candidate causes, and try to disprove them. On plant data the hard part is rarely the method. It is that the evidence has already evaporated, or was never recorded in a form where two systems can be compared.

The Discipline

Seven moves that separate an investigation from a meeting.

State the event precisely“The line stopped” cannot be investigated. “Filler 2 stopped at 14:06:32 on 12 March, 42 minutes lost, third occurrence this quarter, all on product B” already contains three of the clues.
Preserve the volatile evidence firstIn order of how fast it disappears: the people who were there, the failed part and the material, controller and alarm buffers that wrap, high-resolution trends before the historian compresses them, and video. An investigation opened three days late is often an investigation of what people remember.
Build one timeline, in one time baseAlarms, equipment states, operator actions, setpoint and recipe changes, product changeovers, work orders and lab results, merged and ordered. If the systems' clocks are not disciplined to a common source, the first argument will be about which event came first, and nobody will win it.
Generate candidates with structureFive whys suits a simple linear chain with a known mechanism. A cause-and-effect diagram or fault tree suits branching and multi-factor events. Barrier analysis asks a different and often better question: what was supposed to stop this, and why did it not?
Try to disprove each candidateFour tests do most of the work. Timing: did it actually precede the event, allowing for clock skew? Magnitude: is it big enough to explain what happened? Specificity: why not on the identical line, or the other shift? Counterfactual: remove it, and does the event still occur? A candidate that survives all four has earned a mechanism; one that survives none is a story.
Stop at controllable, not at human“Operator error” is a description, not a cause. If a person made a mistake, the useful questions are what made that mistake likely and what should have caught it — the interface, the alarm, the procedure, the workload, the training, the tolerance for the shortcut. That is where a control can actually be placed.
Choose controls by strength, then verifyEliminate the hazard, then engineer it out or interlock it, then detect it, then procedure, then training and awareness — in that order of effectiveness. Set a verification date and name the metric that would show recurrence. An action closed on the day it is implemented has been implemented, not verified.
Common claimWhat would make it evidenceWhat it is without that
“The temperature rose just before the stop”Both records in one disciplined time base, with the clock offset between systems known.A comparison of two clocks that may differ by minutes — an ordering nobody can defend.
“It always happens on night shift”Rate per opportunity, compared at equal product mix and volume.A count, which mostly reflects that night shift runs the long orders.
“It started after the controller update”A change record with a timestamp, plus the same before-and-after comparison on an asset that was not updated.Post hoc reasoning. Changes are memorable, so they attract blame regardless of guilt.
“The new material lot is bad”Lot traceability joining material to defects, with the previous lot still running somewhere as a control.A hypothesis with a conveniently external owner.
“Vibration correlates with the defects”A mechanism that explains why, and a prediction it makes that could fail.A correlation. Search hundreds of tags and some will correlate by chance.
“This has never happened before”Retention that reaches back further than the interval between recurrences.A statement about your archive, not about your plant.

One occurrence, one mechanism

A single event can establish that something is possible and how it happened. It cannot establish a rate, a trend or a comparison between shifts. Investigations lose credibility fastest when a well-evidenced mechanism is stretched into a statistical claim the data cannot carry.

What makes RCA fast

Not analytics. Retained raw events rather than aggregates, one time base, stable asset identity across systems, and enough context that an alarm, a state change and a work order about the same machine can be joined without a spreadsheet. Plants with those four run investigations in hours; plants without them run them in weeks and settle for a plausible story.

Write the causal chain down

Each link stated as a claim, with the evidence for it and its status: supported, refuted, or untested. A fishbone with forty branches and no verdicts is a brainstorm. The written chain is also what makes the investigation reusable when the same event happens on another line.

Why most RCA stalls.

  • Opened days late, after alarm buffers wrapped, trends were compressed, product was consumed and the crew rotated off.
  • Three systems, three unsynchronised clocks, and an unresolvable argument about the order of events.
  • One-minute averages that hide the four-second excursion which actually caused the trip.
  • A cause-and-effect diagram used as a brainstorm and never converted into tests, so nothing is ever eliminated.
  • Stopping at “operator error” or “human factors” because it is a place where the paperwork can be closed.
  • Corrective actions that are entirely retraining and reminders — the weakest controls, chosen because they are the fastest to write.
  • No verification step, so nobody notices the same event recurring under a different description three months later.
  • Retention shorter than the recurrence interval, which makes the pattern permanently invisible — a quarterly problem cannot be seen in a thirty-day archive.
Related Concepts

What an investigation actually runs on.

RCA is the most demanding consumer of the data layer, because it asks questions nobody anticipated when the tags were configured. It needs retained raw events, one time base, and asset context so records from different systems join — which is what contextualization provides and what a Unified Namespace keeps available to every consumer, with ISA-95 giving the equipment one name across OT, quality and maintenance. In practice the raw material arrives from the neighbouring disciplines: the loss breakdown behind an OEE drop, the flood and chattering patterns exposed by alarm rationalization, the condition history behind a maintenance event, and the observations recorded at the shift handover — often the only evidence that something was already abnormal before anyone was looking.

Reason Over The Evidence

The bottleneck is rarely thinking. It is assembling what happened.

Root Cause reads an evidence bundle you supply — event timeline, alarms, equipment states, work orders, quality records — and proposes candidate causal chains with the specific evidence that would confirm or refute each one, so your team spends its time testing rather than collating. It reasons over the export you give it; it does not connect to your control systems. Five free runs, then it is part of the agents plan.

See Root Cause

Frequently asked questions

Are five whys enough?

For a simple linear chain with a known mechanism, often yes. The technique fails on branching or multi-factor events, because it produces one chain and tends to stop at the first plausible answer rather than the best-evidenced one. When more than one condition had to be true at once, a cause-and-effect diagram or a fault tree keeps the alternatives visible instead of quietly discarding them.

How do we tell a real cause from a correlation?

Put every candidate through the same four tests: did it precede the event once clock differences are accounted for; is it large enough to explain the magnitude; why did it not produce the same result on the identical line or the other shift; and would the event still have happened without it. Then ask for the mechanism. A correlation with no mechanism and no failed prediction is a starting point for investigation, never a conclusion.

How long should we keep the raw data?

Longer than the interval between recurrences of the problems you care about. A fault that returns every quarter is invisible in a thirty-day archive, and the investigation will conclude it is new every time. The practical compromise is to keep full-resolution events and states for the assets that actually cost money, and accept aggregation elsewhere — a decision worth making deliberately rather than inheriting from a default setting.