Most investigations stop one question before the useful answer.
Root cause analysis is a structured investigation that works backwards from an event to the conditions that allowed it, and stops only at a cause the evidence supports and the organisation can actually control. The methods — five whys, cause-and-effect diagrams, fault trees, barrier analysis, causal factor charting — are containers for the same discipline: build a timeline from evidence, generate candidate causes, and try to disprove them. On plant data the hard part is rarely the method. It is that the evidence has already evaporated, or was never recorded in a form where two systems can be compared.
Seven moves that separate an investigation from a meeting.
| Common claim | What would make it evidence | What it is without that |
|---|---|---|
| “The temperature rose just before the stop” | Both records in one disciplined time base, with the clock offset between systems known. | A comparison of two clocks that may differ by minutes — an ordering nobody can defend. |
| “It always happens on night shift” | Rate per opportunity, compared at equal product mix and volume. | A count, which mostly reflects that night shift runs the long orders. |
| “It started after the controller update” | A change record with a timestamp, plus the same before-and-after comparison on an asset that was not updated. | Post hoc reasoning. Changes are memorable, so they attract blame regardless of guilt. |
| “The new material lot is bad” | Lot traceability joining material to defects, with the previous lot still running somewhere as a control. | A hypothesis with a conveniently external owner. |
| “Vibration correlates with the defects” | A mechanism that explains why, and a prediction it makes that could fail. | A correlation. Search hundreds of tags and some will correlate by chance. |
| “This has never happened before” | Retention that reaches back further than the interval between recurrences. | A statement about your archive, not about your plant. |
One occurrence, one mechanism
A single event can establish that something is possible and how it happened. It cannot establish a rate, a trend or a comparison between shifts. Investigations lose credibility fastest when a well-evidenced mechanism is stretched into a statistical claim the data cannot carry.
What makes RCA fast
Not analytics. Retained raw events rather than aggregates, one time base, stable asset identity across systems, and enough context that an alarm, a state change and a work order about the same machine can be joined without a spreadsheet. Plants with those four run investigations in hours; plants without them run them in weeks and settle for a plausible story.
Write the causal chain down
Each link stated as a claim, with the evidence for it and its status: supported, refuted, or untested. A fishbone with forty branches and no verdicts is a brainstorm. The written chain is also what makes the investigation reusable when the same event happens on another line.
Why most RCA stalls.
- Opened days late, after alarm buffers wrapped, trends were compressed, product was consumed and the crew rotated off.
- Three systems, three unsynchronised clocks, and an unresolvable argument about the order of events.
- One-minute averages that hide the four-second excursion which actually caused the trip.
- A cause-and-effect diagram used as a brainstorm and never converted into tests, so nothing is ever eliminated.
- Stopping at “operator error” or “human factors” because it is a place where the paperwork can be closed.
- Corrective actions that are entirely retraining and reminders — the weakest controls, chosen because they are the fastest to write.
- No verification step, so nobody notices the same event recurring under a different description three months later.
- Retention shorter than the recurrence interval, which makes the pattern permanently invisible — a quarterly problem cannot be seen in a thirty-day archive.
What an investigation actually runs on.
RCA is the most demanding consumer of the data layer, because it asks questions nobody anticipated when the tags were configured. It needs retained raw events, one time base, and asset context so records from different systems join — which is what contextualization provides and what a Unified Namespace keeps available to every consumer, with ISA-95 giving the equipment one name across OT, quality and maintenance. In practice the raw material arrives from the neighbouring disciplines: the loss breakdown behind an OEE drop, the flood and chattering patterns exposed by alarm rationalization, the condition history behind a maintenance event, and the observations recorded at the shift handover — often the only evidence that something was already abnormal before anyone was looking.
The bottleneck is rarely thinking. It is assembling what happened.
Root Cause reads an evidence bundle you supply — event timeline, alarms, equipment states, work orders, quality records — and proposes candidate causal chains with the specific evidence that would confirm or refute each one, so your team spends its time testing rather than collating. It reasons over the export you give it; it does not connect to your control systems. Five free runs, then it is part of the agents plan.
See Root Cause