Operational root cause analysis
Activated Cloud✓ Officialactivated/operational-root-cause-analysis
Free · MIT
About
Investigates an operational failure or recurring problem without blame: contains the impact first, defines the problem precisely (what, where, when, how much, and what it is not), rebuilds the timeline from records, uses Pareto, fishbone and evidence-checked 5 whys to find causes and contributing factors, verifies them, and sets corrective and preventive actions with an effectiveness check. Use after a missed delivery, a costly error, a complaint spike or a process breakdown. Not for reviewing a whole project: use project-retrospective.
Documentation
Operational root cause analysis
You find out why something went wrong in operations, and fix the conditions that let it happen, so it does not happen again. The standard: the problem is stated with numbers, every causal claim is backed by evidence, the analysis asks how the system allowed the failure rather than who caused it, and every action is owned, dated and later checked to see whether it worked.
When to use
- "Forty orders went out with the wrong items yesterday. What happened?"
- "Why do we keep missing the Friday courier?"
- "Complaints doubled this month, find out why."
- "The payroll file was late again."
- A signal on the operations dashboard (see operations-metrics-dashboard).
What you need
- The records: order, ticket, job or system logs with timestamps, emails and chat about the event, shift rotas, change history (what changed recently in process, people, suppliers, systems). Access through connected apps on the Connections page, your own browser signed in by the owner, or exports.
- The people involved, reached with
ask_teammateorbrief_team, and the owner of the process. - Data on how often the problem happens, for at least several weeks before and after.
Method
- Contain first. If the problem is still causing harm (customers affected, money leaking, safety), propose the immediate containment (hold shipments, add a manual check, pause a process) to the owner. Actions that affect customers, money or staff need the owner's go-ahead. Containment is not the fix; note it as temporary.
- Define the problem with numbers, using "is" and "is not": what is wrong (and what similar thing is fine), where (which site, product, channel, and where it does not happen), when (first seen, how often, which shifts or days, and when it does not happen), how much (count, rate, cost). The "is not" side narrows the causes: if errors happen only on the late shift, causes common to all shifts are unlikely.
- Build the timeline from records, not memory: each event with time, source and who or what acted. Mark the moments where the problem could have been caught.
- Quantify with data (
execute_code): a Pareto of the problem types or locations (a few categories usually account for most cases), and a before-and-after view around recent changes. If a change coincides with the start of the problem, it is a strong lead, not yet a proven cause. - Brainstorm possible causes with a fishbone across People (skills, staffing, workload), Process (steps, handoffs, instructions), Tools and systems (software, equipment, configuration), Inputs (materials, data, information received), Measurement (checks, data, alerts), and Environment (time pressure, layout, external events). Gather views from the people closest to the work individually.
- Drill down with 5 whys on the most likely branches, checking each answer against the records before asking the next why. Stop when you reach a cause the business can change and that, if removed, would have prevented the problem. Expect more than one cause: a root cause plus contributing factors.
- Ask how, not who. Phrase questions as "what made this step easy to get wrong?" and "how did it pass the checks?". "Human error" is never a root cause; it is where the investigation starts. Watch for hindsight bias (it looks obvious only now), confirmation bias (seeking evidence for the first theory), and blaming the last person to touch it.
- Verify the cause. Does it explain every fact in the is and is-not table? Does the data show the problem rises and falls with it? Can it be switched on and off in a safe test? If not, keep investigating.
- Set actions, each with owner, date and expected effect:
- Correction: fix the specific cases (rework the orders, refund, apologise). Customer-facing and money actions need the owner's go-ahead.
- Corrective action: remove the cause (change the process, the system setting, the input).
- Preventive action: stop similar problems elsewhere. Prefer stronger fixes: eliminate the possibility, then automate or mistake-proof, then standardise with a checklist, and only last train or remind.
- Check effectiveness after an agreed period (often 4 to 8 weeks) with the same measure as the problem statement. Schedule it with
cronjob. If the problem has not dropped, the cause was wrong or the fix did not take: reopen.
Worked example: wrong items in parcels
Problem: 41 of 2,310 orders (1.8%) left Warehouse B with the wrong item between 7 and 20 Sep, against a normal rate of 0.3%. Is: late shift, aisles 7 to 9, Mondays and Tuesdays. Is not: Warehouse A, the early shift, other aisles, missing or damaged items. Containment (owner approved, 21 Sep): a second-person check on every late-shift pick from aisles 7 to 9, for two weeks. Timeline from records: 6 Sep, re-slotting job in aisles 7 to 9; 7 Sep, first wrong-item complaint; 14 Sep, scanner battery requests logged; 20 Sep, the complaint spike reaches the owner. Pareto: 38 of the 41 wrong items came from the bin next to the right one. 5 whys, each step checked against a record:
- Why wrong items? Pickers took from the neighbouring bin (pick log against packing photos).
- Why the neighbouring bin? Labels in aisles 7 to 9 were moved during re-slotting and several sit under the wrong bin (walk-through photos, job sheet).
- Why only the late shift? It uses paper pick lists because the scanners run flat by 16:00 (scanner sign-out log).
- Why no scan check? No spare batteries, and scan-to-confirm is not required on paper picks.
- Why not caught at packing? Packing checks the item count, not the item identity. Root cause: pick confirmation depends on a scanner that is unavailable on the late shift, so picks rely on bin labels that the re-slotting misplaced. Contributing: re-slotting had no label check; packing checks count only. Not a cause: picker experience (the same error rate for new and experienced pickers). Actions:
- Correction: re-pick and re-ship the 41 orders with an apology (owner approved, customer-facing).
- Corrective: spare battery packs, and scan-to-confirm required on every pick (mistake-proofing).
- Preventive: a label check step in the re-slotting SOP, and an item-identity check at packing (standardise).
Effectiveness check: the late-shift wrong-item rate, reviewed 4 weeks later; target back under 0.3%. The full write-up follows
references/rca-template.md.
Choosing the strength of a fix
| Strength | What it looks like | When it is right |
|---|---|---|
| Eliminate | Remove the step or the possibility | Whenever the step adds no value |
| Mistake-proof or automate | The system blocks the wrong choice; a scan is required | High-volume, repeatable steps |
| Standardise | A checklist at the point of work | Rare or judgement steps |
| Train or remind | A briefing or a poster | Only alongside something stronger |
| If every action on your list sits in the bottom row, the analysis has not reached a cause the business can change. Go back to the records. |
Asking without blaming
Ask "what made this easy to get wrong?", "what did the check at that step look at?" and "what would have had to be true for this to be caught?". Write conditions, not names: "picks on the late shift relied on misplaced labels", never "Sam picked the wrong item". The wording table in the reference has more examples.
Output
- RCA report (1 to 3 pages, template in
references/rca-template.md): summary, impact, problem statement with is and is-not, timeline, analysis (Pareto, fishbone, 5 whys), root cause and contributing factors with evidence, actions table, effectiveness check date. - A
show_cardwith the impact, the root cause in one sentence, and the actions with owners and dates. - After the check: a one-line update on whether the problem rate fell, with the numbers.
Checks before you finish
- Problem statement has numbers and an is-not side.
- Every causal claim cites a record or data, not an opinion.
- The root cause explains all the facts in the is and is-not table.
- No person is named as the cause; the write-up describes conditions and decisions.
- Each action has an owner, a date and a measurable expected effect, and at least one is stronger than training.
- An effectiveness check is scheduled.
Pitfalls
- Stopping at "human error" or "didn't follow the process". Ask why the process allowed it.
- Fixing the first plausible cause without testing it against the facts.
- One root cause by force. Most failures have a main cause and several contributing factors.
- Actions that are reminders. "Be more careful" changes nothing; change the system.
- Skipping containment while the analysis runs, and letting the harm continue.
- Never checking. Without the effectiveness check you do not know if it worked.
See also
- operations-metrics-dashboard, process-improvement, project-retrospective, process-mapping-and-sops.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
