Resources · Incident reporting
Incident investigation: a practical guide to root cause analysis
Incident investigation is the process of looking past what happened to find why it happened, so you can fix the underlying conditions rather than just the symptom.
Incident investigation is how you turn a report into a fix. Its purpose is not to find someone to blame but to understand the chain of conditions that allowed an event to happen, so that the same chain cannot form again. Root cause analysis is the part of that work where you keep asking “why” until you reach a cause you can actually do something about.
What is an incident investigation?
An incident investigation is a structured review of an event that gathers the facts, reconstructs what happened, and identifies the causes so that corrective action can prevent a recurrence. It applies to incidents that caused harm and, just as usefully, to near misses that did not. The management-system standard ISO 45001:2018 requires organisations to investigate incidents and act on what they learn, and the principle is the same whatever your sector: an event you do not understand is an event you cannot prevent.
The mindset matters as much as the method. A good investigation assumes that people generally try to do the right thing and that failures usually come from the conditions around them: unclear procedures, missing guards, time pressure, poor lighting, confusing handovers. Blaming an individual feels decisive but leaves every one of those conditions exactly where it was.
The aim of an investigation is to fix the system, not to find a culprit. If your finding is “someone was careless”, you have stopped one question too early.
Immediate, underlying and root causes
Most events have causes at three depths, and good investigation works down through all of them.
| Level | What it is | Example |
|---|---|---|
| Immediate cause | The obvious thing that caused the harm. | A worker slipped on a wet floor. |
| Underlying cause | The condition that allowed the immediate cause. | A leak had been left unrepaired and there was no spill procedure. |
| Root cause | The system failure behind it all. | Maintenance requests had no owner and no tracking, so they were lost. |
Fixing only the immediate cause, mopping the floor, solves nothing. The leak returns, the floor gets wet again, and the next person slips. Fixing the root cause, a maintenance system that tracks requests to completion, is what actually removes the risk. This is the same logic that makes near misses so valuable, because they expose these chains before anyone is hurt. See what is a near miss for more on that.
How to run an incident investigation, step by step
A proportionate investigation follows a clear sequence. Match the depth to the severity, or the potential severity, of the event.
- Make the scene safe and gather information. Care for anyone hurt, then preserve the scene and collect evidence while it is fresh: photos, readings, the report itself and, gently, accounts from those involved.
- Build the timeline. Lay out what happened in order. A clear sequence of events is the backbone of every investigation and often reveals gaps on its own.
- Identify the causes. Work down from the immediate cause to the underlying and root causes using a simple analysis method, covered below.
- Agree corrective actions. For each meaningful cause, decide what will change. Aim at conditions and systems, not at telling people to “be more careful”.
- Assign owners and due dates. Every action needs a named owner and a deadline, or it will not happen.
- Verify and close. Check that actions were completed and actually worked, then lock the record as evidence.
- Share the learning. Tell the wider organisation what you found, so other teams benefit and do not relive the same event.
Two simple root cause methods
You do not need elaborate tools to find a root cause. Two simple methods cover most needs.
The Five Whys. Start with the problem and ask “why” until you reach a cause you can fix, usually around five times. For the slip above: the worker slipped (why?) because the floor was wet (why?) because a pipe leaked (why?) because it was never repaired (why?) because the maintenance request was lost (why?) because requests had no owner or tracking. The last answer is something you can change, so it is the root cause. The skill is to keep asking past the first satisfying answer.
The fishbone diagram. Also called an Ishikawa or cause-and-effect diagram, this groups possible causes into categories such as people, equipment, environment, materials, methods and management. It is useful when an event has several contributing factors at once rather than a single clean chain, because it stops you fixating on the first cause you find.
Whichever you use, the test of a real root cause is simple: if you removed it, would the event have been prevented, and is it something within your control to change? If the answer to both is yes, you have found it.
The chain of small failures
Serious incidents are rarely one big failure. They are usually several small, ordinary conditions that happened to line up. The psychologist James Reason described this with his “Swiss cheese” model: every defence you have, training, guards, supervision, checks, is a slice of cheese with holes in it. Most of the time the holes do not align and nothing happens. Occasionally they do, and a hazard passes straight through every layer at once. A good investigation maps those holes. It asks not only “what failed” but “what else would have had to go wrong for this to be much worse”, which is the question that turns a minor event into a serious lesson. For the wider context, see our guide to incident reporting.
Turning findings into prevention
An investigation only pays off if its actions get done and its lessons get used. That means tracking every corrective action to completion, not just agreeing it in a meeting, and watching across many investigations for the patterns that single reports cannot show. If three separate slips, a near miss and a damaged forklift all trace back to lost maintenance requests, the real finding is bigger than any one event. Bringing those threads together is the job of data visualisation, and keeping the whole flow moving from report to closed action is what our health and safety solution is built to do.
Frequently asked questions
What is root cause analysis?
Root cause analysis is the part of an investigation where you keep asking why an event happened until you reach an underlying system cause you can fix. It looks past the obvious immediate cause to the conditions that allowed it, so the fix prevents recurrence rather than just treating the symptom.
What is the difference between an immediate cause and a root cause?
The immediate cause is the obvious thing that caused the harm, such as a wet floor. The root cause is the system failure behind it, such as a maintenance process that loses repair requests. Fixing the immediate cause alone lets the event recur; fixing the root cause removes it.
Should near misses be investigated?
Yes, especially high-potential ones. A near miss reveals the same chain of failures as an accident without the harm, so investigating it lets you fix the risk before anyone is hurt. Match the depth of the investigation to what could have happened, not just to what did.
Is the goal of an investigation to find who is to blame?
No. The goal is to understand and fix the conditions that allowed the event. Blame tends to stop the inquiry early, discourages future reporting and leaves the underlying causes untouched. A just-culture approach produces far more honest information.
What are the Five Whys?
The Five Whys is a simple root cause method where you ask “why” repeatedly, starting from the problem, until you reach a cause you can act on. Five is a rough guide rather than a rule. The point is to keep going past the first easy answer to a cause you can actually change.
Sources
- International Organization for Standardization, ISO 45001:2018 Occupational health and safety management systems. https://www.iso.org/standard/63787.html
- Health and Safety Executive, Investigating accidents and incidents (HSG245). https://www.hse.gov.uk/pubns/priced/hsg245.pdf
- James Reason, “Human error: models and management”, BMJ, 2000;320:768. https://www.bmj.com/content/320/7237/768
From report to root cause to fix
Track investigations and corrective actions to completion, and see the patterns that single events hide.
Book a demo