Incident Investigation & Root Cause Analysis


A few years ago, I was called into an investigation at a client's site, as the investigator representing the supplier side, working alongside the client's own lead investigator. A driver was required to remain in a designated safe area for the duration of a loading operation. That rule existed for a specific reason: a reversing forklift moving a load cannot always see a person on foot nearby, and the yard was not physically separated from the area where drivers were meant to wait.

The driver left that area while the yard was active. A reversing forklift struck him. He did not survive.

I have generalised every detail beyond that description on purpose: the company, the site, the region, and the exact sequence of that day. None of it is identifiable here, and none of it needs to be for the argument to hold. What I can speak about directly is the shape of the investigation that followed, because I sat on that team, and because that shape is the same one I have seen in facility after facility, across three continents, for twenty+ years.


SECTION 1: THE CASE

The finding that writes itself in a case like this is brief. The driver left the designated area when the procedure required him to stay in it. Human error. Corrective action: we will retrain the driver population on the loading procedure, issue a reminder notice, and close the file.

That finding is factually true. It is also the point at which the investigation stopped doing its job.

A root cause explains why a trained, experienced person did something that got him harmed. "He broke the rule" is not that explanation. It is the last visible action before the injury, reported as though it were the reason for it. The real question sits one level back: why was a rule that existed on paper not something the operation was actually enforcing on the ground? Why was there no physical segregation in the yard between the zone where forklifts manoeuvre and the zone where drivers wait? Why did the forklift have no method, mirror, camera, spotter, proximity alarm, to confirm the area behind it was clear before reversing with a load?

None of those are questions about the driver. Every one of them is a question about the organisation, and every one of them was answerable before the incident occurred, not just afterwards.

This scenario is the pattern I see most often when I review closed investigation files. The report identifies an individual action, assigns a corrective action aimed at that individual or his peer group (retraining, a reminder, a disciplinary note), and stops. The organisational conditions that made the individual action likely remain unchanged. The next person in that role, under the same conditions, makes a similar decision, because the conditions never changed.

Recommended Reading: The Case

  • ICAM (Incident Cause Analysis Method), the structured four-level causation model developed for the resources and energy sectors, built specifically to move investigations past individual action into organisational factors; safetyculture.com/library
  • James Reason, Human Error (Cambridge University Press, 1990), the original text behind the Swiss Cheese Model and the distinction between active failures and latent conditions
  • ISO 45001:2018, Clause 10.2, Incident, nonconformity and corrective action, the standard's requirement that organisations determine underlying causes, not only immediate ones; iso.org

SECTION 2: THE EVIDENCE

James Reason's Swiss Cheese Model, developed in 1990 and still a common reference point for serious incident analysis, makes the distinction plainly. Incidents usually result from multiple failures. It is caused by a set of defensive layers, procedures, physical barriers, supervision, and training, each of which has gaps, and an incident occurs when those gaps happen to line up. The driver leaving the designated area was one layer that failed. It was not the only one.

The layer directly behind it was the absence of physical segregation in the yard. A procedure that tells a person to stay in a zone is an administrative control. The Hierarchy of Controls places administrative controls below engineering controls for a reason: a rule depends on someone remembering and choosing to follow it every time, under every condition. A physical barrier, a segregated pedestrian route, or an interlock that prevents forklift movement while the loading zone is occupied does not depend on that. It was not there.

The layer behind that was detection. The forklift operator had no reliable way to confirm that the area behind the vehicle was clear before reversing with a load: no spotter system, no proximity sensor, and no rear camera feeding a monitor at the operator's eye line. In many facilities, this layer is treated as optional because forklift operators are experienced and the yard is "known". Experience is not a detection system. It is a habit that works until the one day it doesn't.

The layer behind that was emergency response. Whether the site's emergency response capability, the training level, the equipment on hand, the response time from the point of injury, was actually matched to the severity of harm this kind of incident can produce was one of the questions the investigation had to confront. A yard with reversing forklifts carrying loaded materials can produce injuries beyond what a basic first aid course prepares someone to manage, and in facility after facility, that gap is assumed rather than tested.

Methodologies exist specifically to force an investigation through all of these layers instead of stopping at the first one. Tripod Beta, developed originally with Shell and grounded directly in Reason's model, categorises Basic Risk Factors, the underlying organisational conditions, design, training, communication, procedures, that sit behind every active failure. ICAM, the Incident Cause Analysis Method used widely across mining, energy, and logistics operations, builds a four-level model that moves from absent or failed defences, through individual and team actions, through task and environmental conditions, to organisational factors. Both methods exist because the person model, stopping at the human action, produces a superficial account that misses the conditions that will produce the next incident.

The data on how incidents actually happen supports this. According to the UK Health and Safety Executive's 2025/26 figures, being struck by a moving vehicle was the second most common cause of fatal workplace injury in Great Britain, 24 of the year's 126 worker deaths, behind only falls from height. In the transportation and storage sector specifically, vehicle-pedestrian interaction is a leading cause of fatality year after year. This is not a rare failure mode that a single distracted driver produced. It is a known, recurring, well-documented risk in any operation where loaded vehicles and people share space, which means the controls for it are not a mystery that could only be identified after the fact. They are documented, available, and frequently not implemented until after someone is killed.

The same requirement exists on the other side of Technique Works' operating footprint, and it is written into the regulation, not left to good practice. In Abu Dhabi, OSHAD-SF, the emirate's occupational safety and health framework, mandates a documented root cause analysis for every reportable incident and near miss under Mechanism 11.0, run under an explicitly no-blame investigation culture, with a preliminary report due within 72 hours and a full investigation report within 30 calendar days. The regulator is not asking whether the driver broke a rule. It is asking what the organisation is doing about the conditions that let the rule get broken. An investigation that stops at "operator failed to follow procedure" satisfies that requirement no better in Abu Dhabi than it satisfies ISO 45001 in Rotterdam.

Recommended Reading: The Evidence

  • HSE (UK), Work-related fatal injuries in Great Britain, 2025/26, the annual national data on causes of workplace fatality, including vehicle-strike incidents; hse.gov.uk/statistics/fatals-overview.htm
  • ADPHC, OSHAD-SF Mechanism 11.0, Incident Notification, Investigation and Reporting, the Abu Dhabi requirement for root cause analysis under a no-blame investigation culture; adphc.gov.ae
  • Tripod Beta, developed by Leiden and Victoria universities with Shell International, the Basic Risk Factor framework for identifying underlying organisational causes; learnfromaccidents.com
  • Health and Safety Executive, Workplace Transport Safety guidance, covering segregation, detection systems, and the hierarchy of controls applied to vehicle-pedestrian interaction; hse.gov.uk

CASE STUDY

Return to the incident in its generalised form and consider what a fuller Tripod Beta or ICAM-style analysis would reveal: a closed file built around one action will never suffice.

The immediate account: a driver left a designated safe area in violation of a written procedure and was struck by a reversing forklift. The easy corrective action: retrain drivers on the procedure, issue a site-wide reminder.

The questions a fuller investigation has to ask sit one level below that. Did the safe-area rule exist as anything more than a written instruction, with real supervision and a verification step confirming compliance before the forklift began moving? Had the yard ever been physically segregated, or had years without incident been read as evidence the layout was safe rather than as evidence the gap simply had not lined up yet? Did the forklift have any detection system fitted, a known and available control for exactly this exposure? And was the site's emergency response capability actually built for trauma consistent with heavy-vehicle impact, or for the minor injuries it had previously handled?

Every one of those questions was answerable before the incident, through a walk-through, a task risk assessment, or a review of near-miss history, if anyone had been looking one level below the procedure document. None of them required the incident to occur in order to be asked.

Where the investigation team pushed into those questions, the corrective actions that followed looked nothing like the first draft: physical segregation between vehicle and pedestrian zones, detection systems on reversing equipment, a verified zero-count procedure confirming the area is clear before any reversing movement with a load, and an emergency response capability reassessed against the severity of harm the site's operations can actually produce, not the severity its existing training assumed.

That is the difference between an investigation that stops at the driver and one that does not. The first produces a training record. The second produces a different yard.

Recommended Reading: Investigation Methodology

  • ICAM Incident Investigation Guide, practical structure for the four-level causation model and interview technique for surfacing organisational factors; safetyculture.com/library
  • Bow-Tie Analysis, the barrier-mapping method that visualises preventive and mitigative controls around a top event, is useful for identifying which specific barrier was absent or failed; CGE Risk Management Solutions, bowtiexp.com
  • UK HSE, Investigating accidents and incidents (HSG245), the practitioner guide to root cause investigation used widely across UK industry; hse.gov.uk/pubns/books/hsg245.htm

SECTION 3: THE PRINCIPLE

June's edition argued that the business continuity plan is for the business, not the department, and that the CEO, not the HSEQ director, is accountable for it once the threshold from operational to organisational is crossed. This month's argument sits directly next to it.

The emergency plan handles the incident. The investigation that follows handles the learning. When a serious event occurs, the investigation determines whether the organisation understands what actually happened or whether it produces a legally defensible finding, operationally comfortable, and structurally wrong.

An investigation that stops at human error is not neutral. It is a decision, made or allowed by leadership, that the organisational conditions behind the incident will not be examined. That decision has a cost that is not visible until the same conditions produce the next incident, because nothing was actually changed.

It also has a cost that is visible immediately, in a courtroom or a regulator's file, and that cost has a number attached to it. Under the UK Sentencing Council's guideline for health and safety offences, amended and in force since June 2025, a large organisation, turnover of £50 million or more, convicted of the most serious category of breach faces a fine with a starting point of £4 million and a range of £2.6 million to £10 million. Two unrelated examples show the range in practice: Decco was fined £2.2 million following an employee's death, and Merlin Entertainments was fined £5 million after a rollercoaster crash caused serious injury. Neither is connected to the case above. These are not outlier figures reserved for chemical majors. They are what a court does with a finding that the organisation knew, or was in a position to know, about the condition that produced the harm, which is precisely the finding an investigation that stops at human error is built to avoid reaching.

"Operator failed to follow procedure, and we retrained our staff" is not a defensible position when the same organisation can be shown to have known, or been in a position to know, that the physical and supervisory conditions made that failure likely. Organisations that stop at human error repeat the incident. Organisations that go deeper, to the systemic condition that made the human failure predictable, stop it, and are also in a materially stronger position if a regulator or a court asks what the organisation did with what it learned.


Recommended Reading: The Principle

  • Sentencing Council, Health and Safety Offences, Corporate Manslaughter and Food Safety and Hygiene Offences Definitive Guideline (amended, in force from 1 June 2025), the fine ranges courts apply to organisations following serious and fatal breaches; sentencingcouncil.org.uk
  • James Reason, Managing the Risks of Organisational Accidents (Routledge, 1997), the extended argument on latent conditions and organisational accountability
  • CCPS (Center for Chemical Process Safety), Guidelines for Investigating Chemical Processes Incidents, the process-safety-specific reference for root cause methodology and organisational learning; aiche.org/ccps
  • ISO 45001:2018, Clause 10.2, on determining causes and taking action to prevent recurrence, the compliance baseline this argument sits above

HSEQ MARKET INSIGHTS: JULY 2026

Four data points shaping the incident investigation conversation in the industrial sector right now.

1. Vehicle strikes remain a leading cause of workplace fatality. HSE's 2025/26 provisional data records 126 worker deaths in Great Britain, with being struck by a moving vehicle the second most common cause at 24 deaths, behind falls from height at 31. This is not a rare or unpredictable failure mode. It is a known, recurring pattern with documented controls. Source: HSE, Work-related fatal injuries in Great Britain, 2025/26.

2. GCC regulators already require going past the human action. OSHAD-SF, the Abu Dhabi occupational safety and health framework, mandates a documented root cause analysis for every reportable incident, run under a no-blame investigation culture, with a full report due within 30 calendar days. Organisations operating in the UAE are already required by their regulator to answer the question this edition is asking. Many investigation files still don't. Source: ADPHC, OSHAD-SF Mechanism 11.0, Incident Notification, Investigation and Reporting.

3. Investigation methodologies built to go past human error exist and are underused. ICAM, Tripod Beta, and Bow-Tie analysis were each developed specifically to move investigations beyond the individual action to the organisational conditions behind it. Their adoption is strongest in energy, resources, and process industries and weakest in logistics, warehousing, and general manufacturing, the sectors where vehicle-pedestrian and struck-by incidents are most common. Of the three, ICAM is the most widely adopted starting point in logistics and manufacturing specifically, the sectors carrying the most struck-by-risk. Source: Technique Works operational observations, 47 facilities, Western Europe and GCC.

4. The finding has a price, and courts are setting it. Under the UK Sentencing Council's guideline, amended and in force since June 2025, a large organisation faces a fine with a starting point of £4 million for the most serious health and safety breaches. In two unrelated cases, Decco was fined £2.2 million following a workplace death and Merlin Entertainments £5 million after a serious injury. ISO 45001:2018, Clause 10.2, and OSHAD-SF both already require organisations to determine underlying causes, not only the immediate action. An investigation that stops at "operator failed to follow procedure" satisfies neither the standard nor, increasingly, the court. Source: Sentencing Council, Health and Safety Offences Definitive Guideline: ISO 45001:2018; ADPHC OSHAD-SF.


QUESTIONS FOR YOU TO CONSIDER

Five questions grounded in this edition's argument. They are not rhetorical. They are diagnostic.

1. Pull your last three closed incident investigation files. How many end with a human action as the stated cause, and how many go one level further to the condition that made that action likely?

2. For your highest-consequence operational area, loading, confined space, high-voltage, working at height, has a physical barrier or engineering control been assessed and rejected in favour of a written procedure, or was the procedure the only control ever considered?

3. When was the last time your organisation changed a physical layout, a piece of equipment, or a supervisory structure because of a near-miss, rather than after an actual injury?

4. Does your emergency response capability match the actual severity of harm your highest-risk operations could produce, or was it built around the injuries you have had rather than the ones your operations are capable of causing?

5. If your last serious incident investigation were read by a regulator or opposing counsel, would it show an organisation that examined its conditions or one that closed the file at the first available individual?

If the honest answer to any of these is "I'd need to check," the investigation quality question is already visible.


PRACTICAL ACTION

Four steps. In sequence. The first one takes fifteen minutes. The last one takes ninety.

Step 1. Pull the most recent incident or near-miss report where the stated cause is a human action, "failed to follow procedure," "didn't notice," "used incorrect technique." Read only the corrective action. If it is retraining, a reminder, or a disciplinary note, and nothing else, the investigation stopped at the first layer.

Step 2. For that same incident, ask what physical or engineering control existed at the point of failure and what would have had to be true for the human action to be prevented rather than corrected after the fact. If the answer is "nothing, the procedure was the only control," that is the finding.

Step 3. Check whether your investigation process specifies a methodology, ICAM, Tripod Beta, Bow-Tie, or an equivalent structured method, or whether it relies on the investigator's individual judgement to decide how far to go. An unstructured process will, by default, stop at the most visible and least uncomfortable explanation.

Step 4. Schedule 90 minutes with your HSEQ director and your legal or compliance lead. Review the last two serious incident files together and ask one question of each: if this were read back to us in a regulatory inquiry, does it show an organisation that understood its own risk, or one that assigned blame and moved on?

If you do not yet have the standing to convene that conversation, start smaller. Run Steps 1 through 3 on a single file, and bring the finding, not the ask, to whoever you report to. A specific gap in one file is easier to act on than a general concern about all of them.


PERSONALISED RECOMMENDATIONS

The same argument lands differently depending on where you sit.

For the CEO / General Manager

You are not expected to run incident investigations. You are expected to know what kind of investigation your organisation produces. An investigation that consistently ends at human error is not a training gap. It is a signal that your organisation has decided, implicitly, not to look at its own conditions. That decision carries legal exposure, priced in the millions once a serious breach reaches a UK court, and it carries repeat incidents, priced in whatever the next one costs you. Before your next board cycle, ask to see the last three closed investigation files and read the stated causes yourself. If the answer is "we retrained the team", the question your board should be asking is what retraining was supposed to fix that a physical barrier would have fixed permanently.

For the COO / Operations Director

Your operational areas are where the layers Reason describes actually exist: procedures, physical barriers, supervision, and detection systems. The most common gap is relying on the procedure alone because it is the cheapest and fastest control to implement. Pull the risk assessment for your highest-consequence operational area and check whether an engineering control was genuinely evaluated, or whether the written procedure was the only option considered.

For the HSEQ Director

This is where you become the person who found the real answer, not the person defending the easy one. A structured methodology forces the investigation through the organisational layers whether or not the investigator is inclined to go there, which protects you as much as it protects the organisation, since the finding is no longer a matter of individual judgement under pressure. ICAM is the most widely adopted starting point in logistics and manufacturing specifically and the easiest to introduce without a lengthy change process. If your current investigation process does not specify a methodology, naming that gap and closing it is a stronger position to be in than waiting for the next serious incident to expose it for you.


NEXT MONTH

August's edition follows the arc directly.

The investigation tells you what happened after the fact. The metrics your board reviews are supposed to tell you before it happens again. Most boardrooms track fatality rates in operations that have never had a fatality. That is not measurement. It is theatre with a spreadsheet attached.

The leading indicators that actually predict whether an incident is coming rarely reach the board table: near-miss reporting trends, barrier verification rates, behavioural observation data, and corrective action closure velocity. The organisations that track these see the next incident before it happens. The organisations that only track the lagging numbers find out the same way everyone else does—after the fact.

August's argument: your board is reviewing the wrong dashboard.


CALL TO ACTION

Each edition of Technique Works HSEQ Insights is published monthly at insights.techniqueworks.com.

If this edition is useful, subscribe to receive it directly. The full archive, including June's edition on business continuity and the complete series from November 2025, is there now.

If it belongs in your GM's inbox, your COO, your legal lead's, or your HSEQ director's, forward it as is. The details have already been generalised beyond identification.

Subscribe → insights.techniqueworks.com


Amador Brinkman · Technique Works info@techniqueworks.com​


Technique Works · HSEQ Insights Newsletter · Edition 11 · July 2026