Root-cause analysis investigations can be streamlined with artificial intelligence (AI), but only if it can be explained by human engineers
Across the process industries, manufacturers are under pressure to improve reliability, efficiency and output while operating with leaner teams and rising operational complexity. Plants must manage aging assets, tighter margins, stricter regulatory expectations and growing pressure to do more with existing resources. Yet one of the most persistent operational challenges remains largely unchanged – when something goes wrong, finding the true root cause can still take far too long.
Whether the issue is an unplanned shutdown, recurring quality deviation, throughput loss, rising energy consumption, abnormal equipment behavior or a chronic process instability, investigations often take days or even weeks. By the time conclusions are reached, production has already been affected, valuable engineering time has been consumed, and in some cases the same issue is already beginning to recur.
At the same time, many organizations are experiencing a quieter structural shift. The tenure of engineers and plant specialists is shortening. Where previous generations spent decades at a single site, today’s careers are more mobile. Mobility brings fresh thinking and broader experience, but it also means plants can no longer rely on deep institutional memory held by a small number of experienced individuals. Troubleshooting knowledge is leaving faster than it is being replaced.
These two trends, slow problem solving and loss of critical experience, are creating a growing operational gap.

Root-cause analysis remains manual
Most facilities today are rich in data. Between plant historians, control systems, laboratory records, maintenance logs and operator shift reports, a modern site captures a remarkably complete picture of what is happening at any moment.
Yet despite this abundance of information, root-cause analysis often remains highly manual. Engineers pull together data from disconnected systems, reconstruct timelines and test competing hypotheses, often against prior events recalled from personal experience. Much of that work depends on good judgement – knowing where to look, which signals matter and which interactions will be meaningful.
The result is slow and inconsistent investigation. Two engineers may approach the same issue differently. Key evidence may be missed. Conclusions may depend more on who is available than on a repeatable method.
For continuously operating plants, delay has a cost. A process upset does not pause while a team investigates it. Lost yield, off-specification product, excess energy use, avoidable maintenance activity and reliability risks can continue accumulating in real time.
Many manufacturers have established digital manufacturing, process excellence or continuous improvement teams specifically to help close this gap. Their mandate is often to improve plant performance through analytics, better workflows and standardized methods. Yet in practice, they frequently encounter the same barriers as frontline engineers: fragmented data sources, limited operational context, difficulty scaling expertise across sites, and tools that generate alerts without explaining causes.
As a result, many digital transformation programs improve visibility, but still struggle to deliver faster root-cause resolution where it matters most – on the plant floor.
The commercial impact is substantial. Industry studies suggest unplanned downtime now costs the process industries roughly 11% of their annual revenue, with hourly losses at large facilities ranging from several hundred thousand dollars to over two million, depending on the industry sector. Every additional hour spent diagnosing a problem translates into lost production, delayed shipments, increased energy use, overtime labor and avoidable waste. Across multi-site organizations, repeated delays in identifying root causes quietly erode margins by millions of dollars per year.
For leadership teams under pressure to improve productivity without increasing headcount, faster and more consistent problem solving is becoming a strategic imperative, rather than simply an engineering objective.
The promise – and problem – of AI
It is no surprise that artificial intelligence is attracting strong interest across all manufacturing sectors. Used appropriately, AI can process large volumes of data quickly, detect patterns and surface relationships that humans may not immediately see. However, many current AI approaches face an important limitation in engineering environments: they operate as “black boxes,” meaning that the inputs and outputs are known, but it is not clear how the AI system carried out the analysis and how it arrived at certain conclusions.
An AI system may suggest that a compressor trip was linked to vibration behavior, for example, or that a quality issue correlates with upstream temperature variation. If the AI model cannot clearly explain how it reached that conclusion, engineers are unlikely to rely on it – nor should they.
In process manufacturing, decisions affect safety, environmental performance, production continuity and product quality. Engineers need to understand evidence, challenge assumptions and verify logic before taking action. A confident answer without a transparent explanation is rarely sufficient. In industrial operations, trust determines whether or not a tool is used.
Many organizations have already learned this lesson with first-generation predictive-maintenance tools. Models may indicate that failure risk is increasing, but they can still leave teams asking the most practical question: what should we do about it?
If a pump is predicted to fail, is the real cause cavitation, seal degradation, process contamination, suction instability, control interaction or changing operating conditions? Without understanding the cause, teams may simply treat symptoms, while the underlying issue remains unresolved.
Prediction undoubtedly has value, but prevention usually depends on understanding causation.

Why explainability matters
The next generation of industrial AI is likely to be defined not only by speed or accuracy, but by explainability.
In engineering contexts, explainability means systems that can show:
- Which variables changed first
- What sequence of events occurred
- Which relationships are likely causal rather than coincidental
- What historical precedents support the conclusion
- What confidence level applies
- What alternative explanations were considered
In other words, the system should help engineers reason more effectively, not simply produce unexplained outputs.
This distinction matters because engineers are trained problem solvers. Their role is not to accept answers blindly, but to assess evidence and make sound decisions under uncertainty. AI should strengthen that capability, not bypass it.
That is especially true in high-consequence industries, such as chemicals, petroleum refining, pharmaceuticals, utilities and energy. If a model cannot explain its reasoning, it will struggle to earn trust in the control room, maintenance shop or technical office.
Explainability also supports organizational learning. When a team can see why a system reached a conclusion, they are more likely to challenge assumptions, improve models and refine operating practices over time. A black-box output may produce a momentary answer, but transparent reasoning can build long-term capability.
From correlation to causation
Many of the AI tools deployed in manufacturing to date have been based on correlation. It identifies statistical relationships in historical data and uses those patterns to forecast future events. Correlation can be useful, but it has limits. In this case, the limits are structural. Historical datasets contain only observations of actions that were previously taken. They do not contain the outcomes of interventions that were never tried. No amount of additional data, nor scaling the size or complexity of the model, can close that gap, because the information is absent from the data by construction.
Complex industrial systems do not fail statistically; they fail physically. Equipment, controls, chemistry and operations interact through real cause-and-effect mechanisms. Fouling may restrict flow. Feed variability may destabilize downstream quality. Pressure changes may influence temperature behavior. Heat-exchanger degradation may affect energy efficiency and throughput. Understanding those relationships is what engineers do every day.
That is why causal approaches are attracting increasing interest. Rather than simply identifying that two variables moved together, causal methods aim to understand which variables influence others, in what direction and under what conditions. The reasoning behind any conclusion can then be followed, questioned and validated against engineering judgment, allowing fast and decisive action.
For the chemical process industries (CPI), this distinction matters, because prevention depends on cause. Prediction may tell you something could go wrong. Causal understanding is what helps you stop it.
This is particularly relevant where operating regimes change. Industrial processes are inherently non-stationary. Models built only on historical patterns can be dangerously confident, yet wrong, when the regime shifts. For example, a regime shift could occur when a feedstock changes, a catalyst ages, or a heat exchanger fouls to the point that historical relationships no longer hold. Causal approaches remain valid across those changes because they align with the physics of the plant, rather than with a snapshot of its history.
This has implications beyond engineering performance. Regulators are increasingly asking for explainable reasoning in high-consequence automation. The European Union AI Act (https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai) classifies industrial control systems as high-risk and requires documented reasoning behind automated decisions. In pharmaceutical manufacturing, similar expectations are extending into process control. A model that cannot defend its reasoning in a formal engineering or regulatory review is, increasingly, a model that cannot be deployed.
From research to plant reality
Only a few years ago, much of the work on explainable AI remained inside university research programs. At The University of Sheffield (UK; www.sheffield.ac.uk), one line of inquiry focused on whether causal AI could help diagnose industrial failures before they escalated – and whether advanced analytics could become genuinely useful in day-to-day operations.
That distinction matters. Many technologies perform well in controlled settings. Far fewer succeed amid noisy plant data, changing conditions, incomplete records and the practical pressures of live manufacturing.
Recent activity has combined academic research with industry operating experience to apply causal reasoning to real engineering problems. The broader lesson is the following: some of the most valuable industrial innovations arise when research excellence is shaped by people who understand how plants actually operate.
That combination of perspectives is increasingly important. Academic teams can push the boundaries of modeling and analytics. Experienced operators and engineers know where time is lost, where decisions stall and where practical constraints determine success or failure.
When those worlds connect effectively, innovation becomes more useful.
There is also a broader geographic dimension. The operational challenges facing process manufacturers are remarkably consistent across regions. Whether in North America, Europe, the Middle East, or Asia, plants are dealing with many of the same pressures, including aging assets, skills transitions, tighter margins and the need to make faster, better decisions using existing data.
Discussions with engineering teams across major industrial hubs reinforce the same point: people with decades of plant experience know what matters. Technology creates the greatest value when it supports that expertise, helps scale it across sites, and makes proven engineering judgement easier to apply under time pressure.
From weeks to minutes
When designed appropriately, explainable AI has the potential to significantly compress investigation time.
Instead of manually searching multiple systems, teams can be presented with an organized timeline of contributing events, ranked hypotheses and supporting evidence. Instead of starting each investigation from scratch, plants can learn systematically from previous incidents, near misses and recurring patterns.
This does not eliminate engineering judgement. Rather, it allows engineers to spend less time gathering information and more time evaluating solutions.
The gains can be meaningful:
- Faster response to process deviations
- Reduced downtime during abnormal events
- Quicker restoration of throughput
- Improved consistency of investigations
- Better retention of troubleshooting knowledge
- Stronger learning across sites and teams
For organizations operating lean teams, these improvements can have disproportionate value. If a small group of specialists can solve more problems, faster and more consistently, capacity increases without adding headcount.
It can also improve morale. Many engineers entered the profession to solve meaningful technical problems, not to spend days assembling spreadsheets or manually reconciling data from disconnected systems.
Capturing knowledge
One of the least discussed benefits of explainable AI may be knowledge preservation. Experienced engineers often recognize failure signatures instinctively. They know which combinations of symptoms indicate fouling, control loop interaction, feed variability, mechanical degradation, instrumentation drift or operating practice issues. But much of that expertise is tacit – built through years of observation rather than fully documented.
As tenure shortens, companies risk losing these insights.
Systems that capture investigation pathways, successful diagnoses and recurring patterns can help transform individual know-how into organizational capability. Newer engineers gain access to reasoning frameworks that previously depended on finding the right colleague at the right time.
This can support workforce development as much as operational performance.
It may also help make engineering careers more attractive. Younger professionals often want modern tools, faster learning and the ability to contribute quickly. Giving them better decision support can improve both effectiveness and engagement.
A tool for engineers, not a replacement
Much of the public conversation around AI focuses on replacement: replacing jobs, replacing decisions, replacing expertise. In engineering, a more realistic and more valuable model is augmentation.
Plants will continue to need engineers who understand process behavior, trade-offs, risk, safety and operations. What may change is the quality and speed of the tools available to them.
If AI can help turn fragmented data into understandable evidence, reduce investigation cycles from weeks to minutes and preserve hard-won operational knowledge, it can become an important part of the engineer’s toolkit.
But only if engineers can understand it.
Looking ahead
Process industries have invested heavily in data collection over the last two decades. The next phase does not focus on collecting more data, but on converting existing data into trusted operational intelligence. The most valuable systems will not simply predict outcomes or generate alerts. They will help engineers understand causes, act faster and improve performance with confidence.
When evaluating any new tool against that standard, three questions tend to cut through the marketing. What decision is it helping us to make, and what is the cost if it is wrong? How does it behave when conditions change, and how was that behavior validated? And when its conclusion disagrees with an experienced operator, can it show its reasoning well enough for the operator to challenge it?
For industries built on evidence, accountability and continuous improvement, that is where innovation becomes genuinely useful.
Edited by Scott Jenkins
Author
Dr. Louis Allen is Founder & CEO at Kausalyze (https://www.kausalyze.com), which develops explainable causal AI tools for root cause analysis and predictive insights in process manufacturing. Kausalyze is a University of Sheffield spinout focused on applying causal AI to root cause analysis and operational improvement in process manufacturing. He holds a first-class degree in Chemical Engineering and completed doctoral research on data-driven modeling and predictive methods for industrial systems. His work focuses on translating advanced analytics into practical tools for engineering decision-making in complex operating environments.