Engagement snapshot
- Mandate
- Shorten post-incident diagnosis while keeping engineering decisions anchored in plant evidence.
- Timeline
- 15 weeks across one flagship facility and two follow-on plants.
- Team shape
- Plant reliability manager, historian specialist, 2 engineers, and 4 maintenance SMEs.
The problem
Reliability engineers spent too long reconstructing failure chains from CMMS tickets, SCADA alarms, and technician shift notes after line stoppages.
What we built
Built a failure investigation assistant that assembled event timelines, summarized likely contributing causes, and linked evidence across historian data and maintenance records.
Operating context
The group already collected data across multiple systems, but post-incident analysis was still slow because the information lived in different tools and in different time formats. Valuable evidence existed, but no one could pull it together quickly.
Key constraints
- The assistant could not imply a root cause without showing the timeline and evidence behind it.
- Historian, alarm, and maintenance data had inconsistent clocks and required normalization.
- Operators needed to challenge conclusions and feed corrections back into the investigation record.
What we built
Temporal event stitching
Aligned alarms, sensor anomalies, maintenance actions, and technician notes into one usable incident timeline.
Evidence-linked reasoning
Generated failure summaries that stayed explicitly tied to the timeline rather than free-form speculation.
Case learning loop
Captured engineer corrections and confirmed causes so future investigations reflected plant-specific patterns.
Delivery path
Incident library build
Collected recent downtime cases and normalized the data needed to reconstruct them reliably.
Engineer review
Used reliability engineers to challenge the system on ambiguous incidents and recurring nuisance alarms.
Plant rollout
Released by line and failure class so teams could compare performance against their current investigation process.
Why it mattered
Diagnosis got faster because engineers stopped spending their first hours reconstructing the event chain. They still owned the conclusion, but the evidence assembly step became materially lighter and more consistent.
Implementation notes
- Time alignment across systems was a bigger engineering challenge than the summarization layer.
- Engineers trusted the assistant only after it showed the exact events that shaped the recommendation.
- A plant-specific incident library improved usefulness faster than generic prompt tuning.