Back to Case Studies

Maintenance Root-Cause Assistant

Process Manufacturing Group · Manufacturing

A temporal retrieval assistant for maintenance logs, historian events, and failure investigation.

The measured result

-36%
Mean Time to Diagnosis
-18%
Unplanned Downtime
-22%
Repeat Failure Incidents

Engagement snapshot

Mandate
Shorten post-incident diagnosis while keeping engineering decisions anchored in plant evidence.
Timeline
15 weeks across one flagship facility and two follow-on plants.
Team shape
Plant reliability manager, historian specialist, 2 engineers, and 4 maintenance SMEs.

The problem

Reliability engineers spent too long reconstructing failure chains from CMMS tickets, SCADA alarms, and technician shift notes after line stoppages.

What we built

Built a failure investigation assistant that assembled event timelines, summarized likely contributing causes, and linked evidence across historian data and maintenance records.

Operating context

The group already collected data across multiple systems, but post-incident analysis was still slow because the information lived in different tools and in different time formats. Valuable evidence existed, but no one could pull it together quickly.

Key constraints

  • The assistant could not imply a root cause without showing the timeline and evidence behind it.
  • Historian, alarm, and maintenance data had inconsistent clocks and required normalization.
  • Operators needed to challenge conclusions and feed corrections back into the investigation record.

What we built

  1. Temporal event stitching

    Aligned alarms, sensor anomalies, maintenance actions, and technician notes into one usable incident timeline.

  2. Evidence-linked reasoning

    Generated failure summaries that stayed explicitly tied to the timeline rather than free-form speculation.

  3. Case learning loop

    Captured engineer corrections and confirmed causes so future investigations reflected plant-specific patterns.

Delivery path

  1. Incident library build

    Collected recent downtime cases and normalized the data needed to reconstruct them reliably.

  2. Engineer review

    Used reliability engineers to challenge the system on ambiguous incidents and recurring nuisance alarms.

  3. Plant rollout

    Released by line and failure class so teams could compare performance against their current investigation process.

Why it mattered

Diagnosis got faster because engineers stopped spending their first hours reconstructing the event chain. They still owned the conclusion, but the evidence assembly step became materially lighter and more consistent.

Implementation notes

  • Time alignment across systems was a bigger engineering challenge than the summarization layer.
  • Engineers trusted the assistant only after it showed the exact events that shaped the recommendation.
  • A plant-specific incident library improved usefulness faster than generic prompt tuning.

Start a Similar Plan

Start a Similar Plan