Service Impact Detection: an AI network analysis tool that went into production in the Network Operations Center

In 2024 I led a team of nine that won the company hackathon with Service Impact Detection (SID). SID decided whether customers were likely affected. It gave the Network Operations Center (NOC) associate on the ticket a summary and prioritized troubleshooting steps.

Summary

Role
Lead Continual Service Improvement Engineer, UScellular, Jan 2021 – Aug 2025. Hackathon team lead, 2024.
Team
A team of nine
Results
Won the 2024 company hackathon. SID went into production in the NOC. Troubleshooting got faster. Several major outages never developed.
Tools and data
Locally hosted open LLM · Utilization and traffic deviations on core network interfaces · 7-day baselines

Problem

  • The NOC associate on the ticket needed to know whether customers were likely affected and what to troubleshoot first.
  • Whether customers were likely affected could be read from utilization and traffic deviations on core network interfaces.
  • Some network problems could develop into major outages.

What I did

  • I led a team of nine at the 2024 company hackathon. We built SID, an AI network analysis tool that ran on a locally hosted open LLM.
  • SID read utilization and traffic deviations on core network interfaces against 7-day baselines. From that, it decided whether customers were likely affected.
  • The NOC associate on the ticket got a summary and prioritized troubleshooting steps.

What changed

  • SID went into production in the NOC.
  • Troubleshooting got faster. Several major outages never developed.

How SID worked

  1. Interface metrics
  2. 7-day baselines
  3. Impact decision
  4. Summary and prioritized troubleshooting steps for the NOC associate on the ticket
The diagram has four stages. SID took utilization and traffic deviations on core network interfaces. It read them against 7-day baselines. It decided whether customers were likely affected. Then it gave the NOC associate on the ticket a summary and prioritized troubleshooting steps.

Related case study: API monitoring wired to PagerDuty. Change validation time fell 85%.

If you're hiring for this kind of work, send a message.