Skip to main content

Fewer alerts, faster answers, a rotation people will hold

Alert volume cut by 78 percent while median time to restore fell from four hours ten minutes to twenty-two minutes across a fleet of 3,400 connected machines.

Industry
Industrial IoT
Location
Stuttgart, Germany
Engagement
9 weeks
Team
2 NordsCode engineers, 3 client engineers
22 minMedian time to restore, down from 4h 10m
Measured across all severity one and two incidents in the six months following handover, compared with the twelve months before the engagement.
78%Fewer alerts reaching an engineer
From an average of 214 alerts per week to 47, of which 11 were severe enough to page. Every remaining alert maps to a customer-visible symptom and carries a runbook link.
3,400Connected machines under a single objective
Telemetry ingestion from the full deployed fleet brought under one availability objective agreed with the operations business, replacing per-component thresholds nobody owned.

What we found when we arrived.

Arcline builds condition monitoring hardware for industrial machinery and sells the data platform behind it as a subscription. Roughly 3,400 machines across 140 customer sites push telemetry into their ingestion pipeline. When ingestion stalled, customers noticed gaps in their dashboards before Arcline did, and the first sign was frequently a support ticket.

There was no shortage of monitoring. Three tools had accumulated over four years, each with its own alerts, and the on-call channel received around 214 alerts per week. Almost none corresponded to anything a customer would notice. The rotation had two willing participants out of a team of nine, and the engineering lead had been paged eleven times in the previous month for a disk threshold on a queue node that auto-recovered every time.

What we did, in the order we did it.

Sequence matters more than tooling on engagements like this one. Each step below existed because the previous one produced something the next one needed.

  1. 01

    Agree what the service owes its customers

    Two weeks of workshops with engineering and the operations business produced three service level objectives in plain numbers, including a freshness target for telemetry that finally made the ingestion delay problem measurable rather than anecdotal.

  2. 02

    Instrument once, portably

    OpenTelemetry across the ingestion pipeline, the API and the processing workers, so a single device message could be traced from the edge gateway through to storage. The stalls turned out to concentrate in one deserialisation path under a specific firmware version.

  3. 03

    Delete more alerts than we created

    Every existing alert was reviewed with the team and either retired, downgraded to a dashboard, or rewritten as a symptom-based alert with an owner and a runbook. Two of the three monitoring tools were switched off, which also removed a EUR 2,100 monthly licence.

  4. 04

    Rehearse an incident before having one

    A deliberate failure injected into the staging ingestion path, run as a real incident with the new severity definitions, escalation path and customer communication template, followed by a blameless review of how the process itself performed.

What changed, and how we know.

Median time to restore across severity one and two incidents fell from four hours ten minutes to twenty-two minutes in the six months after handover. The largest single contributor was tracing: investigations that previously meant correlating timestamps across three tools now start from a trace that shows exactly which stage of the pipeline stalled.

Alert volume fell from 214 per week to 47, with 11 of those severe enough to page a human. The on-call rotation now has seven volunteers out of nine engineers, which the engineering lead describes as the most useful outcome of the engagement. Telemetry costs also fell, because retention tiers and sampling replaced the previous approach of keeping everything at full fidelity for ninety days.

What it was built on

  • OpenTelemetry
  • Grafana Cloud
  • Prometheus
  • Tempo
  • Kubernetes
  • Terraform

Recognise any of this?

Most engagements start with a call describing a situation that sounds a lot like one of these. Tell us yours and we will say plainly whether we can help.