Elevon in ForbesThe moment fault events cross the threshold you set, the system calls the on-call operator and wakes them, with a summary of what is happening. No one has to watch a dashboard through the night.
Client
Our Telco client
Industry
Telco
Solution
Custom AI automation (Elevon suite)
Deployment
Production
Below the threshold, silence. Above it, the phone rings. No one watches a dashboard all night.

Faults happen at all hours, but the night shift is thin. A single event is usually noise; a cluster of events in a short window can be the early signal of a real incident. Telling the two apart at 3 a.m. depends on someone actively watching monitoring, exactly when attention is lowest.
The safe fallback, paging a human for everything, produces alert fatigue and ignored notifications. The unsafe fallback, waiting for the morning shift, means a cluster of faults can grow for hours before anyone acts.
What was missing was a reliable rule: escalate only when the volume of fault events crosses a threshold that genuinely indicates a problem, and escalate in a way that cannot be slept through: a phone call, not another dashboard badge.
The goal was not more alerts. It was the right alert, delivered in a way no one can miss.
Threshold owned by operations
The system counts fault events over a rolling window and escalates only above N, a threshold the ops team sets and can tune. Below it, nothing interrupts the shift.
Escalation that can't be missed
Above the threshold, the system places an actual phone call to the on-call operator, with a concise summary: how many events, of what type, over what window.
Auditable and calibrated
Every escalation is logged with its trigger. Thresholds and the on-call schedule stay with the client, so behaviour is predictable and tunable, not a black box.
“The point was never to alert on everything. It was to make sure that when enough goes wrong at night, the right person actually wakes up.”
The system consumes the fault-event stream and maintains a rolling count per window. When the count crosses the configured threshold, an escalation agent composes a concise summary, event count, dominant type, time window, and triggers a phone call to the operator currently on call.
The on-call schedule and threshold live in configuration the ops team controls. Every trigger, call, and acknowledgement is logged, so the morning shift sees exactly what happened overnight and whether it was actioned.
What this looks like in practice
Through the night, isolated faults pass without interruption. When fault events cross the threshold inside the window, the on-call operator's phone rings with a summary of what is happening. If they don't acknowledge, the escalation continues. Nothing depends on someone staring at a screen at 3 a.m.
Inputs
AI agents
Action
Output
Illustrative reconstruction of the production suite.
Real output format, recreated with blind sample data.
The shift is left alone below the threshold, and cannot be slept through above it.
Escalation only above the configured fault-event threshold; sub-threshold noise no longer interrupts the shift
Above threshold, an automated phone call reaches the on-call operator, not just a dashboard alert
Time-to-human on a real cluster reduced from “next shift” to minutes
Every trigger and acknowledgement logged for the morning handover
Threshold and on-call schedule tuned by operations without engineering changes
minutes
overnight time-to-response, was next shift
↓ risk
faster incident containment
↓ noise
off-threshold pages removed
How we estimate: the value depends on nightly event volume and the threshold N chosen. Every hour a growing cluster goes unactioned overnight is the cost this prevents. Replace with the operator's real event rates and on-call cost to finalize. The figures are illustrative.
“A cluster of faults at night used to wait for the morning. Now it makes the phone ring.”
Overnight monitoring fails in a predictable way: it depends on human vigilance at the hour when vigilance is weakest. Any fix that still requires someone to watch a screen inherits the same failure mode.
Two decisions drove the outcome. Escalation was tied to a threshold the ops team owns, so the system stays quiet on noise and speaks up only when volume indicates a real problem. And the escalation channel is a phone call rather than a dashboard notification, because a call is the one signal that reliably wakes someone. Threshold and schedule stay with the client, so operations can tune sensitivity without an engineering change.
Let's talk about how Elevon can help your team too.
Book consultationWe use essential and analytics cookies by default to ensure proper functionality and understand site usage. Marketing cookies are off unless you opt in. Privacy Policy