AI SRE — Alerting & Escalation

It calls until
someone picks up.

Your monitoring fires a webhook. AI SRE opens the alarm and phones the on-call list in order until somebody presses 1 — while an agent correlates the incident, writes the report and runs the reversible fixes your policy allows.

One incident, replayed — the console's own readout, running a script.

The escalation chain

A page is delivered when a person answers it.

Not when a message is sent. The chain walks your levels in order and only a pressed key stops it.

01
Webhook
Monitoring fires. The alarm opens on the issue key.
02
Level 1 · first member
Called. No answer after 25 seconds.
03
Level 1 · second member
Called. Presses 1.
04
Acknowledged
Chain stops. Level 2 is never paged.

If nobody answers, the chain moves to the next level and then loops — under a loop limit, a monthly call budget and a cap on legs per alarm. A recovery webhook cancels the whole thing mid-flight.

What it does

Everything an on-call rota needs, and nothing it doesn't.

It phones people
Twilio calls the on-call list in order and asks for one digit. No app to install, no notification to sleep through.
Ordered levels, looped
Levels run in sequence and members run in sequence inside them, and the chain loops until someone acknowledges or the brakes stop it.
One webhook per channel
New Relic, Prometheus Alertmanager, Uptime Kuma and Grafana are read natively. Anything else is a JSON path map — no code to run.
Deduplicated by issue key
A flapping monitor is one alarm, not forty. The key is scoped per project and channel, so two teams never collide.
Recovery cancels the chain
When monitoring says it is back, the live chain is cancelled — nobody gets called about an incident that is already over.
An agent works beside it
Correlation, the report and the reversible mitigations your policy allows run alongside the page, never in front of it.
A terminal you can trust
Every action streams live, in order, and survives a reload — detection, each call leg, who acknowledged, what the agent did.
A report per alarm
Detection, every leg dialled, who pressed 1, and each autonomous action — written without anyone staying up to write it.
Where alarms come from

Keep the monitoring you already trust.

AI SRE does not probe your services — that is your monitoring's job, and replacing it would mean rebuilding something that already works. It takes the webhook and turns it into a page that reaches a human.

Uptime Kuma is also mirrored, so monitor status and p95 sit next to the alarms without us running a second set of probes.

New Relic
Native reader
Prometheus Alertmanager
Native reader
Uptime Kuma
Native reader + status mirror
Grafana
Native reader
Anything else
JSON path map, previewable
$ curl -XPOST "…/w/$CHANNEL?dry=1"
→ runs the whole line and returns what the phone would say
Guardrails

The rules it will not break at 3 a.m.

Autonomy is useful exactly as far as it is reversible. Past that line, it wakes you.

Nothing closes an alarm but you
A mitigation that worked still leaves the alarm recovering and waiting on a human confirmation.
Reversible actions only
The agent runs what your policy allows and what can be undone — a rollback, a restart. Never something it cannot take back.
Three independent brakes
Loop limit, monthly call budget and legs per alarm. A misconfigured monitor cannot phone your team all night.
No SMS pretending to be a page
The step is recognised and deliberately skipped. Showing an undelivered page as delivered is more dangerous than not sending it.
One digit, no menu
There is no IVR tree at 3 a.m. Press 1 to acknowledge. That is the entire interaction.
No new dependency on the critical path
The chain keeps working with the AI, the knowledge layer and the monitoring mirror all unreachable.
Use cases

The three ways a night goes wrong.

01

The alert nobody saw

Monitoring already caught it at 03:14 — it went to a channel where everyone had muted notifications.

Outcome — a phone that rings until a person answers, then stops.
02

The on-call rota that only exists in a document

Levels and members are stored in order, and the order is what the engine walks — first, second, then the next level.

Outcome — the rota is executable, not a page someone has to remember.
03

The incident nobody wrote up

Detection, each call leg, the acknowledgement and every autonomous action are recorded as they happen.

Outcome — a post-incident report that exists the next morning without anyone drafting it.
Who it's for

For the people the pager actually reaches.

On-call engineer
The one holding the phone

Wakes to forty notifications and still has to work out what broke and when.

Gets — one call, one digit, and an agent that has already correlated the deploy by the time they open a laptop.
Engineering lead
Owns the rota

Cannot prove a page was delivered, and has no write-up unless somebody sacrifices an hour.

Gets — every leg logged, a report per alarm, and escalation that runs by the rules instead of by memory.
CTO / platform owner
Answers for downtime

Paying per seat for a paging tool, and still unsure whether the chain works at 3 a.m.

Gets — one plan per project, a rehearsal mode that runs the whole chain without ringing anyone, and reversible autonomy under policy.

Find out tonight, not tomorrow.

One plan, priced per project. Bring your monitoring and your on-call list — we will run the whole chain without ringing a single phone.

Get a demo