Six plays that move an operations team from alert noise to one narrow piece of automated remediation, run in sequence rather than in parallel. Each play names the owner, the trigger, the exit criteria and what breaks if you skip it, so the quarter ends with five numbers you can defend in a budget review.
The trap most operations groups walked into: ingest every log, metric and alert on day one including the rules nobody has read since 2021, let anomaly detection run across the lot until it produces a second stream of alerts, then arrive at month six with no baseline and report success as alerts suppressed, which is a number you can move by breaking things.
What the teams who actually get the dividend do: delete the dead alert rules before any vendor conversation, standardise attributes so correlation has something real to join on, and only then let a model group storms, with every play carrying a before-number captured using the query you will re-run afterwards, a re-measure date in the calendar and a rollback that has been tested once.
A play is finished when a stated condition holds, not when the sprint ends. Play two exits above 90% of production services carrying the required attributes, with one real request traced across three services and a queue. Until that holds, play three does not start. The gate is what stops a team running four plays at once and learning nothing from any of them.
No play ships without a number captured before anything changes, using the identical query you will run at the re-measure date. Pages per engineer per week. The member count of the largest storm last quarter. Median time to first useful signal. Skip it and month six becomes a debate about impressions, which the vendor wins because they brought a chart.
The remediation you automate first gets a precondition check that aborts on anything it does not recognise, because refusing to act is the correct default. It also gets one scoped credential: one namespace, one verb set, time bound. Never the platform admin role. And a kill switch any on-call engineer can hit alone, tested by someone who did not build it.
Weeks one and two, owned by the on-call lead. Export every rule from every system into one sheet, join it to 90 days of firing history and to the incident record, then delete whatever fired and was never actioned. Deleted, with the diff in version control, not muted. Publish the noise ratio per team.
Weeks three to six, owned by the platform team. Pin an OpenTelemetry semantic convention version and require service.name, service.version, deployment.environment and an owning team on every metric, log and span. Rename non-conforming attributes at the collector so one team's rollout does not block the rest. Fail the build after the deprecation window.
Weeks seven to ten, owned by SRE tooling with delivery. Group on same service, same alert name, overlapping window, same deploy, and add clustering only for the residue, scored against labelled incident history before it touches a pager. Emit a change event for every deploy, flag flip and config change, then join it to alerts.
Weeks eleven to thirteen, owned by a named engineer. Pick by frequency times runbook determinism, write the runbook as code, run it in shadow for two weeks with a human approving each proposal, then promote behind a rate limit and a kill switch. Ship the first monthly report with the before column filled in.
No, three is enough. Plays one, two and three remove most of the pain: clean alert hygiene, conforming telemetry, working deduplication. A team that stops there has done it without buying anything it cannot switch off. Correlation and remediation are worth more later on that foundation than they are worth now on the current one.
No, it is an argument about order. Machine learning genuinely deduplicates alert storms, correlates across deploys and topology, and drives narrow remediation for failures with a known shape. All three need telemetry that joins. Point a platform at an estate nobody has counted and you have bought a second alert stream at ingest pricing.
Postpone it, not cancel. You cannot scope a tool against an estate you have not counted, and play one costs a day. Walk in with the alert inventory, the noise ratio per team and the member count of the largest storm last quarter, and the conversation moves from a feature list to a scoped pilot.
Treat it sceptically. Root cause explained is statistical proximity dressed as causation, and it is confidently wrong on the novel incidents that actually hurt.
Correlation is still worth having because it shortens the search, so measure median time to first useful signal rather than whether the tool named the right thing.
Heads of SRE and platform leads who own an on-call rota and a budget conversation. It assumes you have alert fatigue, a vendor somewhere in the pipeline, and no baseline anybody could defend in a review. Bring the on-call lead, the platform team and one named engineer, because every play has an owner.
Drop your details and we'll send A Third Of Your Alert Rules Fire And Nobody Acts. AIOps Will Not Fix That. straight to your inbox - no spam, unsubscribe anytime.
Secondary: Book a 30-minute operations review (logiciel.io)
Download the playbook