Observability

Alerting that people still read at 3am

How to build an alerting system with a high proportion of actionable pages: correlation, ownership routing, suppression, and deleting the rules that never produce a decision.

Harshit SengarFounder and platform lead, DevOpsArkPublished 6 May 2026 · Updated 8 August 20268 min read
TL;DR

Alert on symptoms users would notice, not on every cause. Correlate related signals into one incident before notifying. Route by service ownership rather than a generic rota. Suppress disturbance you deliberately caused. Attach the context the responder would gather anyway. Then measure which rules produce action and delete the rest. That last step is the one most teams skip.

Short answer

What makes a good alert?

A good alert is actionable, urgent and attributable: it describes a condition a human needs to do something about now, it reaches the team that owns the affected service, and it arrives with enough context to begin the investigation. If a page can wait until morning it should be a ticket, and if nobody ever acts on it the rule should be deleted.

Alert on symptoms, not causes

The most reliable way to reduce alert volume without reducing coverage is to page on the small number of conditions that mean users are affected, and treat everything else as diagnostic context.

High CPU on one node is a cause. It might produce user impact or it might be entirely fine. Elevated error rate on the checkout endpoint is a symptom, and it is always worth someone looking at. Cause-based alerting produces many pages for each real problem; symptom-based alerting produces roughly one.

The test for a page. Would you want to be woken for this? If the honest answer is no, it is not a page. Make it a ticket, a dashboard or nothing.

Correlate before you notify

A single node failure can produce an alert per pod, per service and per synthetic check. Delivering all of them is worse than delivering one, because the responder now performs the grouping the system should have done, under time pressure.

Correlation means signals sharing a root (the same node, the same deployment, the same downstream dependency) collapse into one incident with the contributing alerts listed inside it. The responder sees one thing to work on and can expand it if they want the detail.

Route by ownership, and suppress the expected

Pages should reach the team that owns the affected service, derived from a service catalogue rather than a separately maintained rota mapping, which is always slightly out of date and never for the service you needed.

Suppression matters just as much. A rolling update causes pod restarts by design. A maintenance window causes hosts to go unreachable by design. If the platform knows a rollout or a window is in progress, the disturbance it causes should not page anyone, while unrelated alerts stay active.

Attach the context the responder would gather anyway

Almost every investigation starts with the same four questions: what deployed recently, what else is alerting, what do the logs say, and has this happened before. All four are answerable by the platform before the notification is sent.

  • Recent deployments and configuration changes to the affected service.
  • Correlated events and other signals that moved in the same window.
  • The log pattern that matches, with its rate compared against baseline.
  • When this alert last fired and what resolved it.
  • A link to the runbook, if one exists.

Prune with evidence

Alert rules accumulate because adding one is easy and deleting one feels risky. The result is a rule set nobody trusts and nobody prunes, defended by the argument that it might catch something one day.

Replace the argument with data. For each rule, record how often it fires, how often it is acknowledged with no action taken, and whether it has ever fired at all. A rule that has fired two hundred times and produced zero actions is not catching anything; it is training people to ignore pages. Delete it.

MetricWhat it tells youHealthy direction
Pages per on-call shiftLoad on the responderLow enough that each is read carefully
Actionable rateShare of pages that led to an actionHigh, most pages should matter
Out-of-hours shareSustainability of the rotaLow, and trending down
Rules never fired in 12 monthsDead weightReviewed and removed
Time to first meaningful actionHow much context the page carriedShort, context should be attached

Key takeaways

  • Page on user-visible symptoms; keep causes as diagnostic context.
  • Correlate related signals into one incident before notifying anyone.
  • Route from a live service catalogue rather than a hand-maintained rota mapping.
  • Suppress the disturbance you deliberately caused, and nothing else.
  • Attach recent changes, correlated signals and matching log patterns to the page itself.
  • Measure the actionable rate per rule and delete rules that never produce a decision.

Frequently asked questions

AlertingSREOn-callObservability

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.