Alert on symptoms users would notice, not on every cause. Correlate related signals into one incident before notifying. Route by service ownership rather than a generic rota. Suppress disturbance you deliberately caused. Attach the context the responder would gather anyway. Then measure which rules produce action and delete the rest. That last step is the one most teams skip.
What makes a good alert?
A good alert is actionable, urgent and attributable: it describes a condition a human needs to do something about now, it reaches the team that owns the affected service, and it arrives with enough context to begin the investigation. If a page can wait until morning it should be a ticket, and if nobody ever acts on it the rule should be deleted.
Alert on symptoms, not causes
The most reliable way to reduce alert volume without reducing coverage is to page on the small number of conditions that mean users are affected, and treat everything else as diagnostic context.
High CPU on one node is a cause. It might produce user impact or it might be entirely fine. Elevated error rate on the checkout endpoint is a symptom, and it is always worth someone looking at. Cause-based alerting produces many pages for each real problem; symptom-based alerting produces roughly one.
Correlate before you notify
A single node failure can produce an alert per pod, per service and per synthetic check. Delivering all of them is worse than delivering one, because the responder now performs the grouping the system should have done, under time pressure.
Correlation means signals sharing a root (the same node, the same deployment, the same downstream dependency) collapse into one incident with the contributing alerts listed inside it. The responder sees one thing to work on and can expand it if they want the detail.
Route by ownership, and suppress the expected
Pages should reach the team that owns the affected service, derived from a service catalogue rather than a separately maintained rota mapping, which is always slightly out of date and never for the service you needed.
Suppression matters just as much. A rolling update causes pod restarts by design. A maintenance window causes hosts to go unreachable by design. If the platform knows a rollout or a window is in progress, the disturbance it causes should not page anyone, while unrelated alerts stay active.
Attach the context the responder would gather anyway
Almost every investigation starts with the same four questions: what deployed recently, what else is alerting, what do the logs say, and has this happened before. All four are answerable by the platform before the notification is sent.
- Recent deployments and configuration changes to the affected service.
- Correlated events and other signals that moved in the same window.
- The log pattern that matches, with its rate compared against baseline.
- When this alert last fired and what resolved it.
- A link to the runbook, if one exists.
Prune with evidence
Alert rules accumulate because adding one is easy and deleting one feels risky. The result is a rule set nobody trusts and nobody prunes, defended by the argument that it might catch something one day.
Replace the argument with data. For each rule, record how often it fires, how often it is acknowledged with no action taken, and whether it has ever fired at all. A rule that has fired two hundred times and produced zero actions is not catching anything; it is training people to ignore pages. Delete it.
| Metric | What it tells you | Healthy direction |
|---|---|---|
| Pages per on-call shift | Load on the responder | Low enough that each is read carefully |
| Actionable rate | Share of pages that led to an action | High, most pages should matter |
| Out-of-hours share | Sustainability of the rota | Low, and trending down |
| Rules never fired in 12 months | Dead weight | Reviewed and removed |
| Time to first meaningful action | How much context the page carried | Short, context should be attached |
Key takeaways
- Page on user-visible symptoms; keep causes as diagnostic context.
- Correlate related signals into one incident before notifying anyone.
- Route from a live service catalogue rather than a hand-maintained rota mapping.
- Suppress the disturbance you deliberately caused, and nothing else.
- Attach recent changes, correlated signals and matching log patterns to the page itself.
- Measure the actionable rate per rule and delete rules that never produce a decision.
Frequently asked questions
Alert fatigue is the loss of responsiveness that follows from receiving too many low-value alerts. It is dangerous precisely because it is rational: when most pages do not matter, ignoring pages is the efficient strategy, and the one that does matter gets ignored with the rest.
Few enough that each is read properly. Many teams aim for a handful per shift and treat a consistently higher number as a signal to fix the rules or the system, not the rota.
Rarely as a page. High CPU may or may not be affecting users. Alert on the user-visible symptom (latency, error rate, saturation of a queue), and keep CPU as context you look at once you are already investigating.
Alerting on conditions that indicate user impact (error rate, latency, availability), rather than on the underlying causes. It produces roughly one alert per real problem instead of one per contributing factor.
Correlate related signals, suppress deliberately caused disturbance, move cause-based rules to diagnostic context, and delete rules with a zero actionable rate. Each of these reduces volume without reducing the set of real problems detected.