Monitoring
Monitoring is the practice of collecting predefined signals from a system and alerting when they leave their expected ranges.
What is monitoring?
Monitoring is the practice of collecting predefined signals from a system and alerting when they leave their expected ranges.
Plain and technical
Monitoring means watching the numbers you decided matter (error rates, response times, memory usage), and being told when one of them looks wrong.
Monitoring covers known failure modes: signals are chosen in advance, thresholds or baselines are established, and evaluation runs continuously. Its limitation is that it can only detect conditions someone anticipated. Baseline-relative alerting, which compares a metric against its own historical behaviour including seasonality, generally outperforms static thresholds because a single number cannot be correct for both a weekday peak and a weekend trough.
What it looks like in practice
Nearby vocabulary
Observability
Observability is the property of a system that allows its internal state to be understood from the signals it emits, including for failures nobody anticipated.
Service level objective
A service level objective is a target for a measurable, user-visible property of a service, such as the share of requests that succeed over a period.
Anomaly detection
Anomaly detection identifies measurements that deviate significantly from a metric learned normal behaviour, rather than comparing against a fixed threshold.
How DevOpsArk handles monitoring
Articles on this subject
Monitoring vs observability: a distinction worth keeping
The difference between monitoring and observability, why the distinction is more than marketing, and what each one is actually for.
Kubernetes monitoring: what to measure and why
A practical guide to Kubernetes monitoring: the cluster, node and workload signals that predict failure, the ones that only look useful, and how to alert on them without drowning.
More definitions
See these concepts in a running system
A 30-minute walkthrough against your own infrastructure rather than a slide about the theory.