Observability
Metrics, logs, traces, alerting design and what observability actually means.
What does the Observability cluster cover?
The Observability cluster on the DevOpsArk blog collects 5 articles on metrics, logs, traces, alerting design and what observability actually means.
Everything in Observability
SIEM vs audit log vs observability: three different jobs, constantly confused
SIEM, audit logs and observability all involve data about your systems, but they answer different questions for different people. How to tell them apart.
Alerting that people still read at 3am
How to build an alerting system with a high proportion of actionable pages: correlation, ownership routing, suppression, and deleting the rules that never produce a decision.
Logs, metrics and traces: which signal answers which question
What each telemetry type is genuinely good at, what it costs, and how to decide where a given piece of information belongs.
Monitoring vs observability: a distinction worth keeping
The difference between monitoring and observability, why the distinction is more than marketing, and what each one is actually for.
What is observability, and how is it different from monitoring?
A definition of observability that does more work than "the three pillars": what property you are actually trying to obtain, and how to tell whether you have it.
What the platform does about this
Observability
Metrics, logs and traces in one plane
Log Management
Centralised logs with structure and retention
AI Log Analysis
Find the line that matters
Anomaly Detection
Detection without hand-written thresholds
Monitoring
Infrastructure and application monitoring
Alerting
Alerts that are worth waking up for
ArkApps
Application inventory and lifecycle
Elsewhere on the blog
DevOps
Fundamentals, automation strategy, incident practice and the platform-versus-toolchain question.
Kubernetes
Architecture, monitoring, deployment strategy, multi-cluster operations and troubleshooting.
AI DevOps
Agentic DevOps, AI log analysis, incident response and where automation should stop.
DevSecOps
Container and Kubernetes security, vulnerability management, secrets handling, audit trails and compliance evidence.
Cloud and cost
Multi-cloud operations, cost optimisation, infrastructure drift and infrastructure hygiene.
Want a structured route through this?
Learning tracks arrange these articles into an ordered path with the glossary terms they depend on.