Observability

What is observability, and how is it different from monitoring?

A definition of observability that does more work than "the three pillars": what property you are actually trying to obtain, and how to tell whether you have it.

DevOpsArk EngineeringEngineering team, DevOpsArkPublished 4 February 2026 · Updated 19 July 20268 min read
TL;DR

Observability is the property of being able to explain what a system is doing from the data it already emits, including for failures nobody anticipated. Metrics, logs and traces are the usual means, not the definition. The practical test is whether you can answer a new question about production without shipping new code, if every novel incident requires adding instrumentation first, the system is not observable.

Short answer

What is observability?

Observability is the property of a system that allows its internal state to be understood from the signals it emits, including for failure modes that were not anticipated in advance. It is typically achieved with metrics, logs and traces, but the defining characteristic is whether you can answer a new question about production behaviour without first adding new instrumentation.

A definition that is actually useful

The word comes from control theory, where a system is observable if its internal state can be determined from its outputs. Carried into software, the useful version is this: can you explain why the system did something, using data it is already producing, for a question you had not thought of in advance?

That last clause is what separates it from monitoring. Monitoring is about known unknowns. You decided in advance that CPU above 90% matters, so you watch for it. Observability is about unknown unknowns: the failure nobody predicted, which by definition has no dashboard.

Why "three pillars" undersells it

Metrics, logs and traces are the standard framing and they are a reasonable inventory of signal types. The framing misleads in two ways. It suggests that having all three means you are done, and it says nothing about the property that actually matters, which is whether they connect.

Three separate systems with three identity models are not observability. They are three data sources and a manual join performed by a tired engineer at 3am. The value appears when a latency metric offers you the traces from that window, the trace offers the logs from each span, and the log line tells you which deployment introduced it.

SignalGood atBad at
MetricsCheap, aggregate, long retention, detectionExplaining a single request; high cardinality gets expensive fast
LogsDetail about a specific moment, arbitrary contextVolume, cost, and being findable when you need one line in millions
TracesWhere time went across services, causalityCoverage (only instrumented paths appear), and sampling decisions

A test you can run this week

Take a genuine question from your last incident, something specific, such as "were the slow requests concentrated in one customer, one region or one code path?" Then try to answer it from data you already have.

  • If you can answer it in a few minutes, the system is observable for that class of question.
  • If you can answer it after an hour of cross-referencing, you have the data but not the connections.
  • If you cannot answer it without deploying new instrumentation, you have monitoring, not observability.

Most teams find they are in the middle category, and the fix is usually not more data. It is a consistent identity model so the data they already collect can be joined.

The cost problem, and the wrong answer to it

Observability spend grows superlinearly with traffic, and the reflex response is to reduce it by sampling more aggressively or dropping log classes. Both work on the bill and both reduce the property you were paying for.

The better levers are cardinality and sampling strategy. High-cardinality labels (user identifiers, request identifiers, full URL paths on metrics) drive most metric cost, and are usually better carried on traces and logs where they belong. And tail-based sampling keeps slow and failing requests in full while discarding routine successful traffic, which is the opposite of what head-based sampling does.

The signal you drop is the one you needed. Head-based sampling at 1% keeps a random 1% of traces. During an incident affecting 0.5% of requests, you will have almost none of the traces that matter. Tail-based sampling decides after the request completes, so it can keep the ones that failed.

Key takeaways

  • Observability is the ability to answer new questions about production without adding instrumentation first.
  • Monitoring covers known unknowns; observability covers the failures nobody predicted.
  • Three signal types with three identity models is not observability. The connections are the property.
  • Test it against a real question from your last incident.
  • Control cost through cardinality and tail-based sampling, not by discarding signal.

Frequently asked questions

ObservabilityMonitoringSRE

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.