Observability is the property of being able to explain what a system is doing from the data it already emits, including for failures nobody anticipated. Metrics, logs and traces are the usual means, not the definition. The practical test is whether you can answer a new question about production without shipping new code, if every novel incident requires adding instrumentation first, the system is not observable.
What is observability?
Observability is the property of a system that allows its internal state to be understood from the signals it emits, including for failure modes that were not anticipated in advance. It is typically achieved with metrics, logs and traces, but the defining characteristic is whether you can answer a new question about production behaviour without first adding new instrumentation.
A definition that is actually useful
The word comes from control theory, where a system is observable if its internal state can be determined from its outputs. Carried into software, the useful version is this: can you explain why the system did something, using data it is already producing, for a question you had not thought of in advance?
That last clause is what separates it from monitoring. Monitoring is about known unknowns. You decided in advance that CPU above 90% matters, so you watch for it. Observability is about unknown unknowns: the failure nobody predicted, which by definition has no dashboard.
Why "three pillars" undersells it
Metrics, logs and traces are the standard framing and they are a reasonable inventory of signal types. The framing misleads in two ways. It suggests that having all three means you are done, and it says nothing about the property that actually matters, which is whether they connect.
Three separate systems with three identity models are not observability. They are three data sources and a manual join performed by a tired engineer at 3am. The value appears when a latency metric offers you the traces from that window, the trace offers the logs from each span, and the log line tells you which deployment introduced it.
| Signal | Good at | Bad at |
|---|---|---|
| Metrics | Cheap, aggregate, long retention, detection | Explaining a single request; high cardinality gets expensive fast |
| Logs | Detail about a specific moment, arbitrary context | Volume, cost, and being findable when you need one line in millions |
| Traces | Where time went across services, causality | Coverage (only instrumented paths appear), and sampling decisions |
A test you can run this week
Take a genuine question from your last incident, something specific, such as "were the slow requests concentrated in one customer, one region or one code path?" Then try to answer it from data you already have.
- If you can answer it in a few minutes, the system is observable for that class of question.
- If you can answer it after an hour of cross-referencing, you have the data but not the connections.
- If you cannot answer it without deploying new instrumentation, you have monitoring, not observability.
Most teams find they are in the middle category, and the fix is usually not more data. It is a consistent identity model so the data they already collect can be joined.
The cost problem, and the wrong answer to it
Observability spend grows superlinearly with traffic, and the reflex response is to reduce it by sampling more aggressively or dropping log classes. Both work on the bill and both reduce the property you were paying for.
The better levers are cardinality and sampling strategy. High-cardinality labels (user identifiers, request identifiers, full URL paths on metrics) drive most metric cost, and are usually better carried on traces and logs where they belong. And tail-based sampling keeps slow and failing requests in full while discarding routine successful traffic, which is the opposite of what head-based sampling does.
Key takeaways
- Observability is the ability to answer new questions about production without adding instrumentation first.
- Monitoring covers known unknowns; observability covers the failures nobody predicted.
- Three signal types with three identity models is not observability. The connections are the property.
- Test it against a real question from your last incident.
- Control cost through cardinality and tail-based sampling, not by discarding signal.
Frequently asked questions
Monitoring watches predefined signals and alerts when one leaves its expected range. Observability is the ability to investigate behaviour you did not anticipate. You need both: monitoring tells you something is wrong, observability lets you find out why.
Metrics, logs and traces. It is a useful inventory of signal types but an incomplete definition, because having all three in separate systems with different identity models does not make a system observable. The connections between them are what matter.
For a distributed system, effectively yes. Once a request crosses several services, metrics and logs can tell you that something is slow but not where the time went. Tracing is the only signal that answers that directly.
OpenTelemetry is a vendor-neutral standard and set of libraries for generating metrics, logs and traces. Instrumenting with it means the telemetry is not tied to one backend, which is the main defence against being locked into an observability vendor.
There is no universal ratio, but if telemetry spend is growing faster than traffic, the cause is usually cardinality rather than volume. Identify the labels driving it before reducing retention or sampling, because those reductions cost you the capability you are paying for.