Observability

Logs, metrics and traces: which signal answers which question

What each telemetry type is genuinely good at, what it costs, and how to decide where a given piece of information belongs.

DevOpsArk EngineeringEngineering team, DevOpsArkPublished 2 April 2026 · Updated 25 July 20267 min read
TL;DR

Metrics are cheap aggregates for detection and trends. Logs are expensive detail for a specific moment. Traces show where time went across services. The most common mistake is putting high-cardinality information on metrics, where it is ruinously expensive, instead of on traces and logs, where it belongs.

Short answer

What is the difference between logs, metrics and traces?

Metrics are numeric measurements aggregated over time, cheap to store and good for detecting change and analysing trends. Logs are discrete timestamped records carrying arbitrary detail about a specific event. Traces follow one request across every service it touches and show where its time was spent. Detection usually starts with metrics; explanation usually needs traces and logs.

What each is genuinely for

Metrics

A metric is a number over time with a small set of labels. Storage cost is roughly proportional to the number of distinct label combinations, not to traffic, which is what makes them cheap, and also what makes them explode if you attach a user identifier. Use them for detection, for trends, for capacity, and for anything you want to retain for a year.

Logs

A log is a record of something that happened, with whatever context the author chose to include. Cost scales with volume, so they get expensive at scale. Use them for the detail that makes a specific event explicable: the exact error, the parameters, the decision the code took.

Traces

A trace follows one request through every service it touches, recording the duration and outcome of each span. It is the only signal that answers "where did the time go" for a distributed request, which is why it is indispensable once you have more than a handful of services.

The cardinality mistake

The single most expensive error in telemetry is putting high-cardinality information on metrics. A counter labelled by user identifier creates one series per user. A histogram labelled by full URL path creates one series per distinct path, and if paths contain identifiers, that is unbounded.

# Expensive: one series per user, per path, per status.
http_requests_total{user_id="8f2c1ad", path="/orders/99213", status="200"}

# Cheap: bounded label set. The identifiers live on the trace and the log.
http_requests_total{route="/orders/:id", method="GET", status="200"}

The information is not lost; it moves to where it costs less. Traces and logs carry arbitrary attributes naturally, and an exemplar link from a metric bucket to a real trace gets you from the aggregate to the specific request without paying metric prices for identifiers.

A decision rule

If you want to knowUse
Is the error rate rising?Metrics
Is this worse than last Tuesday?Metrics with retained history
Which service in the chain is slow?Traces
Why did this particular request fail?Traces, then the logs from the failing span
What exactly did the code do at 03:12?Logs
Has this error ever occurred before?Logs with pattern extraction
How much headroom is left?Metrics

Make them join

The value of all three multiplies when they share identity. Emit trace and span identifiers into your logs, attach exemplars to metric histograms, and label everything with the same service and environment names your inventory uses.

Done consistently, an investigation becomes a path rather than three searches: the metric shows the latency change, an exemplar leads to a slow trace, the trace shows which span consumed the time, and the span links to the log lines emitted while it ran. Done inconsistently, you have three expensive systems and a manual join.

Key takeaways

  • Metrics for detection and trends, logs for detail, traces for where the time went.
  • Metric cost scales with distinct label combinations, so identifiers do not belong on metrics.
  • Move high-cardinality information to traces and logs, and link back with exemplars.
  • Emit trace and span identifiers into logs so the signals join.
  • Consistent service and environment naming across all three is what turns three tools into one investigation.

Frequently asked questions

ObservabilityTelemetryOpenTelemetry

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.