Metrics are cheap aggregates for detection and trends. Logs are expensive detail for a specific moment. Traces show where time went across services. The most common mistake is putting high-cardinality information on metrics, where it is ruinously expensive, instead of on traces and logs, where it belongs.
What is the difference between logs, metrics and traces?
Metrics are numeric measurements aggregated over time, cheap to store and good for detecting change and analysing trends. Logs are discrete timestamped records carrying arbitrary detail about a specific event. Traces follow one request across every service it touches and show where its time was spent. Detection usually starts with metrics; explanation usually needs traces and logs.
What each is genuinely for
Metrics
A metric is a number over time with a small set of labels. Storage cost is roughly proportional to the number of distinct label combinations, not to traffic, which is what makes them cheap, and also what makes them explode if you attach a user identifier. Use them for detection, for trends, for capacity, and for anything you want to retain for a year.
Logs
A log is a record of something that happened, with whatever context the author chose to include. Cost scales with volume, so they get expensive at scale. Use them for the detail that makes a specific event explicable: the exact error, the parameters, the decision the code took.
Traces
A trace follows one request through every service it touches, recording the duration and outcome of each span. It is the only signal that answers "where did the time go" for a distributed request, which is why it is indispensable once you have more than a handful of services.
The cardinality mistake
The single most expensive error in telemetry is putting high-cardinality information on metrics. A counter labelled by user identifier creates one series per user. A histogram labelled by full URL path creates one series per distinct path, and if paths contain identifiers, that is unbounded.
# Expensive: one series per user, per path, per status.
http_requests_total{user_id="8f2c1ad", path="/orders/99213", status="200"}
# Cheap: bounded label set. The identifiers live on the trace and the log.
http_requests_total{route="/orders/:id", method="GET", status="200"}The information is not lost; it moves to where it costs less. Traces and logs carry arbitrary attributes naturally, and an exemplar link from a metric bucket to a real trace gets you from the aggregate to the specific request without paying metric prices for identifiers.
A decision rule
| If you want to know | Use |
|---|---|
| Is the error rate rising? | Metrics |
| Is this worse than last Tuesday? | Metrics with retained history |
| Which service in the chain is slow? | Traces |
| Why did this particular request fail? | Traces, then the logs from the failing span |
| What exactly did the code do at 03:12? | Logs |
| Has this error ever occurred before? | Logs with pattern extraction |
| How much headroom is left? | Metrics |
Make them join
The value of all three multiplies when they share identity. Emit trace and span identifiers into your logs, attach exemplars to metric histograms, and label everything with the same service and environment names your inventory uses.
Done consistently, an investigation becomes a path rather than three searches: the metric shows the latency change, an exemplar leads to a slow trace, the trace shows which span consumed the time, and the span links to the log lines emitted while it ran. Done inconsistently, you have three expensive systems and a manual join.
Key takeaways
- Metrics for detection and trends, logs for detail, traces for where the time went.
- Metric cost scales with distinct label combinations, so identifiers do not belong on metrics.
- Move high-cardinality information to traces and logs, and link back with exemplars.
- Emit trace and span identifiers into logs so the signals join.
- Consistent service and environment naming across all three is what turns three tools into one investigation.
Frequently asked questions
If you want to know how often something happens or how a number trends, emit a metric. If you want to know exactly what happened in one instance, log it. If you need both, do both: they cost differently and answer differently.
Cardinality is the number of distinct combinations of label values for a metric. Each combination is a separate time series with its own storage and memory cost, which is why a label containing user or request identifiers can multiply cost by orders of magnitude.
Structured logging emits records as key-value data, typically JSON, rather than as free text. It makes fields queryable without regular expressions and lets you filter by service, request identifier or user without parsing prose.
An exemplar is a link attached to a metric data point that points at a specific trace which contributed to it. It bridges aggregate and specific: you see a latency spike on a histogram and jump directly to a real slow request from that bucket.
No. Traces show structure and timing; logs carry detail. A span tells you a database call took 4 seconds, and the log line tells you which query it was and what parameters it had.