Distributed tracing
Distributed tracing follows one request across every service it touches, recording the duration and outcome of each step so you see where time went.
What is distributed tracing?
Distributed tracing follows one request across every service it touches, recording the duration and outcome of each step so you see where time went.
Plain and technical
When one user request passes through six services, tracing shows you the whole journey and how long each stage took, so you can find which one was slow.
A trace is a tree of spans sharing a trace identifier, propagated between services through request headers. Each span records a start time, duration, status and attributes. Tail-based sampling (deciding whether to keep a trace after it completes) retains slow and failing requests while discarding routine traffic, which head-based sampling cannot do.
What it looks like in practice
Nearby vocabulary
Observability
Observability is the property of a system that allows its internal state to be understood from the signals it emits, including for failures nobody anticipated.
OpenTelemetry
OpenTelemetry is a vendor-neutral standard and set of libraries for generating and exporting metrics, logs and traces, portable across backends.
Cardinality
Cardinality is the number of distinct combinations of label values for a metric, and it is the primary driver of metric storage cost.
How DevOpsArk handles distributed tracing
Articles on this subject
Logs, metrics and traces: which signal answers which question
What each telemetry type is genuinely good at, what it costs, and how to decide where a given piece of information belongs.
What is observability, and how is it different from monitoring?
A definition of observability that does more work than "the three pillars": what property you are actually trying to obtain, and how to tell whether you have it.
More definitions
See these concepts in a running system
A 30-minute walkthrough against your own infrastructure rather than a slide about the theory.