Glossary

Distributed tracing

Distributed tracing follows one request across every service it touches, recording the duration and outcome of each step so you see where time went.

Short answer

What is distributed tracing?

Distributed tracing follows one request across every service it touches, recording the duration and outcome of each step so you see where time went.

Explained two ways

Plain and technical

In plain terms

When one user request passes through six services, tracing shows you the whole journey and how long each stage took, so you can find which one was slow.

Technically

A trace is a tree of spans sharing a trace identifier, propagated between services through request headers. Each span records a start time, duration, status and attributes. Tail-based sampling (deciding whether to keep a trace after it completes) retains slow and failing requests while discarding routine traffic, which head-based sampling cannot do.

Example

What it looks like in practice

A checkout request takes 4.2 seconds. The trace shows 3.9 of those in a single inventory service span, which in turn made 40 sequential database calls, an N+1 query pattern invisible from metrics alone.
Keep going

More definitions

See these concepts in a running system

A 30-minute walkthrough against your own infrastructure rather than a slide about the theory.