Observability: follow a symptom to its cause without changing tools
Metrics, logs and traces attached to the same service model, so a latency spike leads to the slow span, and the slow span leads to the log line.
What is Observability?
DevOpsArk observability is the module that joins metrics, logs and distributed traces to a single service model so engineers can move from a symptom to its cause across services, clusters and clouds without switching tools.
What Observability is for
The conditions this module removes. If none of these are familiar, you probably do not need it yet.
- Metrics are in one system, logs in another and traces in a third, and the identifiers do not match between them.
- A slow request is visible at the edge but invisible at the seventh service it touches.
- Instrumentation is inconsistent, so half the estate produces traces and half does not.
- Investigations end at "the database looked busy" because nothing connects the query to the request.
- Cardinality explodes, the bill grows, and sampling is turned up until the signal disappears.
Observability ingests metrics, logs and traces (OpenTelemetry natively, plus the common exporters), and attaches every signal to the service and environment it came from using the same identity model as the application catalogue. Because the three signal types share that identity, a metric view offers the traces from the same window, a trace offers the logs emitted during each span, and a log line offers the trace it belongs to. The service dependency map is drawn from observed trace paths, which means it reflects what actually calls what rather than what an architecture diagram claims. Sampling is tail-based where it matters, so slow and failing requests are retained even when the bulk of successful traffic is not.
What Observability does
The 7 capabilities that make up Observability.
Distributed tracing
End-to-end request traces across services, with span timing, errors and attributes, ingested via OpenTelemetry.
Signal correlation
Move from metric to trace to log and back, because all three share the same service and request identity.
Observed service map
The dependency graph is drawn from real trace paths, so it reflects the system as it runs today.
Tail-based sampling
Retain the slow and failing requests that matter for diagnosis rather than a uniform random slice.
Exemplar-linked metrics
A point on a latency histogram links to an actual trace that produced it.
Cardinality control
High-cardinality labels are surfaced with their cost, so the bill is managed deliberately rather than by turning sampling up.
Instrumentation coverage
See which services emit traces, which emit only metrics, and which are dark, so instrumentation gaps are a work list.
How Observability fits together
Outcomes
- Investigations reach a cause instead of stopping at a symptom.
- The dependency map reflects the running system rather than the diagram.
- Slow and failing requests are always available, even under aggressive sampling.
- Instrumentation gaps are visible and can be closed deliberately.
- Telemetry cost is managed by cardinality rather than by discarding signal.
Using Observability, step by step
The path from connecting a source to getting value, in the order it happens.
- 1Instrument
Emit OpenTelemetry from applications, or use auto-instrumentation and mesh telemetry where code changes are not practical.
- 2Ingest and normalise
Signals are attached to services and environments using one identity model.
- 3Sample deliberately
Tail-based sampling retains slow and failing requests in full.
- 4Correlate
Metrics, traces and logs cross-link on the same request and service identity.
- 5Investigate
Follow a symptom to a span to a log line, with Ark summarising the path.
Where teams apply Observability
Trace a latency regression to a span
Follow the p99 spike to the trace, the trace to the slow span, and the span to its log lines.
Find an N+1 query
Read the span waterfall for a slow endpoint and see the repeated database calls.
Close instrumentation gaps
Work through the list of services that emit no traces rather than discovering the gap mid-incident.
Control telemetry cost
Identify the labels driving cardinality and reduce them without losing diagnostic capability.
What Observability works with
Named integrations link to their own page. The rest are supported runtimes and formats.
Observability: frequently asked questions
The 9 questions teams ask most often before adopting Observability.
Observability is the property of a system that lets you understand its internal state from the signals it emits (usually metrics, logs and traces) including for failures you did not anticipate. Monitoring tells you that something is wrong; observability lets you work out why.
Monitoring evaluates predefined signals against expected ranges and alerts when one leaves its range. Observability lets you ask new questions of the system after the fact. Monitoring answers "is it broken"; observability answers "why is it broken".
Metrics are numeric measurements over time, cheap to store and good for detection. Logs are discrete timestamped records, good for detail about a specific moment. Traces follow a single request across services and show where its time was spent. Diagnosis usually needs all three.
Kubernetes observability is the collection and correlation of cluster, workload and application signals (pod metrics, container logs, Kubernetes events and request traces), so you can explain a workload behaviour rather than only observe that it is unhealthy.
Yes. OpenTelemetry is the primary ingestion path for metrics, logs and traces, over OTLP. Existing Prometheus, Loki and Jaeger deployments can also be read rather than replaced.
Tail-based sampling decides whether to keep a trace after the request has finished, so slow and failing requests can be retained in full while most successful ones are discarded. Head-based sampling decides at the start and therefore keeps a random slice, which usually misses the traces you needed.
From observed trace paths, which service actually called which, and how often. It therefore reflects the system as deployed rather than as documented, and it updates as the system changes.
By surfacing the labels and attributes driving cardinality with their cost attached, and by sampling on the tail so volume is reduced without losing the diagnostically valuable traces.
No. DevOpsArk can read from existing Prometheus and Loki deployments and add cross-signal correlation, the service map and retention on top. Replacing them is an option, not a requirement.
See Observability against your own environment
A 30-minute walkthrough with a platform engineer, not a sales deck. Bring a cluster and a problem.