Operate

Observability: follow a symptom to its cause without changing tools

Metrics, logs and traces attached to the same service model, so a latency spike leads to the slow span, and the slow span leads to the log line.

Short answer

What is Observability?

DevOpsArk observability is the module that joins metrics, logs and distributed traces to a single service model so engineers can move from a symptom to its cause across services, clusters and clouds without switching tools.

Why it matters

What Observability is for

The conditions this module removes. If none of these are familiar, you probably do not need it yet.

  • Metrics are in one system, logs in another and traces in a third, and the identifiers do not match between them.
  • A slow request is visible at the edge but invisible at the seventh service it touches.
  • Instrumentation is inconsistent, so half the estate produces traces and half does not.
  • Investigations end at "the database looked busy" because nothing connects the query to the request.
  • Cardinality explodes, the bill grows, and sampling is turned up until the signal disappears.
How it works

Observability ingests metrics, logs and traces (OpenTelemetry natively, plus the common exporters), and attaches every signal to the service and environment it came from using the same identity model as the application catalogue. Because the three signal types share that identity, a metric view offers the traces from the same window, a trace offers the logs emitted during each span, and a log line offers the trace it belongs to. The service dependency map is drawn from observed trace paths, which means it reflects what actually calls what rather than what an architecture diagram claims. Sampling is tail-based where it matters, so slow and failing requests are retained even when the bulk of successful traffic is not.

Capabilities

What Observability does

The 7 capabilities that make up Observability.

Distributed tracing

End-to-end request traces across services, with span timing, errors and attributes, ingested via OpenTelemetry.

Signal correlation

Move from metric to trace to log and back, because all three share the same service and request identity.

Observed service map

The dependency graph is drawn from real trace paths, so it reflects the system as it runs today.

Tail-based sampling

Retain the slow and failing requests that matter for diagnosis rather than a uniform random slice.

Exemplar-linked metrics

A point on a latency histogram links to an actual trace that produced it.

Cardinality control

High-cardinality labels are surfaced with their cost, so the bill is managed deliberately rather than by turning sampling up.

Instrumentation coverage

See which services emit traces, which emit only metrics, and which are dark, so instrumentation gaps are a work list.

Architecture

How Observability fits together

Instrumentation
OpenTelemetry SDKAuto-instrumentationExportersService mesh
Ingest
OTLP collectorTail samplerLabel normaliser
Observability
MetricsLogsTracesService map
Investigation
Correlated viewsArk analysisAlertingSLOs
Observability architecture within the DevOpsArk control plane.

Outcomes

  • Investigations reach a cause instead of stopping at a symptom.
  • The dependency map reflects the running system rather than the diagram.
  • Slow and failing requests are always available, even under aggressive sampling.
  • Instrumentation gaps are visible and can be closed deliberately.
  • Telemetry cost is managed by cardinality rather than by discarding signal.
How to use it

Using Observability, step by step

The path from connecting a source to getting value, in the order it happens.

  1. 1
    Instrument

    Emit OpenTelemetry from applications, or use auto-instrumentation and mesh telemetry where code changes are not practical.

  2. 2
    Ingest and normalise

    Signals are attached to services and environments using one identity model.

  3. 3
    Sample deliberately

    Tail-based sampling retains slow and failing requests in full.

  4. 4
    Correlate

    Metrics, traces and logs cross-link on the same request and service identity.

  5. 5
    Investigate

    Follow a symptom to a span to a log line, with Ark summarising the path.

Use cases

Where teams apply Observability

SRE

Trace a latency regression to a span

Follow the p99 spike to the trace, the trace to the slow span, and the span to its log lines.

Application team

Find an N+1 query

Read the span waterfall for a slow endpoint and see the repeated database calls.

Platform engineering

Close instrumentation gaps

Work through the list of services that emit no traces rather than discovering the gap mid-incident.

FinOps

Control telemetry cost

Identify the labels driving cardinality and reduce them without losing diagnostic capability.

Supported technologies

What Observability works with

Named integrations link to their own page. The rest are supported runtimes and formats.

Do not see your stack? DevOpsArk works over standard interfaces: the Kubernetes API, OCI images, OpenTelemetry and cloud provider APIs, so most environments are supported without a bespoke connector. Ask us about yours.
FAQ

Observability: frequently asked questions

The 9 questions teams ask most often before adopting Observability.

See Observability against your own environment

A 30-minute walkthrough with a platform engineer, not a sales deck. Bring a cluster and a problem.