Monitoring across clusters, servers, clouds and applications
One metric plane for the whole estate, with enough retained history to answer "was it like this last Tuesday" and an agent that reads the graphs with you.
What is Monitoring?
DevOpsArk monitoring is the module that collects, retains and analyses metrics from Kubernetes clusters, servers, cloud services and applications in a single plane, with baselines and agent-led diagnosis on top of the raw signals.
What Monitoring is for
The conditions this module removes. If none of these are familiar, you probably do not need it yet.
- Metrics live in four systems (the cloud provider, Prometheus, an APM tool and a spreadsheet), and none of them agree.
- Retention is short, so questions about last month cannot be answered at all.
- Dashboards proliferate. There are 300 of them and the six that matter are not obvious.
- A metric is out of range and nobody knows whether that is normal for a Tuesday morning.
- Correlating a spike with a deployment means opening two systems and comparing timestamps by eye.
DevOpsArk collects metrics from the Kubernetes metrics API, node exporters, cloud provider monitoring APIs and application endpoints, and stores them in one time-series plane with retention long enough to be useful for capacity and post-incident work. Each series is baselined against its own history, so "high" is measured against how this workload behaves at this hour on this day rather than against a fixed threshold someone guessed. Deployments, releases and configuration changes are overlaid on the same timeline, which turns the most common investigation (did this start when we shipped something) into a glance. When something is wrong, the Ark agent correlates the metric with pod events, container logs and recent changes and reports what it found as a trace you can check, not a verdict you have to trust.
What Monitoring does
The 8 capabilities that make up Monitoring.
Unified metric collection
Kubernetes, servers, AWS, Azure, Google Cloud and application metrics collected into one plane with a consistent label model.
Behavioural baselines
Each series is compared against its own history by hour and day, so a normal Monday peak does not read as an incident.
Change overlay
Deployments, releases and configuration changes are drawn on the metric timeline, so correlation takes a glance.
Service-scoped views
Dashboards derived from the application catalogue, so each service has one canonical view instead of six abandoned ones.
Retention that answers questions
Long enough history for capacity planning, seasonality and post-incident review, not just a live wall display.
Agent-led diagnosis
Ark correlates the metric with events, logs and recent changes, and shows the steps it took to reach its conclusion.
Saturation and capacity signals
CPU throttling, memory pressure, disk headroom, connection pool exhaustion and node allocatable pressure, tracked as first-class signals.
SLO tracking
Define availability and latency objectives per service and track error budget burn rather than only instantaneous health.
How Monitoring fits together
Outcomes
- One plane replaces four, and the numbers reconcile.
- A metric excursion is judged against how the workload actually behaves, not a guessed threshold.
- "Did this start when we deployed" is answered from one timeline.
- Capacity and seasonality questions become answerable because the history exists.
- Diagnosis starts from a correlated summary rather than from a blank dashboard.
Using Monitoring, step by step
The path from connecting a source to getting value, in the order it happens.
- 1Connect sources
Attach clusters, servers and cloud accounts; collection begins without per-host configuration.
- 2Normalise labels
Signals are mapped to services in the application catalogue so views assemble themselves.
- 3Establish baselines
Each series is profiled against its own history by hour and day.
- 4Overlay change
Deployments and configuration changes are drawn on the same timeline.
- 5Alert and analyse
Deviations raise alerts, and Ark correlates them with events, logs and changes.
Where teams apply Monitoring
Investigate a latency regression
Overlay deployments on the latency series and identify the change that introduced it without switching systems.
Plan cluster capacity
Use retained history and saturation signals to size node pools against real demand rather than peak-day guesswork.
Track an SLO
Define an availability objective for a service and watch error budget burn instead of instantaneous status.
Triage at 3am
Start from the agent correlation between the alerting metric, pod events and the last deployment.
What Monitoring works with
Named integrations link to their own page. The rest are supported runtimes and formats.
Monitoring: frequently asked questions
The 9 questions teams ask most often before adopting Monitoring.
Infrastructure monitoring is the continuous collection and analysis of metrics from servers, clusters, networks and cloud services to determine whether they are healthy and to detect when they stop being healthy. It answers what is happening; observability answers why.
Kubernetes monitoring is the collection of cluster, node, workload and pod signals (CPU and memory against requests and limits, restart counts, scheduling failures, node conditions and control-plane health), so cluster problems are detected before workloads fail.
It reads the Kubernetes metrics API and pod events through the cluster API, correlates them with container logs and deployment history, and baselines each series against its own past behaviour. No in-cluster agent is required, though one is available for higher-frequency sampling.
It can, or it can read from an existing Prometheus. Many teams keep Prometheus as a collector and use DevOpsArk for cross-cluster aggregation, retention, baselining and correlation with deployments and logs.
Monitoring tracks known signals and tells you when one leaves its expected range. Observability is the property of a system that lets you ask new questions about it (usually via metrics, logs and traces together), and diagnose a failure you did not anticipate.
Long enough to support capacity planning, seasonality analysis and post-incident review rather than only a live view. Retention is configurable per resolution, with recent data at full resolution and older data downsampled.
A service level objective is a target for a user-visible property of a service, such as 99.9% of requests succeeding over 30 days. Tracking the error budget it implies is more useful than instantaneous up-or-down status, because it shows how much unreliability remains affordable.
Yes. Application metrics exposed over an endpoint or sent via OpenTelemetry are collected alongside infrastructure signals and attached to the same service record in the catalogue.
Yes. AWS, Azure and Google Cloud metrics are collected into the same plane with a consistent label model, so a single view can span providers.
See Monitoring against your own environment
A 30-minute walkthrough with a platform engineer, not a sales deck. Bring a cluster and a problem.