Anomaly detection that understands your Tuesday
Every metric baselined against its own history, including hour-of-day and day-of-week seasonality, so a normal Monday peak is not an incident and a quiet Sunday spike is.
What is Anomaly Detection?
DevOpsArk anomaly detection is the capability that learns the normal behaviour of every metric (including hourly and weekly seasonality), and surfaces statistically significant deviations without hand-written thresholds.
What Anomaly Detection is for
The conditions this module removes. If none of these are familiar, you probably do not need it yet.
- Static thresholds are either too tight, producing noise, or too loose, missing real problems.
- Thresholds are set once at a service launch and never revisited as traffic changes.
- Seasonality is ignored, so every Monday morning peak looks like an incident.
- A metric halving is as significant as it doubling, but only one of them has an alert.
- Cost anomalies are found at the end of the billing period.
Every collected series (infrastructure metrics, application metrics, log pattern rates, deployment frequency and cloud spend) is profiled against its own history with seasonality preserved, so the model knows what this metric does at 09:00 on a Monday and what it does at 03:00 on a Sunday. A deviation is judged against that expectation rather than a fixed number, which catches both directions: a drop in request rate is as interesting as a rise, and is usually more so. Detections are grouped rather than delivered individually, so a single underlying event produces one finding with its affected metrics listed. Each detection carries the expected range, the observed value, the size of the deviation and what else moved at the same time, so an engineer can judge it quickly. Feedback closes the loop: marking a detection as expected teaches the baseline rather than merely silencing the alert.
What Anomaly Detection does
The 7 capabilities that make up Anomaly Detection.
Seasonality-aware baselines
Hour-of-day and day-of-week behaviour is modelled, so normal cycles are not reported as anomalies.
Bidirectional detection
Unexpected drops are surfaced as readily as unexpected rises, which is often where real failures hide.
Grouped detections
Metrics moving together from one underlying event produce one finding, not twelve.
Cost anomaly detection
Spend is baselined like any other series, so an unexpected increase surfaces in a day rather than at month end.
Context on every detection
Expected range, observed value, deviation size and correlated changes attached to each finding.
Feedback that teaches
Marking a detection as expected updates the baseline rather than adding another suppression rule.
Fast adaptation for new services
A newly deployed service gets conservative bounds immediately and tightens them as history accumulates.
How Anomaly Detection fits together
Outcomes
- Detection works without a threshold-tuning project per service.
- Weekly and daily cycles stop generating false positives.
- Silent failures that show as a drop are caught.
- One event produces one finding.
- Feedback improves the system instead of muting it.
Using Anomaly Detection, step by step
The path from connecting a source to getting value, in the order it happens.
- 1Collect
Series arrive from monitoring, logs and billing.
- 2Model
Each series is profiled with hourly and weekly seasonality preserved.
- 3Score
Observations are compared against the expected range in both directions.
- 4Group
Metrics moving together are collapsed into one finding.
- 5Explain and learn
Context is attached, and feedback updates the baseline.
Where teams apply Anomaly Detection
Catch a silent failure
A request rate that drops to a third of normal is surfaced even though nothing crossed an error threshold.
Cover new services automatically
A new deployment is monitored from day one without anyone writing thresholds for it.
Find a cost spike in a day
A spend series that breaks from its trend raises a finding immediately rather than at invoice time.
Notice unusual activity
An unexpected change in access patterns or egress volume surfaces as a deviation.
What Anomaly Detection works with
Named integrations link to their own page. The rest are supported runtimes and formats.
Anomaly Detection: frequently asked questions
The 7 questions teams ask most often before adopting Anomaly Detection.
Anomaly detection identifies measurements that deviate significantly from a metric normal behaviour, learned from its own history, instead of comparing against a fixed threshold someone chose in advance.
A single number cannot be correct for both a Monday morning peak and a Sunday night trough. Set it tight and it produces noise; set it loose and it misses real problems. A baseline that models the metric own cycles avoids that trade-off.
Baselines are built with hour-of-day and day-of-week structure preserved, so the expected range at 09:00 on a Monday differs from the expected range at 03:00 on a Sunday.
It should create fewer. Detections that share an underlying cause are grouped into one finding, and detection replaces the collection of loosely tuned static rules that were generating most of the noise.
Immediately, with conservative bounds derived from comparable workloads. The bounds tighten as the service accumulates its own history, typically becoming reliable within a couple of weekly cycles.
Yes. Cloud spend is treated as a series like any other, so an unexpected increase is surfaced within about a day rather than at the end of the billing period.
The baseline is updated to include that behaviour, so the system learns rather than accumulating suppression rules that hide future genuine deviations.
See Anomaly Detection against your own environment
A 30-minute walkthrough with a platform engineer, not a sales deck. Bring a cluster and a problem.