DevOps

DevOps automation: what to automate, in what order

A practical sequence for automating a delivery path, which step gives the most back first, which automation tends to be regretted, and how to tell when a stage is genuinely done.

DevOpsArk EngineeringEngineering team, DevOpsArkPublished 26 February 2026 · Updated 30 July 20268 min read
TL;DR

Automate in the order the work actually hurts: build and packaging first because it runs on every change, then deployment and rollback because that is where the risk sits, then verification, then operations. The automation people regret is usually the kind that hides a decision: a script that picks a version, resolves a conflict or suppresses an error without telling anyone.

Short answer

What should you automate first in DevOps?

Automate build and packaging first. It runs on every single change, so it has the highest repetition, and its inconsistencies propagate into every later stage. Deployment and rollback come second because that is where the risk concentrates, followed by automated verification and then routine operational tasks.

Order by repetition times pain

The useful heuristic is to rank candidates by how often the task happens multiplied by how much damage it does when performed badly. That ranking almost always puts build and packaging first: it happens on every commit, and an inconsistent build produces problems that surface much later, in a different stage, where they are expensive to diagnose.

It also explains why teams that start with the most visibly painful task (usually a quarterly release process) often see less benefit than expected. A quarterly task automated saves four events a year.

Stage one: build and packaging

The goal is that the same commit produces the same artifact regardless of where it is built. That means pinned base images, resolved lock files, no build-time network calls to anything mutable, and a recorded set of inputs.

This is also the natural place to enforce standards, because it is the one stage every change passes through. Container hardening, dependency scanning and SBOM generation all cost nothing extra here and are expensive to retrofit later.

A done test for this stage. Pick a commit from three months ago and rebuild it. If the resulting image digest differs, or the build fails because something upstream moved, the stage is not finished.

Stage two: deployment and rollback

Most teams automate deployment and stop, leaving rollback as a runbook. This is backwards: the automated part happens when everyone is calm, and the manual part happens under pressure at an inconvenient hour. Rollback is the half that benefits more from automation.

Automated rollback needs a definition of healthy that the system can evaluate, typically success rate and latency compared against the version being replaced. Once that exists, the deployment strategy becomes a per-environment configuration choice rather than a safety mechanism.

Stage three: verification

Verification means tests, scans and policy checks placed where their feedback is still cheap to act on. The ordering principle is to fail fast on the check most likely to catch the problem: linting and unit tests before integration tests, dependency scanning before image build, image scanning before publication.

The trap here is flaky tests. A test suite that fails intermittently trains people to retry rather than read, and once retrying is the reflex, the suite has stopped providing information. Track flakiness explicitly and treat a flaky test as a defect rather than an inconvenience.

Stage four: operations

Certificate renewal, patching waves, backup verification, cost cleanup, right-sizing and access review are all repetitive, well-defined and routinely deferred. They also tend to be the tasks whose omission causes an outage months later, which makes them good automation candidates even though none feels urgent.

Incident response is a partial exception. Automating data gathering (recent deployments, correlated alerts, matching log patterns) is almost pure benefit. Automating the decision about what to do is where teams get uncomfortable, and reasonably so. The approach that works is agents that investigate and propose, with a person approving.

The automation people regret

  • Anything that silently resolves an ambiguity. A script that picks a version when two are plausible will pick wrong eventually, quietly.
  • Retries without limits or logging, which convert a clear failure into an intermittent one.
  • Automation with no dry-run mode, which cannot be safely tested and therefore is not.
  • A pipeline that automatically merges or deploys on a schedule regardless of what changed.
  • Suppression rules added to silence an alert, which accumulate until the alerting system reports nothing.

The common thread is that all of these remove information. Good automation removes work and keeps the information; bad automation removes both, and you only discover which kind you built during an incident.

Key takeaways

  • Rank automation candidates by frequency multiplied by the damage of doing them badly.
  • Build and packaging comes first because it runs on every change and its flaws surface downstream.
  • Automate rollback, not just deployment: the manual half is the one that happens under pressure.
  • Place verification where its feedback is still cheap, and treat flaky tests as defects.
  • Automate incident data gathering; keep the decision with a person.
  • Be suspicious of any automation that resolves ambiguity silently.

Frequently asked questions

DevOpsAutomationCI/CD

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.