DevOps is a way of working in which the people who build software also carry responsibility for running it, supported by automation that makes frequent, small, reversible change safe. It is not a tool, a team name or a job title. The practices that matter are version control for everything, automated build and test, continuous delivery, infrastructure as code, monitoring and observability, and blameless review of failure.
What is DevOps?
DevOps is a set of practices and a working culture that combine software development and IT operations so that teams can build, ship and run software with shared responsibility. Its goal is to shorten the time between a change being written and that change running safely in production, by automating the path between them and by giving the people who write software visibility into how it behaves.
The problem DevOps was invented to solve
For a long time the organisational default was that one group wrote software and another group ran it. The incentives pulled in opposite directions: developers were measured on shipping change, operators on stability, and the fastest route to stability is to ship less. The result was a queue at the boundary, large infrequent releases, and each release carrying enough change that when something broke nobody could say which part caused it.
DevOps is the response: put responsibility for building and running on the same people, and invest in automation so that shipping small changes frequently is safer than shipping large ones rarely. Everything else (the tooling, the practices, the metrics) follows from that.
What DevOps is not
- It is not a job title. A "DevOps engineer" who sits between developers and operations has recreated the boundary the idea was meant to remove.
- It is not a team. A DevOps team that owns all deployment becomes the new queue.
- It is not a tool. Adopting Kubernetes and a CI system without changing who is responsible for production changes very little.
- It is not the same as continuous delivery. Continuous delivery is one practice within it, and probably the most important one, but it is not the whole idea.
- It is not incompatible with a platform team. A platform team that builds a paved path is a good idea; one that executes everyone else deployments is the old operations queue with a new name.
The lifecycle, honestly described
The lifecycle is usually drawn as a loop with eight labels. The labels are fine; the diagram misleads by suggesting the stages are equal in difficulty. In practice most of the pain concentrates in two places: the gap between "it works on my machine" and "it is built the same way every time", and the gap between "it is deployed" and "we know it is healthy".
The loop closes when what you learn from running the system changes what you build next. A team that monitors production but never lets those observations reach the backlog has a lifecycle that is a line, not a loop.
The practices that carry the weight
| Practice | What it means | What it prevents |
|---|---|---|
| Version control for everything | Application code, infrastructure definitions, pipeline configuration and policy all in Git | Undocumented change, and the inability to answer what was different yesterday |
| Automated build and test | Every change built and tested the same way, on every commit | Environment-specific breakage and long feedback loops |
| Continuous delivery | Every change is releasable; releasing is a decision, not a project | Large batches where failure cannot be attributed to a cause |
| Infrastructure as code | Environments described in versioned definitions rather than assembled by hand | Snowflake environments and unreproducible incidents |
| Monitoring and observability | The people who write the code can see how it behaves in production | Learning about failures from users |
| Blameless review | Failure analysis focused on the system, not the individual | Suppressed information, which is the thing that actually causes repeat incidents |
How to tell whether it is working
The four DORA metrics remain the most useful measure, mainly because they pair speed with stability and so resist being gamed in one direction. Deployment frequency and change lead time describe how quickly you can move; change failure rate and time to restore describe whether moving quickly is safe.
The pairing is the point. Deployment frequency alone rewards shipping recklessly. Change failure rate alone rewards shipping nothing. Reported together, they describe a capability rather than an activity.
Where a platform fits
The practical objection to DevOps as originally described is that asking every team to build and operate its own delivery path, observability and security tooling does not scale. Most organisations end up with a platform that provides those as a service, so product teams keep responsibility for their software without each reinventing the machinery.
DevOpsArk is that machinery: build and containerisation, delivery with automatic rollback, monitoring and alerting, security scanning and cost attribution, on one data model, so a team can own its service in production without owning a toolchain.
Key takeaways
- DevOps is shared responsibility for building and running, supported by automation that makes small frequent change safe.
- A DevOps team or a DevOps engineer role usually recreates the boundary the idea was meant to remove.
- Most of the difficulty sits in reproducible builds and in knowing whether a deployment is actually healthy.
- Measure with DORA metrics as a set: speed and stability together resist gaming.
- A platform makes the model scale without each team building its own toolchain.
Frequently asked questions
DevOps means the people who write software also help run it, and that the path from writing a change to it running in production is automated enough that shipping small changes often is safer than shipping large ones rarely.
It should not be. When "DevOps engineer" describes a person who sits between developers and production, the boundary the practice was meant to dissolve has been rebuilt with a new name. Platform engineer is usually the more accurate title for the work involved.
Agile is mainly about how work is planned and delivered up to the point it is written. DevOps is mainly about what happens from there to production and back. They complement each other; Agile without DevOps produces finished work that queues before release.
Deployment frequency, change lead time, change failure rate and time to restore service. They are used together because two describe speed and two describe stability, and optimising either pair alone produces a worse system.
DevOps is a broad set of practices and a culture. Site reliability engineering is a specific implementation, originating at Google, that adds explicit reliability targets, error budgets and a defined split between operational and engineering work. SRE is one way to do DevOps, with more prescription.
No. DevOps predates Kubernetes and works on virtual machines, managed platforms and serverless. Kubernetes helps with certain problems (declarative deployment, scaling, portability), and adds operational complexity of its own.