Site reliability engineering
Site reliability engineering applies software engineering to operations, using reliability objectives and error budgets to balance stability and speed.
What is site reliability engineering?
Site reliability engineering applies software engineering to operations, using reliability objectives and error budgets to balance stability and speed.
Plain and technical
SRE treats keeping systems running as an engineering problem rather than a manual one. It sets a specific reliability target, and how much of that target remains unused determines how much risk the team can take with new releases.
SRE introduces service level objectives and the error budget they imply: if a service targets 99.9% availability over 30 days, the permitted unreliability is the budget. When the budget is being consumed too quickly, the response is to slow feature delivery and invest in reliability. SRE practice also typically caps the proportion of an engineer time spent on operational work, so toil reduction is funded rather than aspirational.
What it looks like in practice
Nearby vocabulary
DevOps
DevOps is a way of working in which the people who build software share responsibility for running it, with automation that makes small change safe.
Service level objective
A service level objective is a target for a measurable, user-visible property of a service, such as the share of requests that succeed over a period.
Monitoring
Monitoring is the practice of collecting predefined signals from a system and alerting when they leave their expected ranges.
Observability
Observability is the property of a system that allows its internal state to be understood from the signals it emits, including for failures nobody anticipated.
How DevOpsArk handles site reliability engineering
Articles on this subject
What is DevOps? A definition that survives contact with practice
What DevOps actually means, where the definition came from, what the lifecycle looks like in practice, and the common misreadings that turn it into a job title instead of a way of working.
Alerting that people still read at 3am
How to build an alerting system with a high proportion of actionable pages: correlation, ownership routing, suppression, and deleting the rules that never produce a decision.
More definitions
See these concepts in a running system
A 30-minute walkthrough against your own infrastructure rather than a slide about the theory.