Glossary

Site reliability engineering

Site reliability engineering applies software engineering to operations, using reliability objectives and error budgets to balance stability and speed.

Short answer

What is site reliability engineering?

Site reliability engineering applies software engineering to operations, using reliability objectives and error budgets to balance stability and speed.

Explained two ways

Plain and technical

In plain terms

SRE treats keeping systems running as an engineering problem rather than a manual one. It sets a specific reliability target, and how much of that target remains unused determines how much risk the team can take with new releases.

Technically

SRE introduces service level objectives and the error budget they imply: if a service targets 99.9% availability over 30 days, the permitted unreliability is the budget. When the budget is being consumed too quickly, the response is to slow feature delivery and invest in reliability. SRE practice also typically caps the proportion of an engineer time spent on operational work, so toil reduction is funded rather than aspirational.

Example

What it looks like in practice

A service targets 99.9% monthly availability. Halfway through the month, two-thirds of the error budget is consumed, so the team pauses risky releases and spends the time on the cause instead.
Keep going

More definitions

See these concepts in a running system

A 30-minute walkthrough against your own infrastructure rather than a slide about the theory.