SRE
Engineering reliability into systems with SLOs, error budgets, and blameless incident learning.
Overview
Read the full SRE guidance
Capabilities
Alerting
Symptom-based, SLO-backed alerts that page a human only when a user-visible promise is burning.
Incident Response
Detecting, coordinating, and resolving unplanned interruptions with defined roles and severities.
SLO Management
Defining, measuring, and governing Service Level Objectives and their error budgets.
Practices
Blameless Postmortems
Incident reviews that fix systems, not people — with tracked action items.
Error Budget Policy
A written contract trading release velocity for reliability when budgets burn.
Toil Budgeting
Measuring and capping manual operational work to fund automation.
Metrics
Availability
Achieved uptime against the defined SLO window.
Unit: %. Good direction: up.
Error Budget Burn Rate
Speed at which the error budget is being consumed.
Unit: × budget/hour. Good direction: down.
MTTR — Mean Time to Restore
Average time from incident start to service restoration.
Unit: hours. Good direction: down.
Architecture patterns
Reference architectures and their trade-offs live in the blueprints library.
Maturity
Maturity for sre is measured, not guessed — every score traces to your answers. See how maturity is scored.
Put it to work
Related Ops disciplines
- Resilience Engineering
Resilience Engineering is related to this area
- Observability Engineering
Observability Engineering enables this area
- DevOps
DevOps overlaps with this area
Sources
- Tier 1 — PrimarySite Reliability Engineering (Google) ↗— Google(verified 2026-09-15)
- Tier 1 — PrimaryThe Site Reliability Workbook (Google) ↗— Google(verified 2026-09-15)