IncidentOps
Running incidents as a discipline: roles, severity, comms, and blameless learning loops.
Overview
Read the full IncidentOps guidance
Capabilities
Incident Response
Detecting, coordinating, and resolving unplanned interruptions with defined roles and severities.
On-Call Management
Sustainable rotations, escalation paths, and compensation for the humans carrying the pager.
Runbook Coverage
Documented procedures for the failure modes that actually recur, kept current.
Practices
Blameless Postmortems
Incident reviews that fix systems, not people — with tracked action items.
Metrics
MTTD — Mean Time to Detect
Average time from fault to detection by a human or automated signal.
Unit: minutes. Good direction: down.
MTTR — Mean Time to Restore
Average time from incident start to service restoration.
Unit: hours. Good direction: down.
Tools
PagerDuty
On-call scheduling, escalation, and incident mobilization.
Visit site ↗
Architecture patterns
Reference architectures and their trade-offs live in the blueprints library.
Maturity
Maturity for incidentops is measured, not guessed — every score traces to your answers. See how maturity is scored.
Put it to work
Related Ops disciplines
- Observability Engineering
Observability Engineering enables this area
Sources
- Tier 1 — PrimaryNIST SP 800-61r2 — Computer Security Incident Handling Guide ↗— NIST(verified 2026-09-15)
- Tier 1 — PrimarySite Reliability Engineering (Google) ↗— Google(verified 2026-09-15)