Observability Engineering
Instrumenting systems so their internal state is interrogatable from the outside.
Overview
Read the full Observability Engineering guidance
Capabilities
Alerting
Symptom-based, SLO-backed alerts that page a human only when a user-visible promise is burning.
Distributed Tracing
End-to-end request traces across service boundaries with consistent context.
Log Management
Structured, centralized, retention-governed logs with correlation identifiers.
Monitoring Coverage
Metrics and health signals across the full estate, including the parts nobody owns.
Metrics
Alert Noise Ratio
Share of pages that required no action.
Unit: %. Good direction: down.
Tools
Grafana
Dashboards and visualization for metrics, logs, and traces.
Visit site ↗OpenTelemetry
Vendor-neutral instrumentation standard for traces, metrics, logs.
Visit site ↗Prometheus
Open-source metrics collection and alerting toolkit.
Visit site ↗
Architecture patterns
Reference architectures and their trade-offs live in the blueprints library.
Maturity
Maturity for observability engineering is measured, not guessed — every score traces to your answers. See how maturity is scored.
Put it to work
Related Ops disciplines
- SRE
SRE is enabled by this area
- IncidentOps
IncidentOps is a dependency of this area
Sources
- Tier 2 — AuthoritativeOpenTelemetry Documentation ↗— CNCF / OTel(verified 2026-09-15)