SRE Principles: Service Level Objectives (SLOs)

A Service Level Objective (SLO) is a target value or range of values for a service level that is measured by an SLI. Setting realistic SLOs aligns engineering velocity with user happiness.
The SLO Definition Workflow
graph TD
Step1["1. Identify Critical User Journeys (CUJs)<br/>(e.g. Checkout, Login, Search)"]
Step2["2. Select & Define SLIs<br/>(e.g. HTTP 2xx rate, P99 Latency < 300ms)"]
Step3["3. Analyze Historical Baseline Data"]
Step4["4. Set Achievable & Aspirational SLO Target<br/>(e.g. 99.9% over a 30-day rolling window)"]
Step5["5. Establish Error Budget Policy & Review Periodically"]
Step1 --> Step2 --> Step3 --> Step4 --> Step5
Types of SLO Targets
1. Achievable Targets (Data-Driven)
Constructed based on historical system metrics during periods when users reported high satisfaction.
2. Aspirational Targets (Business-Driven)
Dictated by market requirements or enterprise customer contracts. If historical performance falls short, dedicated engineering time is allocated to close the reliability gap.
SLO Best Practices
- Never Target 100% Availability: 100% reliability is impossible to achieve and prohibitively expensive. Users' own internet connections and hardware will fail well before the service does.
- Use Rolling Windows: Measure SLO compliance across a rolling 28-day or 30-day window rather than fixed calendar months.
- Document Exclusions: Clearly state which events (such as upstream cloud provider global outages or scheduled maintenance) are explicitly factored into or excluded from the SLO.