Every complex system fails sometimes. What matters is how quickly the team detects and recovers, and whether each incident leaves the system safer. Blameless postmortems and service level objectives are the two most widely used tools for that.
This guide covers the incident lifecycle and its measures, how to set an SLO and compute an error budget, how to run a postmortem that looks at contributing factors instead of blame, and a worked example in which one long incident consumes more than four times a month's budget.
Before You Start
Why Incident Management and Postmortems Matter
Outages Will Happen
Complex systems fail. What separates teams is how fast they detect and recover, and whether they learn.
Blame Hides Causes
If people fear punishment, they hide details, and the real causes stay in place. A blameless approach gets the facts.
Repeat Incidents Are Waste
The same failure twice usually means the first postmortem produced no lasting change.
Reliability Is a Trade-Off
Service level objectives make the trade-off explicit, so teams can balance new features against stability.
The Incident Lifecycle and Its Measures
| Stage | Measure | Meaning |
|---|---|---|
| Detection | Time to detect (MTTD) | From the start of the problem to when it is noticed, ideally by monitoring, not by customers. |
| Response | Time to acknowledge (MTTA) | From the alert to when someone owns it. |
| Mitigation | Time to mitigate | From detection to when customer impact is stopped, even if the root cause is not fixed. |
| Recovery | Time to restore (MTTR) | Total time to restore normal service. |
| Learning | Postmortem completed, actions closed | Whether the organization improved. |
Service Level Objectives and Error Budgets
A service level indicator (SLI) measures some aspect of service, such as the share of requests that succeed. A service level objective (SLO) sets a target for it, for example 99.9% availability over 30 days. The error budget is the allowed unreliability: 100% minus the SLO. The approach is described in Google's Site Reliability Engineering.
When the budget is spent, the team slows feature releases and invests in reliability until it recovers. When plenty of budget remains, it can take more risk. Choose an SLO that reflects what users need; a target higher than users notice is expensive without benefit.
The Blameless Postmortem
A postmortem is a written review of an incident that focuses on how the system and process allowed it, not on who made a mistake. John Allspaw's Blameless PostMortems and a Just Culture (2012) describes the idea: people usually act sensibly given what they knew at the time, so ask what made the action seem reasonable, and what would make the failure harder next time.
- Timeline. Build it from logs, chat, and alerts. Include when the problem started, was detected, was acknowledged, was mitigated, and was resolved.
- Contributing factors. List several, across tooling, process, communication, design, and monitoring. Avoid a single "root cause" that is really a person.
- What went well. Note what worked, such as an alert or a runbook, so it is repeated.
- Action items. Each has an owner, a due date, and a type: prevent, detect, mitigate, or learn. Track them to completion.
Record the analysis in the Blameless Postmortem Template.
Worked Example: A Month Against an SLO
A service has a 99.9% availability SLO over a 30-day month. It had five incidents of 12, 45, 8, 90, and 25 minutes. The numbers are illustrative.
| Measure | Calculation | Result |
|---|---|---|
| Total downtime | 12 + 45 + 8 + 90 + 25 | 180 minutes |
| Availability | 1 − 180 / 43,200 | 99.58% |
| Error budget | 43,200 × 0.001 | 43.2 minutes |
| Budget consumed | 180 / 43.2 | about 4.2 times the budget |
| Mean time to restore | 180 / 5 | 36 minutes |
What the team does. The 90-minute incident accounts for half of the downtime, so it gets a full postmortem. The timeline shows that detection took 22 minutes because the alert fired on CPU load, not on failed requests. Contributing factors also include a deployment without a canary stage and a runbook that was out of date. Actions: add an alert on request success rate (detect), introduce canary releases for this service (mitigate), and update and rehearse the runbook (prevent recurrence). Each has an owner and a date.
Because the budget is spent, the team also pauses non-essential feature releases for two weeks and puts reliability work first. The lesson is not that 99.9% was impossible, but that one long incident can dominate a month, and shortening detection and recovery time is often more valuable than preventing every failure.
Roles During an Incident
During a serious incident, the danger is that everyone works on the same thing, or that nobody coordinates. Clear roles, agreed before an incident, keep the response organized and leave time to think.
Severity levels. Define a small set of severities based on customer impact, and the expected response for each, such as who is paged, how often updates go out, and whether an incident commander is needed. People should be able to declare an incident without fear, and severity can be changed as more is learned.
Prepare in advance. Keep on-call schedules, runbooks, contact lists, and communication templates current. Practice with exercises so that the first time someone is commander is not during a real outage.
Mitigate first. The priority is to restore service, for example by rolling back, failing over, or turning off a feature, then to find the root cause. A perfect diagnosis during an outage is a luxury.
Postmortem Quality: Actions That Get Done
The value of a postmortem lies in what changes afterwards. A well-written document with actions that are never completed teaches nothing.
| Quality of an action | Good | Weak |
|---|---|---|
| Specific | “Add an alert when checkout success rate falls below 99% for five minutes” | “Improve monitoring” |
| Owned and dated | A named owner and a due date | “The team will look into it” |
| Addresses a cause | Linked to a contributing factor | Unrelated to the analysis |
| Strong type | Prevent or detect, with a system change | “Be more careful” or “retrain” only |
| Tracked to completion | Reviewed at a regular meeting | Forgotten after the review |
Look for several contributing factors. Complex incidents rarely have a single root cause. Examine tooling, testing, deployment, monitoring, documentation, communication, and design. Avoid stopping at “human error”; ask what made the error easy to make and hard to notice.
Share widely and look for patterns. Publish postmortems, appropriately edited, where other teams can learn from them. Review several together from time to time to find recurring themes, such as deployment issues or capacity.
Use error budgets. If a service has used up its error budget, agree in advance what changes, for example slowing feature releases and prioritizing reliability work. That makes the trade-off a rule, not a debate in the middle of a crisis.
Measure. Track the time to detect, acknowledge, and restore; the number of repeat incidents; and the completion rate of postmortem actions. See the Blameless Postmortem Template, the MTBF, MTTR, and Availability Guide, and the 5 Whys Guide. Figures in the examples are illustrative.
Self-Assessment Questions
- Do we detect incidents with monitoring before customers report them?
- Do we have SLOs that reflect what users need, and do we track the error budget?
- Do postmortems focus on system and process causes, not on blaming people?
- Does every action item have an owner and a date, and do we close them?
- Do we look for patterns across incidents, not only within one?
Common Mistakes
Blaming an Individual
"Human error" is where the investigation should start, not end. Ask what made the error easy to make.
Postmortems Without Actions
A document nobody acts on teaches nothing. Track actions like any other work.
SLOs Nobody Uses
If the error budget does not change decisions, it is decoration. Agree in advance what happens when it is spent.
Measuring Only MTTR
Mean time to restore hides the spread. Look at each stage of the timeline, and at the longest incidents.
Incident Management and Blameless Postmortems: Frequently Asked Questions
What is a blameless postmortem?
A blameless postmortem is a written review of an incident that focuses on how the system, tools, and processes allowed it, rather than on who made a mistake. It assumes people acted reasonably given what they knew, and asks what would make the failure harder or the recovery faster next time, ending with tracked action items.
What is an error budget?
An error budget is the amount of unreliability a service is allowed under its service level objective, equal to 100% minus the SLO. For a 99.9% SLO over 30 days, it is 43.2 minutes. When the budget is spent, the team prioritizes reliability work over new features until service recovers.
What is the difference between MTTD, MTTA, and MTTR?
MTTD is the mean time to detect a problem from when it starts, MTTA is the mean time to acknowledge an alert, and MTTR is the mean time to restore normal service, though some teams use R for repair or resolve. Definitions vary, so state yours, and look at each stage of the timeline because averages can hide long incidents.
Sources and Further Reading
- Betsy Beyer, Chris Jones, Jennifer Petoff, and Niall Richard Murphy (eds.), Site Reliability Engineering, Google, 2016.
- John Allspaw, "Blameless PostMortems and a Just Culture," Etsy Code as Craft, 2012.
- Sidney Dekker, The Field Guide to Understanding 'Human Error'.
- Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate, on time to restore service.