Written by David Rodgers

Quality and Operations Perspective

Written by David Rodgers, Lean Six Sigma Black Belt and ASQ-certified quality leader. This guide applies quality and process-improvement methods to software and IT operations from a quality and operations perspective. The author is not a software engineer, site reliability engineer, or security professional.

Last editorial review: September 24, 2026. Educational content only: not medical, legal, or regulatory advice. Follow your organization's policies and the requirements that apply to you, and have subject-matter experts review any change to a live process.

  • Lean Six Sigma Black Belt
  • ASQ CQE
  • ASQ CMQ/OE
  • Quality systems and process improvement

Every complex system fails sometimes. What matters is how quickly the team detects and recovers, and whether each incident leaves the system safer. Blameless postmortems and service level objectives are the two most widely used tools for that.

This guide covers the incident lifecycle and its measures, how to set an SLO and compute an error budget, how to run a postmortem that looks at contributing factors instead of blame, and a worked example in which one long incident consumes more than four times a month's budget.

Get the Postmortem Template Open the Reliability Calculator

Before You Start

Educational content. This guide applies quality and process-improvement methods to software delivery and IT operations. It is not security, legal, compliance, or engineering advice. Practices, tools, and risks differ between teams and systems, so have qualified engineers and security professionals review changes to production systems and controls.

Why Incident Management and Postmortems Matter

Outages Will Happen

Complex systems fail. What separates teams is how fast they detect and recover, and whether they learn.

Blame Hides Causes

If people fear punishment, they hide details, and the real causes stay in place. A blameless approach gets the facts.

Repeat Incidents Are Waste

The same failure twice usually means the first postmortem produced no lasting change.

Reliability Is a Trade-Off

Service level objectives make the trade-off explicit, so teams can balance new features against stability.

An on-call engineering team working an incident in an operations room
During an incident, one person coordinates, and everyone else works a defined role.

The Incident Lifecycle and Its Measures

StageMeasureMeaning
DetectionTime to detect (MTTD)From the start of the problem to when it is noticed, ideally by monitoring, not by customers.
ResponseTime to acknowledge (MTTA)From the alert to when someone owns it.
MitigationTime to mitigateFrom detection to when customer impact is stopped, even if the root cause is not fixed.
RecoveryTime to restore (MTTR)Total time to restore normal service.
LearningPostmortem completed, actions closedWhether the organization improved.

Service Level Objectives and Error Budgets

A service level indicator (SLI) measures some aspect of service, such as the share of requests that succeed. A service level objective (SLO) sets a target for it, for example 99.9% availability over 30 days. The error budget is the allowed unreliability: 100% minus the SLO. The approach is described in Google's Site Reliability Engineering.

Error budget (minutes) = period minutes × (1 − SLO). For 30 days and 99.9%: 43,200 × 0.001 = 43.2 minutes

When the budget is spent, the team slows feature releases and invests in reliability until it recovers. When plenty of budget remains, it can take more risk. Choose an SLO that reflects what users need; a target higher than users notice is expensive without benefit.

The Blameless Postmortem

A postmortem is a written review of an incident that focuses on how the system and process allowed it, not on who made a mistake. John Allspaw's Blameless PostMortems and a Just Culture (2012) describes the idea: people usually act sensibly given what they knew at the time, so ask what made the action seem reasonable, and what would make the failure harder next time.

  • Timeline. Build it from logs, chat, and alerts. Include when the problem started, was detected, was acknowledged, was mitigated, and was resolved.
  • Contributing factors. List several, across tooling, process, communication, design, and monitoring. Avoid a single "root cause" that is really a person.
  • What went well. Note what worked, such as an alert or a runbook, so it is repeated.
  • Action items. Each has an owner, a due date, and a type: prevent, detect, mitigate, or learn. Track them to completion.

Record the analysis in the Blameless Postmortem Template.

Worked Example: A Month Against an SLO

A service has a 99.9% availability SLO over a 30-day month. It had five incidents of 12, 45, 8, 90, and 25 minutes. The numbers are illustrative.

MeasureCalculationResult
Total downtime12 + 45 + 8 + 90 + 25180 minutes
Availability1 − 180 / 43,20099.58%
Error budget43,200 × 0.00143.2 minutes
Budget consumed180 / 43.2about 4.2 times the budget
Mean time to restore180 / 536 minutes
0 50 100 150 200 Incident 1 12 min Incident 2 45 min Incident 3 8 min Incident 4 90 min Incident 5 25 min 12 57 65 155 180 Error budget 43.2 min Incident minutes in the month (bars) and cumulative total (line)
The budget is exhausted during the second incident, and by month's end the service has used more than four times what the SLO allowed.

What the team does. The 90-minute incident accounts for half of the downtime, so it gets a full postmortem. The timeline shows that detection took 22 minutes because the alert fired on CPU load, not on failed requests. Contributing factors also include a deployment without a canary stage and a runbook that was out of date. Actions: add an alert on request success rate (detect), introduce canary releases for this service (mitigate), and update and rehearse the runbook (prevent recurrence). Each has an owner and a date.

Because the budget is spent, the team also pauses non-essential feature releases for two weeks and puts reliability work first. The lesson is not that 99.9% was impossible, but that one long incident can dominate a month, and shortening detection and recovery time is often more valuable than preventing every failure.

Engineers reviewing an incident timeline in a blameless postmortem meeting
A postmortem asks how the system made the failure possible, not who made the mistake.

Roles During an Incident

During a serious incident, the danger is that everyone works on the same thing, or that nobody coordinates. Clear roles, agreed before an incident, keep the response organized and leave time to think.

Incident commander Coordinates the response, sets priorities, makes decisions, and delegates; does not fix the problem personally. Technical lead Directs diagnosis and the fix, and brings in specialists as required. Communications lead Keeps stakeholders and customers informed, on a regular cadence, in plain language. Scribe Records actions and times in a shared timeline, for use in the postmortem.
Separating communication and recording from the technical work protects the people who are solving the problem.

Severity levels. Define a small set of severities based on customer impact, and the expected response for each, such as who is paged, how often updates go out, and whether an incident commander is needed. People should be able to declare an incident without fear, and severity can be changed as more is learned.

Prepare in advance. Keep on-call schedules, runbooks, contact lists, and communication templates current. Practice with exercises so that the first time someone is commander is not during a real outage.

Mitigate first. The priority is to restore service, for example by rolling back, failing over, or turning off a feature, then to find the root cause. A perfect diagnosis during an outage is a luxury.

Postmortem Quality: Actions That Get Done

The value of a postmortem lies in what changes afterwards. A well-written document with actions that are never completed teaches nothing.

Quality of an actionGoodWeak
Specific“Add an alert when checkout success rate falls below 99% for five minutes”“Improve monitoring”
Owned and datedA named owner and a due date“The team will look into it”
Addresses a causeLinked to a contributing factorUnrelated to the analysis
Strong typePrevent or detect, with a system change“Be more careful” or “retrain” only
Tracked to completionReviewed at a regular meetingForgotten after the review

Look for several contributing factors. Complex incidents rarely have a single root cause. Examine tooling, testing, deployment, monitoring, documentation, communication, and design. Avoid stopping at “human error”; ask what made the error easy to make and hard to notice.

Share widely and look for patterns. Publish postmortems, appropriately edited, where other teams can learn from them. Review several together from time to time to find recurring themes, such as deployment issues or capacity.

Use error budgets. If a service has used up its error budget, agree in advance what changes, for example slowing feature releases and prioritizing reliability work. That makes the trade-off a rule, not a debate in the middle of a crisis.

Measure. Track the time to detect, acknowledge, and restore; the number of repeat incidents; and the completion rate of postmortem actions. See the Blameless Postmortem Template, the MTBF, MTTR, and Availability Guide, and the 5 Whys Guide. Figures in the examples are illustrative.

Self-Assessment Questions

  • Do we detect incidents with monitoring before customers report them?
  • Do we have SLOs that reflect what users need, and do we track the error budget?
  • Do postmortems focus on system and process causes, not on blaming people?
  • Does every action item have an owner and a date, and do we close them?
  • Do we look for patterns across incidents, not only within one?

Common Mistakes

Blaming an Individual

"Human error" is where the investigation should start, not end. Ask what made the error easy to make.

Postmortems Without Actions

A document nobody acts on teaches nothing. Track actions like any other work.

SLOs Nobody Uses

If the error budget does not change decisions, it is decoration. Agree in advance what happens when it is spent.

Measuring Only MTTR

Mean time to restore hides the spread. Look at each stage of the timeline, and at the longest incidents.

Incident Management and Blameless Postmortems: Frequently Asked Questions

What is a blameless postmortem?

A blameless postmortem is a written review of an incident that focuses on how the system, tools, and processes allowed it, rather than on who made a mistake. It assumes people acted reasonably given what they knew, and asks what would make the failure harder or the recovery faster next time, ending with tracked action items.

What is an error budget?

An error budget is the amount of unreliability a service is allowed under its service level objective, equal to 100% minus the SLO. For a 99.9% SLO over 30 days, it is 43.2 minutes. When the budget is spent, the team prioritizes reliability work over new features until service recovers.

What is the difference between MTTD, MTTA, and MTTR?

MTTD is the mean time to detect a problem from when it starts, MTTA is the mean time to acknowledge an alert, and MTTR is the mean time to restore normal service, though some teams use R for repair or resolve. Definitions vary, so state yours, and look at each stage of the timeline because averages can hide long incidents.

Sources and Further Reading

  • Betsy Beyer, Chris Jones, Jennifer Petoff, and Niall Richard Murphy (eds.), Site Reliability Engineering, Google, 2016.
  • John Allspaw, "Blameless PostMortems and a Just Culture," Etsy Code as Craft, 2012.
  • Sidney Dekker, The Field Guide to Understanding 'Human Error'.
  • Nicole Forsgren, Jez Humble, and Gene Kim, Accelerate, on time to restore service.