Management

Blameless Postmortems: The Engineering Trust Practice That Stops Repeat Failures

Table of Contents:

Almost everyone runs postmortems now. Half of all incidents are still repeats. The gap between those two facts is the whole problem.

At 8:01 on the morning of August 1st, 2012, an internal system at Knight Capital started sending emails.

They referenced SMARS, the firm's order router, and carried an error message: "Power Peg disabled." One arrived roughly every fifty-six seconds. By the time the market opened at 9:30, 97 of them had gone out to a group of Knight personnel.

Nobody looked.

The emails weren't designed to be alerts, so they landed in inboxes as routine noise. And the thing they were describing was real: a deployment the week before had missed one of eight servers, leaving nine-year-old decommissioned code live in production. When the bell rang, that server started trading.

Forty-five minutes later, Knight had turned 212 small retail orders into more than four million executions across 154 stocks, and lost over $460 million.

Here's what I keep coming back to. The alerting worked. The system detected the problem, correctly identified it, and reported it eighty-nine minutes before the market opened. What failed was the organization's capacity to act on information it already had.

That is a bad postmortem in miniature.

Knight Capital Group

The Practice Everyone Adopted and Nobody Finished

Almost every engineering organization runs some form of incident review now. Atlassian's 2024 State of Incident Management survey, fielded by CITE Research across more than 500 US software developers, IT professionals, and IT decision makers, found postmortems close to universal, and consistently so since 2020.

Only 22% said they practise blameless postmortems specifically.

So the meetings happen. Does anything change?

A 2022 benchmark study by Dimensional Research, polling over 300 on-call practitioners and managers responsible for production cloud environments, found that 48% of incidents are straightforward and repetitive. Nearly half of what wakes your engineers at 2am is a variation on something the team has seen before.

That same study found something more pointed: human error caused major incidents five times more frequently than bugs in automation.

Sit with that one, because it's the entire argument for doing this well. If most of your serious incidents involve a person making a decision that turned out badly, then how your team handles those decisions afterward isn't a soft-skills question. It's your reliability strategy.

Meanwhile PagerDuty's 2024 survey of 500 IT leaders at companies with more than 1,000 employees, conducted by Censuswide across the US, UK, and Australia, found 59% reporting that customer-facing incidents had increased, by an average of 43% over the previous twelve months.

Put it together and the picture is uncomfortable. The ritual is universal. The blameless version is rare. Incidents are climbing, and half of them rhyme with something that already happened.

What "Blameless" Was Supposed to Mean

The idea didn't come from software. It came from safety science, where the stakes were counted in lives.

James Reason proposed Just Culture in the 1990s, and his model had three categories, not one:

  1. Honest mistakes. Console the person. Fix the system.
  2. At-risk behaviour, where someone underestimated a danger. Coach the person. Fix the incentives that made the shortcut attractive.
  3. Reckless behaviour, conscious disregard of substantial risk. Discipline is appropriate.

Read that again, because it's the part the industry dropped. Just Culture was never a blanket amnesty. It was a calibrated response that matched consequence to intent.

Sidney Dekker extended it in 2007. His argument: "just" means fair, not permissive. He distinguished retributive justice, which asks who broke a rule and what punishment fits, from restorative justice, which asks who was harmed, what they need, and whose job it is to meet those needs. Accountability was always in the model. It was just pointed forward instead of backward.

John Allspaw brought this into software in 2012 with a post on Etsy's Code as Craft blog called "Blameless PostMortems and a Just Culture." Google codified the approach in Chapter 15 of the SRE book four years later. Within a few years the phrase was industry standard.

And then the machinery fell off.

Reason's three tiers flattened into "don't blame anyone." Dekker's restorative accountability, which demanded someone own the repair, became "no consequences." What survived was the word.

Amy Edmondson has a two-by-two that names the destination precisely. Psychological safety on one axis, accountability on the other. High safety with low accountability produces the comfort zone: people feel safe and nothing improves. High safety with high accountability produces the learning zone.

Most organizations implemented blamelessness and landed squarely in the comfort zone. Everyone feels fine. Nobody changed anything.

We see the same shape in our own data, in a different setting. Teams routinely score well on psychological safety and poorly on the things that turn safety into direction. Our standing advice for that gap is the same prescription this blog is making: don't decrease safety, increase standards.

The Three Ways the Meeting Fails

Three ways the meeting fails

The Stealth Trial

"I'm not blaming anyone, but…"

The questions have a direction. They tighten around one person's decisions while everyone else's go unexamined. Nobody says a name, and everyone knows the name.

What it teachesShip less. Ship later. Route anything risky through someone else's review.

The Process Smokescreen

Every action item is a new checklist.

Tangible, so it feels productive. But process changes are what a team produces when it doesn't want to name the system problem.

Six months laterThe same incident, a different person, and a longer checklist.

The Action Item Graveyard

Twenty-two items. Three owners. Zero dates.

It has a specific time of death: roughly seventy-two hours after the incident resolves, when the next sprint starts and the war room stops feeling recent.

The tell"Improve monitoring" instead of a verb, an owner, and a deadline.

  1. The Stealth Trial. Someone opens with "I'm not blaming anyone, but…" and the room understands immediately. The questions have a direction. They tighten around one person's decisions while everyone else's go unexamined. Nobody says a name, and everyone knows the name. What it teaches: ship less, ship later, route anything risky through someone else's review. You didn't get accountability. You got slower deploys and quieter engineers.
  2. The Process Smokescreen. Every action item is a new checklist, a new required approval, a new field in the deploy template. It feels productive because it's tangible. But process changes are what a team produces when it doesn't want to name the system problem. Six months later the same incident happens to a different person, and the checklist is longer. Knight Capital is the cautionary version. There was no written procedure requiring a second technician to verify the deployment, and the SEC concluded that a simple double-check could have caught the missed server and averted the whole thing. That's a system fix rather than a process one, and the distinction matters: a system change makes the failure impossible or visible, while a process change asks a tired human to remember something.
  3. The Action Item Graveyard. Twenty-two items get logged. Three get owners. Zero get due dates. This one is nearly universal and it has a specific time of death: about seventy-two hours after the incident resolves. For three days people care, because they remember the war room. By day four the sprint has started, there's a demo Friday, and the postmortem items are competing with a backlog that was already overcommitted before production caught fire. Then they wait through that sprint, and the next one. By the time anyone looks again, the context is gone and closing the ticket without doing the work feels like pragmatism.

The clearest standard I know comes from Google's own guidance, published by Lunney, Lueder, and Beyer: action items must be actionable, specific, and bounded. "Improve monitoring" is none of those. "Add latency alerting to the payment service at 500ms p99, owned by the payments team, due in two weeks" is all three. That difference is the difference between something that gets done and something that decays.

The Script

One discussion lead, and it should not be the most senior person in the room. Sixty to ninety minutes.

The structure below is a postmortem-specific version of the After-Action Review, a tool the US military built to debrief operations and one we teach in our leadership work. The AAR runs on five questions: what were our intended results, what were our actual results, what caused our results, what will we do the same next time, and what will we do differently. That's the spine. What follows adapts it for software incidents, where the causal question is the one that needs the most care.

One more thing the military got right, and it's the reason for the rule about the lead: sometimes you run the debrief without the leader in the room, because the team will be more honest. You can't always do that. You can always avoid putting the person whose decisions caused the outage in charge of examining them.

The script
70 minutes of structure. 60 to 90 in the room.

Build the timeline

20 min

Strict chronological order, with timestamps. Signals, not interpretations. A signal is something a log or a graph can prove. An interpretation is a story about why.

The "what did you know when" pass

15 min

At each decision point, two questions: what information did you have, and what was your model of the system? The gap between that model and reality is the fixable part.

Find the contributing conditions

15 min

Not five whys. Ask what conditions had to be true simultaneously for this to happen. List at least three. Then ask which are cheapest to make untrue.

Generate action items

10 min

Maximum six. Each gets a person, a date, and a definition of done. At least one must change the system, not the process.

The "what would you have wanted to know" closer

10 min

Round the room. What did you learn here that you wish you'd known last week? Four minutes, and it surfaces the near-misses in systems that haven't broken yet.

One rule about the lead. Not the most senior person in the room, and never the person whose decisions caused the outage. That isn't a blameless postmortem. It's a hearing where the defendant runs the court. Train two or three people so the role can rotate.

1. Build the timeline (20 minutes)

Strict chronological order, with timestamps, from signals rather than interpretations.

The distinction is the whole craft. A signal is something a log, a graph, a dashboard, or a Slack message can prove. An interpretation is a story about why.

  • Interpretation: "The deploy went out without proper testing."
  • Signal: "14:32, deploy of v2.4.1 to prod. 14:34, error rate on checkout goes from 0.2% to 11%. 14:41, first customer report in #support."

The second version is boring, and boring is the point. You cannot argue with a timestamp. Teams that build the timeline in signals spend the rest of the hour analyzing. Teams that build it in interpretations spend the rest of the hour litigating.

Build it before the meeting, share it beforehand, and open by asking what's missing. Someone always has a Slack thread nobody else saw.

2. The "what did you know when" pass (15 minutes)

Walk the major decision points. At each one, the person who made the call answers two questions: what information did you have, and what was your model of the system?

The second question is the one that surfaces things. People rarely make decisions that are wrong given what they believed. They make correct decisions from an inaccurate mental model, and the gap between that model and reality is a system problem you can actually fix.

This only works if it doesn't feel like a deposition. Some things that help: the lead asks about their own decisions first. The questions stay in the past tense and out of the conditional, so "what did you see" rather than "why didn't you check." And when someone says "I assumed X," that's treated as the most valuable sentence in the meeting, because an assumption held by one engineer is usually an assumption held by five.

3. What caused our results? (15 minutes)

Most templates put Five Whys here, and it's a reasonable place to start. It's simple, it's memorable, and it does push a team past the first answer, which is more than most groups manage on their own.

The limit is that it assumes a chain. Ask why, get an answer, ask why about that answer, and keep going until you reach the bottom. That works beautifully in manufacturing, which is where Toyota built it. A defect in a physical process usually does trace backward through a genuinely linear sequence.

Software incidents rarely do. Richard Cook's "How Complex Systems Fail" makes the point directly: catastrophe in a complex system requires multiple failures, and there is no isolated root cause waiting at the end of the chain.

Look at Knight Capital. There was dead code left callable for nine years. A flag repurposed without removing what it used to trigger. A counter moved in 2005 and never retested. A manual deploy with no second-pair verification. And an alerting channel nobody treated as alerting.

Remove any one of those and the morning goes differently. That's five conditions holding hands, not one root cause with four layers on top. Run Five Whys on it and you'll end up somewhere specific and defensible, but which of the five you land on depends mostly on where the facilitator started.

So keep the AAR's question and read it as plural. What caused our results?

Then make it concrete: what conditions had to be true simultaneously for this to happen? List them. Aim for at least three. Then ask which are cheapest to make untrue.

That last question is what turns analysis into a decision about where the next two weeks go.

4. Generate action items (10 minutes)

Maximum six. Each gets a name, a person, not a team, a due date, and a definition of done.

And at least one must change the system rather than the process.

The test is simple: if the same conditions recurred tomorrow with a different, tired, well-intentioned engineer, would this item prevent the incident, or would it require that engineer to remember something? A required checkbox in the deploy form is process. Automated verification that all eight servers received the build is system. Only one of those works at 3am.

5. The "what would you have wanted to know" closer (10 minutes)

Round the room. Each person answers: what did you learn here that you wish you'd known last week?

It takes four minutes and it's the highest-value part of the meeting, because it converts one team's incident into everyone's knowledge. It also surfaces near-misses. Someone will say "I didn't know that service had no retry logic," and now you know something about a system that hasn't broken yet.

What Makes the Lead's Job Hard

Reframe the verb, in real time. When someone says "the engineer broke prod," the lead says "the deploy caused prod to enter state X." Not as a correction. Just as the next sentence, spoken normally.

  • "She missed the alert" → "the alert didn't reach anyone who could act on it"
  • "He should have rolled back sooner" → "the rollback path wasn't obvious from the runbook"
  • "They pushed without testing" → "the pipeline allowed a deploy without the integration suite"

Nobody announces the rule. The lead just keeps speaking in system language, and within twenty minutes the room is doing it too.

Replace "should have" with "would have helped." "You should have checked the dashboard" is a verdict. "A dashboard link in the alert would have helped" is an action item. Same observation, and only one of them produces work.

Slow down the speed-runners. There's always someone who wants to skip to action items at minute fifteen. They're not being lazy, they're uncomfortable. Sitting in the analysis feels unproductive when you could be fixing things. The lead's job is to hold the room there long enough to find the third and fourth conditions, because the first one is always the obvious one, and the obvious one is rarely the cheapest to fix.

And rotate the role. If the lead is the person whose decisions caused the outage, this isn't a blameless postmortem. It's a hearing where the defendant is running the court. Train two or three people so the role can move.

Why This Is a Psychological Safety Practice

Fahd puts it more bluntly than I would:

"There's one moment that dictates psychological safety. The moment someone screws up. How you react in that moment determines it entirely." That's the whole thing. The values on the wall, the offsite, the engagement survey, all of it gets overwritten by what the team watches happen to the person who broke production. Which brings us back to that Dimensional Research finding: human error causing major incidents five times more often than automation bugs. Your most common serious failure mode is a person. So the question of what happens to that person afterward isn't adjacent to reliability. It is reliability.

Edmondson's original research found something that still surprises people: the better-performing hospital teams reported more errors, not fewer. They weren't making more mistakes. They were surfacing them. The worse teams had the same error rate and a quieter incident log, which is the more dangerous state and the harder one to notice. She tells a story from her interviews about an employee who wouldn't raise problems. Asked why, he said he had four kids in college. Asked how often people actually get fired there, he said this company doesn't fire anyone.

Both things, in the same head, at the same time. That's what an unsafe team looks like from the inside. Not fear of being fired. Fear of being unwelcome, which is quieter and just as effective at keeping people silent.

Psychological safety is the first of the Six Levels of High-Performing Teams for exactly this reason. It's the floor. And a postmortem is one of the few rituals where you can watch it get built or destroyed in real time, in front of everyone, in ninety minutes.

If you want to know where your team currently sits, we've written about how to measure psychological safety rather than guess at it, and about specific exercises for remote engineering teams, where the ambient signals are thinner and this work is harder.

Try This Next Incident

Don't start with the big one. Pick the next non-critical incident. Run the sixty-minute version. Rotate the discussion-lead role to someone who didn't touch it. Build the timeline in signals and share it before the meeting.

Then afterwards, ask each engineer privately: did anything go unsaid?

That question is the actual instrument. What people tell you one-on-one that they wouldn't say in the room is a direct reading of how much safety you have. If the answer is "no, we covered it," you're in good shape. If three people have something, you learned more from asking than from the meeting.

Knight Capital had ninety minutes of warnings and a building full of capable people who never read them. The information was there. What was missing was any structure that turned information into action. Most postmortems fail the same way, for the same reason, at a smaller price.

Team Dynamics Assessment

Want to know where your team's safety actually breaks down?

The Team Dynamics Assessment maps your team against the Six Levels, starting with psychological safety. Real data instead of a guess about whether people are telling you things.

Take the Team Dynamics Assessment

Benchmarked against 78 organizations and 900+ respondents.

Frequently Asked Questions:

Now that you have mastered how to manage conflict - what is your plan of action for making an impact with your team?

Now that you have mastered how to create an environment of empowerment via the 3-P's - what is your plan of action for making an impact with your team?

Developing Your Communication, Empathy and Emotional Intelligence skills is start. What is your plan of action for implementing your learnings within your your team?

Now that you understand the differences in these titles - what is your plan of action for what you learned?

Assessing your team's behaviors is a start - but do you have a plan of action for the results?

Now that you have mastered the art of decision making - what is your plan of action for making an impact with your team?

Download your free leadership guide that outlines the 6 necessary steps you need to acheive in order to develop a high performing team (in weeks, not months).  
Download your free leadership guide that outlines the 6 necessary steps you need to acheive in order to develop a high performing team (in weeks, not months).  
Download your free leadership guide that outlines the 6 necessary steps you need to acheive in order to develop a high performing team (in weeks, not months).  
Download your free leadership guide that outlines the 6 necessary steps you need to acheive in order to develop a high performing team (in weeks, not months).  
Help your managers improve their managing of communication, collaboration and conflict. Download your free leadership guide that outlines the 6 necessary steps you need to achieve in order to develop a high performing team (in weeks, not months).
Download your free leadership guide that outlines the 6 necessary steps you need to acheive in order to develop a high performing team (in weeks, not months).  
Get My Free Leadership Guide Now

A DISC Behaviour Assessment is the best way to understand your team's personalities.

Start by understanding your own behaviour tendencies with a DISC assessment. Learn more about how a DISC Assessment will improve your potential as a leader!

Each DISC Assessment includes a Self Assessment and DISC Style evaluation worksheet
Bill Gates Training Sidebar
Bill Gates Leadership

Curious how to develop into a transformational leader like Bill Gates?

Start by accessing our FREE Training video that outlines six simple steps for creating an environment that will transform YOU, and YOUR TEAM, into Unicorn leaders.

Leadership Training Ad - Sidebar
Leadership Training

Are you ready to transform from just a manager into a Unicorn Leader?

You can access our FREE training that will give you clarity on how to create a successful team in just six steps.

Leadership Training Ad - Sidebar
Leadership Training

Curious on some tips for transforming your managers into leaders?

Access our BEST, free training video that HR leaders are using to inspire real conversations with their managers.

In less than 25 minutes, you can gain clarity on how to turn your teams into centers for growth.

Access FREE Leadership Training Now
Career Conversations Sidebar
Career Conversations Guide

Is Your Company Culture Stuck In A Rut? Sometimes creating an environment for continuous learning can make a huge difference.

We created the best guide for having career development conversations with your teams.

Increase motivation and retain your top talent by following these simple steps.

EQ Leadership Sidebar
EQ Leadership Image

Did you know that EQ is more valuable to a leader than IQ?

We created the BEST, FREE training video to help managers map out a clear path for transformation into a Unicorn Leader.

Are you ready for your leadership transformation?

Access FREE Leadership Training Now
Leadership Workshop Sidebar
Leadership Workshop Image

Empowering your team to make decisions is just the start. Are you also supporting the other key elements for a high-performing team?

An investment in your team's development is an investment in your company's ability to effectively scale.

We've created an experiential, virtual workshop that focuses on developing teams into scalable engines of growth.

Interested in customizing a workshop for your team?

Learn More About The Leadership Workshop

Related posts