DevOps

Post-Mortem Analysis: How to Actually Learn from an Incident

What makes a good incident post-mortem, which questions it should answer, and why placing blame never solves anything.

XALT favicon: white XALT logo on black background with digital light effects.TEAM XALTAtlassian Platinum Partner·24 January 2022·6 min
Incident post-mortem

Errors are human and can lead to minor or even severe incidents. Let's be honest: we can try to avoid errors, but sooner or later they will happen. However, making errors is not the biggest problem. We must ensure that we learn from our mistakes. If you introduce a post-mortem analysis after an incident or error in your projects, you can learn from past mistakes. This article is about introducing a post-mortem culture in your DevOps organisation and why assigning blame is not the solution.

The best way to learn from incidents is to conduct post-mortems.

What is a post-mortem in projects and incidents?

An incident post-mortem or retrospective examines how the incident occurred, what impact it had on business goals and metrics, and how the team fixed the errors.

Why do you need post-mortem analyses in DevOps?

In many companies, whether large or small, major incidents occur at least several times a year. As mentioned earlier, you can work to prevent incidents, reduce their impact, and shorten their consequences for your goals or other key KPIs. But they will still happen, no matter what you do.

Changes to your systems, code, or infrastructure can introduce vulnerabilities that lead to incidents. As a DevOps champion, you are likely releasing new code iterations or updates at a high frequency. This minimises the risk and impact of errors in individual releases. At the same time, the increasing number of releases probably does not lead to a reduction in the number of incidents. However, the likelihood of your entire system failing is drastically reduced.

But what happens when a critical incident occurs?

Instead of assigning blame and pointing fingers at the responsible party, the only thing that matters is finding out what caused the error that led to the incident and, if necessary, mitigating the impact.

Root cause analysis and the implementation of preventive measures are important to ensure that such incidents do not occur too frequently. Otherwise, incidents may become more frequent, and troubleshooting becomes a weekly routine. Sooner or later, your teams will be occupied solely with reacting to incidents. And no one wants that! To get out of this situation, your team must acknowledge the new status quo.

And what does a post-mortem analysis look like?

To learn from past mistakes, we must conduct a precise analysis of events after incidents.

A post-mortem should always be conducted when an incident requires the response of an IT engineer or developer. Typically, a post-mortem analysis captures the following points:

  • What triggered the incident?
  • What were the impacts?
  • How long did it take to detect and contain the incident?
  • What steps were taken to resolve the incident?
  • Did the team conduct a root cause analysis?
  • Can we create a timeline of key events? Summarise the main activities from chat conversations, incident details, and more.
  • What are the lessons learned and next steps? What went well and what did not? How can we prevent this problem from recurring?

In most cases, the analysis is carried out by the team members who responded to the incident and mitigated or investigated the cause.

Example: The Facebook, WhatsApp, and Instagram outage in October 2021

On 4 October 2021, several public Facebook services experienced a global outage lasting almost 24 hours. As Facebook (Meta) explained, configuration changes to the backbone routers, which coordinate network traffic between data centres, caused problems that disrupted communication.

In this article, you will learn more about the service outage and the postmortem at Meta.

What Facebook did seems very simple at first glance: they followed their internal process, "Storm Drills". This is how Facebook ensures they always know exactly what to do next when something happens.

They then made sure to conduct a postmortem to find out what really happened.

Introducing objective postmortems - without blame

Conducting postmortems may seem quite straightforward. However, for many organisations that have never considered incorporating postmortems into their incident response, this could be a challenge they should not take lightly.

Implementing and sustaining the success of a new or modified process requires time and effort at all levels of the organisation.

To facilitate the transition, there are a few key principles you should keep in mind:

  • Make sure to distance yourself from blame and mutual accusations: This is the most important aspect to get things right from the start. If the analysis focuses on blaming the cause of the incident rather than ensuring the team learns and improves, the entire initiative is more likely to cause harm than benefit.
  • Communicate openly and with fault tolerance: Ensure that postmortem meetings are not used to find a scapegoat. It is the only opportunity for teams to learn and improve. This means being honest about what happened and correcting expectations.
  • Introduce a responsible person for the postmortem discussion: An engaged leader ensures that every incident response is concluded with a postmortem. These leaders typically have a comprehensive understanding of all services and DevOps. The leader sets the tone for postmortems and largely determines the collective attitude towards distancing from blame.
  • Collaborate, share information, and foster communication: Ensure that postmortems are well documented on an internal platform (e.g., a Confluence wiki). Each postmortem can then be used as useful training material for your teams.
  • Get top management on board and involve all relevant stakeholders: To engage all team members in the organisation, top-down communication is essential. At the same time, leaders must also know what happened. So make sure you communicate clear figures and outcomes.
  • Make decisions: Ideally, a good, comprehensive postmortem provides preventive suggestions. You must determine who is responsible for approving the recommendations and reviewing the written reports.

BETTER CALL XALT

Ready to build a blameless post-mortem culture?

We help you build incident response processes and post-mortems into Jira Service Management and Confluence.

From the same category

More articles

Title graphic: 5 Steps to Building an Enterprise AI Harness on a dark blue background with network visualization.
DevOps

The Agent Harness: The Real Foundation of Enterprise AI

Why the agent harness – not the AI model or the generated code – determines the success of enterprise AI initiatives, and how to build one in five steps.

Illustration of zero trust and governance for agentic AI
DevOps

Governing Agentic AI Safely: Zero Trust and Compliance for AI Agents

AI agents are transforming the enterprise at speed – and the risk is growing just as fast. How the "in dubio pro securitate" principle, formal verification, and zero trust keep agentic AI safe and compliant.

Laptop displaying code on screen overlaid with a white icon of arrows and gear.
DevOps

Shift Left: Catching Bugs Earlier in Your Development Cycle

Bugs found late cost time, money, and trust. The shift-left approach moves testing and quality assurance as early as possible into the development cycle.