Skip to content
slide-deck.io
BlogGet started free

August 15, 2026

Slide Deck Template for Incident Post-Mortem and Retrospective Presentations

Production incidents are expensive. A major outage can cost hundreds of thousands of dollars in revenue per hour, trigger SLA breach penalties, drive customer churn, and damage the brand reputation that sales and marketing spend millions of dollars building. But the incident itself, as bad as it is, is not the worst possible outcome. The worst outcome is experiencing the same incident twice because the organization failed to learn from the first one.

The post-mortem presentation is the primary mechanism for organizational learning after an incident. Done well, it converts a painful experience into concrete improvements that reduce the probability and impact of future incidents. Done poorly, it becomes a blame exercise that causes engineers to hide problems and incidents go unreported.

This template covers the full post-mortem deck structure based on Google's Site Reliability Engineering blameless post-mortem framework, with specific guidance for each section.

The Blameless Principle

Google's SRE handbook established the blameless post-mortem as the industry standard for reliability engineering, and the principle is worth stating explicitly before every post-mortem process begins: the post-mortem finding is always a systemic failure, not an individual's mistake.

This is not a feel-good statement designed to spare people's feelings. It is a practical engineering principle with empirical support: when individuals fear blame, they hide problems. Hidden problems accumulate until they cause larger incidents. Blameless post-mortems surface root causes rather than scapegoats, which enables the organization to fix the actual problem.

In practice, blameless means: in every place where you might write "Engineer X failed to," rewrite it as "the system did not prevent" or "the runbook did not specify" or "the monitoring did not alert on." The goal is to find which system, process, or tool failed — not which person. If the answer is "a person made a decision based on insufficient information," the systemic question is: why did the system leave them with insufficient information?

Slide 1: Incident Summary

The summary slide gives leadership and stakeholders the complete incident picture in one page. Every person in the review meeting should understand the scope and severity before the detailed discussion begins.

Required elements:

  • Incident timeline: Detected at [time/date], mitigated at [time/date], resolved at [time/date]
  • Duration of customer impact: Not the same as incident duration — customer impact may have started before detection and ended before full resolution
  • Severity classification: Use a standard SEV1-4 classification (SEV1: complete service unavailability, major revenue impact; SEV2: significant degradation affecting many users; SEV3: partial degradation, workaround available; SEV4: minor issue, minimal customer impact)
  • Customer impact: How many users were affected? What functionality was unavailable? Were there geographic or segment-specific patterns?
  • Business impact: Estimated revenue at risk during the outage window, SLA breach determination (yes/no, which customers, what credit obligation), any customer escalations received

A note on detection method: whether the incident was detected by monitoring (best case), internal employees (worse), or customer reports (worst case) is important context. Customer-reported incidents indicate a monitoring gap and have worse trust implications than internally-detected incidents. Always include this.

Slide 2: What Happened — Root Cause and Contributing Factors

This is the most technical slide in the deck. It requires precision because vague root causes lead to ineffective corrective actions. "Infrastructure issue" or "software bug" are not root causes — they are categories. The root cause should be specific enough that someone reading it a year from now could understand exactly what failed and why.

Good root cause statement: "A deployment of service X on [date] introduced a connection pool configuration change that reduced the max connection limit from 500 to 50. Under normal load conditions, the reduced limit was sufficient. During the Tuesday peak traffic window, connection demand exceeded 50, causing connection queue saturation. New requests began failing with 503 errors as the queue filled. The configuration change was not caught in staging because staging load is approximately 15% of production peak."

Bad root cause statement: "Configuration error caused database connection issues leading to service degradation."

Contributing factors explain why the root cause was able to cause the impact it did. These are often as important as the root cause itself. Common contributing factors:

  • Monitoring did not alert on the specific failure mode
  • The runbook for this service did not cover this failure pattern
  • On-call rotation lacked familiarity with this service's architecture
  • Staging environment did not replicate the production load pattern
  • Deployment did not have a canary phase that would have caught the issue at low traffic volume

The incident narrative should describe the chain of events that connected root cause to customer impact, with specific timestamps where helpful.

Slide 3: Incident Timeline and Response

The detailed timeline serves two purposes: it documents the sequence of events for the incident record, and it reveals the response process quality for improvement identification.

Format: a chronological table with timestamp, event/action, and owner (where relevant).

| Time | Event | Owner | |------|-------|-------| | 14:23 | Monitoring alert fires: P99 latency >2s on service X | Automated | | 14:26 | On-call engineer acknowledges alert | On-call SRE | | 14:31 | Incident declared, Slack channel opened | On-call SRE | | 14:38 | Initial hypothesis: database slowdown | On-call SRE | | 14:52 | Hypothesis disproved, investigation widened | SRE lead | | 15:17 | Root cause identified: connection pool exhaustion | SRE lead | | 15:24 | Mitigation deployed: connection limit restored | SRE lead | | 15:27 | Customer impact resolved | — |

Key response quality indicators visible in the timeline:

  • Time from impact start to detection: was monitoring catching problems proactively?
  • Time from detection to hypothesis: was the team prepared to diagnose this type of failure?
  • Time from hypothesis to mitigation: did the team have the right tools and access to act quickly?
  • Communication cadence: were stakeholders updated on a regular basis during the incident?

Slide 4: What Went Well

This is not a performance review of the incident response team designed to balance negative feedback. It is a genuine identification of the response practices that worked — because those practices should be codified and repeated.

Common examples:

  • Monitoring detected the incident before customer reports (vs. customer-reported, this is a win)
  • Incident communication to customers was accurate, timely, and didn't over-promise resolution times
  • Cross-team coordination (e.g., engineering and customer success) was smooth because roles were pre-defined in the incident runbook
  • MTTR was below the team's historical average for this severity level
  • Rollback process was clean and completed in under 5 minutes

The discipline required here is genuine specificity. "The team worked hard" is not a useful observation. "The on-call engineer identified the correct service to investigate within 8 minutes despite ambiguous initial monitoring signals, because they had run a game day exercise for this failure mode two months ago" — that is a useful observation that generates a concrete recommendation (continue game day exercises for critical failure modes).

Slide 5: What Went Wrong

This is the heart of the post-mortem — the honest diagnosis of where the system failed. Apply the blameless principle rigorously in every sentence.

Common systemic failure patterns:

  • Monitoring gaps: The failure mode was not covered by an existing alert. Detection relied on a customer report or manual observation.
  • Runbook gaps: The runbook for this service did not cover this failure scenario, requiring the on-call engineer to improvise during a high-pressure incident
  • Environment parity: Staging did not replicate the load patterns, configuration, or data characteristics of production that were relevant to this failure
  • Deployment process: The deployment lacked a canary phase, progressive rollout, or automated rollback trigger that would have contained the blast radius
  • Communication breakdown: Stakeholder updates were delayed, inaccurate, or inconsistent across channels
  • Escalation delays: The incident was not escalated to additional responders quickly enough, extending MTTR

Each item in "what went wrong" should connect directly to an action item in the next section. If you identify a problem and don't fix it, the post-mortem is theater.

Slide 6: Action Items

Action items are the entire point of the post-mortem. Every action item requires: a specific description (not "improve monitoring"), an owner (a named person, not a team), and a due date.

Categories of action items:

Preventive actions — reduce the probability of recurrence:

  • Add configuration validation to the deployment pipeline that checks connection pool settings against environment-specific limits
  • Add a canary deployment phase to service X deployments that routes 5% of traffic before full rollout

Detective actions — reduce MTTD:

  • Add connection pool utilization monitoring with alert thresholds at 70% and 90% of limit
  • Add connection queue saturation to the SEV2 auto-escalation trigger

Corrective actions — reduce MTTR:

  • Add connection pool exhaustion to the service X runbook with specific diagnostic commands and resolution steps
  • Schedule a game day exercise for connection pool failure scenarios before end of quarter

Track action items from post-mortems in a shared system (JIRA, Linear, or equivalent) with the incident reference, so completion can be tracked and post-mortem quality can be measured by action item completion rate.

Slide 7: SLA and SLO Impact

For incidents that breach customer commitments, the final slide addresses the contractual and trust implications.

SLO (Service Level Objective) breach determination: What is the stated availability or latency SLO? What was the measured performance during the incident window? What is the cumulative SLO performance for the rolling measurement period (typically 30 days)?

SLA (Service Level Agreement) breach determination: Which customer contracts have explicit SLA commitments? Did the incident duration and impact trigger SLA breach for any contracted customers? What credit obligations result?

Customer communication strategy: Was the incident disclosed proactively (status page, email before customers reported) or reactively (in response to customer support tickets)? What follow-up communication is required? For significant incidents, a personalized CEO or VP communication to affected enterprise customers is often warranted — it is the trust-preserving action, even though it is uncomfortable.

Trust recovery actions: Beyond contractual credits, what actions will the company take to rebuild confidence with affected customers? Sharing the post-mortem summary (sanitized for external sharing) demonstrates transparency and commitment to improvement, and is increasingly expected by enterprise customers after significant incidents.

Building This Presentation with slide-deck.io

slide-deck.io builds the full post-mortem deck structure from an incident narrative. Paste the timeline, root cause analysis, what went well and wrong, and action items — the AI formats it into a coherent presentation with proper section structure, the timeline table, and action item formatting. Post-mortems written in slide-deck.io are also easily shared as async artifacts for teams in different time zones or for stakeholders who didn't attend the live review.

The goal of a post-mortem is not the document — it is the changed system. But a well-structured document forces the rigor of analysis that drives real change, and a clear presentation ensures that change gets organizational support.

Build your next presentation with AI

Generate editable .pptx decks in minutes. Free to start — no card required.

Try it free →