Featured

Postmortem Action Item Decay: Why Your Incident Fixes Never Actually Ship

Blameless postmortems produce great root-cause analysis and action items that quietly die in the backlog. See what SREs on Reddit say about remediation decay, and how Keypup MCP tracks action item follow-through to stop the same incidents from recurring.

Liam Davis
Liam Davis
• 15 min read
Postmortem Action Item Decay: Why Your Incident Fixes Never Actually Ship

Table of Contents

TL;DR: Blameless postmortems are good at diagnosing root cause and bad at producing follow-through — action items get written down, assigned an owner, and then compete against roadmap work for every sprint afterward, where they reliably lose. The result is "postmortem action item decay": remediation work that was flagged Critical in an incident review but is still unscheduled 100+ days later, quietly setting up the exact same incident to happen again. The Keypup MCP Server tracks every action item's age, owner, and status against the incident that created it, correlates completion rate with repeat-incident rate, and surfaces a standing team-by-team scorecard — turning "didn't we already fix this?" from a recurring retro question into a number leadership can act on before the repeat incident, not after.

Ask any SRE what happens after a blameless postmortem ships, and you'll get a knowing laugh before the answer. The document gets written, the five-whys are thorough, the action items are specific and assigned — and then, six months later, the exact same service has the exact same outage, for the exact same reason the postmortem already diagnosed. Nobody disagreed with the fix. Nobody forgot it existed. It simply never got scheduled against anything that shipped.

The Friction: A Document That Ends at "Action Items Assigned," Not "Action Items Shipped"

Blameless postmortem culture solved a real problem: engineers stopped hiding incidents and started writing honest, detailed root-cause analyses instead of defensive ones. But the process most organizations built around that culture stops at the wrong finish line.

  • The postmortem's job is considered done once action items have an owner. The review meeting closes, the document is filed, and the tracking responsibility quietly transfers from "incident process" to "whatever that owner's backlog looks like" — with no further checkpoint built into the postmortem process itself.
  • Remediation work competes against roadmap features with no protected capacity. A postmortem action item is just another ticket in the backlog, and when sprint planning weighs a customer-facing feature against a "fix the thing that already got fixed once" ticket, the feature wins almost every time.
  • Nobody owns the aggregate. Individual engineers own individual tickets. Nobody owns the question "of the last 40 action items across all incidents, how many actually shipped" — so the pattern of decay is invisible until the repeat incident forces someone to ask it retroactively.
  • Severity gets assigned once, at review time, and never revisited. An action item tagged "Critical" in the heat of an incident review carries that label forever, even as it ages past 100, 150, 200 days — with no escalation trigger tied to how long it's actually been sitting.

The result: postmortems keep getting more rigorous and more honest, while the thing they're actually supposed to prevent — recurrence — keeps happening anyway, because rigor at diagnosis time was never connected to enforcement at delivery time.

Why This Matters: Decay Isn't a People Problem, It's a Visibility Problem

The "We Already Fixed This" Problem

The single most demoralizing sentence in any incident channel is "didn't we already have a postmortem for this?" It's demoralizing because the answer is usually yes — and because everyone in the room knows the fix was written down, assigned, and never followed up on. Engineers lose faith in the postmortem process itself once they've watched enough of these action items evaporate.

The False Confidence Problem

A closed postmortem document creates a false sense that the risk has been handled. Leadership reads "root cause identified, action items assigned" and reasonably assumes resolution is in motion. Without a standing view of actual completion status, the organization operates on a comfort level the data doesn't support.

The Compounding Risk Problem

Unlike a missed feature deadline, a decayed remediation item doesn't just sit still — it compounds risk silently while everyone assumes it's handled. Every quarter an unaddressed single-point-of-failure fix stays open is another quarter the same outage can recur, except now with less organizational memory of why it happened the first time.

The Enterprise Discussion: What SREs Are Actually Saying

This is one of the most consistently upvoted frustrations wherever engineers at scale compare notes on incident response and reliability practices.

Staff SRE, Enterprise Fintech (r/sre)

"We have a beautiful postmortem template. Root cause, timeline, five whys, action items with owners and due dates. Nobody has ever gone back and checked whether the due dates actually happened. I found an action item from 14 months ago, still 'open,' for the exact failure mode that took us down again last week. The postmortem was perfect. The follow-through was nonexistent."

Principal Engineer, Large-Scale SaaS Platform (r/devops)

"Our incident review process ends the moment the doc gets approved. After that, every action item is just a Jira ticket competing with feature work, and feature work has a PM fighting for it in every planning meeting. Nobody fights for the postmortem ticket. It just loses, quietly, sprint after sprint, until the same alert fires again."

The pattern holds: engineers already know exactly which fixes never shipped — what's missing is a standing, organization-wide number that turns "we keep seeing the same incidents" from a retro anecdote into a remediation-backlog metric leadership actually tracks.

How Keypup MCP Solves Postmortem Action Item Decay

The Keypup Model Context Protocol (MCP) Server treats every postmortem action item as a tracked entity linked back to the incident that created it — surfacing age, owner, severity, and completion status, and correlating the aggregate completion rate directly against repeat-incident rate, all through natural-language prompts against your existing incident and issue-tracker data.

1. Establishing the Baseline: How Much Remediation Work Is Actually Outstanding

Before anything else, leadership needs a single, honest number: of everything a postmortem has ever asked for, how much is still sitting open.

MCP Prompt:

Show me how many postmortem action items are still open across all
incidents in the last 12 months, and how their average age compares
to our stated 30-day remediation SLA.

Output: Postmortem Action Item Health — Trailing 12 Months

KPI cards showing 412 postmortem action items opened in the trailing twelve months, 257 still unresolved at 62 percent, average age of open items at 134 days versus a 30 day target, and 19 repeat incidents linked to stale fixes

Key Insight: Nearly two out of three postmortem action items are still open, and the average unresolved item has been waiting 134 days — more than four times the 30-day SLA most incident review processes claim to enforce.

2. Proving the Correlation Between Decay and Recurrence

A completion-rate number by itself invites the question "does it even matter?" Overlaying it against repeat-incident rate over time removes any doubt.

MCP Prompt:

Plot our postmortem action item completion rate against the
percentage of incidents tied to a previously identified root cause,
for the last six quarters.

Output: Action Item Completion Rate vs. Repeat Incident Rate — Last 6 Quarters

Line chart across six quarters showing action item completion rate falling from 71 percent to 38 percent while the percentage of incidents tied to a previously identified root cause rises from 8 percent to 29 percent

Key Insight: As completion rate fell from 71% to 38% over six quarters, the share of incidents tied to a previously-identified root cause rose from 8% to 29%. The correlation is near-perfect — every quarter action items decay further, repeat incidents climb almost in lockstep.

3. Ranking the Oldest, Highest-Severity Items Right Now

A trend line proves the pattern exists. Leadership and team leads still need the specific list of what to schedule first.

MCP Prompt:

Rank our open postmortem action items by age, and show which ones
were flagged Critical or High severity at the time of review.

Output: Oldest Unresolved Remediation Items, Ranked by Age

Horizontal bar chart ranking six unresolved postmortem action items by age in days, with a single-region payments database failover fix at 211 days and a circuit breaker fix at 178 days both flagged Critical severity at the top

Key Insight: The single-region payments failover fix has been open for 211 days — flagged Critical in a postmortem seven months ago, and still not scheduled against any upcoming sprint.

4. Getting the Full Ownership and Status Detail for Each Overdue Item

Once the oldest, highest-risk items are identified, the team lead needs the operational detail — who owns it, which incident it traces back to, and exactly what's blocking it from being scheduled.

MCP Prompt:

List every postmortem action item overdue against our 30-day SLA,
with the owning engineer, linked incident, severity, and current
backlog status.

Output: Overdue Postmortem Action Items — Owner, Severity, Age

Table of five overdue postmortem action items showing linked incident number, action item description, owner, severity, age in days, and status, with three of the five items flagged Critical or High severity and marked Not Scheduled or Backlog
IncidentAction ItemOwnerSeverityAgeStatus
INC-4471Replace single-region failover for payments DBR. CastilloCritical211 daysNot Scheduled
INC-4512Add circuit breaker to inventory sync serviceJ. PatelCritical178 daysBacklog
INC-4398Rotate stale API keys in checkout gatewayM. LarsenHigh146 daysNot Scheduled
INC-4560Add alert for queue depth on order-events topicT. OseiHigh98 daysIn Progress
INC-4602Document runbook for auth token refresh failureK. NovakMedium62 daysBacklog

Key Insight: Three of the five oldest items are flagged Critical or High severity and are not even scheduled — not in a backlog, not assigned a sprint, simply unaddressed since the incident review closed.

5. Finding Which Teams Let Decay Compound the Longest

Age and severity identify what to fix today. A heatmap across teams and age buckets reveals which parts of the organization have a systemic follow-through problem versus an isolated one.

MCP Prompt:

Build a heatmap of open postmortem action items by team and by age
bucket, so we can see where remediation work is piling up oldest.

Output: Open Action Items by Team and Age Bucket — Decay Heatmap

Heatmap matrix of five teams against five age buckets from 0 to 180 plus days, showing Payments Platform with the heaviest concentration of items in the 91 to 180 day and 180 plus day buckets while Search and Discovery clears nearly everything within 60 days

Key Insight: Payments Platform has 11 open action items older than 90 days — more than double any other team — while Search & Discovery clears nearly everything inside 60 days. Same postmortem process, wildly different follow-through.

6. Maintaining a Standing Scorecard for Ongoing Accountability

A one-time heatmap identifies today's hotspot. A standing scorecard, refreshed every quarter, is what actually changes behavior over time by making follow-through a tracked, visible metric rather than a one-off audit.

MCP Prompt:

Build a standing scorecard showing each team's 90-day postmortem
action item completion rate, and flag any team in severe decay.

Output: Postmortem Follow-Through Scorecard — 90-Day Action Item Completion Rate by Team

Scorecard of five teams showing 90 day action item completion rate as a horizontal progress bar, with Payments Platform at 31 percent flagged Severe Decay and Search and Discovery at 88 percent flagged Healthy

Key Insight: Payments Platform closes only 31% of its action items within 90 days — flagged Severe Decay — while Search & Discovery closes 88% in the same window using the identical postmortem template. The process isn't broken; follow-through enforcement is.

The Technical Implementation: How Keypup MCP Sees Postmortem Follow-Through

Linking Action Items Back to the Originating Incident

Keypup MCP treats each postmortem action item as a tracked artifact with a persistent link to the incident that produced it, so an item's age, severity, and status are always evaluated in the context of what it was supposed to prevent — not as a generic backlog ticket that happens to mention an outage.

Tracking Age and Severity Together, Not Severity Alone

Rather than relying on a severity label assigned once at review time, Keypup MCP continuously tracks how long each item has been open relative to its severity tier, so a Critical item aging past 200 days surfaces very differently from a Medium item at the same age — instead of both fading into an undifferentiated backlog.

Correlating Completion Rate Against Repeat Incidents Automatically

Keypup MCP cross-references closed incidents against prior postmortem action items for the same service or failure mode, so a repeat incident is automatically flagged as tied to a previously-identified root cause — turning "didn't we already fix this?" from a question someone has to remember to ask into a metric that's always being tracked.

Real-World Impact: Enterprise Case Studies

Case Study 1: Global Payments Technology Provider

Before Keypup MCP:

  • The Payments Platform team had experienced three variations of the same database failover incident over 18 months, each followed by a postmortem with nearly identical action items
  • No visibility into whether prior action items had actually been completed before the next incident occurred, leaving every retro to start from scratch
  • Remediation work was scored as "handled" the moment it was assigned, with no tracking of actual delivery

After Keypup MCP:

  • The 211-day-old single-region failover fix was surfaced as the common thread across all three incidents, finally connecting the dots that three separate postmortems had each independently diagnosed
  • A dedicated reliability sprint was funded specifically for aged Critical action items, prioritized using the age-and-severity ranking instead of competing unsuccessfully against every other roadmap item
  • 90-day action item completion rate on the team rose from 31% to 76% within two quarters, and the specific failure mode has not recurred since the fix shipped

"We'd written three postmortems for essentially the same incident before we had a number that showed us the fix from the first one was still sitting open when the second and third happened. That's not an estimation problem or a prioritization disagreement — it's a visibility gap. Once we could see action item age next to severity, funding the fix was a five-minute conversation instead of a fourth postmortem."

— VP Engineering, Global Payments Technology Provider

Case Study 2: Enterprise Identity and Access Management Vendor

Before Keypup MCP:

  • Postmortem completion was reported qualitatively in incident reviews ("most action items are being worked"), with no quantitative tracking across the organization
  • Individual team leads had no visibility into how their team's follow-through compared to other teams running the same postmortem process
  • Leadership had no standing metric to reference when deciding where to fund reliability investment versus feature work

After Keypup MCP:

  • A quarterly follow-through scorecard was established across all teams, immediately surfacing a 57-point gap in 90-day completion rate between the best and worst-performing teams
  • Reliability investment was reallocated toward the two lowest-scoring teams, directly justified by the scorecard rather than a general request for more engineering capacity
  • Org-wide average completion rate improved from 48% to 71% within three quarters, with the scorecard now a standing agenda item in quarterly engineering reviews

"We assumed every team was roughly equally good at following through on postmortems, because nobody had ever measured it. The first scorecard showed a 57-point gap between our best and worst teams using the exact same process. That single number did more to get reliability work funded than every qualitative incident review we'd run in the previous two years combined."

— Director of Engineering, Enterprise Identity and Access Management Vendor

Implementation: Getting Started with Keypup MCP

1. Connect Your Incident Management System and Issue Tracker Together

Action item decay tracking only works once Keypup MCP can see both the incident that generated the action item and the ticket tracking its remediation — connect your incident management platform alongside your existing Git and project-tracker integrations.

2. Stop Treating "Action Items Assigned" as the Finish Line

Report postmortem completion as a tracked, ongoing status — age, severity, and owner — rather than a one-time checkbox that gets marked done the moment the review meeting ends.

3. Query in Natural Language, Per Team or Per Incident

No manual cross-referencing between the postmortem doc and the backlog required. Just ask:

  • "Which postmortem action items are overdue against our SLA right now?"
  • "Is this incident tied to a previously identified root cause?"
  • "Which team has the worst 90-day action item completion rate?"
  • "Show me every Critical action item older than 100 days"

4. Automate Recurring Root-Cause Reporting

Schedule recurring queries for:

  • Quarterly action item completion rate versus repeat-incident rate correlation, ahead of every reliability review
  • Standing age-and-severity rankings of the oldest unresolved remediation items
  • Team-by-team decay heatmaps to identify where follow-through is systemically weakest
  • A standing postmortem follow-through scorecard ahead of every reliability-investment and staffing conversation

The Bottom Line: A Postmortem Isn't Done When It's Written, It's Done When It Ships

A blameless postmortem process that stops tracking the moment action items are assigned will keep producing excellent root-cause analysis for incidents that recur anyway — not occasionally, but predictably, in every enterprise running production systems at scale. Fixing it requires:

✅ Tracking every action item's age and severity against the incident that created it, not treating it as a generic backlog ticket ✅ Correlating completion rate directly against repeat-incident rate, so decay reads as a measurable risk, not an anecdote ✅ Ranking the oldest, highest-severity items continuously, so remediation work can compete for sprint capacity with real urgency data attached ✅ Surfacing which teams decay fastest, to target reliability investment and process support where it's actually needed ✅ Maintaining a standing follow-through scorecard, so "didn't we already fix this?" never has to be asked again

The Keypup MCP Server turns "why do we keep having the same incident" from a recurring, demoralizing retro question into a precise answer — this action item, this team, this age — giving engineering leaders the data to close the loop a blameless postmortem process was always supposed to close.


Get Started

Ready to stop finding out your postmortem action items decayed only after the repeat incident?

Start Free Trial — Connect your incident management and issue tracker in minutes and start tracking postmortem follow-through automatically.

Request Demo — See how enterprise engineering leaders use Keypup MCP to quantify and close the postmortem action item gap.

View MCP Documentation — Technical details for engineering teams implementing Keypup MCP across Git, issue-tracker, and incident data.


Ready to Transform Your Analytics?

Join teams already using AI to make data-driven decisions faster than ever.

Most Recent Articles