Post-incident reviews are where most SOCs quietly fail

DCDaniel C. · Head of Security Operations
Incident Response

Eleven action items assigned to "the team" is what a failed post-incident review program looks like from the inside. Here's the four-check test to find out if yours is real, and the accountability language that quietly stopped meaning anything along the way.

The most expensive document my security operations center (SOC) produced last year was a 12-page incident response report that changed nothing. I sat through its review after a contractor's account had been used to pull data off a file share over a long weekend; my team caught it Tuesday. The meeting ran ninety minutes, the timeline was accurate to the minute, the attack-technique mapping was tidy, and we walked out with 14 action items. Eleven were assigned to "the team." None had a date.

Six months later I sat in a review of a nearly identical incident, same class of account, same missing control, and nobody in the room connected the two until I pulled up the first document. I fund the tooling that feeds these reviews, yet understood the review program least. The meeting gets held, the document gets filed, and its completion goes up the chain as a green box. Nobody reports whether anything closed. That gap stays quiet because completing the review looks exactly like completing the fix.

In Brief:

  • Many SOCs measure whether a review was held, but teams should also audit whether its action items closed; Google's SRE Workbook calls missing ownership a failure mode, while NIST SP 800-61r3 recommends periodically evaluating incident response program performance.
  • "Blameless" describes how you analyze the past. It has been stretched into a reason not to put a name and a date on the future, which is a different thing entirely.
  • Regulators keep documenting the same control gap exploited twice, or three times, at the same organization. Each of those is a review program that failed, often with a fine attached.
  • One afternoon of checks against the last four incident reviews settles whether the program is real.

Special Publication (SP) 800-61r2, now withdrawn and superseded by Revision 3, called for a follow-up document for each resolved incident, useful "for evidentiary purposes" and "for reference in handling future incidents and in training new team members," per its follow-up report guidance. The EU's Digital Operational Resilience Act does something similar for financial entities: Article 19 requires a final report once root cause and actual impact are known, and the applicable reporting standards set that deadline at one month after the last intermediate report. A record can be accurate, on time, and still change nothing in the environment. I've started treating the report as a ledger of commitments instead: who agreed to change what, by when, tracked where, and verified how.

The review that produced a 12-page document and changed nothing

John Allspaw of Adaptive Capacity Labs has a line for this: most incidents are written to be filed, not to be read or learned from. Nora Jones has described the result as a "write only culture" for incident reviews, one that exists to defend the person writing it more than to help the next person reading it. My report had a defensive purpose too: it proved to the board that we'd been thorough.

Nine people sat in the room for ninety minutes, and an analyst spent most of a week writing. The eleven ownerless items sat in an appendix where I found them six months later, untouched. We checked the learning box and skipped the repair, and that class of contractor account stayed unmanaged until the second incident.

Why "blameless" quietly became "no accountability"

John Allspaw articulated Etsy's blameless-postmortem approach in a May 2012 Code as Craft post, drawing on Dekker's human-error work. That original post emphasizes candid accounts without fear of punishment; it doesn't itself spell out accountability for future work, which is a separate governance practice. As one peer-reviewed nursing-management article puts it, "just culture isn't a blame-free culture, rather a culture of balanced accountability." In my SOC, the word got stretched until it meant nobody gets a name next to anything, past or future. One managed detection and response (MDR) provider I evaluated shipped findings reports whose action-item column recommended that the customer review its privileged access policy, with no owner, no date, and no ticket reference.

Blame points backward (an analyst should have had the alert in place); accountability points forward (an analyst will add the alert by Friday). The incident.io framing, vendor-authored but right, is to say out loud in the debrief that "we're assigning ownership of the fix, not responsibility for the incident," per their action-item ownership guidance. Google's SRE team drew the harder line in 2017: "The surest way for a postmortem author to ensure that an action item never gets completed is to leave it without an owner," per their action-item analysis. In my SOC, "the team" owned everything, which meant nobody did.

The same root cause, three incidents in a row, and nobody connected them

The public record includes organizations where the same control gap was exploited twice, or three times, after the weakness had already been flagged internally.

Ownerless action items without follow-through

Capita got three chances. The United Kingdom regulator fined the company £14m in 2025 for a breach where the missing admin-account tiering had been "flagged as a vulnerability on at least three separate occasions but were not remedied," per the regulatory penalty notice. The Federal Trade Commission's (FTC's) Uber complaint tells it twice: in 2014 intruders used an access key an engineer had posted to GitHub. In 2016, per the revised breach complaint, "intruders gained access to the Amazon S3 Datastore using an access key that an Uber engineer had posted to GitHub."

Mandiant's M-Trends 2026 research found that prior compromise became the top initial infection vector in ransomware operations at 30% in 2025, double what it was in 2024. My own repeat cost a week of analyst time re-running an investigation we'd already written up, and a board meeting explaining why the same class of contractor account walked out with data twice.

The items lacked accountable follow-through. Google's SRE Workbook names "Missing ownership" as a failure mode outright and says that "Action items without clear owners are less likely to be resolved" in its missing-ownership guidance. Google's own SRE Prodcast has made a similar point in discussing action-item verification: open-ended monitoring items rarely function as real action items. Left unowned, that remaining work competes for attention against roadmap commitments with a customer's name attached, and tends to lose.

Review completion gets tracked. Action-item completion almost never does.

Every framework I've been audited against requires holding the review and identifying improvements. None prescribes one universal metric for proving a corrective action actually got implemented and stayed effective, and that gap is what a program can hide behind. Special Publication 800-61r3, finalized April 2025, folded the old post-incident phase into the framework's Improvement category. Its ID.IM-01 recommendation is to "periodically evaluate incident response program performance to identify problems and deficiencies that should be corrected," per the current incident-response revision. Controls v8.1 Safeguard 17.8 goes further, asking for reviews that identify "lessons learned and follow-up action" in its post-incident review safeguard.

The Preparation, Identification, Containment, Eradication, Recovery, and Lessons Learned (PICERL) model and the 27035-2:2023 standard call for the same thing: identify improvements, with no built-in mechanism for tracking a named item to closure. Holding the meeting and filing the document can look compliant without that follow-up evidence existing anywhere.

The United Kingdom breach survey for 2025/2026 found that, among organizations identifying breaches or attacks, 61% of businesses and 57% of charities took some action to prevent future incidents; the published results don't report a separate measure of whether that follow-up action closed. The 2026 SOC metrics survey found 70% of SOCs still cite incident count as their top metric.

I went looking for a population-level number on what share of postmortem action items ever close and found none with disclosed methodology. The closest is incident.io's editorial benchmark that below a 50% completion rate, "your team writes post-mortems to satisfy a process, not to change anything," a vendor's rule of thumb rather than research. My board slide reported a 100% post-incident review completion rate, accurate every quarter and silent on whether the environment had changed.

The four-incident test I run before I trust a review program

When I take over a program, and when I evaluate an MDR provider's findings reports, I assess review quality by pulling the last four, plus their tickets, and running the checks below in an afternoon. Every check is a yes or a no, and two failures across the six and I don't trust the program, my own included: your provider's findings report is the review, and your ticket queue is where it closes or dies.

  • Every action item names one individual (not a team), a calendar date, and a verifiable end state. "Improve monitoring" fails. "Add the pool-exhaustion alert to the ops dashboard by March 14" passes. The standard comes from Google's action-item ownership standard.
  • Each item exists as a ticket in the backlog the team actually plans from, linked back to the review. Anything living only inside the review document counts as ownerless.
  • Closed items carry a verification record: a 30/60/90-day check, a monitoring result, a test output, or another documented check. Closed on assertion doesn't count.
  • The created-versus-closed count across the four reviews isn't widening. Google's own action-item research treats a sustained gap between items created and items closed as accumulating technical debt worth tracking over time.
  • Searching incident history for each review's root-cause category turns up no earlier incident with the same cause, or turns up one whose items were still open. A repeat after items were marked closed is a signal to go back and check whether the fix addressed the actual cause, whether it was verified, or whether conditions simply changed.
  • At least one item per review names a detection rule or playbook by identifier, where a detection or playbook gap was part of the finding, and the rule repository or playbook history shows a commit dated after the review. The 11 Strategies guide ties post-incident review quality to detection and response improvement; the commit shows the change was proposed, though it's not proof on its own that the change was deployed, tested, or actually working.

What a review has to produce to count as real

Etsy frames a debrief as a learning exercise in its debriefing facilitation guide. Allspaw goes further, in his incident review observations: "A manager who complains that too few action items were produced has revealed his/her real interest: the reduction of an incident to a manageable discrete list of things to be done." I'm that manager, and I think he's right about the meeting and wrong about the program: the ninety minutes should go to understanding how a contractor account with no multifactor authentication (MFA) sat in scope for two years without anyone owning it. The ledger of who fixes what comes after, and it belongs to someone with backlog authority; Atlassian's own incident-management guidance describes a similar approval model, with approvers expected to prioritize follow-up work in their own backlog.

A real review produces two artifacts, and I fund both. The first is the learning writeup, written to be read by the people who weren't in the room. The second is a table at the top of the incident response report with a name, a date, a ticket link, and a verification step on every line. It's owned by a manager who can bump roadmap work to make room, with detection engineering present to claim the items that are theirs.

The report I sign off on now runs well under 12 pages, and that table is page one. I stopped reporting review completion to the board and started reporting created versus closed; the first quarter I did, the gap was embarrassing enough to get the contractor access project funded.

Frequently asked questions about post-incident reviews

What's the difference between a post-incident review and a post-mortem?

In security operations the terms get used interchangeably, and the difference is mostly lineage. "Postmortem" is a much older term borrowed from medicine and safety engineering; Google's SRE practice popularized a specific blameless, written form of it, defined as a record of an incident's impact, the resolution actions, root causes, and follow-up actions meant to prevent recurrence. "Post-incident review" is the broader term, and it maps loosely onto the U.S. Army's after-action review, which asks what was expected, what actually happened, why any gap existed, and what to sustain or improve.

How do you measure whether post-incident reviews are actually working?

Plot action items created against action items closed over time; a gap that keeps growing means the program is generating debt, not fixes. Confirm each closed item has a verification record, and search incident history for repeats of a root cause after its items were closed. The incident.io vendor benchmarks, 80% closure with high-priority items done inside 30 days and a repeat-incident rate under 30%, are workable targets even though no independent population data exists, per their postmortem benchmark guidance.

Who should own action items from an incident review?

One named individual per item, never a team, assigned during the meeting in the tracker the team plans from. Google SRE's rule is to assign an owner as it is enacted, even if that owner's primary task is to find the best person for the job. A manager with backlog authority approves the full set and answers for prioritizing it against roadmap work. Detection engineers belong in the room so detection and playbook items land with the people who'll ship them.

How often should a SOC revisit open action items from past reviews?

Weekly for high-priority items, folded into whatever retrospective or ops meeting already exists; Rootly's own action-item guidance names weekly ops meetings as one venue for this. Measure closure against a 30-day window, and hold a quarterly leadership review that connects closed action items to the actual reliability trend. Any item still open at the quarter mark gets a new owner and date or gets closed as won't-fix, on the record, so it can't quietly age into irrelevance.


About the author

DCDaniel C. is a security operations leader with over a decade of experience building and scaling SOC capabilities for cloud-native companies. He has led security teams through multiple stages of growth — from early-stage environments with minimal tooling to mature organizations operating 24/7 security operations with distributed teams. His experience includes designing SOC architectures, evaluating and managing MDR providers, and building internal detection and response capabilities. Daniel has been responsible for vendor selection across SIEM, EDR, and XDR platforms, as well as defining SLAs, response models, and escalation frameworks. He has also worked closely with executive leadership on budgeting, board reporting, and aligning security operations with broader business risk. He writes about the practical decisions security leaders face — including build vs buy tradeoffs, how to evaluate security vendors, and what it actually takes to run an effective security operations function at scale

Stay sharp on security operations

Practitioner takes on SOC modernization, detection engineering, threat hunting, and more. No fluff. No product pitches.

Post-incident reviews are where most SOCs quietly fail | Future of SecOps