Security Post-Incident Review
Activated Cloud✓ Officialactivated/security-post-incident-review
Free · MIT
About
Runs a blameless post-incident review once a security incident or near miss is contained: rebuilds the timeline from evidence, measures detection and response times, finds the contributing causes, and turns them into tracked actions with owners and dates. Also prepares a shorter external version for customers, insurers or regulators. Not for handling an incident that is still active (use security-incident-triage).
Documentation
Security Post-Incident Review
You turn a bad day into lasting improvements. The review explains what happened and why it was possible, without blaming people, and produces a small set of actions that are owned, dated and followed up until done. A review whose actions never close was theatre.
When to use
- "The incident is over; let's write it up."
- "The customer / insurer / regulator wants an incident report."
- "We had a near miss with that phishing email; what should we learn?"
- Thirty, sixty and ninety days after a review, to chase open actions.
What you need
- The incident log, evidence register and situation updates from the incident (the security-incident-triage output, or whatever was kept).
- Chat and ticket history from the response, read from connected apps (Slack, Jira, Linear, Gmail) or exports the owner provides.
- Relevant logs to confirm times, especially the earliest malicious activity.
- Time with the people who responded: interviews with
ask_teammate, or a 60-minute review meeting the owner calls.
Method
- Schedule it while memories are fresh. Aim to hold the review meeting within five working days of containment. Name one review owner (often the incident commander) who is accountable for the document.
- Set the tone in writing at the top: the goal is to understand how the system, process and tools allowed this, not who to blame. People acted on the information they had at the time. If people fear blame, they hide facts and the next incident repeats.
- Rebuild the timeline from evidence, not memory, in UTC. Mark each entry with its source. Include the attacker's actions (as far as known), detection, every decision and action, and communications. Note where the evidence is missing.
- Measure the response:
Metric Definition Dwell time Earliest malicious activity to detection Time to acknowledge Detection or alert to a human starting work Time to contain Detection to containment confirmed Time to recover Detection to normal service restored Notification timing Time of awareness to each required notice, against its deadline Compare with any previous incidents to see a trend. - Find contributing factors, not a single root cause. Incidents almost always need several conditions at once. Ask "how did this make sense at the time?" and "what made this possible?" across six areas:
- technical (a missing control, a misconfiguration, an unpatched system)
- detection (why it was not caught earlier; which log or alert was missing)
- process (a step that did not exist or was skipped)
- people and training (unclear ownership, unfamiliar tooling)
- third parties (vendor or supplier factors)
- response (what slowed containment) Use "five whys" only as a prompt, branching wherever more than one answer is true.
- Record what went well and what was lucky. Luck is a hidden risk: if containment worked only because someone happened to be online, that is an action.
- Write actions that close. Each action is specific, has one owner, a due date, and a type:
- prevent (stops it happening again)
- detect (catches it sooner)
- respond (handles it faster or better) Prefer a few high-impact actions over a long wish list. Put them in the team's tracker if it is a connected app; link the tickets in the document.
- Review meeting agenda (60 minutes): read the timeline together (15), discuss contributing factors (20), agree actions and owners (15), what went well (5), wrap-up (5). You take notes and update the document live.
- External version, if needed: a shorter document for customers, insurers or regulators with facts only: what happened, when, what data was affected, what was done, what is changing. It goes through legal review and the owner's approval before it leaves the company. You never send it yourself.
- Follow up. Set a
cronjobto check action status at 30, 60 and 90 days and post the status to the owner. Overdue Highs go to the owner by name.
Questions that surface contributing factors
- What did the people involved know at each decision point, and where did that information come from?
- What made the risky action easy, or the safe action hard?
- Which alert or log would have caught this earlier, and why did it not exist or not fire?
- Who owned the affected system, and did they know they did?
- Was a policy or procedure in place? Was it known, practical and followed?
- Which third party was involved, and what did they tell us and when?
- What slowed containment: access, knowledge, tools, approvals, time zone?
- Has anything similar happened before, here or to peers?
- What would have happened if one lucky break had not occurred?
Blameless wording
| Instead of | Write |
|---|---|
| "Sam carelessly clicked the phishing link." | "A phishing email reached Sam's inbox and the link opened a convincing login page; MFA was not enforced on that account." |
| "Ops forgot to rotate the key." | "Key rotation depended on memory; no reminder or expiry existed." |
| "The developer pushed a secret to GitHub." | "Nothing in the commit process scanned for secrets before push." |
Output
The internal review document using references/pir-template.md; the action list on a show_card (action, owner, due, type, status); and, on request, the external version marked DRAFT for legal review.
Checks before you finish
- Every timeline entry has a UTC time and a source; gaps are marked.
- The document names roles and systems, not people to blame.
- Every contributing factor has at least one action or a recorded decision not to act.
- Every action has one owner and a date.
- The external version contains no speculation and was sent to the owner for legal review.
- Follow-up reminders are scheduled.
Pitfalls
- "Human error" as the cause. It is a starting point. Ask why the error was easy to make and hard to catch.
- One root cause. Picking a single cause hides the other conditions that will combine next time.
- Thirty actions. Nobody does thirty actions. Rank and keep the five that matter most.
- Writing it weeks later. Details fade and logs rotate. Start the draft as soon as the incident closes.
- No follow-up. Track to closure, or the review was wasted.
- Sign-off. The owner approves the internal review and its actions. Anything sent outside the company, especially to regulators, needs legal sign-off first.
Credits for adapted material: references/CREDITS.md.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
