Activated Cloud
← App Store

Incident First Response

Activated Cloud✓ Officialactivated/incident-first-response

No ratings yet6 installsv1.0.0Updated Oct 6, 2026● Unknown

Free · MIT

About

Runs the first hour of a production incident: confirm and size the user impact, declare a severity (when unsure, pick the higher), get the owner and the right people involved, keep one clear coordinator, mitigate before diagnosing (roll back, restart, scale, disable a flag), keep a timestamped log, post regular status updates, confirm recovery, and set up a blameless postmortem. Risky production actions need the owner's go-ahead unless a runbook authorises them. Use when something is down, broken for users, losing data or alerting. Not for routine bugs (use debug-root-cause).

Software Development

Documentation

From SKILL.md · v1.0.0 · what the agent reads when it loads this skill4 files: SKILL.md, references/CREDITS.md, references/incident-templates.md, references/triage-commands.md

Incident First Response

In an incident the goal is to stop the harm fast and safely, not to understand everything first. That means: know the impact, take a clear role, restore service with the least risky action that works, and write everything down as you go so the team can learn afterwards. Calm, structured and loud about what you are doing.

When to use

  • "The site is down", "customers can't log in", "payments are failing", "the queue is backing up", "we're getting 500s".
  • A monitoring alert that shows real user impact, or a credible report of data loss or exposure.
  • Something that will become user-visible soon if nothing is done (disk nearly full, certificate expiring today, replica lag climbing).

What you need

  • Access to the signals: logs, metrics, error tracker, status pages, through the owner's connected apps or the browser the owner signed in, or SSH keys on your computer. If you have none, say so at once and tell the owner exactly what access would help.
  • The owner's fastest channel, and the list of people to involve (memory, team roster). Use brief_team or ask_teammate to reach teammates.
  • Runbooks for the affected service, if the team has them. A runbook step that is pre-authorised can be done without asking; anything else risky in production needs the owner's go-ahead.

Method

First five minutes

  1. Confirm it is real and size it. What do users see? How many (all, a region, a plan, one customer)? Since when? Which features? Check from outside: curl -sS -o /dev/null -w '%{http_code} %{time_total}s\n' https://<host>/<key-path>, the app in the browser, the error tracker, recent alerts.

  2. Declare a severity with this scale, and when unsure between two levels, pick the higher; do not spend time debating it:

    Level Meaning
    SEV1 Critical: most users cannot use the product, or data is being lost or exposed, or money is moving wrongly
    SEV2 Major: a core feature is broken for many users
    SEV3 Partial: a feature degraded for some users, or a risk of becoming SEV2 if nothing is done
    SEV4 Minor: no user impact yet, needs action soon
    SEV1 and SEV2 are major incidents: run the full process below.
  3. Tell the owner now for SEV1 and SEV2, in one message. Do not wait until you understand the cause:

    SEV2: checkout failing for about a third of attempts since 09:05 UTC.
    Started right after the 09:05 release (v2.4.0). I recommend rolling back to v2.3.2 now; approve?
    Next update 09:30 or sooner.
    
  4. Open the incident log (template in references/incident-templates.md) and keep it going: UTC timestamps, observations, actions and their results.

Roles

  1. One person coordinates (the incident commander): sets direction, approves actions, keeps the log and updates flowing. Others investigate or act on request. If the owner or a teammate takes command, follow their instructions and report findings to them; do not take actions on your own initiative during someone else's command. If you are coordinating, do not disappear into debugging: hand the deep investigation to someone else, or explicitly hand over command first.

Mitigate first

  1. Look for the obvious trigger: a deploy, config change, feature flag, migration, cron job, traffic spike, dependency outage (check the provider's status page with web_extract), expired certificate, full disk. git log --since="3 hours ago" --oneline on the deployed branch and the deploy history are quick wins.

  2. Pick the least risky action that stops the harm:

    Situation Mitigation
    Started after a deploy or config change Roll back that change
    A new feature misbehaving Turn its flag off
    Process stuck, leaking memory, wedged Restart it (capture evidence first, step 8)
    Overloaded by traffic Scale out; rate-limit or shed non-critical load; block abusive sources
    A dependency is down Fail over, degrade gracefully, queue work for later, tell users
    Disk full Free space safely (rotate or move logs, clear caches), never delete data
    Data being corrupted or exposed Stop the writer or close the path immediately; protect backups
    Each action: say what you will do, get the go-ahead (owner or runbook), do it, watch the signal for the expected effect, and log the result. If it did not help, consider undoing it before trying the next thing.
  3. Preserve evidence before destroying state. Before restarting or rolling back, grab what you can in a minute: recent logs (docker logs --since 30m <c> > /tmp/inc-<c>.log), a thread dump or py-spy dump, ps, df -h, free -m, the failing request's details. Restarts erase the clues the postmortem needs.

Diagnose with evidence

  1. Use the quick checks in references/triage-commands.md. For resources, look at utilisation, saturation and errors (CPU, memory, disk, network, connections); for services, look at request rate, error rate and duration. Compare with a healthy period, not with zero. Work hypotheses one at a time and log them.
  2. Security incidents (suspected breach, leaked credentials, unknown access): tell the owner privately at once, preserve evidence, rotate the affected credentials only with the owner's go-ahead, do not tip off an attacker in public channels, and never probe or scan systems you are not authorised to touch. A qualified security person or the owner leads.

Communicate

  1. Status updates every 20 to 30 minutes during a major incident, or sooner when something changes, to the owner and the agreed audience: impact now, what we know, what we are doing, next update time (template in references/incident-templates.md). Silence makes people interrupt responders. Customer-facing messages and status-page posts go out only through the owner or the person they designate.
  2. Escalate without hesitation. If you are stuck for 15 minutes or out of your depth, pull in the teammate or vendor who knows the system. Involve only the people needed; let others go when their part is done.

Resolve and hand over

  1. Confirm recovery with the same signals that showed the problem: error rate back to baseline, key flows passing, queues draining, no new errors for a sustained period (at least 15 to 30 minutes for a major incident). Then declare it resolved, say so in the log and to the owner.
  2. List the cleanup: temporary mitigations to undo or make permanent, data to repair, customers to contact, monitoring gaps.
  3. Set up the postmortem within five business days for SEV1 and SEV2 (and any incident where responders were mobilised): draft it from your log using the blameless template in references/incident-templates.md. Focus on contributing factors and system fixes, not on who made a mistake. Every action item gets an owner and an issue.

Output

During the incident: the live log and regular status updates. At the end: a resolution note (impact with numbers, duration, what fixed it, cleanup list) and a postmortem draft with a timeline, contributing factors, what went well and badly, and action items.

Checks before you finish

  • The owner was told promptly, with impact and severity.
  • Every production action was approved (owner or runbook), logged with time and result.
  • Evidence was captured before restarts or rollbacks.
  • Recovery was confirmed with the original symptoms' signals over a sustained window.
  • Cleanup items and a postmortem with owned action items exist.

Pitfalls

  • Debugging while users burn. If a rollback is likely to stop the harm, roll back first and investigate after.
  • Arguing about severity. Pick the higher one and move on; review it in the postmortem.
  • Several people changing production at once. One coordinator approves actions so effects can be attributed.
  • Restarting away the evidence. One minute of capture saves days of guessing later.
  • Going silent. Regular short updates beat one long one at the end.
  • The hero move. Taking every task yourself, or doing someone else's assigned task without telling them, causes collisions. Coordinate.
  • Skipping the postmortem because the cause seems obvious. The obvious cause is usually one of several contributing factors.
  • Blame. People stop reporting problems early when mistakes are punished. Fix the system that allowed the mistake.

Versions

v1.0.0currentOct 6, 2026

Listed from the source repository.

Reviews

No reviews yet. Be the first.

Write a review