Deploy and Roll Back
Activated Cloud✓ Officialactivated/deploy-and-rollback
Free · MIT
About
Ships a change to a running environment with a tested way back: pin exactly what is going out and where, record the current version, check migrations and config for compatibility, deploy through the project's established method, verify with real requests, logs and metrics, watch a bake period against rollback triggers set in advance, and roll back fast when they fire. Production deploys and rollbacks need the owner's explicit go-ahead. Use when asked to deploy, release to staging or production, or undo a bad deploy. Not for publishing a library version (use cut-a-release).
Documentation
Deploy and Roll Back
A deploy is not done when the pipeline turns green; it is done when real requests succeed on the new version and the error rate has not moved. And every deploy starts with knowing how to undo it: the previous version is recorded, the rollback command is written down, and the signals that would trigger it are decided before anything ships.
When to use
- "Deploy this to staging", "ship it to production", "release the fix", "push the new build to the server".
- "Roll back", "the last deploy broke something", "revert production".
- Setting up or reviewing a deployment process.
What you need
- The project's deployment method, from its docs, scripts or pipeline: a CI/CD job, a platform CLI,
docker composeon a server, Kubernetes manifests or Helm, a deploy script. Use the established method; do not invent a new one for production. - Access as the project grants it: the CI system through the owner's connected app (GitHub, GitLab), the hosting provider's dashboard in the browser the owner signed in, or SSH keys the owner installed. If you lack access, ask the owner; never ask for passwords in chat.
- The owner's explicit go-ahead for every production deploy and rollback, unless a written runbook already authorises you for this service. Staging and preview environments follow the owner's standing rules in
memory. - Health signals: a health endpoint, logs, an error tracker or metrics dashboard (connected app or browser).
Method
Pin what and where. Write down the exact commit SHA or image tag being deployed, the target environment, and what changed since the version now running:
git log --oneline <deployed-sha>..<new-sha> git diff --stat <deployed-sha>..<new-sha>Get the currently deployed version from the version endpoint, the platform's release list, the running image tag (
docker inspect,kubectl get deploy -o wide) or the last deploy log. This is your rollback target.Check what makes this deploy risky. Use
references/deploy-checklist.md:- Database migrations: are they backward compatible with the old code (additive columns, no renames or drops in the same release)? Old and new code will run side by side during the rollout, and a rollback runs old code on the new schema.
- Config and secrets: new environment variables set in the target environment before the code that needs them?
- Feature flags: risky behaviour behind a flag, default off?
- Dependencies on other services: does this release need another service deployed first?
- Timing: a low-traffic window, and someone (the owner, or you with authority) available to watch for the next hour. Avoid deploying right before the team goes offline.
Decide rollback triggers in advance. Write them down with the owner before deploying; during a bad deploy is the wrong time to negotiate thresholds. A typical starting set, to adjust per service:
Signal Roll back when Health check Failing for 2 minutes Error rate (5xx or failed jobs) Above twice the pre-deploy baseline for 5 minutes Latency p95 up more than 50 percent for 10 minutes Key flows (login, checkout, main API call) Any scripted check fails twice in a row Data Any sign of wrong or lost writes: roll back at once and tell the owner Get the go-ahead. For production, put a
show_cardin front of the owner: version, changes, migrations, risks, rollback target and command, triggers, and the time window. Proceed only on an explicit yes.Deploy to staging first when one exists, with the same artefact (same image tag, same build). Run the smoke checks from step 7 there.
Deploy with the established method. Trigger the pipeline or run the documented command; watch it to completion (
gh run watch, the platform's deploy log,kubectl rollout status deploy/<name> --timeout=5m). Do not hand-edit files on servers or patch containers in place: that creates drift nobody can reproduce. If an emergency forces a manual change, record exactly what you did and turn it into a proper change afterwards.Verify with real signals.
- The running version is the new one (version endpoint, image tag, release list).
- Health endpoint returns healthy:
curl -fsS https://<host>/health. - Key flows work: scripted requests with expected status and content (
curl -sS -w '\n%{http_code}\n'), and for user-facing apps a quick pass in the browser (browser_navigate,browser_snapshot). - Logs show no new error types:
docker compose logs --since 10m <svc>,kubectl logs deploy/<name> --since=10m,journalctl -u <svc> --since "10 min ago", or the error tracker. - Error rate and latency compared with the pre-deploy baseline.
Bake. Keep watching for the agreed period (often 15 to 30 minutes, longer for low-traffic services or slow-burning effects like queues and caches). Check the triggers at intervals. If you must step away, schedule a check with
cronjobor hand the watch to a teammate explicitly.Roll back when a trigger fires. Do not try a forward fix under pressure unless the cause is known, trivial and safer than rolling back. Rollback paths (exact commands for common setups are in
references/deploy-checklist.md):- redeploy the previous image tag or release through the same pipeline;
kubectl rollout undo deploy/<name>(thenkubectl rollout status);- the platform's "redeploy previous release" action;
git revert <sha>and deploy the revert through the normal pipeline;- turn the feature flag off. Migrations: roll back the code, not the schema, unless the down migration is known to be safe and data-preserving. This is why step 2 insists on backward-compatible migrations. After rolling back, verify the old version is healthy with the same checks, then tell the owner what happened.
Record and report. A deploy log entry: what, when, who approved, result, anything odd. If you rolled back, open an issue with the evidence (errors, graphs, times) so the fix can be investigated calmly.
Output
A short deploy note: environment, version deployed (and previous version), approval, checks run with results, bake outcome, and any follow-ups. For a rollback: the trigger that fired with evidence, the rollback action, verification that the old version is healthy, and the issue opened.
Checks before you finish
- The previous version and the rollback command were recorded before deploying.
- The owner explicitly approved the production deploy or rollback.
- The running version was confirmed, and health, key flows, logs and error rates were checked against the baseline.
- The bake period ran its agreed length, or the watch was explicitly handed over.
- The deploy or rollback is recorded, and the owner was told the outcome.
Pitfalls
- "The pipeline is green, so it's deployed." Pipelines report their steps, not your users' experience. Check the live system.
- Breaking migrations in the same release as the code. A rename or drop makes rollback impossible. Expand first, contract in a later release.
- No recorded rollback target. Finding out which version was running before, mid-incident, wastes the minutes that matter.
- Forward-fixing under pressure. The quick fix often adds a second bug. Roll back, then fix calmly.
- Manual server edits. The next deploy overwrites them, or worse, keeps them invisibly.
- Deploying and walking away. Many failures show up only under real traffic over tens of minutes.
- Deploying several unrelated changes together. When it breaks, you do not know which one did it.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
