Verify Before Claiming Done
Activated Cloud✓ Officialactivated/verify-before-claiming-done
Free · MIT
About
Puts every 'done', 'fixed', 'tests pass' or 'deployed' claim behind fresh evidence: name the command that proves the claim, run it in full after your last edit, read the whole result and the exit code, confirm it tested the right thing, then report exactly what it showed and what you could not check. Use before reporting completion, committing, opening a pull request, ticking a todo item or handing work to a teammate. Not a replacement for writing tests (use write-effective-tests).
Documentation
Verify Before Claiming Done
The most expensive habit a coding agent can have is saying "done" when it is not. The owner then finds the failure, loses trust, and checks everything you do from then on. This skill is the gate in front of every success claim: evidence first, then the claim, and the claim says no more than the evidence shows.
When to use
- You are about to write "done", "fixed", "works", "all tests pass", "deployed", "merged" or anything that implies it.
- Before
git commit, before opening or updating a pull request, before marking atodoitem completed. - When a subagent or teammate reports success and you are about to pass that on.
- After a fix to a bug someone reported: before you tell them it is fixed.
What you need
- The original request and acceptance criteria, re-read now, not remembered.
- The project's real check commands (from CI config or the task runner).
- A way to observe the real behaviour: a running server, a CLI, a browser, logs.
Method
Name the proof. For the claim you are about to make, write down the command or observation that would prove it. Use this table:
Claim Proof required Not proof Tests pass Full test command run after your last edit: 0 failures, exit 0 An earlier run; one test passing; "should pass" Lint or types clean Linter and type checker output after the last edit, exit 0 The editor showing no red; one file checked Build works The build command, exit 0 Tests passing; lint passing Bug fixed The original reproduction now shows the correct behaviour Code changed; a different test passing Regression test guards the bug Test fails with the fix reverted, passes with it restored Test passes once Feature works The acceptance criteria exercised for real (request, CLI run, page) Unit tests alone Deployed The live version or health endpoint reports the new version, key flow works The pipeline said "succeeded" Subagent finished You read its diff and re-ran its checks Its summary saying "done" Requirements met Each criterion checked one by one against evidence "Tests pass" Run it fresh and in full. After your last edit, run the complete command, not a filtered or partial one. Capture the exit code every time:
<command>; echo "exit=$?"Pipes hide failures:
pytest | tail -5returns the exit code oftail. Useset -o pipefail; pytest 2>&1 | tail -40or check${PIPESTATUS[0]}. For suites longer than a few minutes, run withterminal(background=true, notify_on_complete=true)and wait for the notification rather than guessing.Read the result, not just the last line.
- Count: passed, failed, errored, skipped. Compare with the baseline from before your change.
- Check that tests actually ran. pytest exit code 5 means no tests were collected; "0 tests" from any runner is not a pass. A filter (
-k,-t,-run) that matches nothing passes vacuously. - Check skips:
pytest -rslists skip reasons. A test skipped because a service is missing has not verified anything. - Read warnings that mention your files. New deprecation or runtime warnings are often the next bug.
Make sure you tested the code you changed. Common ways to verify stale code:
- a server or worker still running the old code (restart it and confirm the version or PID changed);
- a build output or cache (
dist/,.next/,__pycache__, Docker layer cache, the browser cache) serving old files; - an installed package shadowing the local source (
pip show -f pkg, checkpython -c "import pkg; print(pkg.__file__)"); - the wrong environment (a different virtualenv, Node version or config file).
Prove a regression test can fail. For a bug fix: stash or revert the fix, run the new test and see it fail with the bug's symptom, restore the fix and see it pass.
git stash push -- src/the_fix.py && <test>; echo "exit=$?" # expect failure git stash pop && <test>; echo "exit=$?" # expect passCheck the requirements line by line. Re-read the original request. Make a checklist of every requirement, including the small ones ("and show it in the header", "for both admins and members"). Mark each with its evidence or as not done.
Verify delegated work yourself. For anything a subagent or teammate did: read the diff (
git diff,git log -p -3), confirm it only touched what it should, and re-run the checks it claims passed.Report with the evidence, in plain words.
- Good: "Full suite: 418 passed, 0 failed (
uv run pytest -q, exit 0). Lint and mypy exit 0. Checked the export in the browser: CSV downloads with 3 rows." - Good: "Fixed and verified locally. Not verified in staging: no access to the staging database."
- Not acceptable: "Should work now." "Looks good." "I believe this fixes it." "Done!" with no evidence. If something failed, say so first, with the output.
- Good: "Full suite: 418 passed, 0 failed (
Output
A status line that matches the evidence (done, partly done, blocked, not verified), followed by a compact list of each check: the command or action, and what it showed. Anything not verified is listed with the reason. If the owner is watching live, a show_card with the check table works well.
Checks before you finish
- Every success word in your message is backed by a command run after your final edit.
- Exit codes were read, not inferred from output colour or the last line.
- Test counts are non-zero and compared with the baseline.
- The behaviour was observed on the code you changed, not a stale build or process.
- Each requirement in the original request is marked done with evidence, or listed as not done.
Pitfalls
- Old evidence. A run from before your last edit proves nothing about the current code.
- Partial runs dressed up as full ones. Running one test file and reporting "tests pass".
- Exit codes swallowed by pipes,
|| true, or a script that does not useset -e. - Vacuous passes. Zero tests collected, an over-narrow filter, or everything skipped.
- Trusting summaries. A subagent's "all green" or a CI badge on a different commit.
- Satisfaction words before verification. "Great, that fixed it!" before the check runs. Run the check, then speak.
- Overclaiming scope. Verified locally is not verified in production. Say which environment.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
