Activated Cloud
← App Store

Fix Failing Tests

Activated Cloud✓ Officialactivated/fix-failing-tests

No ratings yet6 installsv1.0.0Updated Oct 6, 2026● Unknown

Free · MIT

About

Gets a red test suite green for the right reason: group failures by cause, reproduce one at a time, decide with evidence whether the code, the test or the environment is wrong, fix that, and stamp out flaky tests by finding the race, shared state, clock or ordering problem behind them. Never deletes, skips or loosens a test to get green without the owner's agreement. Use when tests fail locally or in CI after a change, or pass and fail at random. Not for CI failures that are not test failures (use fix-ci-failures).

Software Development

Documentation

From SKILL.md · v1.0.0 · what the agent reads when it loads this skill3 files: SKILL.md, references/CREDITS.md, references/flaky-test-causes.md

Fix Failing Tests

A failing test is a message: either the code broke, the test is wrong, or the environment differs. The job is to find out which, with evidence, and fix that thing. The cheap moves (delete the test, mark it skipped, widen the assertion, add a retry, raise the timeout) make the suite green and the product worse; they need the owner's say-so and a tracking issue.

When to use

  • "Tests are failing", "CI is red on the test job", "this test is flaky", "fix the suite after the upgrade".
  • After your own change broke tests.
  • A test passes alone and fails in the full run, or fails one run in ten.

What you need

  • The repo and the exact test command CI uses (from the CI config or task runner).
  • The failure output: local run, or the CI log (via the GitHub or GitLab connected app, gh run view <id> --log-failed, or the browser).
  • The baseline: which tests failed before the change in question, if known (git stash, run, git stash pop, or run on the base branch in a worktree).

Method

  1. Collect and group. Run the suite once and list every failing test with its error. Group by error signature: twenty failures with the same ImportError or the same fixture error are one problem. Fix in this order: collection and import errors, then fixture and setup errors, then assertion failures. The first kind often causes the rest.

  2. Reproduce one, narrowly.

    pytest tests/test_orders.py::test_refund -x -vv --tb=long -l
    npx vitest run src/orders.test.ts -t "refund"          # or: npx jest src/orders.test.ts -t "refund" --runInBand
    go test ./orders -run '^TestRefund$' -v -count=1
    cargo test refund -- --nocapture --test-threads=1
    

    If it passes alone but fails in the suite, jump to step 7 (shared state or ordering).

  3. Read the failure fully. The assertion's expected and actual values, the stack trace, the captured output and logs. Find the first line in project code and read it with read_file, along with the test.

  4. Decide who is wrong, with evidence.

    Finding Evidence to look for Fix
    The code regressed The test encodes the agreed behaviour; git log -p shows a recent code change on that path Fix the code
    The behaviour changed on purpose A spec, issue or owner decision says the new behaviour is right; the change was intended in this branch Update the test, and cite the requirement in the commit message
    The test was always wrong The test passes for the wrong reason or fails on correct output (hand-check the expected value) Fix the test
    The environment differs Passes on one machine, fails on another; errors mention missing services, time zone, locale, file paths, versions, network Fix the setup or make the test hermetic
    Use git log -p -S "<assertion value or function>" -- <file> to see when and why the expectation was set. If you cannot tell whether the old or new behaviour is right, ask the owner with clarify and show both.
  5. Fix the cause, minimally. One cause per commit. Re-run the single test, then its file, then the suite.

  6. Snapshot and golden-file tests. Update snapshots only after reading the diff and confirming every change is intended (npx jest -u, npx vitest -u, pytest --snapshot-update, UPDATE_GOLDEN=1, or the project's flag). Never bulk-update to make red go away.

  7. Flaky tests: prove it, find the cause, fix it.

    • Confirm flakiness by repetition: for i in $(seq 50); do <single test cmd> || { echo "failed on run $i"; break; }; done, go test -run X -count=100 -race, pytest --count=50 -x (pytest-repeat).
    • Try random and fixed order, alone and in the suite, serial and parallel, to see which condition triggers it.
    • Match the symptom to the usual causes in references/flaky-test-causes.md (timing and sleeps, order dependence and shared state, the clock, randomness, concurrency, real network, resource leaks, unordered collections, floating point).
    • Fix the cause: wait for a condition rather than sleeping, isolate state per test, freeze the clock, seed randomness, fake the network, sort before comparing.
    • Verify with the same repetition count that showed the flake, and then a full suite run.
  8. When a dependency upgrade broke many tests. Read the dependency's changelog and migration guide (web_extract) for the versions crossed. Fix call sites to the new API rather than pinning back, unless the owner prefers to pin; then open a follow-up issue.

  9. Run everything and compare with the baseline. Full suite, plus the linter and type checker if you changed code. No new failures; the previously failing tests now pass; test counts did not drop (a drop means tests stopped being collected).

  10. What you may not do alone. Deleting a test, adding skip/xfail/.skip/t.Skip/#[ignore], loosening an assertion to accept the current wrong output, adding automatic retries, or raising a timeout without explaining the slowness. If one of these is truly the right call (the feature was removed, the test needs infrastructure you lack), propose it to the owner with the reason and a tracking issue, and do it only on their go-ahead.

Output

A short note per fixed cause: the failing tests, the cause in one or two sentences with file:line, which side was wrong (code, test or environment) and the evidence, the fix, and the verification (single test, repetitions for flakes, full suite against baseline). List anything still failing with what you know about it. For example:

1. 23 failures in tests/api/: one cause. The fixture in `conftest.py` still imported `settings.REDIS_URL`, renamed to `CACHE_URL` in 9c1d2e3. The test side was stale; updated the fixture. All 23 pass.
2. `test_invoice_total_rounds_up`: the code was wrong. `invoice.py:88` switched to banker's rounding in 4ab12f0; the spec (#402) says half-up. Restored half-up; test passes, suite green.
3. `test_export_ordering` (flaky, 3 failures in 50 runs): query had no ORDER BY. Added ordering by id; 0 failures in 200 runs.
Full suite: 418 passed, 0 failed (baseline before the change: 418 passed).

Checks before you finish

  • Each fix addresses a cause you can state, not a symptom.
  • Flaky fixes were verified with at least as many repetitions as were needed to see the flake.
  • The full suite was run after the last change; no new failures against baseline; the test count did not drop.
  • No test was deleted, skipped, loosened or retried without the owner's agreement.
  • Changed expectations cite the requirement or decision that changed them.

Pitfalls

  • Editing the assertion to match the output. That turns a failing test into a test that blesses the bug.
  • Fixing the twentieth failure first. Most of a red wall usually comes from one import, fixture or setup error.
  • "It passed when I reran it." A pass on rerun proves flakiness, not health. Find the cause.
  • Raising timeouts. A test that needs 30 seconds instead of 5 is telling you about a slow or stuck path.
  • Sleep-based fixes. Adding sleep(1) moves the race; it does not remove it.
  • Blaming the environment without proof. Show the difference (versions, time zone, service) before calling it environmental.
  • Bulk snapshot updates. They silently accept every regression in the diff.

Versions

v1.0.0currentOct 6, 2026

Listed from the source repository.

Reviews

No reviews yet. Be the first.

Write a review