Fix Failing Tests
Activated Cloud✓ Officialactivated/fix-failing-tests
Free · MIT
About
Gets a red test suite green for the right reason: group failures by cause, reproduce one at a time, decide with evidence whether the code, the test or the environment is wrong, fix that, and stamp out flaky tests by finding the race, shared state, clock or ordering problem behind them. Never deletes, skips or loosens a test to get green without the owner's agreement. Use when tests fail locally or in CI after a change, or pass and fail at random. Not for CI failures that are not test failures (use fix-ci-failures).
Documentation
Fix Failing Tests
A failing test is a message: either the code broke, the test is wrong, or the environment differs. The job is to find out which, with evidence, and fix that thing. The cheap moves (delete the test, mark it skipped, widen the assertion, add a retry, raise the timeout) make the suite green and the product worse; they need the owner's say-so and a tracking issue.
When to use
- "Tests are failing", "CI is red on the test job", "this test is flaky", "fix the suite after the upgrade".
- After your own change broke tests.
- A test passes alone and fails in the full run, or fails one run in ten.
What you need
- The repo and the exact test command CI uses (from the CI config or task runner).
- The failure output: local run, or the CI log (via the GitHub or GitLab connected app,
gh run view <id> --log-failed, or the browser). - The baseline: which tests failed before the change in question, if known (
git stash, run,git stash pop, or run on the base branch in a worktree).
Method
Collect and group. Run the suite once and list every failing test with its error. Group by error signature: twenty failures with the same
ImportErroror the same fixture error are one problem. Fix in this order: collection and import errors, then fixture and setup errors, then assertion failures. The first kind often causes the rest.Reproduce one, narrowly.
pytest tests/test_orders.py::test_refund -x -vv --tb=long -l npx vitest run src/orders.test.ts -t "refund" # or: npx jest src/orders.test.ts -t "refund" --runInBand go test ./orders -run '^TestRefund$' -v -count=1 cargo test refund -- --nocapture --test-threads=1If it passes alone but fails in the suite, jump to step 7 (shared state or ordering).
Read the failure fully. The assertion's expected and actual values, the stack trace, the captured output and logs. Find the first line in project code and read it with
read_file, along with the test.Decide who is wrong, with evidence.
Finding Evidence to look for Fix The code regressed The test encodes the agreed behaviour; git log -pshows a recent code change on that pathFix the code The behaviour changed on purpose A spec, issue or owner decision says the new behaviour is right; the change was intended in this branch Update the test, and cite the requirement in the commit message The test was always wrong The test passes for the wrong reason or fails on correct output (hand-check the expected value) Fix the test The environment differs Passes on one machine, fails on another; errors mention missing services, time zone, locale, file paths, versions, network Fix the setup or make the test hermetic Use git log -p -S "<assertion value or function>" -- <file>to see when and why the expectation was set. If you cannot tell whether the old or new behaviour is right, ask the owner withclarifyand show both.Fix the cause, minimally. One cause per commit. Re-run the single test, then its file, then the suite.
Snapshot and golden-file tests. Update snapshots only after reading the diff and confirming every change is intended (
npx jest -u,npx vitest -u,pytest --snapshot-update,UPDATE_GOLDEN=1, or the project's flag). Never bulk-update to make red go away.Flaky tests: prove it, find the cause, fix it.
- Confirm flakiness by repetition:
for i in $(seq 50); do <single test cmd> || { echo "failed on run $i"; break; }; done,go test -run X -count=100 -race,pytest --count=50 -x(pytest-repeat). - Try random and fixed order, alone and in the suite, serial and parallel, to see which condition triggers it.
- Match the symptom to the usual causes in
references/flaky-test-causes.md(timing and sleeps, order dependence and shared state, the clock, randomness, concurrency, real network, resource leaks, unordered collections, floating point). - Fix the cause: wait for a condition rather than sleeping, isolate state per test, freeze the clock, seed randomness, fake the network, sort before comparing.
- Verify with the same repetition count that showed the flake, and then a full suite run.
- Confirm flakiness by repetition:
When a dependency upgrade broke many tests. Read the dependency's changelog and migration guide (
web_extract) for the versions crossed. Fix call sites to the new API rather than pinning back, unless the owner prefers to pin; then open a follow-up issue.Run everything and compare with the baseline. Full suite, plus the linter and type checker if you changed code. No new failures; the previously failing tests now pass; test counts did not drop (a drop means tests stopped being collected).
What you may not do alone. Deleting a test, adding
skip/xfail/.skip/t.Skip/#[ignore], loosening an assertion to accept the current wrong output, adding automatic retries, or raising a timeout without explaining the slowness. If one of these is truly the right call (the feature was removed, the test needs infrastructure you lack), propose it to the owner with the reason and a tracking issue, and do it only on their go-ahead.
Output
A short note per fixed cause: the failing tests, the cause in one or two sentences with file:line, which side was wrong (code, test or environment) and the evidence, the fix, and the verification (single test, repetitions for flakes, full suite against baseline). List anything still failing with what you know about it. For example:
1. 23 failures in tests/api/: one cause. The fixture in `conftest.py` still imported `settings.REDIS_URL`, renamed to `CACHE_URL` in 9c1d2e3. The test side was stale; updated the fixture. All 23 pass.
2. `test_invoice_total_rounds_up`: the code was wrong. `invoice.py:88` switched to banker's rounding in 4ab12f0; the spec (#402) says half-up. Restored half-up; test passes, suite green.
3. `test_export_ordering` (flaky, 3 failures in 50 runs): query had no ORDER BY. Added ordering by id; 0 failures in 200 runs.
Full suite: 418 passed, 0 failed (baseline before the change: 418 passed).
Checks before you finish
- Each fix addresses a cause you can state, not a symptom.
- Flaky fixes were verified with at least as many repetitions as were needed to see the flake.
- The full suite was run after the last change; no new failures against baseline; the test count did not drop.
- No test was deleted, skipped, loosened or retried without the owner's agreement.
- Changed expectations cite the requirement or decision that changed them.
Pitfalls
- Editing the assertion to match the output. That turns a failing test into a test that blesses the bug.
- Fixing the twentieth failure first. Most of a red wall usually comes from one import, fixture or setup error.
- "It passed when I reran it." A pass on rerun proves flakiness, not health. Find the cause.
- Raising timeouts. A test that needs 30 seconds instead of 5 is telling you about a slow or stuck path.
- Sleep-based fixes. Adding
sleep(1)moves the race; it does not remove it. - Blaming the environment without proof. Show the difference (versions, time zone, service) before calling it environmental.
- Bulk snapshot updates. They silently accept every regression in the diff.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
