Write Effective Tests
Activated Cloud✓ Officialactivated/write-effective-tests
Free · MIT
About
Writes tests that catch real breaks: work test-first with a red-green-refactor loop, name the bug each test guards against, test behaviour through public interfaces, prefer real code and fakes over mocks, derive expected values by hand, cover the edge cases that actually fail, and prove each test can fail. Includes patterns and commands for pytest, Jest, Vitest, Go and Rust. Use when adding a feature, fixing a bug or covering untested code. Not for repairing a red or flaky suite (use fix-failing-tests).
Documentation
Write Effective Tests
A test earns its place by failing when the code breaks and passing when it works. Tests that mirror the implementation, assert on mocks, or freeze today's values pass happily while bugs ship, then break on every harmless refactor. Write the test first, watch it fail for the right reason, make it pass with the least code, then clean up.
When to use
- "Add tests for X", "make sure this can't break again", "build this test-first".
- Every new behaviour you implement and every bug you fix.
- Before refactoring code that has no tests (pin its current behaviour first).
What you need
- How this project tests: framework, folder layout, fixtures or factories, naming, how to run one test. Read two existing test files in the same area before writing yours, and copy their shape.
- The behaviour to test, stated as observable outcomes, with exact values from the spec.
- Commands for your stack in
references/framework-patterns.md.
Method
Pick the right level. Use the lowest level that exercises the behaviour honestly.
Level Use for Speed Unit Pure logic: calculations, parsing, validation, state transitions Milliseconds Integration Code whose correctness depends on a real collaborator: SQL queries, serialisation, HTTP handlers, file I/O Seconds; use a real local database (container or in-memory) rather than mocking the ORM End to end A few critical user journeys through the running app (sign up, pay, export) Slow; keep them few and stable Many unit tests, fewer integration tests, a handful of end-to-end tests is the usual healthy shape. Name the break. Before writing the body, finish this sentence: "This test fails if ___." The blank must be a realistic bug (wrong branch, missing side effect, off-by-one, wrong rounding, unauthorised access allowed), not "the code changes". If you cannot name one, the test is not worth writing.
Red: write one failing test.
- One behaviour per test. A name that reads as a sentence:
test_expired_coupon_gives_no_discount,it("rejects a coupon past its expiry date"),TestApplyCoupon/expired. - Arrange, act, assert, visibly separated.
- Expected values written as literals you worked out by hand, never computed by the code under test or its helpers.
- Run it and confirm it fails because the behaviour is missing (an assertion failure about the value), not because of a typo, import error or broken fixture. A test that passes on first run is testing something that already exists: fix the test.
- One behaviour per test. A name that reads as a sentence:
Green: the least code that passes. Do not add features the test does not demand. Run the test; then run its file.
Refactor with the tests green. Remove duplication, improve names, extract helpers. Run the tests after each change. If a refactor turns a test red, undo it and take a smaller step.
Repeat in vertical slices. One test, then the code for it, then the next test. Do not write ten tests up front and then all the code: the early tests will be wrong about an interface you have not discovered yet.
Cover the edges that actually break. Go through this list and write the cases that apply:
- empty, one, many; first and last element; exactly at a limit and one past it;
- zero, negative, very large, non-integer, currency rounding;
- missing, null or undefined fields; extra unexpected fields; wrong types;
- unicode, emoji, very long strings, leading and trailing whitespace, case differences;
- dates: time zones, daylight-saving changes, month ends, leap days, the clock exactly at midnight;
- duplicates and repeated calls (idempotency); concurrent calls where the code claims to be safe;
- every error path: the exception type and message, the HTTP status and body, nothing written to the database on failure;
- permissions: the owner can, another user cannot, an anonymous user cannot. Use parametrised or table-driven tests for input variations rather than copy-pasted tests.
Mock only at boundaries you do not own. Network calls to third parties, the clock, randomness, email and payment providers. Prefer a fake (a small working in-memory version) to a mock that returns canned answers. When you must mock:
- mock the slow or external layer, and keep everything your test depends on real;
- make the fake's data match the real shape, all fields, not just the ones you read;
- assert on the outcome, not on the mock existing; assert call arguments only when they are the contract (the email went to the right address);
- freeze time and seed randomness instead of sleeping or hoping.
Do not write these tests.
- Change detectors: asserting a constant's value, an exact list of models, a config version number, or a count that grows whenever someone adds an item. Assert the behaviour that depends on them instead.
- Source-text tests: reading a source file and matching a pattern. Run the code instead.
- Mirror tests: computing the expected value with the same function under test.
- Framework tests: checking that the router calls a registered handler. Test your handler.
- Tests that sleep for a fixed time, depend on test order, share mutable state, or call real external services.
Mutation check. Mentally change the production code in realistic ways (flip a comparison, return early, drop a side effect, swap two arguments, return an empty value) and confirm at least one test fails for each. Where the code is critical (money, auth, data integrity) and the project allows it, run a mutation tool (
mutmut, Stryker,cargo-mutants,go-mutesting) on the module and close the gaps it finds.Prove it, then run everything. For a bug fix, show the test fails with the fix reverted and passes with it in place. Then run the whole suite and compare with the baseline. Check the test count went up and nothing was skipped by accident.
Output
The new or changed test files, plus a short note: which behaviours are now covered (one line each, as "fails if ..." statements), the red-then-green evidence for each new test or for the regression test, and the full-suite result against baseline. Mention any behaviour you chose not to test and why.
Checks before you finish
- Every new test was seen failing for the right reason before it passed.
- Expected values are literals or hand-checked fixtures.
- No test asserts on a mock's existence, on source text or on a value that is expected to change.
- No fixed sleeps, real network calls or order dependence.
- The full suite passes with no new failures and the new tests are counted, not skipped.
Pitfalls
- Tests written after the code that pass first time. You never saw them catch anything. Break the code on purpose and watch them fail.
- Testing the mock. If removing the mock makes the assertion meaningless, the test checks nothing.
- Over-mocking. When mock setup is longer than the test, switch to an integration test with real parts.
- Asserting too much. Snapshotting a whole response makes every harmless field change a failure. Assert the fields that matter.
- Asserting too little.
assert resultor "no exception raised" passes for most wrong answers. - Shared fixtures that mutate. One test edits a shared object and another fails mysteriously later. Build fresh data per test.
- Production code only tests call. Cleanup or reset methods used only by tests belong in test utilities.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
