Debug to the Root Cause
Activated Cloud✓ Officialactivated/debug-root-cause
Free · MIT
About
Finds and fixes the root cause of a bug, crash, wrong output, error log or failing check: build a command that reproduces the exact symptom, read the full stack trace and logs, bisect to the commit or change that broke it, trace the bad value back to where it starts, test one ranked hypothesis at a time, then fix at the source with a regression test and remove the instrumentation. Use for any bug whose cause is not already proven, including 'it worked last week'. Not for sorting incoming reports (use triage-bug-report).
Documentation
Debug to the Root Cause
Guess-and-patch debugging is slow, and it leaves the real bug in place under a symptom fix. The method here is evidence first: a command that goes red on the exact symptom, the full error read line by line, the cause traced back to where the bad state is born, and a fix proven by the same command going green. No fix is proposed until you can say why the bug happens.
When to use
- "This crashes", "the total is wrong", "users get a 500 on checkout", "this test fails and I don't know why".
- An error in the logs, an alert, a stack trace pasted into chat.
- "It worked last week" or "it broke after the upgrade": bisecting is part of this skill.
- A fix you already tried did not work.
What you need
- The symptom in the reporter's words, plus any stack trace, log excerpt, screenshot (
vision_analyze), request ID or timestamp. - The code and a place to run it safely (your computer, a dev or staging environment). Never experiment on production.
- Repository access via the GitHub or GitLab connected app, or a local copy. Log and dashboard access via the owner's connected apps or the browser they signed in; if neither, ask the owner to paste the relevant logs.
Method
Phase 1: Reproduce with a tight loop
- Build one command that shows the symptom: it must fail now and pass only when this bug is fixed. In order of preference: a failing test at the seam that reaches the bug; a
curlagainst a local server; a CLI run with fixture input diffed against expected output; a browser script; replaying a captured request or queue message; a small throwaway harness that calls the failing path. - Make it fast (seconds), deterministic (pin time, seed randomness, isolate files and network) and specific (assert the exact wrong value or error, not "does not crash").
- Intermittent bug: raise the reproduction rate before anything else. Loop it (
for i in $(seq 100); do <cmd> || break; done), add load, run in parallel, narrow timing. A bug that shows up half the time is debuggable; one in a thousand is not yet. - Cannot reproduce at all: collect more data rather than guessing. Compare the reporter's environment with yours (version, config, data, browser, time zone, locale) and ask for the missing pieces.
Phase 2: Read the evidence
- Read the whole error. The exception type and message, every frame of the stack trace, and any chained "caused by" or "during handling of the above" sections. Find the first frame in this project's code: that is where to start reading. Language-specific guidance is in
references/reading-stack-traces.md. - Read the logs around the failure, not just the error line. Filter by request ID or a narrow time window, and look for the first anomaly, which is often a warning minutes before the loud error.
search_filesfor the error message text to find where it is raised, andread_filethat code with its callers.
Phase 3: Narrow down
- What changed?
git log --oneline -20,git diff, lockfile diffs (git diff <good>..HEAD -- '*.lock' package-lock.json go.sum), config and infrastructure changes, data changes, dependency releases. - Known good state? If the behaviour worked at some commit, tag or date, bisect to the change that broke it with
git bisect runand your loop as the check. Method, exit-code rules and test-order bisection are inreferences/bisecting.md. - Working versus broken. Find a similar case that works (another endpoint, another customer, another input) and list every difference. Do not dismiss any difference as "can't matter" until you have tested it.
- Instrument the boundaries. In a multi-part flow (client, API, service, database, queue), log what enters and leaves each boundary on one run to see where good data turns bad. Tag every temporary line with a unique marker such as
DBG-7f3aso you can find and remove them all later. A debugger is often faster than logging: seereferences/debugger-cheatsheet.md. - Trace the bad value upstream. When the failure is deep in the stack, ask what passed the bad value in, then what passed it to that caller, until you reach where it was created. Fix there, not where it exploded.
Phase 4: Hypothesise and test
- Write three to five hypotheses. For each, state a prediction you can check: "If the cache returns a stale price, then clearing the cache makes the loop pass." Rank them by likelihood and by how cheap they are to check.
- Test one at a time with the smallest probe, changing one variable. A hypothesis that survives is not proven until its prediction comes true. If none survive, go back to Phase 2 with what you learned.
- If the owner or a teammate is around and the list is long, share the ranked list: someone who knows the system may reorder it in seconds (
ask_teammate).
Phase 5: Fix at the source
- Write a regression test that reproduces the bug (often your Phase 1 loop, made permanent). Watch it fail.
- Make one fix that addresses the cause. No "while I'm here" changes in the same commit.
- Search for the same mistake elsewhere: the bug class, not just this instance (
search_filesfor the same pattern, sibling call sites, copy-pasted code). - Where cheap, add a guard at the boundary where the bad value entered (validation, a clear error) so the next bug of this kind fails early and loudly.
Phase 6: Verify and clean up
- The loop passes. The regression test fails with the fix reverted and passes with it restored. The full suite, linter and type checker show no new failures against the baseline.
- Remove every temporary log line and breakpoint:
search_filesfor your marker and forbreakpoint(),debugger;,console.log,dbg!,fmt.Printlnyou added.
The rule of three
If three fixes have failed, stop. Each failed fix revealing a new problem somewhere else means the design is the problem, not the line. Write down what you know, what you tried and what happened, and discuss the approach with the owner before attempting a fourth.
Output
A short write-up:
- Symptom: what the user saw.
- Reproduction: the exact command or steps.
- Root cause: what is wrong, at
file:line, and why it produces the symptom. If bisected, the first bad commit. - Fix: what you changed and why there.
- Evidence: the loop red before and green after; the regression test's red-green check; suite results against baseline.
- Same bug elsewhere: sites checked and fixed, or none found.
- Not verified: for example, production behaviour.
Checks before you finish
- You can explain the cause in two sentences, pointing at a line.
- The reproduction fails without the fix and passes with it.
- The full test suite, linter and type checker were run after the fix.
- All temporary instrumentation is gone (
git diffshows none). - Sibling call sites with the same flaw were checked.
Pitfalls
- Fixing where it crashed. The crash site is usually downstream of the cause. Trace back.
- Changing several things at once. You will not know which one mattered, and you may add a bug.
- "Just try this." A change without a hypothesis is a lottery ticket. Each probe should be able to prove you wrong.
- Skipping the reproduction. Without a red loop you cannot know the fix worked.
- Reading half the trace. The useful line is often the chained cause at the bottom or the first frame in project code.
- Debugging in production. Reproduce elsewhere. If production-only data is needed, ask the owner how to get a safe copy.
- Leaving debug code behind. Marked lines make cleanup a single search.
- Silencing the error. A broad
try/except, a default value or optional chaining that hides the failure is not a fix.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
