Activated Cloud
← App Store

Fix CI Failures

Activated Cloud✓ Officialactivated/fix-ci-failures

No ratings yet6 installsv1.0.0Updated Oct 6, 2026● Unknown

Free · MIT

About

Diagnoses and fixes a red CI pipeline: find the failing job and the first real error in its log, classify it (test, lint, type, build, dependency, environment, permission, timeout, flaky infrastructure), check whether main fails too, reproduce locally with the same versions and commands, fix the cause and confirm the rerun is green. Also covers writing and editing GitHub Actions and GitLab CI pipelines well. Use when a check is red on a pull request or the main branch. Not for test failures you can already reproduce locally (use fix-failing-tests).

Software Development

Documentation

From SKILL.md · v1.0.0 · what the agent reads when it loads this skill4 files: SKILL.md, references/CREDITS.md, references/ci-failure-signatures.md, references/workflow-template.md

Fix CI Failures

CI logs are long and the line that matters is rarely the last one. The method: find the first real error, work out what kind of failure it is, prove whether your change caused it, reproduce it locally under CI's conditions, and fix the cause. Turning a check off, adding continue-on-error or merging red is never the fix unless the owner decides so.

When to use

  • "CI is failing on my PR", "the build is red on main", "the deploy pipeline broke", "why did this job fail?"
  • Writing a new workflow or changing an existing pipeline file.
  • A job that fails one run and passes the next.

What you need

  • Access to the CI system: the owner's GitHub or GitLab connected app, the gh or glab CLI authenticated on your computer, or the CI pages in the browser the owner signed in. If none is available, ask the owner to paste the failing job log.
  • The repo checked out at the failing commit.
  • Permission rules from memory: may you re-run jobs, push fixes to this branch, edit workflow files? Workflow and secret changes affect everyone: confirm with the owner before changing permissions, secrets or required checks.

Method

  1. Find the failing run and job.

    gh pr checks <pr>                                   # which checks failed
    gh run list --branch <branch> --limit 5             # recent runs
    gh run view <run-id>                                # jobs and steps
    gh run view <run-id> --log-failed > /tmp/ci-failed.log
    

    GitLab: glab ci status, glab ci view, glab ci trace <job-id> > /tmp/ci-failed.log. In the browser, open the failed job and download the raw log.

  2. Find the first real error. Read the log with read_file or search_files on the saved file. Search for error, Error:, FAILED, ##[error], npm ERR!, Traceback, panic:, exit code. Start from the step that failed and scroll up to the first error in that step: later errors are often consequences. Also note the setup steps' printed versions (runtime, package manager, OS image), because version drift causes many CI-only failures.

  3. Classify it. Match the error against references/ci-failure-signatures.md: test failure, lint or format, type check, build or compile, dependency resolution, missing or wrong environment (env var, service, secret), permissions and tokens, timeout or out of memory, cache corruption, runner or network flake.

  4. Is it your change?

    gh run list --branch main --workflow "<workflow name>" --limit 5
    
    • Main is red with the same error: the failure predates your change. Report it; fix it only if asked or if it blocks you and the fix is small.
    • Main is green: your change (or its interaction with main) caused it. Check whether your branch is behind main (git log --oneline HEAD..origin/main); a stale branch can fail on things already fixed upstream.
    • Suspected flake (network error, runner lost, registry timeout): re-run only the failed jobs once (gh run rerun <run-id> --failed). If it passes, record it as a flake with the error text; repeated flakes are a bug to fix, not something to keep re-running.
  5. Reproduce locally with CI parity. Copy the exact command from the workflow file, not from memory. Match its conditions:

    • same runtime and tool versions as the setup step;
    • clean dependency install from the lockfile (npm ci, uv sync --frozen, pip install -r requirements.txt in a fresh venv, go mod download);
    • same environment (CI=true, TZ=UTC or whatever CI uses, locale), same services (start them with the project's compose file);
    • Linux-specific behaviour: case-sensitive paths, file permissions, line endings. If the job uses a container image, run the command inside that image (docker run --rm -v "$PWD":/w -w /w <image> <command>).
  6. Fix the cause. Apply the matching fix from the signatures reference: fix the code or test, run the formatter, correct the type, add the missing dependency to the manifest through the package manager, update the lockfile, set the missing environment variable in the workflow (never the secret's value in the file), fix paths in the Dockerfile. Run the same command locally until it passes.

  7. Push and watch.

    git add <files> && git commit -m "fix(ci): <what and why>" && git push
    gh run watch                                         # or: gh pr checks <pr> --watch
    

    Confirm every required check is green on the new head commit, not just the one you fixed.

  8. Three strikes. If three different fixes have not turned it green, stop and write down what you know (error, what you tried, what changed each time) and ask the owner or the teammate who owns the pipeline.

  9. Writing or editing pipelines. Follow references/workflow-template.md:

    • least-privilege permissions: at the top; widen per job only when needed;
    • pin third-party actions to a full commit SHA (with the version in a comment) or at least a major version tag, per the repo's policy;
    • dependency caching keyed on the lockfile hash;
    • timeout-minutes on every job; concurrency to cancel superseded runs on the same branch;
    • the same commands developers run locally (a make test target), so local and CI cannot drift;
    • secrets only through the CI secret store, never echoed; fork pull requests do not receive secrets, and pull_request_target with a checkout of untrusted code is dangerous. Validate GitHub workflow files with actionlint (and docker compose config, yamllint where relevant) before pushing.
  10. Never, without the owner's explicit decision: disable or delete a failing check, add continue-on-error: true or allow_failure: true, skip tests in CI only, remove a required status check, merge with red checks, or print secrets to debug.

Output

A short note: the failing job and step, the first real error (quoted, a few lines), the classification, whether main was affected, the cause, the fix (commit link), and the evidence that CI is now green (run link or gh pr checks output). For flakes: the error, the rerun result, and a proposed fix or issue.

Checks before you finish

  • You quoted the first real error, not the last line of the log.
  • You checked main's status for the same workflow.
  • The fix was reproduced and verified locally with CI's command and versions.
  • All required checks are green on the latest commit.
  • No check was disabled, skipped or made non-blocking without the owner's decision.

Pitfalls

  • Reading only the tail. "Process completed with exit code 1" is the effect. The cause is higher up.
  • Fixing the symptom in the workflow. Raising a timeout or pinning an old runtime may hide a real regression.
  • Endless reruns. A job that passes on retry still has a bug.
  • npm install instead of npm ci locally. You test different dependency versions from CI and "cannot reproduce".
  • Debugging secrets by printing them. Logs are often readable by many people; masked values can still leak in transformed forms.
  • Editing shared workflows casually. A change to a reusable workflow or required check affects every branch and every teammate.

Versions

v1.0.0currentOct 6, 2026

Listed from the source repository.

Reviews

No reviews yet. Be the first.

Write a review