SEO Site Audit
Activated Cloud✓ Officialactivated/seo-site-audit
Free · MIT
About
Audits a website's technical and on-page SEO with free tools: crawls it politely with the bundled script, checks indexing, robots, sitemaps, canonicals, redirects, speed and Core Web Vitals, mobile, structured data, titles, headings, content and internal links, then ranks fixes by impact and effort. Use when asked for an SEO audit, why we are not ranking, why traffic dropped, or before a site migration. Not for choosing topics or keywords (use seo-keyword-research).
Documentation
SEO Site Audit
You find what stops search engines from crawling, indexing and ranking the site, and what stops pages from earning the click, then hand over a fix list ordered by impact. Every finding has evidence (a URL, a screenshot, a report row) and a specific fix. You audit only; changes to the live site happen after the owner approves them.
When to use
- "Do an SEO audit of our site."
- "Why aren't we ranking?" or "Why did organic traffic drop?"
- "Google isn't indexing our pages."
- "Check the site before (or after) the migration / redesign."
- "Is the site fast enough?"
What you need
- The site URL and written confirmation from the owner that you may crawl it. Never crawl sites you are not authorised to audit.
- Google Search Console (connected app or the owner's signed-in browser): Pages (indexing) report, Core Web Vitals report, Sitemaps, Manual actions, Performance. Bing Webmaster Tools if set up.
- Google Analytics for traffic trends (connected app or browser).
- Known changes: migrations, redesigns, CMS or plugin changes, and their dates.
- Priority pages and markets.
- Tools on your own computer: Python 3 for
scripts/crawl_site.py; optionally Lighthouse (npm install -g lighthouse, free and open source) for lab speed tests in theterminal. PageSpeed Insights in the browser gives field data from real Chrome users when the site has enough traffic.
Method
Work in this order. A problem higher up blocks everything below it.
- Context and triage. Note the site type, platform, priority pages, recent changes. In Search Console Performance, compare the last 3 months with the previous period and the same period last year. A sudden drop on a date points to a change on the site, a tracking break or a search engine update; check all three.
- Crawl the site. Run
python3 scripts/crawl_site.py https://example.com/ --max-pages 500 --delay 1 --sitemap https://example.com/sitemap.xml --out crawl.csvin theterminal. It obeys robots.txt, waits between requests, and writescrawl.csvpluscrawl_summary.json. It reads static HTML only; check JavaScript-rendered pages in the browser (step 5). - Crawlability and indexing.
- robots.txt: no accidental
Disallowon important paths, CSS or JS; sitemap referenced. - Search Console Pages report: reasons pages are not indexed. Investigate important URLs under "Crawled, currently not indexed", "Discovered, currently not indexed", "Duplicate, Google chose different canonical", "Excluded by noindex", "Soft 404".
- noindex (meta robots or X-Robots-Tag header) on pages that should rank.
- Status codes: internal links to 4xx/5xx; redirect chains longer than one hop; redirect loops; 302s that should be 301s.
- Canonicals: each indexable page canonicalises to itself; no canonical pointing to a redirected, noindexed or 404 URL; HTTP vs HTTPS, www vs non-www and trailing slash consistent.
- XML sitemap: only 200, indexable, canonical URLs; under 50,000 URLs and 50 MB uncompressed per file; submitted in Search Console. Compare sitemap URLs with crawled URLs: in the sitemap but not found by links suggests orphan pages.
- Use
site:searches only as a rough sanity check; Search Console is the source of truth.
- robots.txt: no accidental
- Site structure and internal links. Important pages within about 3 clicks of the home page; no orphan pages; descriptive anchor text; key commercial pages linked from relevant content; pagination and faceted navigation not generating endless crawlable URL variants.
- Rendering and structured data. For key templates, compare the raw HTML (crawler) with the
rendered page (
browser_navigatethenbrowser_snapshot): main content, links and canonical must be present after rendering. Structured data added by JavaScript is invisible to fetch tools; check it in the rendered page or with Google's Rich Results Test before reporting it missing. - Speed and Core Web Vitals. From the Search Console Core Web Vitals report or PageSpeed Insights field data (75th percentile): Largest Contentful Paint good at 2.5 s or less, Interaction to Next Paint good at 200 ms or less, Cumulative Layout Shift good at 0.1 or less. Test each main template, on mobile first. Report the likely causes (heavy images, render-blocking scripts, slow server response, layout shifts from late ads or fonts) and leave implementation to the developer.
- Mobile and security. Responsive layout, viewport set, readable text and tap targets, same content as desktop. HTTPS everywhere, valid certificate, no mixed content, HTTP redirects to HTTPS.
- On-page. For priority pages: one clear primary topic; unique title (aim for roughly 50 to 60 characters so it is not cut off) leading with the topic; unique meta description (about 140 to 160 characters) giving a reason to click; one H1 that matches the intent; logical H2/H3; images with descriptive alt text and sensible file sizes; no two pages competing for the same query (cannibalisation).
- Content quality. Thin pages (the crawler flags under 300 words; judge, do not auto-condemn), duplicate or near-duplicate pages, outdated content, missing author and company information on pages that need trust, and whether pages satisfy the intent of the queries they target.
- International (only if relevant). hreflang on every language version, self-referencing,
reciprocal, valid codes (
en-GB, noten-UK),x-defaultset, targets return 200 and are canonical. Details inreferences/audit-checklist.md. - Prioritise. For each issue: impact (High / Medium / Low, based on how many important pages and how much traffic it affects), effort (S / M / L), evidence, fix, owner. Order: anything blocking indexing of important pages first, then high impact with low effort, then the rest.
- Report and hand over. Write the report, put the top 5 fixes on a
show_card, and offer to brief the developer (brief_team). Do not edit the site, change robots.txt, submit removals or alter redirects without the owner's explicit go-ahead.
Output
- Executive summary: overall state in 3 sentences, top 5 issues, quick wins.
- Findings table: issue, affected URLs (count and examples), evidence, impact, effort, fix, owner.
crawl.csvandcrawl_summary.jsonattached.- Core Web Vitals table per template (field data where available, lab data labelled as lab).
- A 30/60/90-day action plan.
Template in
references/audit-checklist.md.
Checks before you finish
- Every finding has evidence: a URL, a report screenshot or a crawl row.
- "Missing structured data" or "missing content" claims were checked in a rendered page.
- Field data and lab data are labelled separately.
- Each fix says what to change, where, and who should do it.
- The crawl stayed within robots.txt and the authorised site, at a polite rate.
- No changes were made to the live site.
Pitfalls
- Score chasing. A perfect lab speed score is not the goal; real-user Core Web Vitals and indexing are. Prioritise by business impact.
- Reporting tool output as findings. "300 pages under 300 words" is not a problem if they are product pages that serve their purpose. Judge each pattern.
- Trusting raw HTML for JavaScript sites. Check rendered output before reporting missing content, links or schema.
- Ignoring the timeline. Drops usually follow a change. Line up the traffic graph with release dates before guessing.
- Aggressive crawling. High request rates can slow or crash small sites. Keep the delay at 1 second or more and cap the pages.
- Fixing without approval. robots.txt, redirect and canonical changes can remove a site from search. Recommend; the owner and developer decide.
Versions
Listed from the source repository.
Reviews
No reviews yet. Be the first.
