SEO

How to run a free SEO site audit, step by step

In shortAn SEO site audit crawls your website the way a search engine does and lists what stops pages from being found, indexed or understood. Check crawl access, status codes and redirects, canonicals, titles and descriptions, internal links, sitemaps, speed, structured data and AI crawler access, then fix the issues that affect the most pages first.

By the Join Postly teamPublished Updated 7 min read

You don't need an expensive tool to audit a website. You need a crawler, a checklist and a sensible order for fixing things. This guide walks through the checks in the order they matter: first whether search engines can reach your pages at all, then whether they can index and understand them.

What is an SEO site audit?

An SEO site audit is a systematic check of a website for problems that hurt its visibility in search. A crawler visits your pages by following links and sitemaps, records what it finds, and compares it with best practice.

Audits find technical problems well. They can't tell you whether your content is the best answer for a query; that still takes judgement.

Can search engines crawl your site?

Start with robots.txt and server health, because nothing else matters if crawlers can't get in.

  • Open yoursite.com/robots.txt. Make sure no Disallow rule blocks pages or folders you want in search.
  • Under the robots.txt standard (RFC 9309), crawlers treat a 4xx error for robots.txt as "no rules", but a 5xx error as "stay away". A broken robots.txt can stop crawling of the whole site.
  • Check that CSS and JavaScript needed to render pages aren't blocked.
  • If key content only appears after JavaScript runs, check that the raw HTML still contains the main text and links.

Are the right pages indexable?

A page can be crawled and still kept out of search. Check that the pages you want indexed return 200, have no noindex, and don't point their canonical to another URL.

Do the reverse check too: thin, duplicate or private pages (internal search results, filters, staging copies) should be kept out with noindex or removed from links and sitemaps.

Which status codes and redirects need fixing?

Common status codes found in a site audit and what to do
FindingWhat it meansWhat to do
404 / 410Page not found or goneFix or remove internal links to it; redirect only if a real replacement exists
5xxServer errorCheck hosting, timeouts and plugins; these block crawling
Redirect chainA → B → CPoint links and the first redirect straight at the final URL
302 for a permanent moveTemporary redirectUse 301 or 308 when the move is permanent
Soft 404"Not found" page that returns 200Return a real 404 or 410

Also check that http:// redirects to https://, and that the www and non-www versions of your domain redirect to one of them, not serve two copies.

Do canonical tags point the right way?

Each indexable page should have a canonical tag pointing to its own preferred URL. Problems to look for: canonicals pointing to redirects or 404s, chains where A points to B and B points to C, a canonical that disagrees with the sitemap, and pages with a trailing slash and without one both resolving as separate URLs.

Are titles, descriptions and headings in order?

  • Titles: every page has one, it is unique, and it says what the page is about. Very long titles get cut off in results.
  • Meta descriptions: unique summaries that make someone want to click. Search engines may rewrite them, but a good one is often used.
  • Headings: one clear H1 per page, with H2 and H3 headings that reflect the structure.
  • Duplicates: several pages with the same title often means duplicate or competing content.
  • Images: descriptive alt text on images that carry meaning.

Every page you care about should be reachable through normal links within a few clicks of the homepage. Look for broken internal links, orphan pages (in the sitemap but linked from nowhere), and important pages buried deep in the site. Use descriptive anchor text instead of "click here".

Is your XML sitemap clean?

Your sitemap should list only canonical, indexable URLs that return 200. Remove redirects, 404s and noindexed pages from it. One sitemap file can hold up to 50,000 URLs and 50 MB uncompressed under the sitemaps protocol; larger sites use a sitemap index. Reference the sitemap in robots.txt with a Sitemap: line.

How fast is the site?

A crawl can spot causes of slowness: slow server responses, render-blocking scripts in the page head, missing compression and assets without caching headers. Real-user speed is measured differently, with Core Web Vitals.

Google's Core Web Vitals and their good thresholds
MetricMeasures"Good" threshold
LCPLoading of the main content2.5 seconds or less
INPResponse to interactions200 milliseconds or less
CLSUnexpected layout shifts0.1 or less

Google assesses these at the 75th percentile of real visits. Field data comes from real Chrome users, so check the Core Web Vitals report in your Search Console or PageSpeed Insights rather than relying on a crawler's estimate.

Is your structured data valid?

If you use JSON-LD structured data, make sure it is valid, matches what the page shows and uses current types. Common problems are missing required properties, Product markup without price details (an Offer) on a page that sells something, and reviews of your own business marked up as if they were independent. Google's Rich Results Test checks whether a page is eligible for rich results.

Can AI crawlers read your site?

AI assistants and AI search tools use their own crawlers, and robots.txt controls them like any other bot. Some crawlers fetch pages to answer questions and cite sources (for example OAI-SearchBot, Claude-SearchBot and PerplexityBot); others, such as GPTBot and Google-Extended, relate to model training.

Blocking training crawlers is a business choice, but if you block every search-type AI crawler, AI tools can't read your pages to cite them. Also check that your main pages answer their headline question in the first lines, which makes them easier to quote. Some sites add an llms.txt file as well.

How does the Postly Rank Site Audit work?

The Postly Rank Site Audit runs these checks automatically with its own crawler, JoinPostlyBot. For Core Web Vitals field data, still use your Search Console or PageSpeed Insights. On the Free plan it crawls up to 50 pages per audit, 3 audits a day; Rank Pro raises that to 500 pages and 10 a day, and Rank Agency to 2,000 pages and 30 a day.

  1. It reads robots.txt and your sitemaps, then follows same-site links from your start page, treating www and non-www as one site.
  2. It checks up to 50 external links, and probes site-wide behaviour: the HTTPS redirect, the www twin, trailing slashes, a made-up URL (to catch soft 404s) and your llms.txt.
  3. It crawls politely: one request at a time, at least a second apart, honouring robots.txt and Crawl-delay.
  4. The report gives a 0 to 100 score with category scores (technical, content, links, performance, structured data, trust, local, social and accessibility) and a separate AI Search Readiness score.
  5. "Fix these first" lists the five issues that would add the most points, and every issue explains why it matters and how to fix it.

If you'd rather keep JoinPostlyBot out, add User-agent: JoinPostlyBot with Disallow: / to your robots.txt. The JoinPostlyBot page explains how it behaves.

Frequently asked questions

How often should I run an SEO site audit?

After every big change, such as a redesign, a move to a new CMS or a domain change, and otherwise on a regular schedule such as monthly. Paid Postly Rank plans can re-audit a site every week and alert you when new errors appear.

Why does a site audit find pages that don't appear in Google?

An audit crawler follows your links and sitemaps, so it finds every page it can reach. Google decides separately which pages to index. Compare the audit with the page indexing report in your own Google Search Console to see which pages Google left out and why.

Is a site audit score the same as a Google ranking?

No. A site health score summarizes the technical and content issues the audit found. It is a to-do list in number form, not a measure of where Google ranks your pages.

Will the Postly Rank crawler slow down my website?

It shouldn't. JoinPostlyBot fetches one page at a time, at least a second apart, follows any Crawl-delay in robots.txt up to 10 seconds, and only one audit of the same website runs at a time across all users.

What is a soft 404 and why does it matter?

A soft 404 is a missing page that shows a "not found" message but returns a 200 OK status. Search engines may treat it as a real, thin page. A missing page should return 404 or 410, or redirect to a genuinely relevant replacement.

Audit your site for free

Postly Rank's Site Audit crawls up to 50 pages per audit on the Free plan, scores your site from 0 to 100 and explains how to fix every issue.