A technical SEO audit checks whether search engines, and now AI search systems, can crawl your site, render it, index the right pages and serve them quickly. This checklist covers the nine areas a complete audit looks at in 2026, in the order that usually matters most. Every rule below is based on the search engines’ own documentation, linked where it applies, so you can check the source rather than take our word for it.
You need two things to work through it: Google Search Console for what Google actually did with your site, and a crawler for what is on the site. If you do not have one yet, our guides to free SEO audit tools and the best SEO audit tools cover the options.
1. Crawlability
If crawlers cannot reach a page, nothing else on this list matters for it.
- robots.txt blocks nothing you want found. Check every Disallow rule against your important URL patterns. Remember what robots.txt is for: Google says it manages crawling and is not a mechanism for keeping a page out of Google. To keep a page out of the index, use
noindexor password protection instead. - CSS and JavaScript are not blocked. Google renders pages with a headless browser and will not render JavaScript from blocked files.
- The XML sitemap lists only canonical, indexable URLs that return 200. No redirects, 404s or noindexed pages. Keep each sitemap under Google’s limit of 50,000 URLs or 50 MB uncompressed, and reference it in robots.txt and Search Console.
- lastmod is accurate or absent. Google uses
lastmodonly if it is consistently and verifiably accurate, and ignorespriorityandchangefreq, so do not spend time on those two. - Redirects are single hops. Google’s crawlers follow up to 10 redirect hops, but every hop adds a request and slows users down. Point internal links straight at the final URL and remove loops.
- No broken internal links (4xx) and no server errors (5xx) on pages you link to.
- No orphan pages. Every page you want ranked should be reachable through internal links, not only through the sitemap. Important pages should sit a few clicks from the homepage, not deep in pagination.
- No crawl traps. Faceted filters, calendars and session parameters can generate endless URL variations. Our guide to crawl budget on large sites covers how to contain them.
2. Indexability
Start from Search Console’s Page indexing report: it lists the pages Google did not index and the reason for each.
- No important page is noindexed by accident. Check the meta robots tag and the
X-Robots-Tagheader, which templates and plugins sometimes set site-wide. - Canonical tags are consistent. Each indexable page has a self-referencing canonical, and the canonical, the sitemap and your internal links all point to the same URL version.
- Pagination keeps its own canonicals. Google’s guidance is to give each page in a paginated sequence its own canonical URL, not to point them all at page one. Google no longer uses
rel="next"andrel="prev". - One version of every URL. HTTP and HTTPS, www and non-www, trailing slash and no slash, upper and lower case: each pair should resolve to one version with a redirect.
- No soft 404s. A page that says “not found” or is empty but returns 200 is reported by Search Console as a soft 404. Return a real 404 or 410, or redirect to a relevant page.
- Thin and duplicate pages are handled. Consolidate near-duplicates, and improve or remove pages that add nothing. See content pruning.
3. Rendering and mobile
- Content is in the rendered HTML. Google processes JavaScript sites in three phases (crawling, rendering, indexing). Compare the raw HTML with the rendered page for your main templates; anything that only appears after a click or a scroll may be missed. Google still recommends server-side rendering or pre-rendering because not all bots can run JavaScript. See our guide to JavaScript SEO.
- The mobile version has the same content. Google uses the mobile version of a page for indexing and ranking. Text, headings, images, links, titles, meta descriptions and structured data should match the desktop version.
- Primary content is not lazy-loaded behind an interaction. Content that loads only when a user clicks or types will not be seen.
Tip: run this audit right after any migration, redesign, CMS upgrade or template change, not only on a schedule. One template change can affect every page built from it.
4. Performance and Core Web Vitals
Google’s Core Web Vitals are measured on real users at the 75th percentile of page loads, separately for mobile and desktop:
- Largest Contentful Paint (LCP): 2.5 seconds or less.
- Interaction to Next Paint (INP): 200 milliseconds or less.
- Cumulative Layout Shift (CLS): 0.1 or less.
- Time to First Byte (TTFB) is not a Core Web Vital, but web.dev says most sites should aim for 0.8 seconds or less, because a slow server delays everything after it.
How to check and fix:
- Use Search Console’s Core Web Vitals report for real-user data across the site, and PageSpeed Insights or Lighthouse to find the cause on a single URL.
- Fix by template: a slow product template slows every product page.
- Give images and embeds explicit width and height to prevent layout shift, serve modern formats, and do not lazy-load the main image at the top of the page.
- Remove or defer render-blocking CSS and JavaScript, and third-party scripts you do not need.
More detail in Core Web Vitals in 2026.
5. Structured data
- Valid and complete. No errors in Google’s Rich Results Test, and every required property present for each type you use.
- It matches what is on the page. Markup that describes content users cannot see breaks Google’s structured data guidelines.
- Entities are connected. Organization, WebSite, the article and its author as a Person, linked with
@idreferences rather than repeated as separate blocks. - Expectations are current. Since August 2023, FAQ rich results are only shown for well-known, authoritative government and health websites, and How-to rich results are no longer shown. FAQPage markup does no harm, but most sites should not expect a rich result from it. Speakable is still a beta feature for English-language publishers and US users.
- No special schema for AI. Google says there is no special schema.org markup needed for its generative AI features; structured data remains useful for rich results.
See how to add schema markup for the basics.
6. International sites (hreflang)
Skip this section if your site has one language and one country.
- Every version lists itself and all the others. Google says each language version must list itself as well as all other versions; if page B does not link back to page A, the annotations may be ignored. One-way hreflang is the most common mistake.
- Valid codes: an ISO 639-1 language code, optionally with an ISO 3166-1 region (en, en-GB, de-CH). A region on its own is not valid.
- An x-default for users whose language does not match any version, such as a language selector page.
- Canonicals agree with hreflang: each language version is its own canonical, not canonicalised to another language.
7. Security
- HTTPS on every page, with HTTP redirecting to HTTPS and a valid certificate that is not about to expire.
- No mixed content: images, scripts and styles on HTTPS pages also load over HTTPS.
- Nothing in Search Console’s Security issues report, such as hacked content or malware.
- Security headers such as Content-Security-Policy and Strict-Transport-Security protect users. They are good practice for any site, even though they are not a search ranking checklist item.
8. Accessibility
Accessibility work overlaps with SEO: alt text describes images to search engines as well as screen readers, and a clear heading structure helps both. The current standard is WCAG 2.2, a W3C Recommendation.
- Descriptive alt text on meaningful images, empty alt on decorative ones.
- A
langattribute on the page, one H1, and headings in a logical order. - Labels on every form field, sufficient colour contrast, and a page you can use with the keyboard alone.
9. AI search readiness
Google’s own guide says that optimizing for its generative AI features is still SEO: everything above already counts. On top of it, check:
- Eligibility in Google: pages must be indexed and eligible to show with a snippet, and the site must be included in Search generative AI features in Search Console.
- The right AI crawlers are allowed. Separate search bots from training bots. OpenAI uses OAI-SearchBot to show sites in ChatGPT search and GPTBot for model training; you can allow the first and block the second. Google-Extended controls use for Gemini training and grounding and, per Google, does not affect inclusion or ranking in Google Search. See our robots.txt guide.
- llms.txt is optional. Google says Google Search ignores llms.txt files; publishing one for other services neither helps nor harms you in Google. See what llms.txt is.
- Pages are worth quoting: clear answers near the top, original information, a named author.
- Measure it: Search Console’s Generative AI performance report shows how your content performs in Google’s AI features. For other engines, see our guide to AI visibility tools.
Prioritising what you find
A full audit of a real site usually finds more issues than anyone can fix at once. Work in this order:
- Anything that blocks crawling or indexing of important pages: robots.txt rules, noindex, wrong canonicals, server errors.
- Problems on templates, because one fix repairs every page built from them.
- Issues on the pages with the most traffic or revenue; Search Console shows which those are.
- Everything else, in batches, re-crawling after each batch to confirm the fix and catch anything it broke.
The checklist at a glance
- robots.txt blocks nothing important, CSS and JS crawlable
- Sitemap: only canonical 200 URLs, accurate lastmod, submitted
- Single-hop redirects, no loops, no broken internal links
- No orphan pages, no crawl traps
- No accidental noindex; consistent canonicals; one URL version
- Paginated pages keep their own canonicals; no soft 404s
- Content present in rendered HTML and on mobile
- LCP ≤ 2.5 s, INP ≤ 200 ms, CLS ≤ 0.1 at the 75th percentile; TTFB ≤ 0.8 s
- Valid structured data that matches the page; realistic rich result expectations
- hreflang with return links and x-default (if international)
- HTTPS everywhere, no mixed content, no security issues
- Alt text, lang, heading order, labels, contrast (WCAG 2.2)
- AI search crawlers allowed, generative AI eligibility on, performance measured