Technical SEO: a practical guide

·4 min read

Technical SEO is the work that lets a search engine fetch your pages, read their content, store them in its index and show them for the right queries. Think of it as four steps, crawl, render, index and serve, and ask at each step what could stop the page. Most of the gains come from a few blockers like a stray noindex, a robots.txt rule, a wrong canonical or content that only exists after JavaScript runs. The same plumbing now decides what AI crawlers from OpenAI, Anthropic and Perplexity can read.

This post is the mental model. For the item-by-item list with a test for each, use the technical SEO checklist.

How technical SEO works: crawl, render, index, serve

Google describes Search in three stages in how Search works: crawling, indexing and serving. Rendering sits between the first two and breaks often enough to count as its own step.

Step What the engine does What breaks it
Crawl Finds the URL through links or a sitemap and requests it robots.txt blocks, 5xx errors, firewalls that challenge bots, orphan pages
Render Runs the page's JavaScript to get the final HTML Content, links or metadata that only appear after scripts run
Index Decides whether to store the page and which URL represents it noindex, a canonical pointing elsewhere, near-duplicates, thin pages
Serve Picks indexed pages for a query and shows them Slow or broken mobile pages, missing titles, weak relevance

When a page is missing from search, work down the table and stop at the first step that fails. A page that is never crawled cannot have a canonical problem.

What matters most in technical SEO

A small set of signals can remove a page from search entirely. Get these right before anything else.

  1. Status codes. Pages you want found return 200. Moved pages return a single 301 to the final URL, not a chain. See HTTP status codes for SEO.
  2. Crawl access. One Disallow: / left over from a staging site wipes out everything. So does a firewall that serves bots a challenge page.
  3. Indexing directives. A noindex in a meta tag or an X-Robots-Tag header overrides every other effort on the page.
  4. Canonicals. Each page should point its canonical at itself or at the version you want ranked, and that target must return 200 and be indexable. The canonical tag guide shows the common mistakes.
  5. Content in the HTML. If the main copy loads from an API after the page mounts, you depend on Google's render step and get nothing from crawlers that skip it.
  6. Internal links. A page no other page links to is hard to find and looks unimportant. Real <a href> links, not click handlers.

On the sites I have fixed, one of these explains most "why is this page not ranking" questions.

What is overrated in technical SEO

Plenty of audit findings are true and still not worth your week.

Core Web Vitals as a ranking lever

Google's page experience documentation says it shows the most relevant result even when page experience is poor. Fix a slow page for users, not for a ranking jump.

Crawl budget on small sites

Google's crawl budget guide is written for sites with around a million pages, or tens of thousands that change daily. On a 300-page site, unindexed pages point to quality or links, not budget.

Perfect audit scores

A 61-character title is a real finding that changes nothing. Sort by impact, not by count.

HTML validation and text-to-code ratio

Search engines parse messy HTML fine. Neither number shows up in Google's documentation as something it uses.

llms.txt as a fix

Google says Search ignores it. It does not replace server-rendered content or open crawler access.

Technical SEO for AI crawlers

The four steps apply to AI search too, with one big difference at the render step. Googlebot renders JavaScript with a current version of Chromium, per Google's JavaScript SEO basics. Most AI crawlers, including GPTBot, ClaudeBot and PerplexityBot, read the HTML your server returns and do not run scripts. A client-rendered pricing page can rank in Google and still be empty to ChatGPT. The JavaScript SEO post walks through how to test and fix that.

Crawl access also gets more complicated. OpenAI alone runs separate agents for training, search and user-initiated fetches, listed in its crawler documentation, and you can allow one while blocking another. A CDN rule written to stop scrapers may block the ones you want.

There is no separate AI technical stack to build. A page that returns 200, has its content in the server HTML, is indexable and is open to the crawlers you choose is in good shape for Google, Bing and the AI engines at once.

Check your technical SEO with our tools

Our SEO tools run each step of this model on your own pages. Three of them cover most of it.

The Indexing & Canonical Checker (8 credits per URL) reports the page's status, the robots.txt rules for Google and Bing, meta robots and X-Robots-Tag, the canonical and whether its target is indexable, and snippet controls, with the tag or header to change. It does not confirm that Google has the URL in its index, because it cannot see your Search Console, and it does not render JavaScript. The Robots.txt & AI Crawler Checker (10 credits) shows which search and AI crawlers your robots.txt allows and which rule matches a given path. AI Crawler View (8 credits) compares the raw HTML with the page rendered in a headless browser and lists the content, links and structured data that only appear after JavaScript runs.

Keep reading