Orphan pages: how to find and fix them

·4 min read

Orphan pages are pages on your site that no other page on your site links to. Search engines and visitors can only reach them through a sitemap, a bookmark, an ad or a link from another site. To find them, take every URL you know exists, crawl your site by following internal links, and compare the two lists. Then link each orphan you want ranked, redirect the ones that duplicate a better page, and remove or noindex the rest.

What is an orphan page?

An orphan page is a live URL with zero internal links pointing at it. It can still return 200, sit in your XML sitemap and even rank, but you cannot click your way to it from the homepage.

Two near relatives are worth separating:

Orphan page Deep page
Internal links pointing at it None At least one
Reachable by following links No Yes, after several clicks
Usual fix Add a link from a relevant page Link it from a page closer to the homepage

A page linked only through JavaScript that runs after load can behave like an orphan too. Google says it can't reliably extract URLs from <a> elements that lack an href attribute, so a click handler on a <div> or a bare <a> hides the link from crawlers.

Why orphan pages hurt SEO

Orphan pages hurt SEO because links are how search engines find pages and judge which ones matter. Google's link guidelines state that it uses links "to find new pages to crawl" and as a relevancy signal, and that every page you care about should have a link from at least one other page on your site.

An orphan loses on each count.

Discovery. Google's How Search works guide names two ways it finds URLs, extracting a link from a known page or reading a sitemap you submitted. An orphan missing from the sitemap has neither. One that is listed gets found, but Google's sitemap overview says a sitemap doesn't guarantee that every item in it will be crawled and indexed. Orphans are common among URLs stuck in Discovered, currently not indexed.

Internal link signals. Every link carries anchor text and tells Google the linking page thinks the target is worth a visit. An orphan gets none of that, so it competes on its own content alone.

AI crawlers. Bots that follow links to build AI answers never reach a page nothing links to.

How to find orphan pages

You find orphan pages by comparing a list of URLs you know exist with the list a crawler reaches from your homepage. Anything in the first list and missing from the second is an orphan.

  1. Build the list of known URLs. Combine your XML sitemap, a CMS export of published pages, landing pages from your analytics and URLs that appear in your server logs. Each source catches orphans the others miss. Logs catch old pages that still get hits, and log file analysis shows how to pull them.
  2. Crawl the site from the homepage. Use a crawler that follows only <a href> links, respects robots.txt and records every internal URL it reaches.
  3. Normalize both lists. Match protocol, host, trailing slashes and case the way your site does, and strip tracking parameters. Otherwise https://example.com/pricing and https://example.com/pricing/ show up as two pages and one looks orphaned.
  4. Subtract. Known URLs minus crawled URLs gives your orphan candidates.
  5. Check each candidate. Confirm it returns 200 and isn't a redirect, a noindex page or a URL robots.txt blocks. Those belong in a sitemap cleanup, not on your orphan list.

How to fix orphan pages

Fix each orphan by deciding whether it deserves to exist, then picking one of four actions.

Situation Action
Useful page you want ranked Link it from a relevant page, hub or category
Thin page that overlaps a stronger one Merge the content and 301 redirect to the stronger page
Page that must stay live but shouldn't rank, such as a campaign thank-you page Add noindex and drop it from the sitemap
Outdated page with no traffic and no backlinks Remove it and return 404 or 410

For pages you keep, skip the footer link. Link from the page a reader would be on right before they want this one, with anchor text that names the target. Our post on internal linking covers how to build those hub pages.

Before deleting anything, check backlinks. An orphan with links from other sites deserves a redirect or an internal link.

Then fix what created them. Redesigns that drop a navigation level and category pages that paginate old posts out of view are the usual causes. A related-content block in your CMS templates stops new orphans appearing.

Find orphan pages on your site with one crawl

The Orphan Page Finder crawls up to 150 pages from the URL you enter, following same-host links breadth-first and skipping URLs robots.txt blocks for Googlebot. It doesn't follow links on pages marked nofollow. In parallel it reads the sitemaps your robots.txt declares, or tries /sitemap.xml, /sitemap_index.xml and /wp-sitemap.xml, up to 10 sitemap files and 10,000 URLs. Sitemap URLs no crawled page links to come back as orphans. The same run lists indexable pages missing from the sitemap, pages four or more clicks deep, broken internal links and sitemap URLs that redirect, fail, carry noindex or point their canonical elsewhere.

Its limits matter. It finds only orphans listed in your sitemap, so pages known only to your CMS or logs need the manual comparison above. It doesn't see links added by JavaScript after load or links from other sites. If the crawl stops at 150 pages before covering the site, unlinked sitemap URLs are reported as unconfirmed, and starting from a section hub gives a fuller answer for that section. Each run costs 20 credits.

Keep reading