How to check hreflang tags for errors

·4 min read

A hreflang checker has to confirm six things. Every alternate links back, every code is a valid language and region, every target is the final canonical URL, no target is noindex, each page lists itself, and the HTML, header and sitemap declarations agree. You can check all six by hand with view source, curl -I and your sitemap file. Start with missing return links, because they cancel everything else.

This post assumes you already know what the tags look like. If you don't, read what hreflang is and how to write it first.

Hreflang checker list of the six errors

Error What Google does Where to look
Missing return link Ignores the pair Source of each alternate
Invalid code, such as en-UK Ignores the annotation Source, headers, sitemap
Target redirects or is canonicalized away Return link may not match, set may collapse curl -I, canonical tag of each target
Target is noindex Version can't show in search Meta robots, X-Robots-Tag
No self-reference Set is incomplete Source of the page itself
Methods disagree Unpredictable Compare all three

Google's localized versions guide is blunt about this. If two pages don't both point to each other, it ignores the tags. Page A listing page B does nothing until page B lists page A.

To check, open the English page, copy every hreflang URL, then open each one and search its source for the English URL. The match has to be exact. If the German page links back to https://example.com/en and the English page lives at https://example.com/en/, that is a different URL.

Return links usually break when one language runs on a different template, or when someone adds a new market to one template and forgets the others.

Check the self-reference while you are in the source. Google says each language version must list itself as well as every other version. CMS plugins that build the list from "other translations" often skip the current page.

2. Invalid language and region codes

Google supports only ISO 639-1 language codes with an optional ISO 3166-1 alpha-2 region. The classic mistake is en-UK. The region code for the United Kingdom is GB, and Google's guide lists UK among the reserved codes that have no effect in Search. Use en-GB.

Other codes I see on real sites:

  • jp for Japanese. It should be ja, since jp is the country.
  • cn alone. Write zh-CN, or zh-Hans for the script.
  • se and dk for Swedish and Danish. The language codes are sv and da.
  • en_US. The separator must be a hyphen, so en-US.
  • gb on its own. A region needs a language in front of it.

A country code alone is never valid. The value always starts with a language.

3. Redirected, noindexed or non-canonical targets

Each hreflang URL should be the final URL that returns 200 and declares itself canonical. Check the status with curl:

curl -sI https://example.com/de/ | grep -iE "^HTTP|^location"

A 301 or 302 with a Location header means the annotation points at a redirect. Swap in the destination. Redirects also break the return-link match from error 1, because the alternate links back to one URL and the page lives at another. The redirect chains guide covers tracing longer hops.

Then check the canonical tag on each target. If /de/ has <link rel="canonical" href="https://example.com/en/">, you are telling Google the German page is a duplicate of the English one, which cancels the hreflang relationship. Every language version should be canonical to itself. The canonical tag guide explains how Google treats that signal.

A language version that is noindex can't appear in search, so the annotation pointing at it is wasted. Check two places, because noindex hides in both:

curl -sI https://example.com/fr/ | grep -i x-robots-tag
curl -s https://example.com/fr/ | grep -i '<meta name="robots"'

A common cause is a staging leftover, where a new market launches with the noindex it had in development. See the X-Robots-Tag post for the header version.

4. Conflicting declarations across methods

Google treats the HTML head, the HTTP Link header and the XML sitemap as equivalent and says using more than one has no benefit. The risk is that two methods drift apart, for example the sitemap says de-AT maps to /at/ while the page head says /de-at/. Google does not document which one wins.

Compare all three for one page:

  1. View source and list every <link rel="alternate" hreflang> in the head.
  2. Run curl -sI <url> | grep -i "^link" to see header annotations, which PDFs and other non-HTML files depend on.
  3. Search your sitemap for the page's <loc> and read its <xhtml:link> entries.

If more than one method is active, pick one and remove the others. The XML sitemap guide shows the sitemap format if that is the method you keep. Also confirm each code points to a single URL. Two URLs under the same hreflang value is a conflict inside one method, and Google can't tell which one serves that audience.

Check your hreflang tags in one run

Our Hreflang Checker reads the hreflang links in the page's head and its Link header, then fetches up to 20 language versions they list. It flags invalid codes with the fix (en-UK becomes en-GB), duplicate codes, relative URLs, a missing self-reference, a missing x-default and a page canonicalized elsewhere. For each alternate it reports whether it loads, redirects, is noindex, is canonicalized away or fails to link back.

It does not read hreflang from your XML sitemap and it checks one page and its alternates, not the whole site, so run the sitemap comparison above by hand. If an alternate's bot protection blocks the checker, the report says that version could not be verified instead of guessing. Each run costs 10 credits.

Keep reading