How to check hreflang tags for errors
·4 min read
A hreflang checker has to confirm six things. Every alternate links back, every code is a valid language and region, every target is the final canonical URL, no target is noindex, each page lists itself, and the HTML, header and sitemap declarations agree. You can check all six by hand with view source, curl -I and your sitemap file. Start with missing return links, because they cancel everything else.
This post assumes you already know what the tags look like. If you don't, read what hreflang is and how to write it first.
Hreflang checker list of the six errors
| Error | What Google does | Where to look |
|---|---|---|
| Missing return link | Ignores the pair | Source of each alternate |
| Invalid code, such as en-UK | Ignores the annotation | Source, headers, sitemap |
| Target redirects or is canonicalized away | Return link may not match, set may collapse | curl -I, canonical tag of each target |
| Target is noindex | Version can't show in search | Meta robots, X-Robots-Tag |
| No self-reference | Set is incomplete | Source of the page itself |
| Methods disagree | Unpredictable | Compare all three |
1. Missing return links
Google's localized versions guide is blunt about this. If two pages don't both point to each other, it ignores the tags. Page A listing page B does nothing until page B lists page A.
To check, open the English page, copy every hreflang URL, then open each one and search its source for the English URL. The match has to be exact. If the German page links back to https://example.com/en and the English page lives at https://example.com/en/, that is a different URL.
Return links usually break when one language runs on a different template, or when someone adds a new market to one template and forgets the others.
Check the self-reference while you are in the source. Google says each language version must list itself as well as every other version. CMS plugins that build the list from "other translations" often skip the current page.
2. Invalid language and region codes
Google supports only ISO 639-1 language codes with an optional ISO 3166-1 alpha-2 region. The classic mistake is en-UK. The region code for the United Kingdom is GB, and Google's guide lists UK among the reserved codes that have no effect in Search. Use en-GB.
Other codes I see on real sites:
jpfor Japanese. It should beja, sincejpis the country.cnalone. Writezh-CN, orzh-Hansfor the script.seanddkfor Swedish and Danish. The language codes aresvandda.en_US. The separator must be a hyphen, soen-US.gbon its own. A region needs a language in front of it.
A country code alone is never valid. The value always starts with a language.
3. Redirected, noindexed or non-canonical targets
Each hreflang URL should be the final URL that returns 200 and declares itself canonical. Check the status with curl:
curl -sI https://example.com/de/ | grep -iE "^HTTP|^location"A 301 or 302 with a Location header means the annotation points at a redirect. Swap in the destination. Redirects also break the return-link match from error 1, because the alternate links back to one URL and the page lives at another. The redirect chains guide covers tracing longer hops.
Then check the canonical tag on each target. If /de/ has <link rel="canonical" href="https://example.com/en/">, you are telling Google the German page is a duplicate of the English one, which cancels the hreflang relationship. Every language version should be canonical to itself. The canonical tag guide explains how Google treats that signal.
A language version that is noindex can't appear in search, so the annotation pointing at it is wasted. Check two places, because noindex hides in both:
curl -sI https://example.com/fr/ | grep -i x-robots-tag
curl -s https://example.com/fr/ | grep -i '<meta name="robots"'A common cause is a staging leftover, where a new market launches with the noindex it had in development. See the X-Robots-Tag post for the header version.
4. Conflicting declarations across methods
Google treats the HTML head, the HTTP Link header and the XML sitemap as equivalent and says using more than one has no benefit. The risk is that two methods drift apart, for example the sitemap says de-AT maps to /at/ while the page head says /de-at/. Google does not document which one wins.
Compare all three for one page:
- View source and list every
<link rel="alternate" hreflang>in the head. - Run
curl -sI <url> | grep -i "^link"to see header annotations, which PDFs and other non-HTML files depend on. - Search your sitemap for the page's
<loc>and read its<xhtml:link>entries.
If more than one method is active, pick one and remove the others. The XML sitemap guide shows the sitemap format if that is the method you keep. Also confirm each code points to a single URL. Two URLs under the same hreflang value is a conflict inside one method, and Google can't tell which one serves that audience.
Check your hreflang tags in one run
Our Hreflang Checker reads the hreflang links in the page's head and its Link header, then fetches up to 20 language versions they list. It flags invalid codes with the fix (en-UK becomes en-GB), duplicate codes, relative URLs, a missing self-reference, a missing x-default and a page canonicalized elsewhere. For each alternate it reports whether it loads, redirects, is noindex, is canonicalized away or fails to link back.
It does not read hreflang from your XML sitemap and it checks one page and its alternates, not the whole site, so run the sitemap comparison above by hand. If an alternate's bot protection blocks the checker, the report says that version could not be verified instead of guessing. Each run costs 10 credits.