Status code 500: internal server errors and SEO

·7 min read

Status code 500, Internal Server Error, means the server ran into an unexpected condition and could not finish the request. It names no cause, so the real error is in your server and application logs, never in the browser. A few minutes of 500s do no lasting SEO harm. Google slows its crawling when it sees them, though, and URLs that keep returning a 500 are eventually dropped from the index.

What does status code 500 mean?

RFC 9110, the HTTP standard, says a 500 means the server "encountered an unexpected condition that prevented it from fulfilling the request." The visitor sees a generic error page, while the stack trace that explains it goes to a log file.

500 vs 502, 503 and 504

The four codes differ in who failed. A 500 usually comes from your application itself. A 502 or 504 comes from a proxy in front of it, reporting that the app behind it gave a bad answer or no answer in time.

Code Name What it means Who usually sends it Look first at
500 Internal Server Error The server hit an unexpected condition and could not finish Your application or web server Application and error logs
502 Bad Gateway A proxy got an invalid response from the server behind it nginx, a load balancer or a CDN Whether the app process is running
503 Service Unavailable The server is overloaded or in maintenance and expects to recover Your server, on purpose or under load Maintenance mode, load, rate limits
504 Gateway Timeout A proxy waited too long for the server behind it nginx, a load balancer or a CDN Slow queries and long requests

For the proxy side, see our guide to the 502 Bad Gateway error.

What causes a 500 internal server error?

The pattern of failures is the fastest clue to the cause. One URL failing points to bad data or a code path only that page hits. Every URL failing at once points to configuration, a deploy or the database.

An unhandled exception

A template assumes every product has a price. One product does not, and the code throws TypeError: Cannot read properties of null (reading 'toFixed'). Nothing catches it, so the framework turns it into a 500. Only the pages with the odd data fail, which is why these bugs survive a quick test of the homepage.

A bad .htaccess directive

Apache reads .htaccess on every request, and one line it cannot parse breaks every URL in that directory and below it. The error log names the problem:

/var/www/html/.htaccess: Invalid command 'RewriteEngine', perhaps misspelled or defined by a module not included in the server configuration

That message means mod_rewrite is not enabled. Undo the last edit to .htaccess before you debug anything else.

PHP running out of memory

PHP Fatal error:  Allowed memory size of 134217728 bytes exhausted (tried to allocate 20480 bytes)

134,217,728 bytes is 128 MB, PHP's default memory_limit. A heavy plugin or a large export pushes one request over it and PHP dies mid-request. Raising the limit brings the page back. Then find the plugin or query that needs that much, because it will hit the new ceiling too.

A failed deploy or a missing environment variable

If every URL returns 500 minutes after a release, suspect the deploy. The new code reads an environment variable that exists in staging but not in production, or queries a column whose migration never ran. Roll back, then debug. If the app process fails to start at all, the proxy in front of it usually answers 502 instead.

A database that is down

When the database refuses connections or fills its disk, every dynamic page fails together. WordPress prints "Error establishing a database connection" on every URL, including the robots.txt it generates. A failing robots.txt stops Google from crawling the whole site for 12 hours.

How to find the real error behind an HTTP 500

The real error is in your web server's error log or your application's log, not in the browser.

  1. Note the exact URL and time of a failure. For intermittent errors, take the timestamp from your access log or uptime monitor.
  2. Read the web server's error log. On Debian and Ubuntu that is /var/log/apache2/error.log or /var/log/nginx/error.log. RHEL-based servers use /var/log/httpd/error_log.
  3. Read the application log, such as the PHP-FPM log or your Node process output. On WordPress, set WP_DEBUG and WP_DEBUG_LOG to true and WP_DEBUG_DISPLAY to false in wp-config.php, and errors go to wp-content/debug.log.
  4. Search the access log to see how widespread the failures are and whether Googlebot hit them.
# URLs that returned 5xx, most frequent first (combined log format)
awk '$9 ~ /^5/ {print $9, $7}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -20

# The same, limited to requests that claim to be Googlebot
grep Googlebot /var/log/nginx/access.log | awk '$9 ~ /^5/ {print $9, $7}' | sort | uniq -c | sort -rn | head -20

Do not switch on display_errors in production to read the error in the browser. It shows file paths and query details to every visitor, Googlebot included.

Do 500 errors hurt SEO?

Occasional 500s do not hurt SEO. Persistent ones do. Google's HTTP status code documentation describes how its crawlers react:

  • 5xx responses make Google's crawlers slow down, and the slowdown is proportionate to how many URLs return a server error.
  • Google ignores any content it receives with a 5xx status, so a friendly error page full of links gives it nothing to follow.
  • URLs already in the index stay there at first. URLs that keep returning a server error are eventually removed.
  • Once the server returns 2xx again, Google raises the crawl rate gradually.

Google publishes no timeline for "eventually". A bad deploy you roll back in ten minutes is unlikely to matter. A template bug that breaks 3,000 product pages for a week is a real risk. Those pages can drop out, and the slower crawl means Google takes longer to see the fix.

The opposite mistake hurts too. An error page that returns 200 tells Google the error message is the page's content. Keep the real status code and customize only the body.

Return 503, not 500, for planned downtime

For maintenance you plan, send 503 Service Unavailable with a Retry-After header. Google's guide to pausing a website recommends exactly that for an outage of one or two days and calls a few days the limit. For longer closures it recommends an indexable home page that returns 200.

Google's crawling documentation lists 500, 502 and 503 together, so a 503 earns no special crawl treatment. The case for it is clarity. A 503 says the outage is deliberate and temporary, Retry-After suggests when to come back, and your logs keep 500 for real bugs.

On Apache, this .htaccess block puts the site in maintenance mode and lets your own IP through:

RewriteEngine On
RewriteCond %{REMOTE_ADDR} !=203.0.113.10
RewriteCond %{REQUEST_URI} !=/maintenance.html
RewriteRule ^ - [R=503,L]
ErrorDocument 503 /maintenance.html
Header always set Retry-After "7200"

Retry-After takes a number of seconds or an HTTP date, and Header needs mod_headers. Remove the block when you finish. A 503 left in place for weeks becomes the persistent server error that Google drops URLs for.

What happens when robots.txt returns a 500?

Google stops crawling your whole site for 12 hours, even if every page works. Google's robots.txt specification sets out what happens next:

  1. For 12 hours, Google crawls nothing on the site and keeps retrying robots.txt.
  2. For the next 30 days, it uses the last good copy it fetched. With no cached copy, it assumes there are no restrictions.
  3. After 30 days, if the rest of the site is available, Google acts as if there is no robots.txt. If the site is generally unavailable, Google stops crawling but keeps checking robots.txt.

That makes a generated robots.txt a weak point. WordPress builds its virtual robots.txt in PHP after connecting to the database, so a fatal PHP error or a database outage takes it down with the pages. A missing robots.txt is safer than a failing one, because Google treats a 404 there as permission to crawl everything. The WordPress robots.txt guide shows how to replace the virtual file with a real one.

How to find 500 errors in Search Console

Search Console shows the 500s Googlebot actually hit, in three places.

Page indexing report. Under Indexing, then Pages, look for the reason "Server error (5xx)" and click it for example URLs. Fix them, then click Validate fix. Google says validation typically takes up to about two weeks.

Crawl stats. Open Settings, then Crawl stats. The breakdown by response shows "Server error (5XX)" over time, and host status flags robots.txt fetch failures, DNS failures and connectivity problems from the past 90 days. Google offers the Crawl stats report only for Domain properties and root-level URL-prefix properties.

URL Inspection. Test live URL fetches the page now. Google warns that server errors can be transient, so a passing live test does not prove the page was fine when Googlebot crawled it.

Check a list of URLs for 500 errors

Our HTTP Status Bulk Checker requests each URL in a pasted list, one per line. It checks up to 20 URLs per run, follows up to five redirects and gives each URL eight seconds to answer. For every URL it returns the final status code, the first redirect and where it points, the number of hops and the response time, plus totals for server errors, client errors and redirects. A URL that answers with a rate limit, a bot check or a 403 is listed as "did not answer", because bot protection sent that response, not the page.

It does not read page bodies, crawl your site or expand a sitemap, so you choose the URLs and an error page served with a 200 shows up as 200. Each run sends one request per URL, which means a 500 that fires only under load or only for Googlebot still needs your logs. Put your robots.txt URL in the list next to one URL from each template. A run costs 20 credits.

Keep reading