Introducing Highlights: the context that matters

HTTP 404 Not Found: What It Means and How to Fix It

HTTP 404 Not Found is a status code meaning the origin server did not find a current representation of the requested URL, or is not willing to say that one exists. It does not say whether the absence is temporary or permanent. Caches may store a 404 by default, and crawlers treat it as a signal to drop the URL.

Code
404
Name
Not Found
Class
4xx client error
Retry?
No, fix the cause first

What causes a 404 error?

  • →The page was deleted or moved without a redirect.
  • →A typo, wrong case, or trailing-slash mismatch in the URL.
  • →A stale link from a sitemap, an old crawl, or a model that guessed the URL.
  • →A server hiding a forbidden resource behind a 404 instead of a 403.

How do you fix a 404 error when web scraping?

  • →Refresh your URL list from the site’s sitemap instead of reusing an old crawl. A sitemap extractor returns the URLs a site currently publishes.
  • →Normalize URLs before fetching: lowercase the host, keep the path case, and follow the site’s trailing-slash convention.
  • →Do not retry a 404. Record it, drop the URL from the queue, and re-check it on the next full crawl.
  • →Watch for soft 404s: pages that return 200 with "not found" text. Check the content, not just the status.

How do you fix a 404 error on your own server?

  • →Redirect moved pages with a 301 to the closest equivalent.
  • →Return a real 404 status for missing pages rather than a 200 with an error message.
  • →Use 410 Gone when a page was removed on purpose and is not coming back.

How do you handle a 404 error in a retry loop?

404 is not in RETRYABLE, so raise_for_status() raises on the first response instead of spending retries on a request that will fail the same way. Fix the cause, then send the request again.

import random
import time

import requests

RETRYABLE = {408, 429, 500, 502, 503, 504, 520, 521, 522, 523, 524}


def fetch(url: str, max_attempts: int = 5) -> requests.Response:
    for attempt in range(max_attempts):
        try:
            response = requests.get(url, timeout=(10, 60))
        except requests.Timeout:
            time.sleep(2**attempt + random.uniform(0, 1))
            continue
        if response.status_code not in RETRYABLE:
            response.raise_for_status()
            return response
        retry_after = response.headers.get("Retry-After", "")
        backoff = 2**attempt + random.uniform(0, 1)
        time.sleep(min(int(retry_after) if retry_after.isdigit() else backoff, 60))
    raise RuntimeError(f"Gave up on {url} after {max_attempts} attempts")

How does Context.dev handle a 404 error?

When a target page returns 404, the scrape output fails with NOT_FOUND. It is the one failure Context.dev still bills, at the request’s base price (1 credit for a standard scrape), because the fetch itself succeeded. Use Map URLs to get a current URL list before a large job.

See what the web scraping API does on every request, or read how to fix HTTP errors in web scraping for a longer walkthrough.

Frequently asked questions about a 404 error

What does 404 mean in HTTP?

The server found nothing at that URL, or will not say whether something is there. The request reached the server successfully; only the resource is missing.

What is the difference between 404 and 410?

A 404 makes no claim about the future. A 410 says the removal is deliberate and permanent, which search engines treat as a stronger signal to drop the URL.

Do 404 errors hurt SEO?

A 404 on a page that should not exist is normal. The problem is internal links and sitemaps that point at 404s, which waste crawl budget and send visitors to dead ends.

Which status codes are related to 404?

Sources

Last reviewed

Ship an agent that actually knows things.

Free tier, 10-minute integration, and the same API powering agents at Mintlify, daily.dev, and Propane. No credit card to start.