HTTP 403 Forbidden: What It Means and How to Fix It
HTTP 403 Forbidden is a status code meaning the server understood the request but refuses to fulfill it, and sending the same credentials again will not help. For scrapers, a 403 usually comes from a bot-management layer such as Cloudflare, which blocks requests whose IP address, headers, or browser fingerprint look automated, rather than from the page itself.
- Code
- 403
- Name
- Forbidden
- Class
- 4xx client error
- Retry?
- No, fix the cause first
What causes a 403 error?
- →A bot-management or WAF rule (Cloudflare, Akamai, and others) that flags the request as automated. Cloudflare returns 403 for most of its security blocks.
- →A datacenter IP range, a country, or an IP address the site has blocked.
- →Headers that do not look like a browser: a library default
User-Agent, or noAccept-Language. - →Valid credentials without permission for this resource, or a directory listing the server will not show.
How do you fix a 403 error when web scraping?
- →Look at the response body and headers first. A Cloudflare block has
server: cloudflareand acf-rayheader, and error 1020 means a specific firewall rule matched. - →Send the headers a real browser sends, and render the page in a headless browser when the site checks JavaScript execution.
- →Slow down and spread requests over time. Many 403s are rate limits that escalated after a run of 429s.
- →Respect explicit refusals. If robots.txt or the terms disallow automated access, use an official API or ask the site owner. The guide to fixing 403 errors when web scraping covers the options in order.
How do you fix a 403 error on your own server?
- →Check WAF and bot-management logs for the Ray ID or request ID before changing rules.
- →Return 403 only when credentials or permissions are the problem, and 401 when credentials are missing.
- →If you block bots on purpose, publish a robots.txt policy and a contact address so legitimate crawlers can ask for access.
How do you handle a 403 error in a retry loop?
403 is not in RETRYABLE, so raise_for_status() raises on the first response instead of spending retries on a request that will fail the same way. Fix the cause, then send the request again.
import random
import time
import requests
RETRYABLE = {408, 429, 500, 502, 503, 504, 520, 521, 522, 523, 524}
def fetch(url: str, max_attempts: int = 5) -> requests.Response:
for attempt in range(max_attempts):
try:
response = requests.get(url, timeout=(10, 60))
except requests.Timeout:
time.sleep(2**attempt + random.uniform(0, 1))
continue
if response.status_code not in RETRYABLE:
response.raise_for_status()
return response
retry_after = response.headers.get("Retry-After", "")
backoff = 2**attempt + random.uniform(0, 1)
time.sleep(min(int(retry_after) if retry_after.isdigit() else backoff, 60))
raise RuntimeError(f"Gave up on {url} after {max_attempts} attempts")
How does Context.dev handle a 403 error?
Context.dev runs every scrape in a browser with managed proxies and handles proxy selection and fetch retries for you. If the site still serves an anti-bot challenge or access wall, the output fails with WEBSITE_BLOCKED inside an HTTP 200 response, and a request where every output fails is not charged. Retrying later or from another sharedParams.country sometimes succeeds. A 403 from the Context.dev API itself means browser actions need a paid plan, zero data retention is not enabled, or the API key lacks permission.
See what the web scraping API does on every request, or read how to fix HTTP errors in web scraping for a longer walkthrough.
Frequently asked questions about a 403 error
Why do I get 403 Forbidden from Cloudflare when scraping?
Cloudflare scored the request as automated or matched a firewall rule. The usual triggers are datacenter IPs, non-browser headers, missing JavaScript execution, or request rates no person produces.
Is a 403 permanent?
Not always. A 403 from a bot wall can clear when you slow down or change network, while a 403 for missing permissions stays until access is granted. Retrying the identical request immediately rarely helps.
What is the difference between 403 and 404?
A 403 says the resource exists and you may not have it. RFC 9110 lets servers return 404 instead when they do not want to reveal that a forbidden resource exists.
Which status codes are related to 403?
Sources
Last reviewed