403, 429 and 503 errors when scraping — causes and how to handle them properly (with Python code)
In this article
When a scraper starts getting errors, the most common reflex is "change the IP". But each HTTP status code says something different, and rotating is right in only a few cases. Reading the code properly lets you fix the right thing — and avoid turning a temporary limit into a lasting block. This guide covers the errors you'll meet most when scraping through a proxy, with Python code that handles them properly.
Read the code before rotating
| Code | Meaning | What to do |
|---|---|---|
| 429 | Too many requests | Wait per Retry-After, slow down |
| 403 | Refused | Find the cause; don't blindly retry |
| 503 | Overloaded or a challenge page | Back off; read the response body |
| 401 | The site needs a login | The data isn't public — reconsider |
| 407 | The proxy rejects your login | Fix the proxy username and password |
| 404 | No such page | Fix the URL; don't retry |
| Connection error, timeout | Network trouble or the proxy reconnecting | Retry after a few seconds |
The key distinction: 407 comes from the proxy, while 401, 403, 429 and 503 come from the website. Proxy-side errors are covered in common proxy errors.
429 Too Many Requests
The site is saying plainly that you've sent too many requests in a given time. It's the most "polite" error — the site hasn't blocked you, it's asking you to slow down.
- Read the
Retry-Afterheader. It says how long to wait: a number of seconds (Retry-After: 120) or a point in time. - No header? Back off progressively — wait 5, 10, 20, 40 seconds… with a little randomness.
- Lower your overall pace after a 429, not just for the failed request.
Rotating to dodge a 429 only moves the problem: your pace is unchanged, and you'll hit 429 again on the new IP — sometimes with a harsher block.
403 Forbidden
The site understood the request but refuses to serve it. Common causes:
- The IP is in a server range and the site blocks the whole range — try the same URL through a 4G proxy.
- Your program doesn't look like a browser — missing
User-Agent,Accept,Accept-Language. - Content restricted by country — the IP is in the wrong country.
- The site forbids automated access to that part.
A quick diagnosis: open that exact URL in a normal browser through the same proxy. The browser gets in but the script doesn't → the problem is the script. The browser gets 403 too → the problem is the IP or country. If the site clearly forbids automated access, stop — see scraping anti-bot websites.
503 and challenge pages
A 503 can mean two very different things:
- The site is overloaded or under maintenance — often with
Retry-After. Back off, and consider slowing down, since you might be part of the load. - An anti-bot challenge page — the body is a "checking your browser" page rather than data.
Log part of the error response body to tell the two apart. Repeated challenges mean the site doesn't want automated traffic; retrying faster only makes things worse.
401 and 407
- 401 Unauthorized — the site needs a login. That data isn't public; collecting it automatically usually breaks the terms.
- 407 Proxy Authentication Required — the proxy doesn't accept your username and password. Recheck the proxy details and encode special characters in the password. Details in the 407 error.
Connection errors and timeouts
With a 4G proxy, brief connection errors are normal right after a rotation: the device needs a few seconds to reconnect to the mobile network. The odd error when signal is weak can happen too.
- Set a sensible timeout (20–30 seconds) instead of waiting forever.
- Retry after a few seconds on connection errors, with a cap on attempts.
- After rotating, wait until the proxy is back before sending requests — sample code is in rotating proxy IPs via API.
When rotating does make sense
- The current IP is flagged because of someone else. Mobile IPs are shared with many subscribers, so one is sometimes limited before you've even started.
- Starting a new session — a new run, a new profile.
- Prolonged flaky connectivity from the device.
Rotating is not the fix for a 429 caused by going too fast, or a 403 because the site forbids automation. With 65Proxy, each proxy rotates at most once every 5 minutes — a limit that happens to suit using it properly.
Sample code: handling errors properly
The Python below waits per Retry-After on 429 or 503, backs off on connection errors, and does not blindly retry a 403 or 404. Install with pip install requests and set the PROXY environment variable:
import os
import random
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
import requests
PROXY = os.environ["PROXY"] # http://USER:PASS@HOST:PORT
session = requests.Session()
session.proxies = {"http": PROXY, "https": PROXY}
session.headers["User-Agent"] = "my-crawler/1.0 (+mailto:[email protected])"
def retry_after(resp):
"""Seconds to wait per the Retry-After header, or None."""
value = resp.headers.get("Retry-After", "").strip()
if value.isdigit():
return int(value)
try:
when = parsedate_to_datetime(value)
except (TypeError, ValueError):
return None
if when.tzinfo is None:
when = when.replace(tzinfo=timezone.utc)
return max(0, (when - datetime.now(timezone.utc)).total_seconds())
def fetch(url, max_tries=5):
delay = 5
for _ in range(max_tries):
try:
resp = session.get(url, timeout=30)
except requests.RequestException as err:
print(f"connection error: {err}")
else:
if resp.status_code < 400:
return resp
if resp.status_code not in (429, 503):
resp.raise_for_status() # 403, 404…: stop, don't retry
wait = retry_after(resp)
if wait is not None:
print(f"{resp.status_code}: waiting {wait:.0f}s per Retry-After")
time.sleep(wait)
continue
time.sleep(delay + random.uniform(0, delay / 2))
delay = min(delay * 2, 300)
raise RuntimeError(f"giving up on {url} after {max_tries} tries")
for url in ["https://example.com/page/1", "https://example.com/page/2"]:
print(url, fetch(url).status_code)
time.sleep(random.uniform(3, 6)) # pause between pagesIf 429s keep coming, don't raise the retry count — lengthen the pause between pages.
Frequently asked questions
Can I just rotate the IP on a 429?
Technically yes, but you shouldn't. A 429 means you're going too fast; a new IP at the same pace will hit 429 again, and constantly rotating to dodge limits is a pattern sites recognise quickly.
Why does the browser get in when my script gets 403?
Because the site can tell the script isn't a browser — from missing headers, the TLS fingerprint, or because it doesn't run JavaScript. See how websites detect proxies.
How many times should I retry?
Three to five times for temporary errors (429, 503, connection errors), with growing waits. Don't retry 403, 404 or 401 — they won't go away on their own.