Using a proxy with Scrapy — authenticated setup, polite throttling, error fixes and rotating between runs
In this article
Scrapy is Python's most popular scraping framework, and it supports proxies out of the box — the documentation is just scattered. This guide gathers what you need: two ways to add an authenticated proxy, a "polite" throttling setup that keeps you from getting blocked, how to check the proxy is in use, the usual errors, and how to rotate a 4G proxy's IP without breaking a run.
How Scrapy uses a proxy
Scrapy ships with HttpProxyMiddleware, enabled by default. It takes the proxy from two places:
- the environment variables
http_proxyandhttps_proxy, read once when Scrapy starts, - each request's
request.meta["proxy"]— which takes priority over the environment.
The username and password go straight into the proxy address (http://USER:PASS@HOST:PORT) — Scrapy splits them out and sends the proxy an authentication header.
Note: Scrapy only supports HTTP/HTTPS proxies, not SOCKS5. With 65Proxy, use the HTTP connection details. The difference between the two is in SOCKS5 vs HTTP proxies.
Way 1: environment variables
The quickest way, applying to every request. Both variables start with http://, even https_proxy, because Scrapy talks to the proxy over HTTP:
export https_proxy="http://USER:PASS@HOST:PORT"
export http_proxy="$https_proxy"
scrapy crawl prices -O prices.jsonOn Windows PowerShell:
$env:HTTPS_PROXY = "http://USER:PASS@HOST:PORT"
$env:HTTP_PROXY = $env:HTTPS_PROXY
scrapy crawl prices -O prices.jsonWay 2: your own middleware
When you want control in code — say, using the proxy only for some domains — write a small middleware that sets meta["proxy"]. In middlewares.py:
import os
class ProxyMiddleware:
def __init__(self):
self.proxy = os.environ["PROXY"] # http://USER:PASS@HOST:PORT
def process_request(self, request, spider=None):
request.meta.setdefault("proxy", self.proxy)Then enable it in settings.py with a number below 750 so it runs before HttpProxyMiddleware:
DOWNLOADER_MIDDLEWARES = {
"myproject.middlewares.ProxyMiddleware": 350,
}Replace myproject with your project's name. Reading the proxy from an environment variable keeps the username and password out of your code.
Polite throttling
With a 4G proxy, every request goes out from one IP — so pace matters more than ever. A sensible starting point for settings.py:
ROBOTSTXT_OBEY = True
USER_AGENT = "my-crawler/1.0 (+mailto:[email protected])"
CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_DELAY = 2 # seconds; Scrapy varies it 0.5–1.5x
DOWNLOAD_TIMEOUT = 30 # the 180-second default is too long
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 2
AUTOTHROTTLE_MAX_DELAY = 60
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0
RETRY_TIMES = 3
RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429]
HTTPCACHE_ENABLED = True # while developing: don't re-download pages- AutoThrottle slows down automatically when the site responds slowly.
- The default retry middleware doesn't wait per
Retry-After. If you see 429s often, raiseDOWNLOAD_DELAYand lowerCONCURRENT_REQUESTS_PER_DOMAINrather than adding retries. What each error code means is in 403, 429 and 503 errors when scraping. - The HTTP cache while developing lets you fix your parsing code without hitting the target again — turn it off for real runs if you need fresh data.
Check the proxy is in use
scrapy fetch downloads a page with the current settings and prints it:
scrapy fetch --nolog https://api.ipify.orgThe result should be the proxy's IP, not your home IP. To see the network and country as well:
scrapy fetch --nolog https://ipinfo.io/jsonRotating between runs
With a 4G proxy, rotating means calling the proxy's rotation link and waiting for the device to reconnect. In Scrapy, the cleanest approach is to rotate between runs, not between requests:
- rotating mid-run leaves cookies and sessions out of step with the IP,
- in-flight requests fail while the device reconnects,
- and each 65Proxy proxy rotates at most once every 5 minutes.
A script to run before each crawl (the rotation link lives in the ROTATE_URL environment variable):
export https_proxy="http://USER:PASS@HOST:PORT"
export http_proxy="$https_proxy"
old=$(curl -s -m 15 -x "$https_proxy" https://api.ipify.org)
curl -s -m 60 "$ROTATE_URL" > /dev/null
for i in $(seq 1 20); do
sleep 3
new=$(curl -s -m 10 -x "$https_proxy" https://api.ipify.org)
if [ -n "$new" ] && [ "$new" != "$old" ]; then break; fi
done
echo "IP: $old -> $new"
scrapy crawl prices -O prices.jsonPython and Node.js versions of this wait loop are in rotating proxy IPs via API.
Common errors
- 407 Proxy Authentication Required — wrong username or password. If the password has special characters, encode them in the proxy address:
@becomes%40,:becomes%3A. - TunnelError: Could not open CONNECT tunnel — the proxy refused to open a tunnel for HTTPS; usually wrong credentials, the wrong port, or a SOCKS5 port by mistake.
- Errors with
socks5://— Scrapy doesn't support SOCKS; switch to the HTTP port. - A burst of timeouts — usually while the proxy is rotating or signal is weak; sensible
RETRY_TIMESandDOWNLOAD_TIMEOUThandle it. - The IP is still your home IP — the environment variable was set after Scrapy started, or in a different terminal window.
General proxy errors are covered in common proxy errors.
65Proxy's Singapore 4G proxies for Scrapy
Our proxies have an HTTP port with a username and password — exactly what Scrapy needs — plus an on-demand rotation link (at most once every 5 minutes). Singapore mobile-carrier IPs help you reach public data on sites that block server IPs, or data shown only to Singapore visitors. 200GB per plan, priced 1 day $4, 7 days $13, 1 month $40. Use them with the throttling setup above and respect each site's rules — see is web scraping legal and pricing.
Frequently asked questions
Is one 4G proxy enough for Scrapy?
For polite scraping — a few requests a second or fewer per site — it usually is. If you need much more speed, ask the target site about an API before thinking about more IPs.
Is there a rotating-proxy middleware for Scrapy?
There are libraries that cycle through a list of proxies. With a 4G proxy that rotates by link you don't need them — just rotate between runs as above.
Can Scrapy run JavaScript?
Scrapy itself can't. For pages that need JavaScript, look for the API the page calls, or use a headless browser — see scraping JavaScript sites.