Skip to content
Esc

Articles

What is a 4G proxy? How mobile proxies work, pros, cons and when to use oneRotating proxies explained — rotation types and how to change a proxy IP with a linkSOCKS5 vs HTTP proxies — what's the difference and which should you use?Proxy vs VPN — what's the difference, and when should you use each?Static vs rotating proxies — how they differ, pros and cons, and when to use eachDedicated vs shared proxies — how they differ, and what "dedicated" means for a 4G proxyIPv4 vs IPv6 proxies — how they differ, why IPv6 proxies are cheap, and when they don't workHow do websites detect proxies? 8 signals anti-bot systems look atProxies for Facebook — managing several accounts more reliably with 4G IPsProxies for TikTok — when you need a Singapore IP and how to use it properlyProxies for advertising — checking and managing ads aimed at SingaporeProxies for Shopee and Lazada Singapore — prices, competitors and running your storeProxies for web scraping — which type to choose, how to rotate, and how to play fairDo you need a proxy to buy tickets on ThaiTicketMajor, Ticketmelon or SISTIC?Scraping Google and Bing search results with a proxy — and when to use the API insteadHow to scrape without getting blocked — pace, sessions and clean IPsProxies for price monitoring — tracking competitor and retail prices across regionsScraping anti-bot and anti-proxy websites (Cloudflare, Akamai, DataDome) — what works, and the limitsIs web scraping legal? Robots.txt, terms, personal data and copyright (Vietnam, Singapore and beyond)How to set up a proxy on Windows, macOS, iPhone and Android (with a username and password)How to add a proxy to GoLogin, AdsPower, Hidemyacc, GenLogin and MultiloginRotate a proxy IP automatically with an API — curl, Python and Node.js examplesHow to check a proxy — IP, country, carrier, speed and DNS and WebRTC leaksCommon proxy errors and how to fix them — 407, timeouts, ERR_PROXY_CONNECTION_FAILEDScraping JavaScript-rendered sites with a headless browser and a proxy403, 429 and 503 errors when scraping — causes and how to handle them properly (with Python code)Using a proxy with Scrapy — authenticated setup, polite throttling, error fixes and rotating between runs4G mobile vs residential vs datacenter proxies — a detailed comparison and how to chooseSingapore proxies — when you need a Singapore IP and which type to chooseClean proxies and blacklists: what a clean IP means, and why "never blocked" is a mythRotating mobile proxies (rotating 4G proxies) — what "clean rotation" means and how to choose oneHow to tell a real 4G proxy from a fake — check the carrier, IP type, latency and rotationFree vs paid proxies — the real risks of free proxies, and when paying is worth itDedicated 4G proxies, not shared — why a proxy only you use matters, and how to check it yourselfCheap 4G proxies — $4 gets you 200GB, and how to compare mobile proxy prices properly before you buyRefer a friend and earn 10% on every proxy order — how the 65Proxy referral program worksGuides to 4G mobile proxies

Using a proxy with Scrapy — authenticated setup, polite throttling, error fixes and rotating between runs

Setup and use 4 min read

In this article
  1. 01 How Scrapy uses a proxy
  2. 02 Way 1: environment variables
  3. 03 Way 2: your own middleware
  4. 04 Polite throttling
  5. 05 Check the proxy is in use
  6. 06 Rotating between runs
  7. 07 Common errors
  8. 08 65Proxy's Singapore 4G proxies for Scrapy
  9. 09 Frequently asked questions

Scrapy is Python's most popular scraping framework, and it supports proxies out of the box — the documentation is just scattered. This guide gathers what you need: two ways to add an authenticated proxy, a "polite" throttling setup that keeps you from getting blocked, how to check the proxy is in use, the usual errors, and how to rotate a 4G proxy's IP without breaking a run.

How Scrapy uses a proxy

Scrapy ships with HttpProxyMiddleware, enabled by default. It takes the proxy from two places:

  • the environment variables http_proxy and https_proxy, read once when Scrapy starts,
  • each request's request.meta["proxy"] — which takes priority over the environment.

The username and password go straight into the proxy address (http://USER:PASS@HOST:PORT) — Scrapy splits them out and sends the proxy an authentication header.

Note: Scrapy only supports HTTP/HTTPS proxies, not SOCKS5. With 65Proxy, use the HTTP connection details. The difference between the two is in SOCKS5 vs HTTP proxies.

Way 1: environment variables

The quickest way, applying to every request. Both variables start with http://, even https_proxy, because Scrapy talks to the proxy over HTTP:

Terminal
export https_proxy="http://USER:PASS@HOST:PORT"
export http_proxy="$https_proxy"
scrapy crawl prices -O prices.json

On Windows PowerShell:

PowerShell
$env:HTTPS_PROXY = "http://USER:PASS@HOST:PORT"
$env:HTTP_PROXY  = $env:HTTPS_PROXY
scrapy crawl prices -O prices.json

Way 2: your own middleware

When you want control in code — say, using the proxy only for some domains — write a small middleware that sets meta["proxy"]. In middlewares.py:

Python
import os


class ProxyMiddleware:
    def __init__(self):
        self.proxy = os.environ["PROXY"]  # http://USER:PASS@HOST:PORT

    def process_request(self, request, spider=None):
        request.meta.setdefault("proxy", self.proxy)

Then enable it in settings.py with a number below 750 so it runs before HttpProxyMiddleware:

Python
DOWNLOADER_MIDDLEWARES = {
    "myproject.middlewares.ProxyMiddleware": 350,
}

Replace myproject with your project's name. Reading the proxy from an environment variable keeps the username and password out of your code.

Polite throttling

With a 4G proxy, every request goes out from one IP — so pace matters more than ever. A sensible starting point for settings.py:

Python
ROBOTSTXT_OBEY = True
USER_AGENT = "my-crawler/1.0 (+mailto:[email protected])"

CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_DELAY = 2              # seconds; Scrapy varies it 0.5–1.5x
DOWNLOAD_TIMEOUT = 30           # the 180-second default is too long

AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 2
AUTOTHROTTLE_MAX_DELAY = 60
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0

RETRY_TIMES = 3
RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429]

HTTPCACHE_ENABLED = True        # while developing: don't re-download pages
  • AutoThrottle slows down automatically when the site responds slowly.
  • The default retry middleware doesn't wait per Retry-After. If you see 429s often, raise DOWNLOAD_DELAY and lower CONCURRENT_REQUESTS_PER_DOMAIN rather than adding retries. What each error code means is in 403, 429 and 503 errors when scraping.
  • The HTTP cache while developing lets you fix your parsing code without hitting the target again — turn it off for real runs if you need fresh data.

Check the proxy is in use

scrapy fetch downloads a page with the current settings and prints it:

Terminal
scrapy fetch --nolog https://api.ipify.org

The result should be the proxy's IP, not your home IP. To see the network and country as well:

Terminal
scrapy fetch --nolog https://ipinfo.io/json

Rotating between runs

With a 4G proxy, rotating means calling the proxy's rotation link and waiting for the device to reconnect. In Scrapy, the cleanest approach is to rotate between runs, not between requests:

  • rotating mid-run leaves cookies and sessions out of step with the IP,
  • in-flight requests fail while the device reconnects,
  • and each 65Proxy proxy rotates at most once every 5 minutes.

A script to run before each crawl (the rotation link lives in the ROTATE_URL environment variable):

Terminal
export https_proxy="http://USER:PASS@HOST:PORT"
export http_proxy="$https_proxy"

old=$(curl -s -m 15 -x "$https_proxy" https://api.ipify.org)
curl -s -m 60 "$ROTATE_URL" > /dev/null
for i in $(seq 1 20); do
  sleep 3
  new=$(curl -s -m 10 -x "$https_proxy" https://api.ipify.org)
  if [ -n "$new" ] && [ "$new" != "$old" ]; then break; fi
done
echo "IP: $old -> $new"

scrapy crawl prices -O prices.json

Python and Node.js versions of this wait loop are in rotating proxy IPs via API.

Common errors

  • 407 Proxy Authentication Required — wrong username or password. If the password has special characters, encode them in the proxy address: @ becomes %40, : becomes %3A.
  • TunnelError: Could not open CONNECT tunnel — the proxy refused to open a tunnel for HTTPS; usually wrong credentials, the wrong port, or a SOCKS5 port by mistake.
  • Errors with socks5:// — Scrapy doesn't support SOCKS; switch to the HTTP port.
  • A burst of timeouts — usually while the proxy is rotating or signal is weak; sensible RETRY_TIMES and DOWNLOAD_TIMEOUT handle it.
  • The IP is still your home IP — the environment variable was set after Scrapy started, or in a different terminal window.

General proxy errors are covered in common proxy errors.

65Proxy's Singapore 4G proxies for Scrapy

Our proxies have an HTTP port with a username and password — exactly what Scrapy needs — plus an on-demand rotation link (at most once every 5 minutes). Singapore mobile-carrier IPs help you reach public data on sites that block server IPs, or data shown only to Singapore visitors. 200GB per plan, priced 1 day $4, 7 days $13, 1 month $40. Use them with the throttling setup above and respect each site's rules — see is web scraping legal and pricing.

Frequently asked questions

Is one 4G proxy enough for Scrapy?

For polite scraping — a few requests a second or fewer per site — it usually is. If you need much more speed, ask the target site about an API before thinking about more IPs.

Is there a rotating-proxy middleware for Scrapy?

There are libraries that cycle through a list of proxies. With a 4G proxy that rotates by link you don't need them — just rotate between runs as above.

Can Scrapy run JavaScript?

Scrapy itself can't. For pages that need JavaScript, look for the API the page calls, or use a headless browser — see scraping JavaScript sites.