Proxies for web scraping — which type to choose, how to rotate, and how to play fair
In this article
Web scraping means using a program to read public data from websites: product prices, listings, search results. As your request volume grows, sites start limiting or blocking your IP — that's when people think of proxies. But a proxy isn't always needed, and the wrong type costs money and still gets blocked. This guide helps you decide whether you need a proxy, which type, how to rotate, and how to collect data responsibly.
First: do you need a proxy at all?
Often not:
- The site has an API or open data: an official API is always more stable and legitimate than parsing HTML.
- Small volumes: a few hundred pages a day at a gentle pace usually work fine from your own IP.
- Other sources exist: RSS, sitemaps, ready-made exports.
A proxy is needed when you want country-specific data (prices or search results in Singapore), or when your reasonable volume still exceeds what one IP is allowed.
Choosing a proxy for the job
| Need | Proxy type | Why |
|---|---|---|
| High volume, sites that don't block server IPs | Datacenter | Cheap, fast, many IPs |
| Many pages and countries, varied IPs | Residential, rotating per request | Very large IP pools |
| Singapore data, sites that check IPs closely, moderate volume | 4G mobile | Trusted IPs on a real mobile network |
A 4G proxy's strength is trust, not quantity: at any moment one proxy is one IP. The full comparison is in 4G vs residential vs datacenter proxies.
Rotating the IP sensibly
- Don't rotate for no reason: a new IP for every request doesn't make you "invisible"; it just fragments your sessions.
- Rotate when limited: when a site starts returning 429s or captchas, slow down first, then rotate.
- Keep cookies with their IP: one session (cookies) should stay on one IP; a new IP means a new session.
- With 65Proxy's 4G proxies: rotate with each proxy's own link, at most once every 5 minutes. How to call the link and wait for the new IP from code is in rotating the IP with an API.
Rate limits and retries
Sustainable scraping rests on a reasonable pace more than on the number of IPs:
- Pause between requests (a few seconds, with some randomness) and cap parallel connections.
- Honour
Retry-Afterwhen a site returns 429 or 503; without it, back off progressively (5, 10, 20 seconds…). - Cache: don't re-download pages that haven't changed; use
If-Modified-Since/ETagwhere the site supports them. - Run off-peak for the site, where you can.
Playing fair
- Read the site's terms and
robots.txt; respect what they don't allow. - Only collect public data: don't log in to scrape behind a login, don't bypass captchas.
- Be careful with personal data: names, phone numbers and emails are protected by personal-data laws such as Singapore's PDPA; don't collect them without a lawful basis.
- Don't overload sites: if your activity slows someone else's site down, you're harming them.
Example: Python with a proxy and a rate limit
The code below checks robots.txt, goes through the proxy, pauses between requests and backs off when limited. Install with pip install requests and set the PROXY environment variable to http://USER:PASS@HOST:PORT:
import os
import random
import time
from urllib import robotparser
import requests
AGENT = "price-monitor/1.0 (contact: [email protected])"
PROXY = os.environ["PROXY"]
session = requests.Session()
session.proxies = {"http": PROXY, "https": PROXY}
session.headers["User-Agent"] = AGENT
robots = robotparser.RobotFileParser("https://example.com/robots.txt")
robots.read()
def wait_after(response, attempt):
try:
return int(response.headers["Retry-After"])
except (KeyError, ValueError):
return 5 * 2 ** attempt
def fetch(url, tries=4):
for attempt in range(tries):
try:
r = session.get(url, timeout=30)
except requests.RequestException:
time.sleep(5 * 2 ** attempt) # a slow network, or the proxy is rotating
continue
if r.status_code in (429, 503):
time.sleep(wait_after(r, attempt))
continue
r.raise_for_status()
return r.text
raise RuntimeError("giving up after several tries: " + url)
for url in ["https://example.com/p/1", "https://example.com/p/2"]:
if not robots.can_fetch(AGENT, url):
continue
html = fetch(url)
# ... read the data you need from html ...
time.sleep(random.uniform(2, 5)) # one request every few secondsAre 65Proxy's Singapore 4G proxies a fit?
They fit when you need data as Singapore users see it, from sites that check IPs closely, at moderate volume — say, tracking a few hundred products on a Singapore marketplace each day. Every plan includes 200GB, priced 1 day $4, 7 days $13, 1 month $40.
If you need millions of pages a day, datacenter or per-request residential proxies are usually more economical — we say so plainly so you pick the right tool.
Frequently asked questions
Is web scraping legal?
It depends on the data, how you collect it and the site's terms. Public data, collected at a reasonable pace and free of personal data, usually carries less risk; for large projects or anything involving personal data, get legal advice.
How much bandwidth is enough?
It depends on page size. Measure a typical page (your browser's Network tab, or the total downloaded in your script), multiply by pages per day — then compare with 200GB per plan.
Should I use a headless browser (Playwright)?
Only when the page needs JavaScript to show the data. A browser uses far more bandwidth and resources than plain requests. Using a proxy with Playwright is covered in rotating the IP with an API.