Skip to content
Esc

Proxies for web scraping — which type to choose, how to rotate, and how to play fair

By use case 4 min read

In this article
  1. 01 First: do you need a proxy at all?
  2. 02 Choosing a proxy for the job
  3. 03 Rotating the IP sensibly
  4. 04 Rate limits and retries
  5. 05 Playing fair
  6. 06 Example: Python with a proxy and a rate limit
  7. 07 Are 65Proxy's Singapore 4G proxies a fit?
  8. 08 Frequently asked questions

Web scraping means using a program to read public data from websites: product prices, listings, search results. As your request volume grows, sites start limiting or blocking your IP — that's when people think of proxies. But a proxy isn't always needed, and the wrong type costs money and still gets blocked. This guide helps you decide whether you need a proxy, which type, how to rotate, and how to collect data responsibly.

First: do you need a proxy at all?

Often not:

  • The site has an API or open data: an official API is always more stable and legitimate than parsing HTML.
  • Small volumes: a few hundred pages a day at a gentle pace usually work fine from your own IP.
  • Other sources exist: RSS, sitemaps, ready-made exports.

A proxy is needed when you want country-specific data (prices or search results in Singapore), or when your reasonable volume still exceeds what one IP is allowed.

Choosing a proxy for the job

Need Proxy type Why
High volume, sites that don't block server IPs Datacenter Cheap, fast, many IPs
Many pages and countries, varied IPs Residential, rotating per request Very large IP pools
Singapore data, sites that check IPs closely, moderate volume 4G mobile Trusted IPs on a real mobile network

A 4G proxy's strength is trust, not quantity: at any moment one proxy is one IP. The full comparison is in 4G vs residential vs datacenter proxies.

Rotating the IP sensibly

  • Don't rotate for no reason: a new IP for every request doesn't make you "invisible"; it just fragments your sessions.
  • Rotate when limited: when a site starts returning 429s or captchas, slow down first, then rotate.
  • Keep cookies with their IP: one session (cookies) should stay on one IP; a new IP means a new session.
  • With 65Proxy's 4G proxies: rotate with each proxy's own link, at most once every 5 minutes. How to call the link and wait for the new IP from code is in rotating the IP with an API.

Rate limits and retries

Sustainable scraping rests on a reasonable pace more than on the number of IPs:

  • Pause between requests (a few seconds, with some randomness) and cap parallel connections.
  • Honour Retry-After when a site returns 429 or 503; without it, back off progressively (5, 10, 20 seconds…).
  • Cache: don't re-download pages that haven't changed; use If-Modified-Since / ETag where the site supports them.
  • Run off-peak for the site, where you can.

Playing fair

  • Read the site's terms and robots.txt; respect what they don't allow.
  • Only collect public data: don't log in to scrape behind a login, don't bypass captchas.
  • Be careful with personal data: names, phone numbers and emails are protected by personal-data laws such as Singapore's PDPA; don't collect them without a lawful basis.
  • Don't overload sites: if your activity slows someone else's site down, you're harming them.

Example: Python with a proxy and a rate limit

The code below checks robots.txt, goes through the proxy, pauses between requests and backs off when limited. Install with pip install requests and set the PROXY environment variable to http://USER:PASS@HOST:PORT:

Python
import os
import random
import time
from urllib import robotparser

import requests

AGENT = "price-monitor/1.0 (contact: [email protected])"
PROXY = os.environ["PROXY"]

session = requests.Session()
session.proxies = {"http": PROXY, "https": PROXY}
session.headers["User-Agent"] = AGENT

robots = robotparser.RobotFileParser("https://example.com/robots.txt")
robots.read()


def wait_after(response, attempt):
    try:
        return int(response.headers["Retry-After"])
    except (KeyError, ValueError):
        return 5 * 2 ** attempt


def fetch(url, tries=4):
    for attempt in range(tries):
        try:
            r = session.get(url, timeout=30)
        except requests.RequestException:
            time.sleep(5 * 2 ** attempt)  # a slow network, or the proxy is rotating
            continue
        if r.status_code in (429, 503):
            time.sleep(wait_after(r, attempt))
            continue
        r.raise_for_status()
        return r.text
    raise RuntimeError("giving up after several tries: " + url)


for url in ["https://example.com/p/1", "https://example.com/p/2"]:
    if not robots.can_fetch(AGENT, url):
        continue
    html = fetch(url)
    # ... read the data you need from html ...
    time.sleep(random.uniform(2, 5))  # one request every few seconds

Are 65Proxy's Singapore 4G proxies a fit?

They fit when you need data as Singapore users see it, from sites that check IPs closely, at moderate volume — say, tracking a few hundred products on a Singapore marketplace each day. Every plan includes 200GB, priced 1 day $4, 7 days $13, 1 month $40.

If you need millions of pages a day, datacenter or per-request residential proxies are usually more economical — we say so plainly so you pick the right tool.

Frequently asked questions

It depends on the data, how you collect it and the site's terms. Public data, collected at a reasonable pace and free of personal data, usually carries less risk; for large projects or anything involving personal data, get legal advice.

How much bandwidth is enough?

It depends on page size. Measure a typical page (your browser's Network tab, or the total downloaded in your script), multiply by pages per day — then compare with 200GB per plan.

Should I use a headless browser (Playwright)?

Only when the page needs JavaScript to show the data. A browser uses far more bandwidth and resources than plain requests. Using a proxy with Playwright is covered in rotating the IP with an API.