Scraping anti-bot and anti-proxy websites (Cloudflare, Akamai, DataDome) — what works, and the limits
In this article
- 01 The layers you'll run into
- 02 Read the signals: what the site is telling you
- 03 Step 1: find a way that doesn't need scraping
- 04 Step 2: if you still scrape — shrink your footprint
- 05 What a proxy can and can't do
- 06 Lines not to cross
- 07 A sensible workflow
- 08 65Proxy's Singapore 4G proxies
- 09 Frequently asked questions
More and more sites put anti-bot services such as Cloudflare, Akamai or DataDome in front of their content. If a script that worked yesterday now only gets a "Checking your browser" page or an endless CAPTCHA loop, this guide is for you. We won't promise a trick that "gets past every protection" — there isn't one, and many of those protections exist precisely to say "no" to automation. Instead: how these layers work, how to read the signals a site is sending, sensible routes to the data, and what a proxy really does.
The layers you'll run into
A protected site usually stacks several layers, from cheap to expensive:
| Layer | What it does | What you see |
|---|---|---|
| Rate limiting | Counts requests per IP and session | 429 errors, pages slowing down |
| IP reputation | Scores the IP type and history | 403 on the very first request from a server IP |
| Browser challenge | Runs JavaScript to verify a real browser | A "checking your browser" page |
| CAPTCHA | Asks a human to confirm | A CAPTCHA box, again and again |
| Device fingerprint | Matches browser, TLS, operating system | Blocked despite a clean IP |
| Behaviour scoring | Watches speed, paths, interaction | Blocked after running for a while |
| Login wall | Shows data only to accounts | A login page |
Each signal is covered in how websites detect proxies.
Read the signals: what the site is telling you
Every response is a message. Reading it right lets you fix the right thing:
- 429 Too Many Requests — you're going too fast. Slow down and honour
Retry-After. - 403 on the first request, while a normal browser on the same IP gets in — your program doesn't look like a browser, or lacks basic headers.
- 403 for every program from a server IP — the site blocks datacenter ranges.
- A JavaScript challenge page — the site suspects automation and wants to verify a browser.
- CAPTCHA after CAPTCHA — the site has decided it doesn't want this traffic. That's the moment to stop, not to look for a way round.
A detailed error table and handling code are in 403, 429 and 503 errors when scraping.
Step 1: find a way that doesn't need scraping
Before taking on an anti-bot system, check these routes — they're usually faster, cheaper and more durable:
- An official API, even a paid one. Clean, structured data that doesn't get blocked.
- Public datasets — many agencies, marketplaces and organisations publish open data.
- Sitemaps and RSS — lists of pages and new content that the site provides precisely for machines.
- Ask for access. A short email explaining what you need and how often sometimes earns a dedicated access route or a data export.
- Buy licensed data from a data provider.
Step 2: if you still scrape — shrink your footprint
If the data is public, the terms and robots.txt don't forbid it, and none of the options above exist:
- Fetch only what you need — no images, videos or unrelated pages.
- Cache — don't re-download unchanged pages; use
If-Modified-SinceorETagwhen the site supports them. - Go slowly — a few seconds per request, few parallel connections, during the site's off-peak hours.
- Identify yourself when appropriate — a User-Agent with your tool's name and a contact lets an admin reach you instead of blocking you.
- Use a real browser when the site needs JavaScript, at human pace — see scraping JavaScript sites.
Pacing and sessions are covered in detail in how to scrape without getting blocked.
What a proxy can and can't do
A proxy can help:
- When a site blocks whole server IP ranges. Many sites block datacenter IPs but still serve visitors on mobile networks. A 4G IP gives you the access an ordinary phone user has.
- When content differs by country — prices, products and search results in Singapore are only visible from a Singapore IP.
- When the current IP is flagged because of someone else — mobile IPs are shared, so one is occasionally limited for a while; a fresh IP fixes that.
A proxy can't help with:
- CAPTCHAs and JavaScript challenges — they check your program and behaviour, not just your IP.
- Fingerprints — TLS and the browser belong to the program you run.
- Behaviour — going too fast through the cleanest IP still gets you limited.
- Login walls — a proxy doesn't change what you're allowed to access.
An "anti-proxy" site struggles to block mobile IPs without blocking real customers, but it still rate-limits per IP. A 4G proxy lets you start as an ordinary visitor; the rest depends on how you behave.
Lines not to cross
These are things we don't teach and advise you not to do:
- Using CAPTCHA-solving services to get past CAPTCHAs at scale — CAPTCHAs exist to stop exactly that.
- Using proxies to carry on after a clear block or after being asked to stop. In some countries, courts have treated this as unauthorised access — see is web scraping legal.
- Taking data from behind a login against the terms, or creating fake accounts to do it.
- Collecting personal data — names, phone numbers, emails — without a legal basis.
- Sending enough traffic to slow the site down. That's harming someone else, not collecting data.
A sensible workflow
- Look for an API, open data or sitemaps; ask for access if needed.
- Read robots.txt and the terms; collect public data only, not personal data.
- Try each page by hand through the proxy before writing code.
- Write a slow scraper with caching and handling for 429 and
Retry-After. - Watch the error rate; when challenges or CAPTCHAs repeat, stop and rethink instead of speeding up or adding IPs.
65Proxy's Singapore 4G proxies
Our proxies use Singapore mobile-carrier IPs, rotated on demand by link or API (at most once every 5 minutes), 200GB per plan, priced 1 day $4, 7 days $13, 1 month $40. They suit accessing public data on sites that block server IPs, and data that only Singapore visitors see. They're not a tool for beating CAPTCHAs or blocks — and we won't advertise them as one. See proxies for web scraping and pricing.
Frequently asked questions
Can a 4G proxy get past Cloudflare?
No proxy "gets past" Cloudflare. Mobile IPs usually meet fewer challenges than server IPs during normal browsing, but if a site has switched on challenges for automated traffic, that's the owner's decision and no IP changes it.
Should I use a CAPTCHA-solving service?
We don't recommend it. A CAPTCHA is the site owner saying clearly that they don't want automation there. Getting round it raises your legal risk and almost certainly breaks the site's terms.
The site has a challenge but the data is public — what now?
Ask the owner about an API or access, find another data source, or fetch at a very low rate with a real browser, like a user would. If the site still blocks you, respect that decision.