Scraping JavaScript-rendered sites with a headless browser and a proxy
In this article
Some pages show nothing useful in their raw HTML — the prices, listings or reviews only appear after JavaScript runs in a browser. Plain requests can't see that content, so people reach for a headless browser like Playwright or Puppeteer. It works, but a browser is heavy: it uses far more bandwidth and is easier to fingerprint. This guide covers when you actually need one, how to wire a proxy into it, and how to keep the cost and the blocks down.
Do you actually need a browser?
Check before you reach for one, because the alternative is much cheaper:
- View the raw HTML (your browser's "View source"). If the data is already there, plain
requestsis enough. - Look at the Network tab while the page loads. Many sites fetch their data from a JSON endpoint you can call directly — far lighter than rendering the page.
- Use a browser only when the content truly needs JavaScript or user interaction to appear.
The basics of plain-request scraping are in proxies for web scraping.
The cost of running a browser
A headless browser loads the page like a real one: HTML, CSS, JavaScript, images and fonts. On a 4G proxy that eats your 200GB many times faster than a single request would. Two habits keep it down:
- Block what you don't need — images, fonts and media rarely matter for scraping.
- Reuse one browser context for a run instead of launching a fresh browser per page.
Setting a proxy in Playwright
Playwright takes the proxy when you launch the browser. Install with pip install playwright and playwright install chromium, set HOST, PORT, USER and PASS to your proxy, and block heavy resources to save bandwidth:
from playwright.sync_api import sync_playwright
PROXY = {
"server": "http://HOST:PORT", # or socks5://HOST:PORT
"username": "USER",
"password": "PASS",
}
with sync_playwright() as p:
browser = p.chromium.launch(proxy=PROXY, headless=True)
page = browser.new_page()
# Save bandwidth: drop images, fonts and media
page.route(
"**/*",
lambda route: route.abort()
if route.request.resource_type in ("image", "font", "media")
else route.continue_(),
)
page.goto("https://example.com/product/1", wait_until="domcontentloaded")
page.wait_for_selector(".price") # wait for the data, not a fixed sleep
print(page.inner_text(".price"))
browser.close()Puppeteer (Node) is the same idea: pass --proxy-server on launch and the credentials with page.authenticate.
Wait for the content, not a fixed timer
The most common mistake is sleep(5) and hope. It's slow when the page is quick and flaky when it's slow. Instead:
wait_for_selectorfor the element that holds your data, or- wait for network to go idle when the page keeps loading pieces.
This is faster and far more reliable, and it matters doubly over a 4G proxy where latency varies.
Browsers still get blocked
A real browser engine helps, but automation leaves traces, and a browser carries a full fingerprint to match:
- Keep one profile per IP and don't run many at once through a single address.
- Rotate sensibly, not every page — calling the rotation link from a script and waiting for the new IP is covered in rotating the IP with an API.
- For managing several identities, a purpose-built antidetect browser is steadier than raw automation — see adding a proxy to antidetect browsers.
- The behaviour rules in how to scrape without getting blocked apply to browsers too.
Singapore 4G proxies from 65Proxy
Our Singapore mobile IPs work with Playwright, Puppeteer and any antidetect browser, over HTTP or SOCKS5, with rotation by link (the 5-minute rule applies) and 200GB per plan, priced 1 day $4, 7 days $13, 1 month $40. Because a browser uses bandwidth quickly, block images and fonts and prefer a JSON endpoint where one exists. See pricing.
Frequently asked questions
Playwright or plain requests — which should I use?
Use plain requests whenever the data is in the HTML or a JSON endpoint; it's faster and far lighter. Reach for Playwright only when the content needs JavaScript to appear.
Does a headless browser really use that much more bandwidth?
Yes — it downloads the whole page, scripts and assets included. Blocking images, fonts and media, and reusing one context, cuts it sharply. Measure a run against 200GB per plan.
Can I rotate the IP on every page?
You can, but you usually shouldn't — rotating mid-session breaks logins and looks suspicious, and each 65Proxy proxy rotates at most once every 5 minutes. Keep a session on one IP and rotate only when limited.