Skip to content
Esc

Articles

What is a 4G proxy? How mobile proxies work, pros, cons and when to use oneRotating proxies explained — rotation types and how to change a proxy IP with a linkSOCKS5 vs HTTP proxies — what's the difference and which should you use?Proxy vs VPN — what's the difference, and when should you use each?Static vs rotating proxies — how they differ, pros and cons, and when to use eachDedicated vs shared proxies — how they differ, and what "dedicated" means for a 4G proxyIPv4 vs IPv6 proxies — how they differ, why IPv6 proxies are cheap, and when they don't workHow do websites detect proxies? 8 signals anti-bot systems look atProxies for Facebook — managing several accounts more reliably with 4G IPsProxies for TikTok — when you need a Singapore IP and how to use it properlyProxies for advertising — checking and managing ads aimed at SingaporeProxies for Shopee and Lazada Singapore — prices, competitors and running your storeProxies for web scraping — which type to choose, how to rotate, and how to play fairDo you need a proxy to buy tickets on ThaiTicketMajor, Ticketmelon or SISTIC?Scraping Google and Bing search results with a proxy — and when to use the API insteadHow to scrape without getting blocked — pace, sessions and clean IPsProxies for price monitoring — tracking competitor and retail prices across regionsScraping anti-bot and anti-proxy websites (Cloudflare, Akamai, DataDome) — what works, and the limitsIs web scraping legal? Robots.txt, terms, personal data and copyright (Vietnam, Singapore and beyond)How to set up a proxy on Windows, macOS, iPhone and Android (with a username and password)How to add a proxy to GoLogin, AdsPower, Hidemyacc, GenLogin and MultiloginRotate a proxy IP automatically with an API — curl, Python and Node.js examplesHow to check a proxy — IP, country, carrier, speed and DNS and WebRTC leaksCommon proxy errors and how to fix them — 407, timeouts, ERR_PROXY_CONNECTION_FAILEDScraping JavaScript-rendered sites with a headless browser and a proxy403, 429 and 503 errors when scraping — causes and how to handle them properly (with Python code)Using a proxy with Scrapy — authenticated setup, polite throttling, error fixes and rotating between runs4G mobile vs residential vs datacenter proxies — a detailed comparison and how to chooseSingapore proxies — when you need a Singapore IP and which type to chooseClean proxies and blacklists: what a clean IP means, and why "never blocked" is a mythRotating mobile proxies (rotating 4G proxies) — what "clean rotation" means and how to choose oneHow to tell a real 4G proxy from a fake — check the carrier, IP type, latency and rotationFree vs paid proxies — the real risks of free proxies, and when paying is worth itDedicated 4G proxies, not shared — why a proxy only you use matters, and how to check it yourselfCheap 4G proxies — $4 gets you 200GB, and how to compare mobile proxy prices properly before you buyRefer a friend and earn 10% on every proxy order — how the 65Proxy referral program worksGuides to 4G mobile proxies

Is web scraping legal? Robots.txt, terms, personal data and copyright (Vietnam, Singapore and beyond)

By use case 5 min read

In this article
  1. 01 The short answer
  2. 02 Public data and data behind a login
  3. 03 Terms of service and robots.txt
  4. 04 Personal data: PDPA, Vietnam's rules and GDPR
  5. 05 Copyright and databases
  6. 06 Circumvention is the riskiest zone
  7. 07 Checklist before you scrape
  8. 08 Where a proxy fits in all this
  9. 09 Frequently asked questions

"Is scraping illegal?" has no yes-or-no answer. The same code can be entirely ordinary when it collects public product prices, yet risky when it collects customers' phone numbers or keeps running after a site has blocked you. This guide sums up the factors that decide the risk, with examples from Vietnam, Singapore and a few well-known cases abroad.

Note: this is general information to help you ask the right questions, not legal advice. For a high-stakes project, or anything involving personal data, talk to a lawyer where you operate.

The short answer

The risk depends on four questions:

  1. What data? Public, non-personal data (prices, product descriptions, event listings) is the lowest risk. Personal data and copyrighted content carry more.
  2. How do you get it? Visiting like an ordinary user is very different from logging in, beating CAPTCHAs or getting round a block.
  3. What do you do with it? Internal analysis differs from republishing it verbatim or reselling the data.
  4. Where? The laws of your country, the site's country and the country of the people in the data may all apply.

Public data and data behind a login

  • Data anyone can see without an account is the lowest-risk zone.
  • Data behind a login comes with the terms you accepted when you signed up. Most platforms ban automated collection in their terms, and breaking that can mean a banned account or a breach-of-contract claim.
  • Creating fake accounts to get data almost certainly breaks the terms, and makes everything else riskier.

In the US, hiQ v. LinkedIn is often cited: the appeals court found that scraping public profiles was unlikely to be "unauthorised access" under the computer intrusion law. But hiQ later lost on breaching LinkedIn's user agreement and had to stop collecting. The lesson: "public" doesn't mean "anything goes".

Terms of service and robots.txt

  • Terms of service are a contract between you and the site (most clearly once you've registered an account). If they forbid automated collection, you're taking on contractual risk by doing it.
  • robots.txt is the file a site uses to tell automated programs what they may and may not access. In many places it isn't law, but it's clear evidence of the owner's wishes — and ignoring it looks bad in any dispute.

The simple, safe approach: read both, and follow them.

Personal data: PDPA, Vietnam's rules and GDPR

This is the biggest risk area, and "it was public" is not a shield:

  • Vietnam: Decree 13/2023/ND-CP on personal data protection, and the Personal Data Protection Law passed in 2025, in force from 1 January 2026. Collecting and processing personal data needs a lawful basis, and trading in personal data is prohibited.
  • Singapore: the Personal Data Protection Act (PDPA) governs how organisations collect, use and disclose personal data. There are exceptions for publicly available data, but they're not a free pass.
  • European Union: under the GDPR, public personal data is still personal data, and scraping it needs a lawful basis. Many data protection authorities have issued a joint statement warning about scraping personal data from social media.

The practical rule: if you don't truly need personal data, don't collect it. Filter out names, phone numbers and emails at the point of collection.

  • Facts and figures (prices, dates, specifications) generally aren't protected by copyright.
  • Creative content — articles, photos, videos, user reviews — is. Republishing it verbatim is a copyright issue whether you copied it by hand or by code.
  • Some countries have exceptions for computational analysis. Singapore, for example, has a computational data analysis exception in its Copyright Act 2021 — provided you access the work lawfully.
  • The European Union also protects databases whose makers invested substantially in building them.

Analysis, statistics and short, credited quotes carry far less risk than copying and republishing.

Circumvention is the riskiest zone

Once a site has actively stopped you — blocked your IP, put up a CAPTCHA, required a login, sent a demand to stop — carrying on by getting round those barriers changes the picture completely.

The case usually cited is Craigslist v. 3Taps in the US: after Craigslist sent a cease-and-desist letter and blocked its IPs, 3Taps kept accessing the site through proxies, and the court found that could amount to unauthorised access. In Vietnam, the Criminal Code has an offence of illegally infiltrating another party's computer network or electronic devices (Article 289), covering getting past warnings, access codes or firewalls.

Put plainly, as a proxy seller: using proxies to carry on after being blocked or asked to stop is exactly the kind of thing that loses lawsuits. A proxy doesn't make a risky activity legal.

Checklist before you scrape

  • Is there an API or open data for this?
  • Is the data public, or does it need a login?
  • What do the terms and robots.txt say?
  • Is there personal data? If so, do you really need it, and do you have a lawful basis?
  • Will you republish the content, or only analyse it?
  • Could your pace affect the site?
  • Is the site blocking you, or has it asked you to stop? If so — stop.

Where a proxy fits in all this

A proxy is a network tool, like a browser or a VPN. Its legitimate uses in data collection:

  • seeing data as a user in a given country — prices and results in Singapore,
  • reaching public data on sites that block whole server IP ranges without forbidding access,
  • spreading load gently for work that's already permitted.

A proxy doesn't change what you're allowed to do. See how to do it properly in proxies for web scraping and scraping anti-bot websites.

Frequently asked questions

Is scraping competitors' product prices legal?

Public prices are usually the lowest-risk zone: they're facts, not personal data, and need no login. Still read the terms, go slowly and don't republish their content. See proxies for price monitoring.

If data is public on social media, can I take it freely?

No. Public names, photos and phone numbers are still personal data under Vietnam's rules, the PDPA or the GDPR, and platform terms usually ban automated collection.

Does using a proxy protect me from liability?

No. A proxy changes your IP, not who's responsible. Using one to get round a block can even be read as evidence that you knew you weren't allowed.