Is web scraping legal? Robots.txt, terms, personal data and copyright (Vietnam, Singapore and beyond)
In this article
"Is scraping illegal?" has no yes-or-no answer. The same code can be entirely ordinary when it collects public product prices, yet risky when it collects customers' phone numbers or keeps running after a site has blocked you. This guide sums up the factors that decide the risk, with examples from Vietnam, Singapore and a few well-known cases abroad.
Note: this is general information to help you ask the right questions, not legal advice. For a high-stakes project, or anything involving personal data, talk to a lawyer where you operate.
The short answer
The risk depends on four questions:
- What data? Public, non-personal data (prices, product descriptions, event listings) is the lowest risk. Personal data and copyrighted content carry more.
- How do you get it? Visiting like an ordinary user is very different from logging in, beating CAPTCHAs or getting round a block.
- What do you do with it? Internal analysis differs from republishing it verbatim or reselling the data.
- Where? The laws of your country, the site's country and the country of the people in the data may all apply.
Public data and data behind a login
- Data anyone can see without an account is the lowest-risk zone.
- Data behind a login comes with the terms you accepted when you signed up. Most platforms ban automated collection in their terms, and breaking that can mean a banned account or a breach-of-contract claim.
- Creating fake accounts to get data almost certainly breaks the terms, and makes everything else riskier.
In the US, hiQ v. LinkedIn is often cited: the appeals court found that scraping public profiles was unlikely to be "unauthorised access" under the computer intrusion law. But hiQ later lost on breaching LinkedIn's user agreement and had to stop collecting. The lesson: "public" doesn't mean "anything goes".
Terms of service and robots.txt
- Terms of service are a contract between you and the site (most clearly once you've registered an account). If they forbid automated collection, you're taking on contractual risk by doing it.
- robots.txt is the file a site uses to tell automated programs what they may and may not access. In many places it isn't law, but it's clear evidence of the owner's wishes — and ignoring it looks bad in any dispute.
The simple, safe approach: read both, and follow them.
Personal data: PDPA, Vietnam's rules and GDPR
This is the biggest risk area, and "it was public" is not a shield:
- Vietnam: Decree 13/2023/ND-CP on personal data protection, and the Personal Data Protection Law passed in 2025, in force from 1 January 2026. Collecting and processing personal data needs a lawful basis, and trading in personal data is prohibited.
- Singapore: the Personal Data Protection Act (PDPA) governs how organisations collect, use and disclose personal data. There are exceptions for publicly available data, but they're not a free pass.
- European Union: under the GDPR, public personal data is still personal data, and scraping it needs a lawful basis. Many data protection authorities have issued a joint statement warning about scraping personal data from social media.
The practical rule: if you don't truly need personal data, don't collect it. Filter out names, phone numbers and emails at the point of collection.
Copyright and databases
- Facts and figures (prices, dates, specifications) generally aren't protected by copyright.
- Creative content — articles, photos, videos, user reviews — is. Republishing it verbatim is a copyright issue whether you copied it by hand or by code.
- Some countries have exceptions for computational analysis. Singapore, for example, has a computational data analysis exception in its Copyright Act 2021 — provided you access the work lawfully.
- The European Union also protects databases whose makers invested substantially in building them.
Analysis, statistics and short, credited quotes carry far less risk than copying and republishing.
Circumvention is the riskiest zone
Once a site has actively stopped you — blocked your IP, put up a CAPTCHA, required a login, sent a demand to stop — carrying on by getting round those barriers changes the picture completely.
The case usually cited is Craigslist v. 3Taps in the US: after Craigslist sent a cease-and-desist letter and blocked its IPs, 3Taps kept accessing the site through proxies, and the court found that could amount to unauthorised access. In Vietnam, the Criminal Code has an offence of illegally infiltrating another party's computer network or electronic devices (Article 289), covering getting past warnings, access codes or firewalls.
Put plainly, as a proxy seller: using proxies to carry on after being blocked or asked to stop is exactly the kind of thing that loses lawsuits. A proxy doesn't make a risky activity legal.
Checklist before you scrape
- Is there an API or open data for this?
- Is the data public, or does it need a login?
- What do the terms and robots.txt say?
- Is there personal data? If so, do you really need it, and do you have a lawful basis?
- Will you republish the content, or only analyse it?
- Could your pace affect the site?
- Is the site blocking you, or has it asked you to stop? If so — stop.
Where a proxy fits in all this
A proxy is a network tool, like a browser or a VPN. Its legitimate uses in data collection:
- seeing data as a user in a given country — prices and results in Singapore,
- reaching public data on sites that block whole server IP ranges without forbidding access,
- spreading load gently for work that's already permitted.
A proxy doesn't change what you're allowed to do. See how to do it properly in proxies for web scraping and scraping anti-bot websites.
Frequently asked questions
Is scraping competitors' product prices legal?
Public prices are usually the lowest-risk zone: they're facts, not personal data, and need no login. Still read the terms, go slowly and don't republish their content. See proxies for price monitoring.
If data is public on social media, can I take it freely?
No. Public names, photos and phone numbers are still personal data under Vietnam's rules, the PDPA or the GDPR, and platform terms usually ban automated collection.
Does using a proxy protect me from liability?
No. A proxy changes your IP, not who's responsible. Using one to get round a block can even be read as evidence that you knew you weren't allowed.