Web Scraping Ethics in the AI Era
Patrick Bushe
June 20, 2024 · 5 min read
Web scraping, using software to collect information from websites automatically, powers price comparison, research, archives and search engines. It also powers AI models trained on huge amounts of text and images taken from the web, often without asking. That has turned an old technical practice into a live ethical and legal debate.
What the law says, roughly
This isn't legal advice, and rules vary by country.
- Public data in the US: in hiQ Labs v. LinkedIn, the Ninth Circuit court of appeals found that scraping publicly available profiles likely didn't violate the main US anti-hacking law. LinkedIn later won on breach of its user agreement.
- Anti-hacking law narrowed: in Van Buren v. United States (2021), the US Supreme Court read the Computer Fraud and Abuse Act narrowly, focusing on whether someone accessed areas they weren't allowed into.
- Personal data in Europe: scraping personal information is still processing under GDPR. Data protection authorities fined Clearview AI, which scraped billions of facial images, including €20 million in Italy in 2022 and €30.5 million in the Netherlands in 2024.
- Copyright: copying content can raise copyright issues regardless of how it was collected.
The AI training dispute
AI companies argue that training on public web content is lawful and transformative. Authors, artists and publishers argue it copies their work without permission or payment, and lets AI compete with them. Lawsuits, including The New York Times against OpenAI and Microsoft, are testing these questions. See the attribution problem.
robots.txt: a request, not a lock
robots.txt lets site owners tell crawlers what not to collect. It became an internet standard (RFC 9309) in 2022, but following it is voluntary. Many AI companies now publish crawler names that can be blocked. See whether AI bots respect robots.txt.
A practical code of conduct for scrapers
- Respect robots.txt and the site's terms.
- Identify your bot with a clear user agent and contact details.
- Go slowly: limit request rates so you don't burden the site.
- Prefer official APIs or data downloads when they exist.
- Avoid personal data unless you have a lawful basis and a real need.
- Don't bypass sign-ins, paywalls or technical blocks.
- Credit and link to sources when you publish what you found.
- Ask when in doubt; many site owners will share data if asked.
For site owners
Decide what you're comfortable sharing. You can block specific AI crawlers in robots.txt, use your hosting provider's bot controls, and keep sensitive data behind sign-in. See blocking AI crawlers.
For readers
If you'd rather go straight to original sources than AI summaries built from them, Search Cleaner hides Google's AI Overviews and the AI Mode tab.