Start free trial of Lex HR →

EDPB opens consultation on web‑scraping for generative AI

The European Data Protection Board published Draft Guidelines 03/2026 on web‑scraping for generative AI and opened a public consultation for stakeholders.

2 September 2026

The European Data Protection Board has published Draft Guidelines 03/2026 on web‑scraping in the context of generative AI and opened a public consultation inviting stakeholders to comment on how GDPR applies to large‑scale data collection for model training.

Issued in June 2026, the draft guidance sets out the board’s reading of legal and technical limits on scraping personal data from the web for the purpose of building generative models and explains how data‑subject rights and core GDPR obligations should be respected in that context. The consultation is hosted on the EDPB’s online portal and invites submissions from industry, civil society and supervisory authorities.

At the centre of the draft is a practical framing of common training‑data practices: large‑scale automated harvesting, reuse of publicly accessible profiles and posts, and the mixing of scraped material with internal or proprietary datasets. The EDPB maps those activities to GDPR concepts such as lawfulness of processing, purpose limitation, data minimisation and the responsibilities of controllers and processors. It emphasises that simply being publicly accessible does not remove personal data from the scope of the regulation.

The guidance devotes attention to data‑subject rights where scraped material is used to train models. It addresses how rights to access, rectification and erasure interact with model training and outputs, and flags the particular challenges posed by models that ingest content at scale. The board also outlines when a data protection impact assessment will be required, and underscores that processors and model trainers must put in place appropriate technical and organisational measures to mitigate risks to individuals.

For HR teams, talent vendors and model trainers that reuse scraped candidate profiles, workplace posts or employee‑related content, the draft is directly relevant. The EDPB cautions that scraping professional networks or public forums to build training sets can create compliance exposures unless the legal basis for processing is clear, retention is limited and individuals’ rights can be operationally upheld. Employers embedding generative models into talent products are urged to reassess data flows, contractual terms with vendors and consent practices where personal data is used.

The publication sits alongside a broader wave of regulator activity in Europe focused on training data for AI systems. National data protection authorities have already issued opinions and enforcement actions related to large‑scale data collection, and the EDPB’s draft is intended to harmonise supervisory practice across the bloc by offering interpretive guidance rather than new legal rules.

What the draft does not do is set bright‑line numerical thresholds or technical recipes for anonymisation that would remove data from GDPR’s scope. It also stops short of prescribing specific engineering standards for provenance labelling or dataset auditability, leaving those technical details to sectoral practice and national authorities. The EDPB provides principles and examples, but firms will still need to translate them into operational controls and may face differing interpretations in member‑state enforcement.

Responses to the consultation will shape the final guidelines and, in the interim, organisations that handle scraped personal data should expect closer scrutiny. HR vendors and in‑house teams that rely on scraped profiles or workplace content to train or fine‑tune models should prioritise data inventories, contractual clarity with upstream data suppliers, and demonstrable compliance measures such as DPIAs, retention limits and mechanisms to respect individual rights. The final guidance is likely to influence supervisory priorities and contractual risk in talent‑tech procurement across Europe.

Sources
  1. Guidelines 03/2026 on web‑scraping in the context of generative AI — public consultation
  2. Regulation (EU) 2016/679 (General Data Protection Regulation)