This browser is not actively supported anymore. For the best passle experience, we strongly recommend you upgrade your browser.
The Lens

Digital developments in focus

| 3 minute read

Caught in the web: EDPB guidelines on scraping for generative AI

As regulators and legislators across Europe continue to grapple with the intersection of AI development and data protection, the EDPB has published new draft guidelines on one of the most contentious data collection practices underpinning generative AI (genAI) – web scraping. The guidelines confirm that web scraping falls squarely within the GDPR's scope whenever personal data is collected - and given the nature of the internet, it almost invariably will be. Aligning with the EDPB’s recent push to support innovation (discussed here), the guidelines set out practical steps developers can take to help ensure scraping is lawful, but the bar remains high. 

Legal basis

The guidelines confirm legitimate interests (LI) as the most likely legal basis for web scraping. Consent is effectively ruled out, as scrapers of third-party data have no direct relationship with individuals and cannot obtain consent at scale. 

Consistent with its previous Opinion 2024/12 (discussed here), the EDPB recognises that controllers may rely on LI for web scraping where the usual three-part test can be satisfied. Detailed guidance is provided on the application of the test to web scraping, including: 

  • on the articulation of the developer’s LI in connection with the development/improvement of general-purpose AI models where the end-use is yet to be decided (a point not addressed by the EDPB’s previous Opinion). The EDPB recommends that reference should be made to the objective pursued by the model development, e.g. whether it is commercial or scientific research, for the organisation itself or a third party.
  • on the necessity limb, it is suggested that narrowing the collection criteria to exclude unnecessary personal data and avoid scraping “a wide part of the internet” may be crucial. Alternatively, use of synthetic or pseudonymised data should be considered as less intrusive alternatives. More generally, the guidance recommends that certain types of site and content are excluded from scraping, e.g. particularly sensitive data types, sites relating to vulnerable populations, or that appear available only to logged-in users - this is likely to be challenging to implement in practice.
  • on the balancing test, it is recognised that large scale data collection poses risks to individuals, including a “chilling effect” on freedom of expression. However, the EDPB provides new guidance to help organisations navigate these risks and data subjects’ varying expectations. For example, it notes that developments around genAI in recent years may mean people are now aware the data they publish online may be used by third parties, although expectations will vary by website type and will be impacted by whether sites include anti-scraping features (e.g. CAPTCHAs or robots.txt). The guidance places significant emphasis on these features, in this context and also as part of data minimisation. As such, organisations hosting personal data may want to revisit their own websites' use of them.

Transparency

The Article 14(5)(b) privacy notice exemption can apply in the context of web scraping where providing information proves impossible or would require disproportionate effort, but this should not be treated as a blanket exemption but evaluated for each data set. Data age and the volume of data subjects impacted are key considerations. Where the exemption is relied on, controllers must publish detailed privacy information on their websites, including whether scraping sources are public or private and (if public) details of the relevant the crawler. A list of sources should also be provided to the greatest extent possible. 

Special category data

Incidental collection of special category data (SCD) is practically unavoidable when scraping at scale. Although the EDPB emphasises that processing SCD remains prohibited under the GDPR, the guidelines add a new nuance. The EDPB makes a distinction between intentional processing of SCD (where an Article 9(2) derogation is required) and non-intentional “residual” processing. Relying on the CJEU’s ruling in GC & Others (C-136/17), the EDPB says that such residual processing is not necessarily unlawful absent an Article 9 ground, where the reasoning in the case applies and the controller takes appropriate measures throughout the processing lifecycle. The EU’s Digital Omnibus proposals also potentially address residual processing of SCD. If enacted they would create an explicit EU GDPR legal basis for such processing, subject to requiring robust safeguards to be in place. The measures in the guidelines may serve as a reference point for what those should look like in practice.

Conclusion

While largely consistent with existing EDPB policy positions, the guidelines address areas of challenge and uncertainty that have surfaced since the EDPB’s previous AI opinion. With new AI guidance also being published by the Dutch DPA (on generative AI) and French DPA (on agentic systems) this month, the direction of travel is clear: GDPR should neither be seen as a blocker to AI development nor as something to be ignored. AI developers need to keep grasping the nettle.

The consultation on the guidelines closes on 30 October 2026.

Sign up to receive the latest insights. Click here to subscribe to The Lens Blog.

Tags

dp, ai