The article of the EU Artificial Intelligence Act (Regulation (EU) 2024/1689) that imposes obligations on providers of general-purpose AI models: maintaining technical documentation, supplying information to downstream integrators, putting in place a policy to comply with EU copyright law including text-and-data-mining opt-outs, and publishing a sufficiently detailed public summary of the content used for training. It is the principal legal hook connecting web scraping and training-data practices to EU regulatory enforcement.

Semantic Classification

Content

Definition

Article 53 of the EU AI Act sets out the baseline obligations that apply to every provider of a general-purpose AI (GPAI) model placed on the EU market, regardless of whether the model is classified as posing systemic risk. It sits in Chapter V of Regulation (EU) 2024/1689 and became applicable on 2 August 2025, one year after the Act entered into force.

The article imposes four core duties. Providers must (a) draw up and keep up to date technical documentation of the model, including its training and testing process and evaluation results, for supply to the AI Office and national authorities on request; (b) make information and documentation available to downstream providers who integrate the model into their own AI systems, so those providers can meet their own obligations; (c) put in place a policy to comply with Union copyright law — in particular to identify and respect rights reservations (text-and-data-mining opt-outs) expressed under Article 4(3) of the Copyright in the Digital Single Market Directive; and (d) publish a sufficiently detailed summary of the content used to train the model, following a template issued by the AI Office.

Article 53 is distinct from the Act’s transparency provisions for AI systems (such as Article 50’s disclosure duties for chatbots and synthetic media): its subject is the upstream model provider and the provenance of Training Data. Because duties (c) and (d) reach directly into how corpora are assembled, the article effectively regulates the conduct of AI Scrapers — a scraper-fed training pipeline must now honour machine-readable opt-outs and be documentable in a public summary.

Current Landscape

Operationalisation arrived through the GPAI Code of Practice (July 2025), whose Transparency and Copyright chapters give signatories a presumption-of-conformity route to Article 53 compliance, and the AI Office’s training-content-summary template published the same month. Open-source model providers enjoy a partial exemption from duties (a) and (b) — but not from the copyright policy or the training-content summary — unless their model is designated as carrying systemic risk.

The AI Office’s formal enforcement powers, including the ability to impose fines of up to 3% of global annual turnover or 15 million euros for GPAI infringements, activated on 2 August 2026 — the obligations themselves had applied since 2 August 2025, so the intervening year was compliance without penalty exposure. Notably, the 2026 “Digital Omnibus on AI”, which deferred the Act’s standalone high-risk (Annex III) obligations to 2 December 2027 and embedded (Annex I product) high-risk obligations to 2 August 2028, explicitly left the GPAI obligations of Articles 51-56 — Article 53 among them — and their enforcement dates untouched. The article’s extraterritorial pull is significant: any provider whose model reaches the EU market must apply an EU-grade copyright-reservation policy to training conducted anywhere in the world, making Article 53 a de facto global standard for training-data governance and a template debated in other jurisdictions.

Sources: