EU AI Case Law Watch: Copyright, Training Data, and the Limits of Text & Data Mining

·

The copyright fight over AI isn’t theoretical anymore.

It’s moving from “is training fair?” to “what is the legally relevant act of copying—and who pays when models reliably reproduce protected text?”

In the EU, that question gets answered in courtrooms, not in policy PDFs. And the most important development to watch is how judges treat memorisation: when training data appears to be embedded in model parameters and can be extracted through simple prompts.

At a high level, EU copyright disputes in generative AI tend to collapse into three operational questions:

  1. Reproduction: Was the protected work fixed or copied “in any manner and in any form”? (InfoSoc Directive framing)
  2. Communication / making available: Did the system make the work available to the public through outputs?
  3. Exceptions: If copying occurred, does a text and data mining (TDM) exception apply—and if so, to which acts?

Executives often hear this as “training might be illegal.” That’s too blunt. Courts don’t rule on vibes. They rule on:

  • what the system did
  • what the defendant controlled
  • what a user can reliably extract
  • whether the conduct fits inside a statutory exception

Landmark Signal: GEMA v OpenAI (Munich Regional Court)

In GEMA v OpenAI, Germany’s collecting society sued OpenAI over the reproduction of song lyrics. The Munich I Regional Court (42nd Civil Chamber) essentially upheld claims for injunctive relief, information, and damages related to both (a) reproduction of lyrics in the language models and (b) reproduction in the chatbot outputs. The court rejected a separate personal-rights claim about incorrect attribution of modified lyrics.
See: the unofficial translated press release PDF circulated by IFRRO: German court press release (unofficial translation).

Two themes in the court’s reasoning are worth treating as “new baseline assumptions” for EU disputes:

1) “Memorisation” can be reproduction

The court accepted that training data can become embedded in model parameters and remain retrievable. It treated that embed-and-retrieve phenomenon as relevant “fixation” (even if encoded as probability values), and therefore potentially a reproduction within the meaning of copyright law when the work can be perceived with technical means.
See discussion and framing in: MediaLaws case note and Bird & Bird’s analysis: GEMA v OpenAI ruling overview.

This matters because it shifts the legal conversation. It is no longer limited to “outputs are generated, not copied.” If a court accepts that training led to durable embodiment of protected expression in model weights, then training itself may be litigated as a reproduction event, not only outputs.

2) TDM exceptions have a purpose boundary

The defendants argued that relevant acts were covered by text and data mining limitations. The court’s view (as described in the press release and subsequent analyses) is that TDM is designed to permit reproductions that are necessary for mining and analysis—not permanent reproduction of the work itself inside the model when that reproduction interferes with rights holders’ exploitation interests.
See: IFRRO press release PDF, MediaLaws case note, and Bird & Bird analysis.

Whether other courts adopt the same “purpose boundary” logic is a major watch item. But as a risk posture, you should assume that “TDM applies to everything” is not a stable compliance theory when memorisation is in the fact pattern.

Key takeaway: Courts are increasingly willing to treat memorised, extractable training data as a legally relevant form of copying—both in-model and in outputs.

What “Memorisation” Means in Practice (Operationally)

Most governance discussions are too abstract. Here is the operational version:

  • If a user can produce near-verbatim excerpts with simple prompts, you have an “extractability” problem, not just a licensing problem.
  • If the model can output protected text reliably, a court may infer that it is not coincidence and that the work is embodied in the system’s parameters (as Munich did).
  • If you can’t demonstrate suppression controls, injunctive relief becomes realistic—even if removal is technically hard.

This is why copyright risk is now an engineering-and-legal joint issue.

How This Interacts With the AI Act (Especially GPAI)

The EU AI Act is not a copyright statute, but it explicitly pushes GPAI providers toward copyright-aware governance (especially around transparency and copyright-related rules). The Commission’s AI Act overview is a good high-level summary of how the Act treats GPAI models and transparency obligations.
See: European Commission AI Act overview and the authoritative text: Regulation (EU) 2024/1689 (EUR‑Lex).

The key strategic point: copyright compliance is converging on evidence—what you can prove about sources, opt-outs, and output controls—not on statements of intent.

Practical Implications (Three Audiences)

For model developers (providers)

Treat “copyright” as a lifecycle control surface:

  • Dataset governance: provenance, licensing status, opt-out signals, and retention rules need to be auditable.
  • Memorisation testing: build a repeatable eval suite for near-verbatim reproduction risk (lyrics, books, news, code, standards).
  • Output mitigation: layered controls (filters, refusal policies, safety tuning, retrieval suppression) and a documented incident response path when extraction is reported.
  • Litigation readiness: be able to explain—simply—how the model works, what it stores, and what controls exist. Courts are already engaging with architecture arguments (transformers, weights, extraction).

Bird & Bird notes a hard reality: once memorised content is embedded, compliance may require mitigation measures, licensing, or retraining—technical difficulty is not a complete defence.
See: GEMA v OpenAI ruling overview.

For rights holders

The playbook is getting clearer:

  • Evidence matters: show that simple prompts reproduce substantial parts near verbatim.
  • Remedies are expanding: injunction + information/disclosure demands + damages theories appear in early rulings.
  • Strategic pressure: suits can force negotiation leverage around licensing, filtering, and future training.

The IFRRO summary is useful for the remedy shape (injunctive relief + information + damages) and the court’s emphasis on both in-model and output reproductions.
See: IFRRO press release PDF.

For downstream users (deployers and enterprises)

Many organisations think “copyright is the vendor’s problem.” That assumption is fragile.

Downstream exposure comes through:

  • reputational risk (your brand publishes AI output that reproduces protected material)
  • contractual risk (indemnities that don’t cover your use patterns, or require compliance with vendor policies you can’t implement)
  • regulatory risk (transparency obligations and misleading-disclosure issues intersect with AI Act transparency duties in August 2026)

Operationally, downstream users should have:

  • usage policies that prohibit prompting for protected text
  • logging sufficient to investigate extraction claims
  • escalation routes to the vendor and legal counsel

Key takeaway: Even if your vendor carries primary infringement risk, your organisation carries publication and governance risk.

What to Watch Next (EU and Member States)

The Munich decision is first-instance and appeal paths may reshape the reasoning. But as a case-law signal, it raises questions you should monitor across jurisdictions:

  • Will other courts adopt “memorisation = reproduction” logic?
  • Will TDM exceptions be interpreted narrowly when extractability is shown?
  • Will courts differentiate between training on lawfully accessed content vs unlawfully scraped sources—or treat output extractability as the decisive factor?
  • Will EU-level guidance on GPAI transparency and copyright compliance affect how judges view “reasonable measures”?

Closing

The EU copyright-and-AI battlefield is moving fast. The core shift is simple:

Training data governance and memorisation controls are becoming litigated facts.

If you build or deploy generative AI in the EU, the safe posture is no longer “we rely on an exception.” It’s:

  • we can prove provenance and licensing posture
  • we test for extractability and near-verbatim reproduction
  • we have layered mitigation controls
  • we can respond quickly when a rights holder shows a reproducible prompt

That’s how you turn this from a headline risk into an operationally managed one.


Suggested reading (sources)

  • Munich Regional Court / GEMA v OpenAI press release (unofficial translation): https://ifrro.org/resources/documents/General/German_Court_OpenAI_Memory_Output_Infringe_Copyright_NOV25.pdf
  • Case note: https://www.medialaws.eu/gema-v-openai-decision-of-the-munich-regional-court/
  • Bird & Bird analysis: https://www.twobirds.com/en/insights/2025/landmark-ruling-of-the-munich-regional-court-(gema-v-openai)-on-copyright-and-ai-training
  • EU AI Act overview: https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai
  • EU AI Act text (EUR‑Lex): https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng