# September 2026 AI Training Data Copyright Ruling and Global Industry Impact

> Explore how generative AI copyright disputes, US fair use battles, and EU regulations are reshaping data rights and intellectual property.

Published: 2026-09-19T06:34:25.283Z
Updated: 2026-09-19T06:34:25.283Z
URL: /en/article/ai-copyright-training-2026

# Data Rights and Algorithmic Boundaries: Generative AI Copyright Disputes and Regulatory Turning Points

As generative artificial intelligence (AI) transitions from experimental research into mission-critical industrial infrastructure, legal and ethical debates over the data used to train large language models (LLMs) have reached a critical juncture. The era of unchecked data harvesting—where the web was treated as an open commons and text and images were scraped indiscriminately as "the new oil"—has effectively come to an end. Today, judicial precedents and statutory regulations are actively establishing institutional mechanisms to scrutinize the legitimacy of AI training datasets.

The development of fair use jurisprudence in the United States and the codification of regulatory frameworks in the European Union represent the two primary pillars of emerging global data governance. Intersecting with technical infrastructure controls and commercial data licensing agreements, a new framework is forming to balance intellectual property protection with artificial intelligence innovation.

---

## Background

The rapid advancement of generative AI models relies fundamentally on pre-training across vast repositories of human creative work scraped from the open web. However, the systematic omission of explicit creator consent and fair remuneration has inevitably led to widespread litigation. While initial controversies centered on claims of unauthorized digital reproduction, the legal focus has since deepened into structural questions: Does AI model training constitute "market substitution" that undermines the creative economy, or does it qualify as "transformative use" by adding distinct utility and meaning to original works?

In response, the United States and the European Union have pursued diverging paths. The U.S. continues to rely on flexible, ex-post judicial interpretations through case law, whereas the EU has opted for ex-ante statutory mandates enforcing transparency and copyright opt-outs. These distinct approaches are reshaping not only how Big Tech corporations approach AI research and development, but also the fundamental business models of digital media publishers, creative professionals, and rights holders worldwide.

---

## Core Issues

Global AI copyright disputes currently center on two primary axes: American judicial determinations and the European Union’s legislative framework.

### 1. Fair Use Scrutiny and the Limits of Platform Liability in U.S. Courts

In U.S. federal courts, the scope of the fair use doctrine under Section 107 of the Copyright Act has become the central battleground. In *The New York Times Co. v. Microsoft Corp. and OpenAI*, the federal district court denied the defendants' motion to dismiss core infringement claims, permitting the case to proceed to the merits. This ruling indicates that courts are unwilling to grant blanket fair use exemptions to AI developers without a granular, evidentiary evaluation of alleged market harm. Conversely, in September 2026, the U.S. Department of Justice (DOJ) filed an official statement of interest suggesting that utilizing protected works as training data may qualify as fair use, highlighting growing friction between administrative policy priorities and the intellectual property claims of creators and media organizations.

At the same time, the judiciary has maintained limits against overbroad infringement theories. The U.S. Court of Appeals for the Ninth Circuit rejected plaintiffs' expansive interpretations of Section 1202 of the Digital Millennium Copyright Act (DMCA)—which governs the removal or alteration of Copyright Management Information (CMI)—in litigation involving OpenAI, Microsoft, and GitHub. Similarly, in a class-action suit against Meta, the court dismissed claims where the plaintiffs failed to demonstrate concrete market dilution or commercial injury resulting from model training.

Nevertheless, internal communications surfaced during the discovery process have intensified public scrutiny. Unsealed documents revealed a senior Microsoft executive internally characterizing the mass scraping of web content for LLMs as "the largest labor theft in human history." This disclosure has amplified criticisms that major technology firms proceeded with large-scale ingestion while privately acknowledging the structural legal vulnerabilities of their training pipelines.

### 2. The EU AI Act and Codification of TDM Exceptions

In contrast to the case-law-driven approach of the United States, the European Union has established an explicit statutory regime by integrating the EU AI Act with the Directive on Copyright in the Digital Single Market (CDSM Directive). To enforce Articles 53 and 55 of the AI Act, the European Commission introduced the General-Purpose AI (GPAI) Code of Practice.

This regulatory framework imposes tiered transparency obligations scaled to model compute capacity:
* **Models trained with $\ge 10^{23}$ FLOPs**: Developers must publish comprehensive summaries of their training datasets and implement rigorous compliance protocols aligned with EU copyright law.
* **Models trained with $> 10^{25}$ FLOPs**: Classified as carrying "systemic risk," these frontier architectures face heightened technical audits, adversarial testing, and direct governance oversight.

A vital component of this framework is the application of the Text and Data Mining (TDM) exception. Open-source models classified as presenting systemic risks receive no exemption from these obligations and must strictly respect machine-readable opt-outs exercised by rights holders under Article 4(3) of the CDSM Directive. While entities such as OpenAI, Google, and Microsoft have expressed commitments to comply with the Code of Practice, Meta and several major Chinese AI developers have declined to sign, signaling growing geopolitical and regulatory fragmentation across the global AI ecosystem.

---

## Multifaceted Analysis

Judicial rulings and legislative mandates are triggering systemic shifts across the technology stack and the broader digital economy. This transformation extends beyond legal theory into digital business models, network infrastructure, and industrial competitiveness.

### 1. Restructuring Commercial Compensation and Data Licensing Models

Confronted with mounting legal exposure, AI developers are actively deploying commercial compensation structures to secure uninterrupted access to premium, high-integrity training data. Google has introduced an "AI Contribution Pilot" alongside pay-per-value licensing frameworks, designed to remunerate publishers based on the extent to which their content informs generative responses in AI Overviews and Gemini. Concurrently, OpenAI has tested hyperlinked, cost-per-click (CPC) attribution models within ChatGPT to compensate news organizations and digital publishers facing declining referral traffic.

This shift marks a structural evolution away from traditional search indexing—where platforms and publishers coexisted via outbound web traffic—toward a "zero-click" generative paradigm. In this environment, raw web data is increasingly commodified through formal bilateral licensing rather than informal, uncompensated harvesting.

### 2. Proliferation of Infrastructure-Level Web Governance Technologies

To automate compliance and enforce rights retention programmatically, network-level infrastructure is evolving rapidly. Cloudflare, for instance, has deployed features such as "Accountable" and "Bot Preference Sync." These protocols allow web publishers to remain indexable for standard search engines while selectively blocking automated scrapers deployed for AI model training.

This technological evolution moves the web beyond the static, non-binding conventions of `robots.txt` files toward dynamic, cryptographically verifiable protocols. Publishers can now exercise granular sovereignty over whether their digital assets are utilized strictly for discovery or repurposed as algorithmic training inputs.

### 3. Counterarguments: Overregulation and Market Entrenchment

Despite these protections, critics argue that excessive regulatory overhead and mandatory licensing fees present substantial economic risks. Highly restrictive opt-out frameworks and aggressive compensation demands may disincentivize the deployment of state-of-the-art foundation models within strictly regulated jurisdictions, accelerating regional technological divergence.

Furthermore, these compliance burdens may inadvertently cement market concentration. Because only well-capitalized tech conglomerates possess the financial reserves required to absorb astronomical licensing agreements and institutional compliance overhead, these mandates risk creating prohibitive barriers to entry. This dynamic could sideline early-stage AI startups and entrench a closed, oligopolistic market dominated by incumbent platforms.

---

## Outlook

The intersection of generative AI and intellectual property has fundamentally transitioned from an era of unmonitored data ingestion to one defined by commercial negotiation, technical governance, and legal verification. The trajectory of this ecosystem will likely be shaped by three critical factors:

First, the definitive legal resolution of fair use in the United States. As cases like *The New York Times v. OpenAI* advance to trial, the judicial balance struck between the transformative value of algorithmic synthesis and direct market substitution will establish critical precedent. Given the Ninth Circuit’s recent rulings and executive branch interest in preserving national competitiveness, U.S. courts will likely calibrate platform liabilities carefully to avoid stifling domestic innovation.

Second, regulatory divergence and jurisdictional market fragmentation. While the compute thresholds ($10^{23}$ and $10^{25}$ FLOPs) and mandatory TDM opt-outs under the EU AI Act aim to establish a global benchmark, the refusal of non-EU developers to adopt these standards foreshadows a split market. This division may yield marked regional differences in model capabilities, data diversity, and commercial availability.

Third, the institutionalization of a value-based data economy. If programmatic compensation systems—such as Google's contribution-based models, OpenAI's click-attribution pilots, and Cloudflare’s granular blocking tools—achieve mainstream market adoption, online content will no longer serve as a free, extractable commodity. Instead, it will be formalized as a priced, trackable digital asset within the AI supply chain.

Establishing an intellectual property framework for the artificial intelligence era requires balancing equitable compensation for human creators with broader technological advancement. The degree to which legal doctrines and network infrastructure can collaboratively forge this equilibrium will define the productivity, architecture, and governance of the digital economy for decades to come.

## Claims

- EU AI Act 실천강령 하에서 10^23 FLOPs를 초과하는 연산량으로 훈련된 GPAI 모델은 투명성 및 저작권 정책 요건을 준수해야 한다. (Supported)

## Forecasts

- 75% — 미국 연방법원의 AI 훈련 데이터 1차 공정이용 판결 도출 (2027년 상반기). Signal: 뉴욕타임스 대 오픈AI 약식판결 심리 결과 및 미 법무부 의견서가 법원 판단에 반영되는지 여부
- 65% — 빅테크 기업의 EU GPAI 규제 대응에 따른 파운데이션 모델 역내 출시 지연 (2026년 하반기). Signal: Meta 등 비서명 기업의 EU AI Act 제53조 투명성 요건 불응에 따른 집행위원회 조사 착수 여부

## Sources

- [Earl Spencer defends Diana book claims about King Charles - BBC News](https://www.bbc.co.uk/news/articles/cmqxvd1drd35o?at_medium=RSS&amp;at_campaign=rss) — bbc.co.uk, 2026-09-19
- [Flight chaos caused by millisecond software defect, says air traffic control body - BBC News](https://www.bbc.co.uk/news/articles/cw0kl1571lpmo?at_medium=RSS&amp;at_campaign=rss) — bbc.co.uk, 2026-09-19
- [Billionaire Manchester United owner Sir Jim Ratcliffe says he has lost confidence in UK - BBC News](https://www.bbc.co.uk/news/articles/cm0463619r1no?at_medium=RSS&amp;at_campaign=rss) — bbc.co.uk, 2026-09-19
- [Google&#x27;s Gemini AI hacked three companies in security test - BBC News](https://www.bbc.co.uk/news/articles/c607l0k72rlvo?at_medium=RSS&amp;at_campaign=rss) — bbc.co.uk, 2026-09-19
- [All smiles in Strasbourg but uncertainty clouds Canada’s EU membership plan | European Union | The Guardian](https://www.theguardian.com/world/2026/sep/17/smiles-strasbourg-uncertainty-canada-eu-membership-plan-mark-carney) — theguardian.com, 2026-09-19
- [AI타임스](https://www.aitimes.com/) — aitimes.com, 2026-09-19
- [News and Analysis on Cryptocurrencies, Blockchain and Decentralized Finance - Cryptonomist](https://en.cryptonomist.ch/) — en.cryptonomist.ch, 2026-09-19
- [IPDaily | 지식재산 전문 미디어](https://www.ipdaily.co.kr/) — ipdaily.co.kr, 2026-09-19
- [Home - AI Insider](https://theaiinsider.tech/) — theaiinsider.tech, 2026-09-19
- [Generative AI - Wikipedia](https://en.wikipedia.org/wiki/Generative_AI) — en.wikipedia.org, 2026-09-19
- [OpenAI - Wikipedia](https://en.wikipedia.org/wiki/OpenAI) — en.wikipedia.org, 2026-09-19
- [Clearview AI - Wikipedia](https://en.wikipedia.org/wiki/Clearview_AI) — en.wikipedia.org, 2026-09-19
- [Trained AI models are *arguably* fair use under the &quot;transformative&quot; prong (whic... | Hacker News](https://news.ycombinator.com/item?id=33276345) — news.ycombinator.com, 2026-09-19
- [&gt; Any time you use something with &quot;fair use&quot; in mind, it is the equivalent of sa... | Hacker News](https://news.ycombinator.com/item?id=39075282) — news.ycombinator.com, 2026-09-19
- [This analogy doesn&#x27;t work. Fair use is an affirmative defense to copyright infri... | Hacker News](https://news.ycombinator.com/item?id=38815958) — news.ycombinator.com, 2026-09-19
- [&gt; isn&#x27;t this a slam dunk case? Meta literally published a paper where they said ... | Hacker News](https://news.ycombinator.com/item?id=42974403) — news.ycombinator.com, 2026-09-19
- [한국데이터경제신문](https://www.dataeconomy.co.kr/) — dataeconomy.co.kr, 2026-09-19
- [Davis Wright Tremaine](https://www.dwt.com/) — dwt.com, 2026-09-19
- [The New York Times v. Microsoft and OpenAI - Wikipedia](https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsoft_and_OpenAI) — en.wikipedia.org, 2026-09-19
- [Fair use - Wikipedia](https://en.wikipedia.org/wiki/Fair_use) — en.wikipedia.org, 2026-09-19
- [OpenAI | Research & Deployment](https://openai.com/) — openai.com, 2026-09-19
- [General-Purpose AI Code of Practice - Wikipedia](https://en.wikipedia.org/wiki/General-Purpose_AI_Code_of_Practice) — en.wikipedia.org, 2026-09-19
- [&gt; Nvidia wouldn’t say where this training data came from, but at least one repor... | Hacker News](https://news.ycombinator.com/item?id=42644416) — news.ycombinator.com, 2026-09-19
- [&gt; Also explains why Apple isn’t launching in EU I don&#x27;t buy it. From the AI act:... | Hacker News](https://news.ycombinator.com/item?id=41003909) — news.ycombinator.com, 2026-09-19
- [&gt; the EU still doesn&#x27;t have a clear definition of open source AI One can debate ... | Hacker News](https://news.ycombinator.com/item?id=41957772) — news.ycombinator.com, 2026-09-19
- [The New York Times Amends Lawsuit Against OpenAI and Microsoft - The New York Times](https://www.nytimes.com/) — The New York Times, 2026-09-19
- [Getty Images largely loses landmark UK lawsuit over AI image generator - Reuters](https://www.reuters.com/) — Reuters, 2026-09-19
- [Exclusive | Meta Approaches Media Companies About AI Content-Licensing Deals - WSJ](https://www.wsj.com/) — WSJ, 2026-09-19
- ["초5인데 국가대표라고?" 11살 日 게임 천재, 'AI급' 실력으로 金 노린다..."메달 따면 AG 역대 최연소 대기록"](https://www.chosun.com/sports/sports_general/2026/09/19/HAYTQZRWMJQTOZLEGBSGGNZWGI/) — chosun.com, 2026-09-19
- [When Trump and Xi meet they will discuss AI. &apos;Track Two&apos; talks are already buzzing](https://www.npr.org/2026/09/18/nx-s1-5971481/trump-xi-meeting-ai-track-two-talks) — npr.org, 2026-09-19
- [Will even one of the U.N.&apos;s 17 &apos;sustainable development goals&apos; be met by 2030?](https://www.npr.org/2026/09/18/g-s1-143738/united-nations-sustainable-development-goals-hunger-climate-gender) — npr.org, 2026-09-19
- [OpenAI Ordered to Hand Over 20M ChatGPT Logs in NYT Copyright Case - Decrypt](https://decrypt.co/) — Decrypt, 2026-09-19
- [News generative AI deals revealed: Who is suing, who is signing? - Press Gazette](https://pressgazette.co.uk/) — Press Gazette, 2026-09-19
- [California District Court upholds transparency requirements for generative AI training data - Norton Rose Fulbright](https://www.nortonrosefulbright.com/) — Norton Rose Fulbright, 2026-09-19
- [How Big AI Developers are Skirting a Mandate for Training Data Transparency - Tech Policy Press](https://techpolicy.press/) — Tech Policy Press, 2026-09-19
- [Researchers have trouble finding AI training data summaries - euractiv.com](https://www.euractiv.com/) — euractiv.com, 2026-09-19
- [Historic NYT v. OpenAI copyright battle heats up - axios.com](https://www.axios.com/) — axios.com, 2026-09-19
- [AI Companies Prevail in Path-Breaking Decisions on Fair Use - crowell.com](https://www.crowell.com/) — crowell.com, 2026-09-19
- [Two California district judges rule that using books to train AI is fair use - whitecase.com](https://www.whitecase.com/) — whitecase.com, 2026-09-19
- [Britannica Sues OpenAI Over ChatGPT 'Memorizing' Content - The Tech Buzz](https://www.techbuzz.ai/) — The Tech Buzz, 2026-09-19
- [Rightsholders Take the Lead: GEMA v. OpenAI - Taylor Wessing](https://www.taylorwessing.com/) — Taylor Wessing, 2026-09-19
- [Google Strikes Tough Negotiating Stance With Publishers on AI Licensing - The Information](https://www.theinformation.com/) — The Information, 2026-09-19
- [The price of AI training data, from $5M to $250M - qz.com](https://qz.com/) — qz.com, 2026-09-19
- [Here are the companies OpenAI has made deals with to train ChatGPT - Fast Company](https://www.fastcompany.com/) — Fast Company, 2026-09-19
- ["AI학습데이터 활용 시의 저작권, 법적 불확실성 해결해야" - 대학지성 In&Out](https://www.unipress.co.kr/) — 대학지성 In&amp;Out, 2026-09-19
- [온신협·신문협 '생성형 AI 뉴스 저작권 침해' 공동 대응 - 한국기자협회](http://m.journalist.or.kr/) — 한국기자협회, 2026-09-19
- [AI로 교과서 학습해 문제집 팔면 저작권 침해…정부, AI 학습 가이드라인 발표 - 조선일보](https://www.chosun.com/) — 조선일보, 2026-09-19
- [DOJ urges judge to rule for OpenAI, Microsoft in N.Y. Times lawsuit - The Washington Post](https://www.washingtonpost.com/) — The Washington Post, 2026-09-19
- [Stability AI challenges Getty copyright claim over watermark intent standard - Daily Journal](https://www.dailyjournal.com/) — Daily Journal, 2026-09-19
- [Getty Images vs. Stability AI copyright infringement lawsuit](https://copyrightlately.com/pdfviewer/getty-images-v-stability-ai-complaint/) — copyrightlately.com, 2026-09-19
- [OpenAI strikes ‘first of its kind’ deal with publisher Axel Springer to use its news content in ChatGPT | CNN Business - CNN](https://www.cnn.com/) — CNN, 2026-09-19
- [[AI 시대의 저작권] ③ 학습 데이터 전쟁 격화…창작자 보상 시스템 시급 - 매일경제 마켓](https://stock.mk.co.kr/) — 매일경제 마켓, 2026-09-19
- [학습하고 싶은 'AI' VS 거부하는 '창작자'[AI시대와 저작권①] - 네이트](https://news.nate.com/) — 네이트, 2026-09-19
- [인공지능 학습을 위한 저작권법 개정이라니 [왜냐면] - 한겨레](https://www.hani.co.kr/) — 한겨레, 2026-09-19
- [What are the copyright challenges of AI training? - coe.int](https://www.coe.int/) — coe.int, 2026-09-19
- [European Parliament Proposes Changes to Copyright Protection in the Age of Generative AI - Global Policy Watch](https://www.globalpolicywatch.com/) — Global Policy Watch, 2026-09-19