Ntụziaka ọha

Fair Use and AI Training Data

Fair use is the US copyright doctrine that courts are using to decide whether training AI models on copyrighted works without permission is lawful.

  • 4 min read
  • Emelitere ikpeazụ
Na ibe a4 min read
  1. Nchịkọta
  2. Ime miri emi
  3. Mmetụta atụmatụ
  4. The Future of Fair Use and AI Training Data
  5. Mmejuputa n'ezie n'ụwa
  6. Ihe ize ndụ & okporo ụzọ nche
  7. Map mmejuputa
  8. Nọgide na-eme nchọpụta
  9. Ajụjụ a na-ajụkarị

Nchịkọta

Judges weigh four factors: the purpose of the use, the nature of the work, how much was copied, and the effect on the market for the original. The first 2025 rulings split. Training was found highly transformative in some generative AI cases, while pirated data and non-generative copying drew liability. The answer affects every AI developer and every creator whose work is in training data.

Ime miri emi

Section 107 of the Copyright Act lists four factors. Factor one asks whether the use is transformative and commercial. The Supreme Court's 2023 Warhol v. Goldsmith decision emphasized that a different purpose matters more than adding new meaning. Factor two considers whether the work is creative or factual. Factor three looks at how much was taken. Factor four asks whether the use harms the market for the original, and courts often treat it as the most important. The first rulings came in 2025. In February, Judge Stephanos Bibas held in Thomson Reuters v. Ross that Ross's copying of Westlaw headnotes to train a legal search tool was not fair use. Ross was building a direct competitor, and the AI was not generative. The case went to the Third Circuit on interlocutory appeal. In June, Judge William Alsup ruled in Bartz v. Anthropic that training on books was 'exceedingly transformative' and fair use, and that scanning purchased print books was also fair. He held that building a central library from pirate sites was a separate use that fair use did not cover. After class certification, Anthropic agreed to a settlement reported at $1.5 billion, about $3,000 per covered work. Days later, Judge Vince Chhabria ruled for Meta in Kadrey v. Meta, but only because the authors did not develop evidence of market harm. He stressed that 'market dilution' from floods of AI-generated competing works could weigh heavily against fair use in a better-argued case. A common misconception is that courts have declared AI training fair use across the board. These are district court rulings on specific facts. Appeals, other cases and the method of acquiring data all matter.

Mmetụta atụmatụ

Ihe ize ndụ na nchekwa

Ọdachi na mmerụ AI kwa ụbọchị dabere na onye ghọtara ihe egwu dị na onye nwere ike ime ihe.

Mkpebi doro anya

mmuta nke ọha na nke ọkachamara na-akpụzi ma amụma nchekwa siri ike ọ ga-ekwe omume na ndọrọ ndọrọ ọchịchị.

Ịcha site hype

Nkọwa doro anya na-ebelata njide site na hype, ụlọ nyocha PR na ụlọ ihe nkiri na-edoghị anya.

The Future of Fair Use and AI Training Data

Appellate decisions, including the Third Circuit's review of Thomson Reuters v. Ross, are likely to shape the doctrine more than any single trial ruling. Expect courts to keep separating how data was acquired from how it was used, and to look closely at market harm evidence. Licensing deals will probably keep growing, since they reduce legal risk regardless of outcome. Congress could legislate, and the Copyright Office has published its own analysis of generative AI training. Until then the law remains unsettled and depends on the facts of each case.

Mmejuputa n'ezie n'ụwa

In Bartz v. Anthropic, a judge found training Claude on lawfully purchased and scanned books was fair use, but downloading millions of pirated books into a library was not excused.

In Thomson Reuters v. Ross Intelligence, a court rejected fair use for a legal research startup that used Westlaw headnotes to build a competing, non-generative search tool.

In Kadrey v. Meta, authors lost at summary judgment because they did not prove market harm, even though the judge suggested such harm could exist in other cases.

A news publisher signs a paid licensing deal with an AI company, which both earns revenue and supports arguments that a training license market exists.

Ihe ize ndụ & okporo ụzọ nche

  • Ịgwọ ihe egwu dị adị dị ka sci-fi mgbe ike ogige.

  • Nchekwa ngwaahịa elu na-agbagwoju anya yana itinye n'okpuru ikike dị elu.

  • Hapụ ndị na-abụghị ndị bekee na ndị ọkachamara nwere naanị isi mmalite dị ala.

Map mmejuputa

  1. Mmebi ngwaahịa dị iche iche, iji ya eme ihe na enweghị njikwa / ihe egwu adịghị mma.

  2. Jụọ ihe akaebe ga-agbanwe echiche gị na usoro iheomume na ịdị njọ.

  3. Na-ahọrọ isi mmalite na nyocha pụtara ìhè karịa nzọrọ ahịa.

  4. Chọpụta otu ụzọ omume: ọrụ, amụma, ego, ma ọ bụ nka - ọ bụghị naanị mmata.

Nọgide na-eme nchọpụta

Free newsletter

Get the daily AI briefing

Three verified AI stories every weekday morning, written in plain English. Free forever, no ads.

One email each weekday. Unsubscribe in one click. We never sell or share your address.

Test yourself

Take the Fair Use and AI Training Data quiz

Instant feedback on every answer, and a shareable certificate with a verifiable ID once you pass a course.

Malite ajụjụ

Support free AI education. AI Understanding is a 501(c)(3) nonprofit — no ads, no paywall, ever. Make a donation

Ajụjụ a na-ajụkarị

What is Fair Use and AI Training Data?

Fair use is the US copyright doctrine that courts are using to decide whether training AI models on copyrighted works without permission is lawful. Judges weigh four factors: the purpose of the use, the nature of the work, how much was copied, and the effect on the market for the original. The first 2025 rulings split. Training was found highly transformative in some generative AI cases, while pirated data and non-generative copying drew liability. The answer affects every AI developer and every creator whose work is in training data.

Which fair use factor asks about harm to the market for the original work?

Factor four considers the effect on the potential market for or value of the original, and courts often treat it as the most important.

Why did the court reject fair use in Thomson Reuters v. Ross Intelligence?

Judge Bibas found Ross's use was not transformative enough and served as a direct market substitute for Westlaw.

In Bartz v. Anthropic, what did Judge Alsup find was NOT covered by fair use?

Alsup held training and scanning lawfully bought books were fair use, but acquiring and keeping pirated books in a library was a separate, unexcused use.

What settlement did Anthropic reach in Bartz, as reported?

After class certification, the reported settlement was about $1.5 billion, approximately $3,000 per covered work.

Why did Meta win in Kadrey v. Meta?

Judge Chhabria ruled narrowly because the plaintiffs failed to show market harm, while suggesting a better-argued case could succeed.