Expertise

Copyright & Content Law

Copyright counsel for companies that train models, generate content, or host what users upload. In the United States the training question has turned on how the corpus was acquired. In the EU it turns on the text and data mining opt-out and the AI Act's copyright duties. We work on both.

Training Data: What the Courts Have Decided So Far

The decisions handed down since 2025 have not produced a single rule, but they point in a consistent direction. Courts have been willing to treat training on lawfully obtained works as fair use, and unwilling to excuse how a corpus was acquired: the $1.5 billion settlement approved in the Anthropic class action in July 2026 concerned acquisition from pirate sources rather than training itself. Where the output competed with the source, fair use has failed. The New York Times case against OpenAI and Microsoft remains pending. We advise on the position as it stands and identify plainly where it is still unsettled.

  • Provenance records for every corpus, because acquisition is the fact the cases have turned on
  • Fair use analysis by model architecture and by output behavior, rather than as a single answer for the company
  • Licensing where the fair use position is too thin to rely on
  • Collection practices, including terms of use, robots exclusion and the contract and CFAA claims that sit alongside copyright
  • Risk allocation between you, your vendors and your customers, drafted to match the analysis rather than to sound reassuring

Sources17 U.S.C. § 107·US Copyright Office on AI

The EU: The TDM Opt-Out and Article 53

European law reaches the same question from the other direction. The text and data mining exceptions in the 2019 Copyright Directive permit training, but Article 4(3) allows rightholders to reserve their works, and a reservation has to be respected. Since August 2, 2025 the AI Act has required providers of general-purpose models to operate a copyright policy honoring those reservations and to publish a summary of training content on the Commission's template. Those duties became enforceable by the AI Office on August 2, 2026. They follow the model into the EU market regardless of where training took place.

  • A copyright policy that meets Article 53(1)(c), rather than a statement of intent
  • Detection of and compliance with machine-readable reservations under Article 4(3) of the Copyright Directive
  • The public training-content summary, on the AI Office template rather than a format of your own
  • Annex XI technical documentation, and the information downstream providers are entitled to receive

SourcesCopyright Directive (EU) 2019/790·AI Act·Training-content summary template·GPAI Code of Practice

Ownership of AI-Generated Output

In the United States this is settled at the level that matters. The DC Circuit held in Thaler v. Perlmutter that human authorship is required for copyright, and the Supreme Court declined to review the decision on March 2, 2026. Output generated without human authorship is not protected, although a work created with AI assistance may be, depending on what the human contributed. That has direct consequences for what you can promise a customer and for how your terms should be written.

  • Terms that describe accurately what the customer receives, which is often a license and a disclaimer rather than ownership
  • Records of human contribution where protection is going to be claimed
  • Warranties and indemnities scoped to what the law will support, so the indemnity is worth giving
  • Attribution and disclosure duties, including Article 50 of the AI Act for synthetic content

SourcesThaler v. Perlmutter (D.C. Cir.)·Supreme Court docket 25-449·US Copyright Office on AI·AI Act Article 50 guidance

User-Generated Content and the DMCA

If your platform hosts what users upload, the section 512 safe harbor is worth having and straightforward to lose. It requires a designated agent on the Copyright Office register, a notice and takedown process that meets the statute, and a repeat infringer policy that is applied rather than merely published. Safe harbor is usually lost on the last of those.

  • Designated agent registration, and the renewal that falls due every three years
  • Notice, counter-notice and takedown handling in the form the statute requires
  • A repeat infringer policy with evidence that it is enforced
  • Where generated output sits in the analysis, which is not always where user uploads sit

Sources17 U.S.C. § 512·Copyright Office on section 512·Designated agent directory

Content Licensing

Products that aggregate, transform and redistribute content at scale need licenses permitting each of those acts, including the ones a standard stock license does not cover. Training and synthetic derivation are rarely granted by default. We negotiate the licenses and structure the library so that ordinary operation does not need a legal review every time.

  • Stock, creator and influencer agreements that permit the uses the product actually makes
  • Training and fine-tuning rights, stated expressly rather than inferred
  • Territory, term and sublicensing terms that match how you distribute
  • Records that let you show the chain of rights for any asset on demand

Where This Currently Stands

The dates below are the fixed points in a field that is otherwise still moving. We will tell you which of them bear on your position and which do not.

  • August 2, 2025: AI Act copyright policy and training-content summary duties began to apply to providers of general-purpose models
  • March 2, 2026: the Supreme Court declined to review Thaler, leaving the human authorship requirement in place
  • August 2, 2026: the AI Act copyright duties became enforceable by the Commission's AI Office
  • December 2, 2026: marking and detection of synthetic content under Article 50 reaches systems placed on the market before August 2026

SourcesAI Act·Article 50 guidance·Thaler (D.C. Cir.)

Key Questions We Help You Answer

  • ?Can we train on data we collected ourselves, and what turns on how we obtained it?
  • ?Do the AI Act copyright duties reach us if we train outside the EU?
  • ?Who owns what our model generates, and what can we honestly promise a customer?
  • ?Are we relying on DMCA safe harbor, and would it survive a challenge?
  • ?Do our content licenses actually permit training and synthetic derivation?
  • ?Are we liable if a customer uses our tool to infringe?
Your plan

Project-based legal advice

Customscoped

Scoped engagements on training-data provenance, AI Act copyright duties, output ownership and terms, DMCA safe harbor, content licensing and IP assignments.

See pricing
Start now

Online intake takes about five minutes.

Prefer a human first? Book a call

Discovery Call

30 min • Video call

Pacific Time (PT) • San Francisco