Muskeology
Frontier tech, minus the hype

AI & Compute

Training data, copyright and the unresolved question

Models are trained on text and images their creators did not license, several jurisdictions disagree about whether that matters, and the courts are still working.

Row of similar lockers with various optic fiber cables in modern data server room
Row of similar lockers with various optic fiber cables in modern data server room · Photo via Pexels

Large models are trained on very large corpora assembled largely from the public web. Whether that constitutes infringement is genuinely unsettled, and the answer differs by jurisdiction.

What is actually being argued

Several distinct questions get conflated.

Is copying for training an infringement? Training requires reproducing works into a dataset and processing them. Whether that reproduction is permitted is the first question.

Is the model itself an infringing derivative work? A model's weights are not a copy of any input in any ordinary sense, and they were derived from them.

Are outputs infringing? Where a model reproduces substantial portions of a work — which has been demonstrated for memorised text and images — that is a more conventional infringement question.

Is there a separate harm from market substitution? If a model trained on a publication's archive produces summaries that reduce readership of that publication, the harm is economic even where no copying is apparent in the output.

Courts have been treating these differently, which is why headlines announcing that a case has settled the question are usually reporting a narrower ruling.

The jurisdictional split

United States. The argument centres on fair use, a four-factor test weighing the purpose of the use, the nature of the work, the amount taken, and the effect on the market.

Transformative use has historically been treated favourably. Whether training is transformative, and how much weight the market-effect factor carries, is being litigated across multiple cases with differing outcomes.

European Union. Text and data mining exceptions exist, with a mechanism allowing rights holders to reserve their rights — effectively an opt-out.

The practical difficulty is that the reservation must be expressed in a machine-readable way, and the standards for doing so are still settling.

United Kingdom. A narrower research-only exception, with proposals to broaden it having met substantial opposition from creative industries.

Japan has a relatively permissive provision for machine learning, subject to conditions.

The result is that the same training run may be lawful in one jurisdiction and not in another, which is an awkward position for a global service.

Memorisation

The technical fact underlying much of the legal argument.

Models demonstrably memorise portions of their training data, particularly text that appears many times in the corpus or that is unusual enough to be encoded distinctly.

Research has shown extraction of verbatim passages, and image models have been shown to reproduce near-copies of training images in some conditions.

Rates are low relative to total output and are not zero, and mitigations — deduplicating training data, output filtering — reduce rather than eliminate it.

Its existence undermines the cleanest version of the argument that a model contains no copies.

What has changed in practice

Regardless of how the law settles, behaviour has shifted.

Licensing deals. Several model developers have signed agreements with publishers, image libraries and forums for training access. These are now a meaningful cost line.

Robots.txt and crawler declarations. Publishers increasingly block AI crawlers specifically, and major developers have begun honouring those signals.

Whether historical training data collected before those blocks is affected is a separate question.

Provenance and dataset documentation. Growing pressure, and regulatory requirement in some jurisdictions, to disclose training data sources at least in summary.

Indemnification. Several providers now indemnify enterprise customers against copyright claims arising from outputs, which is a commercial answer to a legal uncertainty.

The adjacent questions

Personal data. Separate from copyright and equally live. Training corpora contain personal information, and data protection law grants rights — including erasure — that are difficult to honour once information is encoded in model weights.

Consent and compensation for creators, which is a policy question rather than a purely legal one, and where the loudest disagreement is.

Output ownership. Whether AI-generated material can be copyrighted at all. Several jurisdictions have held that works lacking human authorship are not protected, with the boundary depending on the degree of human contribution.

How to read developments

A ruling in one case, in one jurisdiction, on one of the four questions above, is not a settlement of the field.

Look for what was actually decided: whether training was addressed, whether outputs were, and whether the decision turned on facts specific to that dataset.

And note that the commercial arrangements are moving faster than the law, which may mean the question is answered by licensing markets before it is answered by courts.

copyrightdatalawlicensing
Tobias Nkemelu
AI & Compute, Muskeology

Tobias builds and breaks machine learning systems for a living, which makes him a difficult audience for benchmark announcements.

More from Tobias →

Also by Tobias Nkemelu

AI & Compute

What to actually worry about with AI

A great deal of the risk discussion is about scenarios, and a shorter list of problems is already causing measurable harm.

Tobias Nkemelu··3 min read

AI & Compute

Where AI systems actually fail in production

Not in the model. In the data pipeline, the distribution shift, the feedback loop and the assumption that the world stays still.

Tobias Nkemelu··3 min read

AI & Compute

Evaluating an AI product claim

A short checklist for reading announcements, which mostly consists of asking what was measured and against what.

Tobias Nkemelu··3 min read

Space

Space law and who owns anything up there

A treaty framework from the 1960s, written for two states and no companies, now governing a commercial industry it did not anticipate.

Ravi Shankaran··3 min read

Neurotech

Neural data and who gets to hold it

Recordings from the brain are unusually intimate, largely unregulated as a category, and several jurisdictions have started legislating specifically.

Lena Brandt··3 min read

Robotics

Autonomous mobile robots outside the warehouse

Hospitals, hotels, factories and pavements — where wheeled autonomy has spread, and the specific reasons each environment is harder than a warehouse.

Lena Brandt··3 min read