Home/Blog/AI and copyright: the three questions people confuse

AI and copyright: the three questions people confuse

Whether training on protected work is lawful, whether AI output can be owned or infringes, and whether any of it is fair to creators are three separate questions with different rules and different answers. Most of the argument consists of people answering different ones at each other.

Almost every argument about AI and copyright is conducted as though there were one question with one answer. There is not. The dispute contains at least three distinct questions, governed by different legal doctrines, decided by different facts, moving in different directions, and in one case probably not a legal question at all. Whether training a model on protected work is lawful copying, whether the output of a model can be owned or can infringe, and whether the whole arrangement is fair to the people whose work was used are three separate questions, and because an answer to any one of them tells you nothing about the others, most of the public argument consists of people answering different questions at each other.

This guide separates them, sets out how the training question is handled across the major legal traditions and why they diverge so sharply, explains why the output question is governed by entirely different rules, and describes the fairness argument on its own terms rather than collapsing it into the legal ones. One necessary caveat before starting: this is an explanation of how the questions are structured, not legal advice, the law differs by country and is actively changing, and anyone with a real exposure should consult a qualified lawyer in their jurisdiction.

Question one: is training on protected work lawful?

Start with what training physically involves, because the legal analysis follows from it. Building a model means acquiring large quantities of material, copying it, processing it, and using it to adjust parameters. Copying protected work is the thing copyright regulates, so the question is not whether copying occurred but whether this particular copying is permitted.

This applies across the board, from text models to the systems that generate images, where the objection from visual artists has been loudest. Two features make it unusual. The copying is instrumental rather than expressive: nobody reads the copies, and the output of the process is a set of statistical parameters rather than a reproduction. And it is massive, involving quantities of work no licensing negotiation has historically contemplated. Legal systems have responded to that combination in two structurally different ways.

Two models of regulation

The first model relies on flexible judicial doctrine. The United States is the main example, applying its fair use test, which weighs four factors: the purpose and character of the use, including whether it is transformative; the nature of the work used; the amount and substantiality of what was taken; and the effect on the market for the original. There is no statutory provision written for AI training, so courts assess it case by case against a standard designed for other situations. Several other jurisdictions, including the United Kingdom, work from comparable but narrower doctrines of fair dealing, which typically require the use to fall within an enumerated purpose rather than being assessed openly.

The second model uses explicit statutory exceptions for text and data mining, meaning automated analysis of large bodies of material. The European Union created such an exception in its copyright directive, permitting mining broadly for scientific research while allowing rights holders to reserve their rights against commercial use, which in practice means honouring machine-readable signals that say the material is not available for this purpose. Japan took a more permissive route, allowing use of protected work for information analysis where the use is not directed at the expressive content itself, a distinction between processing a work and enjoying it. Singapore adopted a broad exception for computational data analysis.

The practical difference is large. Under the statutory-exception model, a developer knows in advance whether an activity is permitted and what conditions attach. Under the flexible-doctrine model, the answer emerges from litigation, sometimes years later, and can differ between courts. Neither model is obviously better: the first offers certainty at the cost of flexibility, and the second adapts to circumstances nobody anticipated at the cost of leaving everyone guessing.

What the arguments actually are

Within the flexible model, the substantive dispute has a recognisable shape and it is worth understanding both sides properly.

Developers argue that training is transformative, meaning the use serves a fundamentally different purpose from the original. A novel is written to be read; a model derives statistical relationships from it and produces no substitute for reading it. On this view the copying is a technical step toward something new rather than an appropriation of expression, which is the classic profile of a permitted use. There is a supporting argument about consequences, and it connects to the wider problem of where training material comes from as human data runs short, which is why synthetic data has become central: restricting training to licensed material would leave developers with whatever data is cheapest to clear, which skews toward low-quality and unrepresentative sources and may make systems worse and more biased rather than fairer.

Rights holders argue that the analysis cannot stop at transformation. Even a transformative use can fail if it harms the market for the original, and they contend the harm is real: models trained on a body of work can produce material that competes with it, displacing demand for the very work that made the system possible. They also point out that the scale is unprecedented, that no consent was sought, and that a use being technically novel does not make it costless to the people whose work it consumed.

Courts working through these arguments have shown some convergence on the transformation point, treating the training of general-purpose systems as substantially transformative, while disagreeing sharply on other elements, particularly market effect and on how the material was obtained in the first place, which is a separate matter from what was done with it once acquired. The area remains unsettled, and confident predictions in either direction should be treated with suspicion.

Why the divergence between countries matters more than it looks

Copyright is territorial: each country's law applies to acts within its borders, and there is no global answer. That has become one of the more consequential facts in the field.

A developer can train where an explicit exception makes the activity clearly lawful and deploy where a different regime applies. The two acts are assessed separately, which means the relevant question is not "is AI training legal" but "which acts occurred where, and under whose rules." A system trained under a permissive exception and offered commercially in a jurisdiction with rights-reservation requirements faces a compliance question at the point of deployment that has nothing to do with the lawfulness of the training. The interaction between independently designed regimes, rather than the content of any one of them, is now the harder problem, and it is producing pressure toward licensing arrangements simply because a licence is the one thing that works everywhere.

Question two: the output side

Now the second question, which is governed by different law and has different answers. It splits into two parts that also get confused with each other.

Can AI-generated output be owned? In several major jurisdictions, copyright requires human authorship, so material produced without meaningful human creative contribution is not protectable. The practical consequence is a spectrum rather than a rule: output generated from a brief instruction with no further involvement generally attracts little or no protection, while work in which a person makes substantial creative choices, selects and arranges, edits significantly, or incorporates generated elements into a larger original work can be protected to the extent of that human contribution. Jurisdictions differ, and some have taken different positions on computer-generated works, so this is another place where the answer depends on where you are.

Can AI output infringe? Yes, and this is entirely separate from whether training was lawful. If a generated work is substantially similar to a protected work, ordinary infringement analysis applies, exactly as it would if a person had produced it. The AI-specific wrinkle is that this can happen without anyone intending it, because models can memorise portions of their training data and reproduce them, particularly material that appeared many times. That is a real and studied phenomenon rather than a hypothetical, and it is why output filtering and similarity checking have become standard parts of serious deployments.

Why the answers do not transfer

This is the heart of the confusion, so it is worth making explicit with the two cases that break the intuitive link.

A model can be trained entirely lawfully, under a clear statutory exception or fully licensed data, and still produce an output that infringes, if what it generates is substantially similar to a protected work. Lawful training does not immunise output.

A model can be trained on material acquired in ways that turn out to be unlawful, and still produce output that is completely original and infringes nothing, because the output resembles no particular work.

The two questions have different legal tests, turn on different facts, and produce different remedies.

1 · INPUT Is training on protected work lawful copying? fair use / fair dealing, or statutory text-and-data-mining answer varies by country 2 · OUTPUT Can it be owned? Can it infringe? human authorship, and substantial similarity entirely different doctrines 3 · FAIRNESS Should it be allowed, and on what terms? consent, compensation, displacement policy, not doctrine No answer transfers. Lawful training can produce infringing output; unlawful training can produce original output; and a use can be perfectly legal while the objection to it remains reasonable.
Three questions, not one. Each is governed by different rules, decided on different facts, and moving in a different direction, which is why a ruling about training settles nothing about output and a finding of legality does not address the fairness objection.
Anyone who tells you that a ruling on training settles the output question, or the reverse, has collapsed two distinct legal analyses into one. The litigation itself reflects this, having begun with training and increasingly moved toward output, which are separate battles rather than stages of one.

Question three: is it fair?

The third question is the one people usually care about most, and it is not really a legal question at all, which is why legal answers keep failing to satisfy anyone.

When a writer or illustrator objects that their work was used without permission to build a system that now competes with them, they are making a claim about consent, compensation, and economic displacement. Copyright is a poor instrument for that claim, for several reasons that are worth stating plainly. Copyright protects particular expression, not style, so a system that learns to produce work in the manner of an artist without copying any specific piece is doing something copyright was never designed to prevent. Several jurisdictions have exceptions specifically permitting analytical use of the kind training performs. And copyright's remedies address copying rather than market displacement, which is the actual grievance.

So it is entirely possible for a use to be lawful and for the objection to it to be reasonable. Those are compatible positions, and recognising that dissolves a great deal of pointless argument. Someone saying "this is legal" and someone saying "this is not right" are frequently both correct, because they are answering different questions. Whether the arrangement should be permitted, and on what terms, is a policy question about how the benefits of a technology built from collective creative output ought to be distributed. That question can be answered through licensing markets, statutory levies, opt-out or opt-in defaults, disclosure requirements, or by leaving things as they are, and the choice is political rather than doctrinal.

The counter-position deserves equal statement: knowledge and style have always built on prior work, human creators learn by studying protected material without licensing it, and a rule requiring permission for statistical analysis of published work would be difficult to limit and might restrict research and competition in ways that concentrate capability among whoever can afford the largest licensing deals. That argument is serious, and treating either side as obviously self-serving is a failure of attention rather than a conclusion.

Where things are actually heading

Without predicting outcomes, several directions are visible in the structure of the situation.

Licensing is expanding, driven less by legal defeat than by the interaction problem: a licence is valid everywhere, while a jurisdictional exception is not, so paying is increasingly the cheapest route to operating globally. Transparency obligations are growing, with several regimes moving toward requiring disclosure of training data sources, which shifts the practical question from what is permitted to what must be documented, and which connects directly to the debate over what counts as an open model, since data disclosure is precisely the component most releases withhold. Attention is moving from input to output, with more effort spent on preventing regurgitation, filtering, and provenance than on the training question. And technical provenance infrastructure, including machine-readable reservation signals, watermarking, and content credentials, is becoming the mechanism through which policy is implemented, which matters because a right that cannot be expressed in a form a crawler respects is difficult to exercise at scale.

What this means in practice

For someone creating work, the available protections vary and are imperfect. Where rights-reservation mechanisms exist, using them is the way to express a preference in a form that regimes built around opt-out will recognise, though the effect is limited to jurisdictions that honour them and to actors that comply. Understanding the distinction above also matters practically: objecting to training and objecting to output-level copying are different complaints with different footings.

For someone building with these systems, the exposures are separable. Training exposure depends on data sources and on where activities occur. Output exposure depends on whether generated material can resemble protected work, which is addressable through filtering, similarity checking, and retrieval that grounds output in licensed material. Ownership of what you produce depends on how much human creative contribution went into it, which is worth knowing before assuming a generated asset can be protected. Many providers now offer indemnities, and the scope of those matters more than their existence.

And for anyone following the argument, the most useful habit is to ask which of the three questions a given claim is actually about. A ruling about training tells you nothing about ownership of output. An argument about fairness is not refuted by a finding of legality. A statement about one country's law is not a statement about the world's.

The short version

AI and copyright is not one question but three. The first is whether training on protected work is lawful, which turns on the fact that training involves copying but for an instrumental rather than expressive purpose. Jurisdictions handle it in two structurally different ways: flexible judicial doctrines such as fair use in the United States, where courts assess transformation, market effect and other factors case by case, and explicit statutory exceptions for text and data mining, as in the European Union, which permits research broadly while allowing rights holders to reserve commercial rights, Japan, which allows use for information analysis not directed at the expressive content, and Singapore. Because copyright is territorial, where training and deployment occur are separate legal questions, and the interaction between independently designed regimes is now the hardest part, pushing developers toward licensing because a licence works everywhere. The second question concerns output, and splits into whether AI-generated material can be owned, which in several jurisdictions requires meaningful human creative contribution, and whether output can infringe, which it can if substantially similar to a protected work, including through memorisation of training data. These answers do not transfer: lawful training can produce infringing output, and unlawful training can produce entirely original output. The third question, whether the arrangement is fair to creators, is a policy question about consent and compensation that copyright is poorly suited to answer, which is why a use can be lawful while the objection to it remains reasonable.

The idea to hold onto is that these are three separate questions with different rules, different facts, and different answers, so a ruling about training settles nothing about output, a finding of legality does not address the fairness objection, and one country's position is not the world's. Most of the argument is people talking past each other, and the first useful move in any discussion is asking which question is on the table.

Common questions

Is it legal to train AI on copyrighted material? It depends on the jurisdiction, and there is no global answer because copyright is territorial. Some countries have explicit statutory exceptions permitting text and data mining, including the European Union for research with a commercial opt-out for rights holders, Japan for information analysis not directed at expressive content, and Singapore for computational data analysis. Others, including the United States, have no purpose-built provision and assess training under general doctrines such as fair use, which weighs transformation, the nature of the work, the amount used, and market effect. In those jurisdictions the question is being resolved through litigation and remains unsettled.

What is the difference between the input and output copyright questions? The input question asks whether the copying involved in training a model on protected work is lawful, and is governed by exceptions such as fair use or text and data mining provisions. The output question asks whether what a model produces can be owned and whether it can infringe, and is governed by authorship requirements and ordinary substantial-similarity analysis. They are separate: a model trained entirely lawfully can still produce an infringing output, and a model trained on improperly obtained material can produce output that infringes nothing. A ruling on one does not settle the other.

Can AI-generated content be copyrighted? In several major jurisdictions copyright requires human authorship, so purely machine-generated material with no meaningful human creative contribution generally receives little or no protection. The practical position is a spectrum: output from a short instruction with no further involvement is weakly protected at best, while work involving substantial human creative choices, significant editing, or incorporation of generated elements into a larger original work can be protected to the extent of that human contribution. Jurisdictions differ, and some treat computer-generated works differently, so the answer depends on where protection is sought.

Can AI output infringe copyright? Yes, and independently of whether the training was lawful. If generated material is substantially similar to a protected work, ordinary infringement analysis applies exactly as it would to human-produced material. This can happen unintentionally, because models can memorise portions of their training data and reproduce them, especially content that appeared many times. This is a documented phenomenon rather than a theoretical risk, which is why output filtering, similarity checking, and grounding generation in licensed sources have become standard practice in production systems.

Why do countries have such different rules on AI training? Because copyright is territorial and legal traditions differ, and because governments have taken different positions on the trade-off between supporting AI development and protecting creative industries. Two broad approaches have emerged. Some jurisdictions rely on flexible judicial doctrines like fair use or fair dealing, which adapt to new situations but leave outcomes uncertain until litigated. Others enacted explicit statutory exceptions for text and data mining, which give developers advance certainty but fix conditions that may not suit unanticipated cases. The result is that the same activity can be clearly permitted in one country and contested in another.

If AI training is legal, does that mean it is fair? Not necessarily, and conflating the two produces most of the unproductive argument in this area. Legality and fairness are different questions. Copyright protects particular expression rather than style, several jurisdictions permit analytical uses of the kind training performs, and copyright remedies address copying rather than economic displacement, which is the grievance many creators actually have. So a use can be lawful while the objection to it remains reasonable. Whether the arrangement should be permitted, and on what terms, is a policy question about consent and compensation that can be settled through licensing, disclosure requirements, or changed defaults rather than through doctrine.

What can creators do to stop their work being used for AI training? The options are limited and vary by jurisdiction. Where regimes are built around rights reservation, as in the European Union's commercial text and data mining framework, expressing a machine-readable reservation is the recognised way to signal that material is not available for that purpose, using mechanisms that crawlers can detect. The limitations are real: the effect extends only to jurisdictions that honour reservations and to actors who comply, and it does not affect copies already made. It is also worth distinguishing complaints, since objecting to training and objecting to output that resembles your work are separate claims with different legal footing.

What should someone building with AI worry about? The exposures separate along the same lines as the questions. Training exposure depends on where data came from and where activities took place, which mainly concerns those building models rather than those using them. Output exposure applies to everyone and depends on whether generated material can resemble protected work, which is addressable through filtering, similarity checking, and grounding output in licensed sources. Ownership of what you produce depends on the degree of human creative contribution, which matters before assuming a generated asset can be protected. Provider indemnities are common, and their scope and exclusions matter more than their existence.

Sources & further reading

The primary literature behind the claims above, drawn from the concept entries this post links to, so a claim carries the same source here as it does there.

  • Longpre et al. (2023), The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI — the audit; most licence tags on popular datasets are wrong or missing. Data Provenance
  • Gebru et al. (2018), Datasheets for Datasets — the template; the good sections are the ones people skip. Data Provenance
  • Bourtoule et al. (2021), Machine Unlearning — why you can't take it back, and what the approximate methods don't guarantee. Data Provenance

Related articles

Learn the concepts

← All posts