THIS WEEK IN CULTURE + THE CULTURE BUSINESS
News Briefing: AI + fair use developments

In the ongoing debate over copyright and AI, the big question in the US is whether or not using existing creative works to train an AI model is ‘fair use’ under American copyright law. If it is AI companies could use those works without getting permission from the relevant creators or rightsholders.
There were two big rulings in the US courts last week that centred on that question involving Meta, Anthropic and a group of authors, including Charles Graeber, Andrea Bartz, Richard Kadrey and comedian Sarah Silverman.
In this TW News Briefing, we explain what those cases are about and why the rulings favoured the AI companies over the creators, but may nevertheless prove useful for the creative industries in the bigger picture debate over copyright and AI.
THE BIG DEBATE: A RECAP
To train a generative AI model – which can then generate text, images, audio, video and/or music – you need a training dataset, which basically means a large quantity of existing content.
This means that when tech companies train generative AI models, they need to copy large quantities of existing and usually copyright protected content onto their servers.
Copyright law gives copyright owners control over the reproduction – or copying – of their creative works. Which means, if any third party wants to make a copy, they need to get the copyright owner’s permission. The copyright owner would usually charge for this permission, which is how copyright enables creators to make money.
However, copyright law also includes some exceptions, scenarios when third parties can make use of copyright protected works without getting the copyright owner’s permission. This includes things like critical analysis and parody.
Under US law there is a similar but more ambiguous concept called fair use – if a third party’s use of a copyright protected work is ‘fair use’, permission is not required.
Many AI companies argue that AI training is – or should be – covered by a copyright exception. Or, in the context of US law, that AI training is fair use.
The creative industries are adamant that that is not the case, that AI companies must get permission before using existing content to train their models, and if they don’t that’s copyright infringement.
If AI training is covered by an exception – or is fair use – the AI companies can make use of existing content without seeking permission or paying a fee.
If it isn’t, creators and copyright owners can stop their content from being used, or – more likely – negotiate a licensing deal with the AI companies and secure payment.
WHAT’S HAPPENING IN COURT?
There are currently numerous lawsuits working their way through the US courts where a group of creators or rightsholders has sued an AI company for copyright infringement for making use of large quantities of existing content without permission when training their respective models.
Although each lawsuit often has multiple elements to it, at the heart of every one of these legal battles is that key question: is AI training fair use?
Because there are so many lawsuits that swing on that simple question, rulings in the first cases are considered to be very important, in that they could set a precedent that impacts on all the other lawsuits – and on the evolution of the generative AI business more generally.
That said, while all these lawsuits have been going through the motions, the US Copyright Office published a report which said AI training may be fair use in certain circumstances, but is likely not fair use in other circumstances.
Which means whether or not AI training is fair use may well depend on various specifics regarding what content is being used and in what way.
Nevertheless, these early rulings could still provide some clarity on the copyright obligations of AI companies.
FAIR USE FACTORS
Although fair use is a somewhat abstract concept, there are four factors judges and juries are meant to consider when assessing whether any one use of a copyright protected work is fair. They are as follows:
1. The purpose and character of the use – including whether the original work has been ‘transformed’ into something new. The more ‘transformative’ the use the more likely it is to be fair.
2. The nature of the copyrighted work – including whether the original work is based on facts and information, or if is fictional or more creative. Copying the former is more likely to be fair use.
3. The amount and substantiality of the portion taken – the less you take the more likely it is to be fair use.
4. The effect of the use upon the potential market – does the use deprive the creator or owner of the original work income, or undermine a new or potential market for the original work? If the use causes ‘market dilution’, it is less likely to be fair.
WHAT HAPPENED LAST WEEK?
We got rulings in two cases last week both involving a group of authors who sued over text-based AI models which were trained with those writers’ works without permission.
In the first case authors including Charles Graeber and Andrea Bartz sued AI company Anthropic over its AI model Claude.
In the second authors including Richard Kadrey and comedian Sarah Silverman sued Facebook and Instagram owner Meta over its AI model Llama.
In both cases the judges made a summary judgement on whether or not the AI companies training their respective models with copyright protected books was fair use. In both cases they sided with the AI companies.
Those rulings centred on the fact that Anthropic and Meta’s use of the authors’ books was transformative, in that – when Claude and Llama both output new text – that new text is nothing like the books written by the authors.
The judge in the Anthropic case described the AI company’s use of the authors books as being “spectacularly transformative”.
Therefore – based on fair use factor one – the judges ruled that Anthropic and Meta’s AI training was fair use.
IS THERE ANY GOOD NEWS FOR CREATORS?
Despite both of last week’s rulings finding in favour of the AI companies, they were in fact mixed bag judgements that had positives and negatives for both sides.
In the Anthropic case, the judge said that AI training is fair use providing an AI company sources the works it copies from legitimate places. If pirated content is used in the training process, it is unlikely to be fair use.
Anthropic initially sourced digital copies of more than seven millions books from piracy sites, before later legitimately buying and digitising about a million physical books. The judge said “Anthropic had no entitlement to use pirated copies”. As a result, the authors’ lawsuit will proceed in relation to the pirated books.
Given, under US law, rightsholders can sue for so called ‘statuary damages’ of up to $150,000 per infringement – and Anthropic pirated millions of books – it is arguably in the AI company’s interests to now seek to negotiate some sort to settlement, despite in theory winning in court last week.
In the Meta case, the judge said that – while Meta’s use of the authors’ works was clearly very transformative – he felt that the Llama generative AI would have a significant negative impact on the potential market for the authors’ books.
And therefore the fair use defence could fail based on the fourth factor – “the effect of the use upon the potential market”.
However, he felt that lawyers working for the authors had presented weak arguments regarding market dilution.
He then set out what kind of arguments he thought would be stronger – basically that the existence of AI models like Llama will have a general negative impact on the market for books, and that in turn will have a general negative impact on the authors whose books were used to train the model.
The judge in the Meta case also rejected any arguments by AI companies that having to secure licences from creators and rightsholders to train their models would stop the development of generative AI models. After all, he said, these are billion dollar companies pursuing a trillion dollar opportunity.
“The suggestion that adverse copyright rulings would stop this technology in its tracks is ridiculous”, he wrote.
“These products are expected to generate billions, even trillions, of dollars for the companies that are developing them. If using copyrighted works to train the models is as necessary as the companies say, they will figure out a way to compensate copyright holders for it”.
WHAT NEXT?
It’s likely the authors in both these lawsuits will appeal the rulings against them. So this is only phase one of these specific legal battles. However, even if the fair use rulings stand, there are positives for creators and rightsholders.
First, it’s thought many AI companies have scraped the internet for content to train their models, meaning they have likely used pirated content. And the Anthropic ruling was clear that that is not acceptable and opens up AI companies to liability for copyright infringement and mega-damages.
Second, the ruling in the Meta case sets out arguments that could be employed in future AI lawsuits regarding market dilution that might allow other creators and rightsholders to successfully defeat the fair use defence.
Plus, remember the Copyright Office’s position that whether or not AI training is fair use will depend on various specifics about what content is being used and in what way – meaning these rulings don’t necessarily set an industry-wide precedent. The Meta ruling in particular seemed to back up this position.
Optimists in the creative community hope that if there are enough ambiguities as to whether or not fair use applies in any one kind of AI training, that provides an incentive for AI companies to secure licensing deals, just to be certain they might not end up losing a big lawsuit and paying mega-damages down the line.
Though, of course, creators and rightsholders would prefer it if there were clear cut rulings in court that AI training is never fair use.
Plus there are concerns that a big old report on AI currently being prepared by Donald Trump’s administration might come out in favour of big tech and disagree with the position of the Copyright Office, whose boss was recently sacked by the President.
So, two big interesting rulings last week, but plenty of uncertainty remains, even within the US. And, of course, the copyright debates around AI are different in other parts of the world, even if the US lawsuits seem particularly significant.
