Adobe hit with proposed class-action, accused of misusing authors’ work in AI training
Adobe, like many tech companies, has aggressively integrated AI technologies in recent years, launching several AI-based services including its Firefly media-generation suite. However, a recent lawsuit suggests Adobe's use of AI has raised legal issues regarding unauthorized use of copyrighted materials.
The Lawsuit Details
A proposed class-action lawsuit filed on behalf of Oregon author Elizabeth Lyon accuses Adobe of using pirated books—including her own works—to train its SlimLM AI model. Adobe describes SlimLM as a compact language model designed for mobile document assistance, developed using the SlimPajama-627B dataset, which Adobe claims is a "deduplicated, multi-corpora, open-source dataset" released by Cerebras in mid-2023.
Lyon contends that Adobe’s SlimLM training data includes a derivative of the RedPajama dataset, specifically containing the Books3 dataset, which has been central to multiple copyright infringement claims. Her lawsuit argues that these works were copied and used in violation of copyright protections.
Background on Datasets and Legal Context
Books3, featuring around 191,000 books, is widely known as a major data source for training generative AI models but has been the subject of numerous copyright lawsuits. Similarly, the RedPajama dataset has been implicated in litigation against other tech giants.
- In September, Apple faced a lawsuit alleging unauthorized use of copyrighted books to train its Apple Intelligence model.
- Salesforce was sued in October over similar claims related to RedPajama usage.
- In a notable case, Anthropic agreed to a $1.5 billion settlement with authors accusing it of using pirated work to train its AI chatbot Claude, setting an important legal precedent.
Implications for the AI Industry
These legal challenges point to growing concerns over the use of copyrighted materials in AI training datasets. While the AI sector continues to expand, the controversy around dataset sourcing and intellectual property rights remains a critical issue for developers and companies.
The Adobe lawsuit represents the latest chapter in ongoing disputes about how AI models are trained and the responsibilities tech companies have regarding the content they use.
For more details, see the original article on TechCrunch.