Copyright Subsistence in Training Data: The Legal Battle Over Whether Scraping Constitutes Infringement

Introduction

Generative AI systems, from large language models to image synthesis engines, owe their capabilities to training on vast datasets assembled by scraping content from the internet and other digital repositories. This process, which involves systematically copying, processing, and storing text, images, audio, and video created by human authors, has triggered one of the most consequential copyright disputes of the present decade. The central legal question is deceptively simple: does the act of using a copyrighted work to train an AI system constitute copyright infringement? The answer varies by jurisdiction in ways that will have enormous consequences for the global AI industry, for the creative communities whose work provides the raw material for AI training, and for the legal infrastructure that governs information use in the digital economy.

India sits at an interesting intersection of these tensions. As both a major producer of creative content and a major exporter of AI software services, India has significant economic interests on both sides of the debate. Yet India’s Copyright Act 1957, drafted in an era of physical reproduction and subsequently amended to address digital networks, contains no provision specifically addressing text and data mining (TDM), and its exceptions to copyright are narrower than those available in the European Union or the United States. The result is a statutory gap that courts have begun to address through litigation but that only legislative reform can properly resolve.

Legal Framework

Copyright Subsistence and Exclusive Rights Under the Copyright Act 1957

The Copyright Act 1957 vests copyright in “original literary, dramatic, musical and artistic works” (Section 13), with originality assessed under the standard endorsed by the Supreme Court in Eastern Book Company v. D.B. Modak (2008), which aligned India with the “skill and judgment” standard derived from CCH Canadian Ltd. v. Law Society of Upper Canada rather than the more stringent requirement of creativity associated with Feist Publications v. Rural Telephone Service in the United States.

Section 14 of the Copyright Act sets out the exclusive rights of copyright owners. For literary works, these include the rights to reproduce the work, to issue copies to the public, to perform the work, and to make any translation or adaptation. Reproduction is defined broadly in Section 2(m) as including the storage of a work in any medium by electronic or other means. On a plain reading, the temporary storage of a copyrighted work in computer memory during the web scraping process, and the more permanent storage of that work in a training dataset, both constitute reproduction within the meaning of the Act.

The training of an AI model on a copyrighted dataset involves further acts that may infringe additional exclusive rights. The creation of numerical vector representations of textual or visual content for the purpose of machine learning might constitute adaptation under Section 14(a)(vi), which grants the copyright owner exclusive rights over adaptations of their work. Whether vector embeddings constitute adaptations is an unsettled question, but the argument is not implausible given the breadth of the statutory definition.

The Transient Copies Exception: Section 52(1)(b)

Section 52(1)(b) of the Copyright Act provides a fair dealing exception for “the making of copies or adaptation of a computer programme by the lawful possessor of a copy of such computer programme, from such copy in order to utilise the computer programme for the purpose for which it was supplied.” This provision, modelled on software-specific exceptions in other jurisdictions, is clearly designed for end-user software use rather than for large-scale training data processing.

More relevant is the general transient copies exception that was introduced by the Copyright (Amendment) Act 2012, which added Section 52(1)(b) in its current form and also addressed issues of internet intermediary liability. However, the 2012 amendment did not introduce a standalone transient copies exception for the kind of intermediate copying that occurs during web scraping. The European Union’s 2001 Information Society Directive (Article 5(1)) and the UK Copyright Act’s Section 28A both contain explicit exceptions for transient or incidental copies that are an integral part of a technological process and that have no independent economic significance. India has no direct analogue to these provisions, meaning that the intermediate copies made during the scraping process are not clearly protected by any existing statutory exception.

The Absence of a Text and Data Mining Exception

The EU’s Directive on Copyright in the Digital Single Market (DSM Directive, 2019) introduced two TDM exceptions in Articles 3 and 4. Article 3 creates a mandatory exception for TDM by research organisations and cultural heritage institutions for scientific research. Article 4 creates a general TDM exception that allows anyone to engage in TDM for any purpose, subject to a rights-holder opt-out mechanism. The UK retained its own TDM exception (Section 29A of the Copyright, Designs and Patents Act 1988) after leaving the EU, though the UK’s Intellectual Property Office proposed in 2022 to extend that exception to all TDM purposes and subsequently retreated from that position following creator lobbying.

India has no equivalent of either provision. The Copyright Act 1957’s research and private study exception (Section 52(1)(a)) permits fair dealing with any work for purposes of research or private study, criticism or review, or reporting of current events. However, “fair dealing” in India is assessed against a multi-factor standard that weighs the purpose of the use, the nature of the work, the amount of the work used, and the effect on the potential market for the original. Large-scale systematic scraping of copyrighted content for commercial AI development is unlikely to qualify as fair dealing under any plausible application of these factors, given that it involves copying vast quantities of protected works for a purpose that directly substitutes for or competes with markets in which copyright owners have legitimate commercial interests.

Judicial Developments

ANI Media Pvt. Ltd. v. OpenAI Inc. (Delhi High Court, 2024)

The most significant Indian litigation on training data copyright is ANI Media’s suit against OpenAI before the Delhi High Court. ANI, the Asian News International news agency, alleged that OpenAI had used its copyrighted news articles to train ChatGPT without authorisation or compensation. OpenAI filed applications challenging the territorial jurisdiction of the Delhi High Court on the ground that OpenAI Inc. is a US entity with no relevant activities in India. The Delhi HC’s orders on jurisdiction, delivered in 2024, confirmed that the Court could exercise jurisdiction where the cause of action partly arises in India, including by virtue of alleged infringement of Indian copyright in content accessible in India. The case has not yet proceeded to the merits of the infringement question, but it has established that Indian courts will accept jurisdiction over international AI companies in copyright disputes involving Indian content.

The significance of ANI v. OpenAI for Indian copyright law extends beyond the immediate parties. The case has prompted the Ministry of Information and Broadcasting to commission a review of copyright implications for news publishers in the AI context, and has accelerated discussions within the Department for Promotion of Industry and Internal Trade (DPIIT) about whether a TDM exception should be introduced into the Copyright Act.

International Cases with Persuasive Relevance

While not binding on Indian courts, several international cases are of significant persuasive value. In The New York Times Company v. Microsoft Corporation and OpenAI (S.D.N.Y., filed 2023), the NYT alleges wholesale verbatim reproduction of its articles in ChatGPT outputs, a claim that goes beyond the training data question to the generation question. The case is important because it separates two analytically distinct copyright issues: infringement in the training process and infringement in the output. Indian courts will eventually need to address both dimensions independently.

Getty Images’ suit against Stability AI in the UK High Court (filed 2023) raises the specific question of whether scraping images, including embedded metadata and watermarks, for training image-generation models constitutes infringement. The UK proceedings are being watched closely by rights holders in India’s creative industries, including Bollywood studios and photography agencies.

The US authors’ collective actions, including Kadrey v. Meta Platforms (N.D. Cal., 2023) concerning Meta’s LLaMA model training, raise questions about whether the “fair use” analysis under Section 107 of the US Copyright Act protects AI training. The four-factor fair use analysis in the United States is considerably more flexible than India’s fair dealing framework, yet even in the US the outcome of these cases is uncertain. The transformative use doctrine, which has been the primary vehicle for expanding fair use in digital contexts since Authors Guild v. Google (2d Cir. 2015), is being tested by arguments that AI training consumes the original work without transforming it in the way that a search index or a critical commentary does.

Contemporary Issues and Analysis

The core conceptual difficulty in the training data copyright debate is that AI training involves a form of “reading” that is simultaneously like and unlike the kind of reading that human learners engage in. When a human scholar reads thousands of articles, acquires knowledge from them, and then produces original scholarship influenced by what they have read, copyright law does not regard the reading as infringement. The knowledge gained from reading is not protected by copyright; only the expression in the original works is protected. AI training is structurally similar in that the trained model does not literally store the training texts but rather encodes statistical patterns derived from those texts.

The difference, and it is legally significant, is that AI training involves the actual reproduction of the original works, even if only temporarily. Copyright protects the act of reproduction, not only the retention of reproduced copies. Human reading does not reproduce the text; the human’s neurons encode meaning, not copyright-protected expression. The AI training pipeline, by contrast, creates digital copies of the original works at multiple stages: scraping and storing raw HTML, processing and tokenising text, and encoding embeddings. These intermediate reproductions may not resemble the original work in form, but they are legally copies within the meaning of the Copyright Act.

The WIPO Framework

WIPO’s Standing Committee on Copyright and Related Rights (SCCR) has addressed AI and copyright in multiple sessions since 2019. The Revised Issues Paper on Artificial Intelligence and Intellectual Property Policy (2020) identifies three approaches taken or proposed by member states: no change to existing law (relying on exceptions and contractual arrangements), introduction of specific TDM exceptions, and introduction of remuneration rights for authors whose works are used in AI training. India has participated in these discussions and has generally taken a position sympathetic to authors’ rights, consistent with its status as a major cultural exporting nation.

Comparative and International Perspective

The EU’s Article 4 TDM exception, which allows scraping for any purpose subject to a rights-holder opt-out, represents the most sophisticated legislative attempt to balance AI development interests with copyright protection. The opt-out mechanism is implemented through machine-readable reservations, typically expressed in robots.txt files or site terms of service, giving rights holders practical control over whether their content is used for AI training while ensuring that accessible, unrestricted content is available for the general AI training market.

Japan’s approach is more permissive: the 2018 amendment to Japan’s Copyright Act (Article 30-4) created a broad exception allowing use of copyrighted works for AI training without consent or compensation, on the theory that informational consumption of creative works for machine learning purposes does not substitute for the market for the works. Japan’s approach has attracted criticism from creator communities but is attractive to the AI industry.

The United States relies on its open-textured fair use doctrine to assess AI training cases individually, without legislative guidance. The outcomes of pending litigation will substantially shape US law on this point, and Indian courts and policy makers will need to monitor those outcomes carefully.

Practical and Policy Implications

For Indian AI companies such as those developing large language models for Indian languages (including models using Indic script training data), the absence of a TDM exception creates acute legal uncertainty. Scraping publicly available Indian language content from government websites, newspapers, and cultural institutions is potentially infringing if those websites assert copyright in their content. The practical reality is that large-scale enforcement is unlikely against domestic AI developers given the commercial and policy importance of the AI sector, but legal uncertainty is itself an economic cost that deters investment and complicates licensing negotiations.

For Indian content creators, including authors, journalists, musicians, and visual artists, the absence of a compensation mechanism for AI training use represents a real economic loss. The growth of AI-generated content that competes with human creative output, trained on human creative output without compensation, is a structural threat to creative labour markets that existing copyright law is not designed to address.

Suggestions and Reforms

India should introduce a TDM exception into the Copyright Act 1957 through amendment, structured along the following lines.

A mandatory TDM exception, not subject to opt-out, should be available for non-commercial research and for cultural heritage preservation, modelled on EU Article 3. This exception serves clear public interest purposes and is unlikely to harm copyright owners’ commercial interests materially.

A commercial TDM exception should be introduced, subject to rights-holder opt-out through machine-readable reservation mechanisms. The opt-out should be enforceable and technically standardised, perhaps through CGPDTM-issued technical specifications, to ensure that rights holders can meaningfully exercise the right to withhold their content from AI training without requiring individual litigation.

A remuneration mechanism should be created through collective rights management, administered by existing copyright societies such as the Indian Performing Right Society (IPRS) and the Phonographic Performance Limited (PPL) or a new AI-specific copyright society. AI developers who use the commercial TDM exception should pay a statutory licence fee, collected and distributed to rights holders through the collective rights management infrastructure. This mechanism balances the economic interests of AI developers (certainty and predictable cost) with those of creators (compensation without requiring individual enforcement).

The Copyright Act should also be amended to address AI-generated outputs, clarifying whether copyright subsists in content produced by AI systems without meaningful human creative input. The current position under Section 13 is that copyright requires human authorship, but the Act is silent on AI-generated works, creating uncertainty that affects both the AI industry and the creative economy.

Conclusion

The copyright dimensions of AI training data are among the most consequential and least resolved questions in contemporary intellectual property law. India’s Copyright Act 1957, despite its 2012 amendments, was not designed to address the scale and nature of digital reproduction involved in large-scale AI training. The ANI v. OpenAI litigation marks the beginning of a period of judicial engagement with these questions that will produce important precedents, but litigation alone cannot provide the systemic clarity that the AI industry, the creative community, and the public interest require.

A legislative TDM exception, calibrated with a rights-holder opt-out and a collective remuneration mechanism, offers the most coherent path forward. Such reform would position India as a jurisdiction that supports AI development while respecting the rights of the authors, journalists, and artists whose creative labour provides the foundational resource on which AI systems depend. The alternative, prolonged legal uncertainty resolved through unpredictable individual litigation, serves neither the AI industry nor the creative community and risks deterring both investment in AI development and investment in the creative industries that generate the content AI systems require.

About the Author

Leave a Reply

Your email address will not be published. Required fields are marked *

You may also like these

✶ Message sent! We'll get back to you shortly.