AI Training Data Copyright in India

AI Data

Introduction : The development of generative artificial intelligence depends on large quantities of text, images, recordings, software, databases and other forms of information. AI systems process this material during training to identify patterns and generate responses. The legal difficulty is that training may require copying, storing, filtering and transforming works that are protected by copyright.

In India, this issue has become particularly significant because the Copyright Act 1957 contains a fair-dealing framework but does not expressly create a comprehensive text-and-data-mining exception. The law protects authors and copyright owners through exclusive reproduction rights, while Section 52 recognises limited situations in which use of a work will not constitute infringement. The question is whether AI training can fit within those existing exceptions.

The Delhi High Court’s decision in ANI Media Pvt Ltd v OpenAI Opco LLC has made the issue more important. In an interim ruling delivered on 24 July 2026, the Court declined to grant ANI an injunction and held, on a prima facie basis, that OpenAI’s storage and use of ANI’s literary works to train the models underlying ChatGPT could fall within the fair-dealing exception for private or personal use, including research. The decision does not finally settle Indian copyright law, but it provides the most important judicial indication yet of how courts may analyse AI training.

The emerging Indian position is therefore neither a complete licence for AI companies nor an automatic requirement to obtain permission for every training use. It is a fact-sensitive framework in which purpose, method, output, market harm, technological necessity and the nature of the works will all matter.

Copyright related Rights Engaged by AI Training

Section 14 of the Copyright Act gives copyright owners exclusive rights over their works. These include reproduction, issuing copies, communication to the public, adaptation and other forms of exploitation depending on the category of work. Section 51 treats unauthorised exercise of rights protected by Section 14 as infringement, subject to the exceptions in the Act.

AI training may engage the reproduction right in several ways. A system may download or ingest works, create temporary or permanent copies, convert files into machine-readable formats, store them in a dataset, and create intermediate representations such as tokens, vectors or embeddings. Each stage may raise a different legal question. A court may distinguish between copying the expressive content of a work and making a technical copy that allows a system to identify statistical relationships without reproducing the work for human consumption.

Training datasets may also themselves attract protection. Indian copyright law recognises compilations and computer databases as literary works where originality requirements are satisfied. The copyright may not extend to every individual fact within the dataset, but it may protect the selection or arrangement produced through intellectual effort. An AI company may therefore infringe not only by copying individual works, but also by reproducing a protected database or a substantial part of its arrangement. 

This distinction is important for licensing. A company that obtains rights to use individual articles may still need to consider whether it has lawfully acquired the database, archive or platform through which those articles were organised. Conversely, the existence of copyright in a compilation does not mean that every underlying fact is protected.

Section 52 and Fair Dealing

Section 52(1)(a) provides that fair dealing with a work, other than a computer programme, for private or personal use, including research, criticism or review, or reporting current events and current affairs, does not constitute infringement. The provision is purpose-specific. It does not create a broad exception for every socially beneficial or technologically innovative use. 

The phrase “including research” is particularly relevant to AI training. Training a model may be described as a research and development activity, especially when the model is being tested, improved or evaluated. But commercial purpose remains relevant. An AI company may be conducting research while also building a product from which it earns revenue. The question is whether commerciality automatically defeats the exception or whether it is one factor in determining fairness.

The recent ANI Media ruling indicates that commercial operation does not necessarily prevent an AI developer from relying on Section 52. The Delhi High Court treated the training process as potentially falling within private or personal use, including research, and did not consider the fact that OpenAI is a commercial organisation sufficient by itself to exclude the defence. This is significant because a contrary approach would make the research limb unavailable to almost every private technology company. 

However, the ruling was made at the interim stage. It did not establish that all commercial AI training is fair dealing. The Court’s reasoning must be understood in relation to the facts before it, including the nature of the training process, the storage of works, the allegations concerning outputs and the evidence available at that stage.

Indian courts have traditionally assessed fair dealing through a contextual inquiry rather than a rigid mathematical test. Relevant considerations may include the purpose of the use, the amount and substantiality of the material taken, the manner of use, the degree of transformation, and the effect on the copyright owner’s market. These factors are not applied mechanically. The analysis asks whether the defendant’s use competes with or substitutes the original work, and whether the use is consistent with the purpose of the statutory exception.

The ANI Media v. OpenAI Decision

ANI alleged that OpenAI had copied and stored its news content for training and that ChatGPT could reproduce or communicate ANI’s articles in response to user prompts. The case therefore involved two distinct issues. The first concerned the training process. The second concerned outputs generated after training.

The Delhi High Court treated the storage and use of ANI’s literary works for model training as capable, prima facie, of satisfying the purpose requirement under Section 52(1)(a). The Court also found that the outputs did not amount, on the material before it, to substantial reproduction of ANI’s articles.

The distinction between training and output is legally important. Training may involve internal processing of works without presenting them to the public in their original form. An output, by contrast, is delivered to a user and may reproduce protected expression. A finding that training is fair dealing cannot automatically protect an output that reproduces an entire article, a substantial passage, a distinctive image or a protected software component.

The decision therefore suggests a two-stage infringement analysis:

  1. Whether the acts involved in collecting, storing and processing works for training are protected by Section 52.
  2. Whether the system’s outputs reproduce a substantial part of a protected work or otherwise infringe a copyright owner’s rights.

This separation is analytically sound. It prevents the legality of the training process from immunising downstream copying. It also prevents the existence of harmful outputs from automatically proving that every intermediate training act was unlawful.

The ruling should nevertheless be treated cautiously. It was not a final determination after a complete trial. The Court considered the evidence and legal arguments in the context of interim relief. The final decision may require closer analysis of the dataset, the extent of copying, the method of access, the contractual terms governing the content, the nature of the outputs and the commercial effect on ANI.

Licensing as the Safer Route

Even if fair dealing may protect some training uses, licensing remains the safest route where the dataset is commercially valuable, curated, paywalled or likely to generate direct market substitution. A licence can address questions that statutory exceptions do not resolve clearly.

A training-data licence should identify the works covered, the permitted modes of access, whether copies may be stored, the duration of use, whether the model may be commercialised, and whether the licence extends to fine-tuning, evaluation, retrieval-augmented generation or future models. It should also specify whether the licensee may permit affiliates, contractors and cloud providers to access the data.

The parties should distinguish between training rights and output rights. A licence may allow use of works for model development but impose restrictions on generating outputs that reproduce or closely imitate the licensed material. It may require filtering, attribution, opt-out processing, takedown systems, audit rights and technical measures to reduce memorisation.

Payment models may include fixed fees, usage-based royalties, revenue sharing, collective licensing or a combination of these structures. News publishers may prefer licences tied to the number of articles, queries, subscribers or revenue generated by AI services. Academic and public-interest datasets may adopt different terms based on research purpose and non-commercial access.

Licensing may also reduce evidentiary uncertainty. If a dispute reaches court, a written agreement establishes the scope of consent and limits arguments over implied permission. It can also allocate responsibility for third-party claims and specify which jurisdiction and dispute-resolution mechanism will apply.

Infringement Risk for AI Developers

AI developers face several distinct risks in India. The first is unauthorised copying. If a company obtains works through scraping, circumvents access controls or copies a protected database, the owner may argue that the conduct falls outside fair dealing.

The second is excessive reliance on the research purpose. A company should not assume that describing a commercial product as research will automatically satisfy Section 52. Courts may examine whether the research purpose is genuine, whether the process is connected to public or private study, and whether the commercial exploitation goes beyond what the exception can reasonably accommodate.

The third is output memorisation. A model that repeatedly produces substantial portions of protected works may expose its operator to infringement claims even if the training process is treated favourably. Technical safeguards should therefore be part of legal compliance, not an afterthought.

The fourth concerns contractual restrictions. A website’s terms of use, database licence or subscription agreement may prohibit automated extraction, even if copyright infringement is uncertain. Contractual liability may arise separately from copyright liability.

The fifth involves confidential information and personal data. Copyright clearance does not authorise disclosure of trade secrets or processing of personal information. AI developers must evaluate confidentiality, privacy and cybersecurity obligations alongside copyright.

Implications for Copyright Owners

Copyright owners should improve their evidence and licensing position. They should maintain records showing ownership, publication history, database structure, access conditions, licensing revenue and instances of alleged output reproduction. A general assertion that an AI system used “online content” may be insufficient to prove copying or market harm.

Publishers and archives should also clarify website terms, crawler permissions and machine-readable licensing signals. They may choose to create separate licensing products for training, search, retrieval and output display rather than treating all digital access as one category.

The ANI Media decision makes output monitoring especially important. Owners should test whether a system reproduces their works in response to prompts, whether it attributes sources accurately, and whether it provides summaries that compete with the original service. This evidence may support a stronger claim than simply showing that a work appeared somewhere in a training dataset.

Conclusion

Indian copyright law is entering a period of significant adjustment as courts apply a pre-existing fair-dealing framework to generative AI. Section 14 protects reproduction rights, Section 51 identifies infringement, and Section 52 may protect certain research-oriented uses. The Delhi High Court’s ruling in ANI Media v OpenAI indicates that model training and generated outputs should be examined separately and that commercial AI development is not automatically excluded from the research limb of fair dealing.

The decision does not create a general exemption for AI companies. Training remains exposed where copying is excessive, access is unauthorised, database rights are violated, contractual restrictions are breached, or the system generates substantial reproductions. Licensing remains the most secure option for high-value, curated and commercially sensitive datasets.

Author:- Amrita Pradhanin case of any queries please contact/write back to us at support@ipandlegalfilings.com or   IP & Legal Filing.

References

  1. ANI Media Pvt Ltd v. OpenAI Opco LLC, CS(COMM) 1028/2024.
  2. Copyright Act, 1957, Section 14.
  3. Copyright Act, 1957, Section 51.
  4. Copyright Act, 1957, Section 2(o)
  5. Copyright Act, 1957, Section 13(1)(a).
  6. Copyright Act, 1957, Section 52(1)(a)(i).
  7. The Indian Express, ‘How India Proposes to Deal with Legal Challenges Posed by AI to Copyright Law’ (20 December 2025) https://indianexpress.com/article/explained/explained-law/how-india-proposes-to-deal-with-legal-challenges-posed-by-ai-to-copyright-law-10430219/
  8. Copyright Act, 1957, Section 52(1)(a)(ii) and (iii).
  9. Copyright Act, 1957, Section 57 and 63.
  10. R G Anand v. Deluxe Films, (1978) 4 SCC 118.
  11. The Chancellor, Masters and Scholars of the University of Oxford v Rameshwari Photocopy Services, 2016 SCC OnLine Del 5128.
  12. Copyright Act, 1957, Section 2(k).
  13. NLS International Journal of Law and Technology, ‘The Copyright Play: AI, Section 52, and the Acts of Fair Dealing’ (13 May 2026) https://forum.nls.ac.in/ijlt-blog-post/the-copyright-play-ai-section-52-and-the-acts-of-fair-dealing/