[codicts-css-switcher id=”346″]

Global Law Experts Logo
can ai be trained

Our Expert in India

Can AI Be Trained on Copyrighted Books in India? (the ANI V. Openai Litigation)

By Global Law Experts
– posted 1 hour ago

Can AI Be Trained on Copyrighted Books in India?

The legal position after the Delhi High Court’s decision in ANI v. OpenAI
The Delhi High Court has significantly shifted the Indian copyright debate on AI training. Its July 2026 ruling suggests that storing copyrighted material for LLM training can, in appropriate cases, qualify as fair dealing for “private or personal use, including research”. But the judgment is not a blanket license to ingest protected content: the source material, market impact and—most importantly—the model’s outputs remain critical.

A recurring question for businesses building domain-specific AI tools is deceptively simple: can an AI model be trained on copyrighted books without obtaining a license from the copyright owner? Consider an AI platform that scans specialist technical reference books, trains a model, deletes the scanned files after training, and then gives users consolidated answers in its own words with book and page references. Until recently, the conservative position under Indian law was that this sat in a legal grey area leaning towards infringement. The Delhi High Court’s decision in ANI Media Pvt. Ltd. v. Open AI OpCo LLC (24 July 2026) has materially changed that risk analysis.

  • Why AI training raises a copyright issue in the first place

Under Section 14 of the Copyright Act, 1957, the copyright owner of a literary work has the exclusive right to reproduce the work, including storing it electronically. Scanning a textbook or copying digital content into a training dataset therefore engages the reproduction right at the outset. Deleting the scan after training does not undo the fact that electronic storage occurred.

The difficulty has always been Section 52. India follows a purpose-specific “fair dealing” regime rather than the broader US-style fair use doctrine. Section 52(1)(a) protects fair dealing for specified purposes, including “private or personal use, including research”. Indian law still does not contain an express statutory exception specifically labelled for AI training or text-and-data mining. The core question is therefore whether AI training can fit within the existing fair dealing framework.

  • ANI v. OpenAI: the key change in the Indian position

In July 2026, the Delhi High Court refused ANI’s application for an interim injunction against OpenAI. At the prima facie stage, the Court held that storage of ANI’s original literary works for training LLMs could fall within Section 52(1)(a) and would therefore not amount to copyright infringement.

 Three aspects of the ruling are particularly important for AI businesses:

  • Commercial use is not automatically disqualified. The Court declined to read a non-commercial limitation into Section 52(1)(a). A business does not lose the fair dealing defense merely because the model is ultimately monetized.
  • Machine learning can amount to “research”. The Court treated LLM training—screening, organizing and analyzing stored literary works to train a model—as a form of research and applied an updating interpretation of the statute to modern technology.
  • “Private” use can include closed model training. The Court emphasized that the training material was used in a closed environment and was not itself made available to the public. That internal use was capable of satisfying the private-use limb, subject to fairness.

The Court then applied a broader fairness analysis, looking at the purpose and character of the use, the extent to which the training use differed from the original use, market substitution, and public interest. On the facts before it, the Court found the requirements of fair dealing satisfied at the interim stage.

  • Training and output are two different copyright questions

This distinction is arguably the most useful part of the judgment. Even if the training process is protected as fair dealing, the model’s output can independently infringe copyright if it reproduces a substantial part of the protected expression.

In ANI, the Court was not satisfied that the challenged ChatGPT responses were substantially similar to ANI’s original articles. It also distinguished copyright in expression from the underlying facts or information.

However, the Court made clear that regurgitation or extraction of training data that results in substantial reproduction of the original work can amount to infringement.

For a domain-specific AI product trained on technical books, this means the technical design of the output layer matters enormously. A system that synthesizes information, answers in genuinely original language and points the user to the source is in a materially better position than one capable of reproducing paragraphs, tables, explanations or distinctive formulations from the book.

Citations help—but they are not a copyright defense. Providing the book title, page or paragraph reference may support the argument that the AI is a research-assistive tool rather than a replacement for the source. It can also reduce the likelihood of market substitution by directing users back to the original. But attribution does not legalize copying that would otherwise infringe copyright.

  • Does the ruling give a green light to train on technical books? Not quite.

ANI is a major pro-AI development, but it should not be read too broadly. The ruling was made on an interim injunction application, the observations are prima facie, and ANI has already challenged the order in appeal. The final legal position therefore remains capable of change.

More importantly, the facts of a technical-book case may be harder than the facts of ANI. News articles are largely factual and, in ANI, were freely available online. A specialist engineering reference book is a paid, structured work whose commercial purpose may be closer to the purpose served by a domain-specific question-answering tool. If the AI product enables users to obtain the same practical value that they would otherwise obtain by purchasing or consulting the book, a publisher may have a stronger market-substitution argument.

Equally, scanning an entire book is a significant act of copying. The July 2026 judgment makes such copying more defensible where it is confined to a closed training process and the ultimate outputs are non-substitutive, but the analysis remains fact-sensitive. A license from the publisher therefore remains the lowest-risk route for core proprietary content, particularly where a small number of high-value works are central to the product.

  • Practical guardrails for businesses building AI on copyrighted content

    • Use lawfully accessed source material and avoid pirated repositories, shadow libraries or circumvention of paywalls. Source provenance should be documented.
    • Keep training copies in a controlled environment and retain them only for as long as technically necessary. Document deletion and access controls.
    • Build anti-regurgitation safeguards. Test the model specifically for verbatim or near-verbatim extraction, including through adversarial prompts.
    • Design outputs to synthesize rather than substitute. Short answers, multiple-source synthesis and source references are safer than reproducing lengthy explanations from a single work.
    • Assess market substitution work-by-work. Where the product is likely to replace the need to buy or consult a core publication, licensing or a publisher partnership should be seriously considered.
  • The takeaway

The Indian position has moved considerably in favour of AI developers. After ANI v. OpenAI, it is no longer accurate to say that commercial AI training on copyrighted material necessarily falls outside Section 52 merely because a company is doing the training or intends to monetize the resulting model. The Delhi High Court has recognized, at least prima facie, that machine learning can be “research” and that closed training use can be “private”.

But the judgment is not a blanket text-and-data-mining exception. The safest legal analysis still separates three questions: how the source material was obtained; what happens to it during training; and what the model ultimately gives the user. For businesses developing specialized AI tools, that distinction is likely to determine whether the product looks like a legitimate research technology—or an unlicensed substitute for the copyrighted work.

Note: This article is for general information only and does not constitute legal advice. The discussion reflects the legal position and publicly reported procedural status as at 9 September 2026.

FAQs

Can AI models legally be trained on copyrighted books in India?
There is no settled, categorical answer under Indian law, and the ANI v. OpenAI litigation remains a live matter. In principle, storing and processing copyrighted works to train a closed model might be argued to qualify as “private or personal use, including research” under the Copyright Act in appropriate cases, provided access is controlled and outputs are non-substitutive. Any such position is conditional, fact-specific, and not a blanket permission.
Not necessarily automatically. Commerciality is a relevant factor, but the key questions are whether the use serves a permitted research purpose and whether it substitutes for the original in the market. Closed training with non-substitutive outputs may be defensible despite a commercial context, though this depends on the facts.
Potentially yes. Reproducing a substantial part of a protected work, assessed qualitatively as well as quantitatively, can be infringement. Any fair dealing shelter that might apply to training does not automatically protect a regurgitated output, which is judged as a fresh act of reproduction.
Immediate priorities include: a data-provenance register, training-specific licences where available, a genuinely closed model with access controls, anti-regurgitation and output-filtering controls, and red-team memorisation testing with documented results. Contractual indemnities and takedown processes should follow.
Notices alone rarely halt training, but they establish a record. Rights-holders with evidence of regurgitation or market substitution can seek injunctions and damages, and can pursue licensing negotiations. The strength of any remedy depends on documented, forensically sound evidence.
Arguments favouring developers tend to turn heavily on model closure and access controls. Open models that redistribute weights or corpora face a materially different, and generally higher, risk profile, because the closure and control indicators are weaker or absent.
Honouring reasoned deletion requests from rights-holders is good practice and can support a fair dealing position, though the precise obligation depends on contractual terms and the legal basis for the request. A documented deletion policy and clear escalation pathway are advisable.
For advisory work, compliance audits and litigation strategy on AI and copyright in India, seek counsel experienced in corporate and intellectual property matters. Global Law Experts’ corporate practice in India can assist with tailored guidance.
qfc company formation qatar
By Jonathon Richards

posted 26 minutes ago

complete due diligence
By Global Law Experts

posted 1 hour ago

Find the right Legal Expert for your business

The premier guide to leading legal professionals throughout the world

Specialism
Country
Practice Area
LAWYERS RECOGNIZED
0
EVALUATIONS OF LAWYERS BY THEIR PEERS
0 m+
PRACTICE AREAS
0
COUNTRIES AROUND THE WORLD
0
Lawyer Profile Page - Lead Capture
GLE-Logo-White
Lawyer Profile Page - Lead Capture

Can AI Be Trained on Copyrighted Books in India? (the ANI V. Openai Litigation)

Send welcome message

Custom Message