Permission at the Point of Extraction: The Access Question After ANI v. OpenAI

ANI Media v. OpenAI leaves unresolved a logically prior question to copyright use - when may publicly accessible material be lawfully extracted for AI systems? This article argues that Indian law now reflects opt-out reasoning but no opt-out regime. Drawing on Section 43(b) of the Information Technology Act and a twenty-four-domain RightSignal audit, it proposes a purpose-specific framework based on technical conduct, declared function, communicated scope, and collection within scope, distinguishing public accessibility from unrestricted permission.

Manraj Singh Chandpuri

October 10, 2026 12 min read
Share:

Introduction

Artificial intelligence (“AI”) related copyright disputes raise two distinct questions. The ‘use’ question asks whether storing works for model training or reproducing them in outputs infringes copyright. The logically prior ‘access’ question asks whether the developer was permitted to collect or scrape those works in the first place. In ANI Media Pvt. Ltd. v. OpenAI OpCo LLC (“ANI matter”), the Delhi High Court (“DHC”) declined interim relief, holding prima facie that storage for training fell within Section 52(1)(a) of the Copyright Act, 1957 (“the Act”), and that ANI had not shown substantially similar or memorised outputs [¶¶256, 271]. 

Access was not among the four issues framed for interim determination [¶9]. The DHC nevertheless noted that ANI had not alleged acquisition from unauthorised sources or circumvention of a paywall, and later relied on ANI’s failure to block crawlers when assessing the balance of convenience [¶¶200, 262–263]. Those observations establish non-prevention, not necessarily permission and provides the Indian AI industry a temporary defence and not an invariant rule. The DHC did not consider Section 43(b) of the Information Technology Act, 2000 (“IT Act”), which addresses the extraction of data without permission. 

This article argues that the ANI matter leaves Indian law with opt-out reasoning but no opt-out regime, because the DHC treated the availability and non-exercise of an “opting-out option” as legally relevant, without defining the rules that would make such an opt-out system operational. If failure to block crawlers carries legal significance, the law must determine what counts as an effective reservation, which automated purpose it governs, and how silence or competing signals should be treated. Comparatively, the European Union’s Text and Data Mining (“EU TDM”) framework does so through an express statutory reservation mechanism, while India has no equivalent framework. Using RightSignal, an instrument developed by the author, the article audits twenty-four public domains and proposes a four-part Section 43(b) inquiry into technical conduct, declared function, communicated scope, and compliance with that scope. Ergo public accessibility of a work may evidence permission, but does not by itself establish permission for every automated use. 

The Access and Permission Question ANI did not decide

This article does not contest the DHC’s provisional conclusion on the use question. ANI argued that the Explanation to Section 52(1)(a) confines the exception to non-infringing copies, so that an unlawfully acquired copy could not attract the defence. The DHC rejected that reading and held that the limitation applies only to incidental storage of a computer programme and noted that Sections 52(1)(aa), 52(1)(ab) and 52(1)(ad), unlike Section 52(1)(a), expressly refer to a lawful possessor or a legally obtained copy [¶¶194–196, 201]. On that construction, lawful acquisition was not made to be an element of the Section 52(1)(a) defence.

That conclusion does not decide whether collection may attract an independent legal rule. Paragraph 200 records, in the alternative, that ANI had not alleged that OpenAI obtained an infringing copy from an unauthorised source. Paragraphs 262 and 263 concern the balance of convenience and the conduct of a party seeking discretionary relief. Neither passage asks whether ANI permitted extraction for model training. The DHC had already treated that question as unnecessary to its construction of Section 52(1)(a) of the Act.

Paragraph 262 matters significantly because the DHC expressly referred to ANI’s “opting-out option” and treated its non-exercise as relevant. That reasoning resembles an opt-out rule, where automated collection is treated as permissible unless the publisher says otherwise. A workable opt-out, however, needs more than the abstract ability to block a crawler. It must identify how a reservation is expressed, what content and purpose it covers, and what happens when a website publishes more than one instruction.

The Copyright Act contains no comparable condition in Section 52(1)(a). By contrast, Section 52(1)(c) expressly makes its exemption depend on the rightsholder not having prohibited the relevant links, access or integration. This contrast showcases the Parliament’s awareness of construing a non-prohibition part of an exception when it chooses to do so.

The EU TDM illustrates the missing machinery without providing the Indian answer. Article 4(3) of the CDSM Directive creates a general TDM exception for lawfully accessible works only where the use has not been expressly reserved. For content made publicly available online, Article 4(3) recognises reservations made through appropriate means, including machine-readable means. India need not adopt that default, but once silence is given legal effect, the law must specify what counts as an effective signal.

Section 43(b) of the IT Act supplies the vocabulary most directly concerned with extraction. It applies where a person, without the permission of the owner or person in charge of a computer resource, downloads, copies or extracts data or information. The relevant permission is therefore permission concerning the computer resource, which need not always be identical to a copyright license from the owner of the work. This route has been already identified by the executive. Answering Rajya Sabha Unstarred Question No. 558, the MeitY stated that “web scraping of any publicly available user data … for training AI models or any other purpose is regulated under the … IT Act..” and referred to Section 43’s consequences for unauthorised access. The answer is not binding and does not decide its application to publisher content, but identifies a plausible statutory route for the access question left outside the ANI matter.

What Operators Publish

Automated collection is ordinarily performed by a crawler, which is a software that requests webpages at scale. Before or while serving those requests, a website may publish instructions directed at crawlers. The best-known location is the robots.txt, a small file at the root of a domain that can address named crawlers and particular paths. Other instructions      may appear in HTTP response headers, page metadata, dedicated TDM reservation policies, or machine-readable license links. These mechanisms were created for different purposes and do not, by themselves, determine legal permission. They are evidence of what a site communicated about an agent, path or use.

RightSignal was developed to make that evidence inspectable. A user enters a domain, the instrument retrieves published signal locations, preserves the source text, translates recognised instructions into four purposes – search indexing, AI training, generative-AI training and commercial TDM – and applies disclosed resolution rules. It is deterministic, meaning that the same evidence produces the same result, without an external language model deciding the classification. A later live audit may differ because this website or server response may have changed. Most importantly, the output is an evidentiary record, not a conclusion that scraping is lawful or unlawful.

The underlying empirical record (“ER”), preserved on the Internet Archive, documents a purposive pilot, rather than a prevalence study. Twenty-four domains were selected before the audit to span different publishers and jurisdictions; every selected domain was retained. Across ninety-six domain-use observations, the instrument returned different outcomes for different purposes      on fifteen domains. Those differences were not equivalent. Five domains expressly differentiated between automated agents or purposes. On nine, one purpose was restricted while another encountered silence, so the difference came from the instrument’s default rather than an affirmative statement by the operator. One arose from a mapping artefact. Every permissive resolver result was produced by silence. Manual review identified only one affirmative permission.

That permission appeared on Sveriges Television’s domain [ER, §§3, 6; pp. 3-4, 249]. Its robots.txt stated in ordinary language that the site’s journalism was available for public search indexing and real-time retrieval, but not for training AI models (Figure 1). 

 

Figure 1. Sveriges Television’s robots.txt, admitting OpenAI’s retrieval and search agents while refusing AI training (captured 25 July 2026).

It also published a machine-readable Content-Signal instruction:

Published Instruction Plain meaning
search=yes search indexing permitted
ai-input=yes AI input or retrieval permitted
ai-train=no AI training refused

The file allowed OpenAI’s OAI-SearchBot and ChatGPT-User at the site root while maintaining the training refusal. At the level of published instructions, the site therefore did not choose between blocking OpenAI and permitting every OpenAI use. It admitted particular retrieval and search functions while refusing training. The robots.txt is not a binding protocol. The legal significance is not that it necessarily binds an Indian court. It is that a bare finding that “the crawler was not blocked” cannot reveal the purpose for which access was communicated.

ANI’s robots.txt (Figure 2) demonstrates the problem of incomplete evidence. The file disallows fourth paths, publishes two sitemaps, including a Google News sitemap, and names no AI crawler. A sitemap tells search engines where content can be found. It can support an inference that the operator contemplated search discovery, but it is not an express legal permission and says nothing directly about model training.

Figure 2. ANI’s robots.txt, retrieved manually. The audit’s automated request for the same file returned HTTP 403.

The manner in which the file was retrieved adds another layer. RightSignal’s automated request returned HTTP 403, a server response refusing that particular request, although the file could be opened manually. The response does not reveal whether the cause was a content-delivery-network rule, user-agent filtering, rate control or bot mitigation. Nor does an audit conducted on 25 July 2026 establish the site’s configuration when OpenAI allegedly collected the material. Because no readable carrier supplied a qualifying instruction, the resolver produced permissive outputs through its silence default; the ER correctly reclassified that observation as indeterminate rather than affirmative permission.

These examples yield three legal propositions. First, websites can communicate permission and refusal by purpose rather than for “access” in the abstract. Second, silence, a failed request and an express instruction are different forms of evidence and should not be collapsed into the same conclusion. Third, technical output requires interpretation, as a transparent instrument can preserve and organise the evidence, but cannot determine the legal scope of permission. Those propositions supply the bride to Section 43(b) of the IT Act.

A Purpose-Specific Inquiry under Section 43(b)

Section 43 of the IT Act does not prescribe a particular form in which permission must be given. Applied to the open web, permission may therefore have to be inferred from conduct and context. Publicly positing a page ordinarily invites human retrieval. Depending on the circumstances, it may also support an inference permitting conventional search indexing and the incidental copying necessary to provide it. The difficult question is how far that inference extends.

A court considering whether extraction occurred without permission should assess four non-exhaustive matters:

  1. Technical Conduct: Did the collector defeat authentication, a paywall, a meaningful crawler block, a rate limit or another boundary? Circumvention strongly weakens an inference of permission. Its absence is relevant, but not conclusive.
  2. Declared Function: Did the crawler identify itself as operating for search, user-directed retrieval, training or commercial extraction? A declared identity is evidence, not a formal prerequisite, because identifiers may be missing, inaccurate or imitated.
  3. Communicated Scope: Did the operator publish a reasonably discoverable instruction relevant to the crawler, path, content or purpose? Until Indian law recognises one controlling protocol, the inquiry should examine the available carriers together rather than treating robots.txt, a header or a policy page as automatically decisive.
  4. Collection within Scope: Did the collector remain within the purpose, paths, rate and other limits reasonably communicated? A discovered, relevant restriction should weigh against an inference of unrestricted permission. Silence may preserve an inference drawn from public accessibility, but it should not be redescribed as affirmative consent.

The Sveriges Television example primarily informs the second and third considerations. Its published instructions differentiate search and retrieval agents from training. ANI’s file, by contrast, supports a narrower inference relating to search discovery and leaves training unaddressed. The unexplained HTTP 403 response and the audit’s later date prevent a stronger conclusion. Nothing in the pilot establishes whether OpenAI had or lacked permission on the facts of the lawsuit, and the article does not ask a technical file to decide that issue.

The framework instead identifies the evidence a court would need. It preserves the ordinary inference arising from an open website, avoids requiring a negotiated licence for every public page, and still gives legal relevance to purpose-specific instructions where they are reasonably discoverable. It also keeps the access question distinct from copyright infringement. A collector may have permission to retrieve content without having permission to reproduce it in every later use, and a copyright exception may apply to a use without resolving an independent complain about the manner of extraction.

The institutional implication is limited, yet important. Reservation signals are being standardised abroad. The European Commission consulted between December 2025 and January 2026 on machine-readable protocols for reserving rights against TDM, in support of Article 53(1)(c) of the AI Act and the Copyright section of the General-Purpose AI Code of Practice, to which OpenAI is a signatory, and will publish an agreed list of such protocols. India faces the prior question whether non-blocking or usage of such protocols should carry that consequence at all. Courts can begin by treating published signals as evidence under Section 43(b). The Parliament or the appropriate authorities may later define recognised carriers.

Conclusion

The interim ruling gives OpenAI a substantial victory on the use question, but does not settle the logically prior access question. The DHC’s observations on paywalls and crawler blocking were made in alternative reasoning and on the balance of convenience. They should not become a complete permission test by implication. Section 43(b) of the IT Act offers a more direct interpretation of extraction without permission. Applied to public websites, that interpretation should distinguish technical non-prevention from permission and should ask what purpose, scope and limits were communicated. The four considerations proposed here provide a workable starting point without covering every open page into licensed content or every technical instruction into law. Until a court or Parliament supplies a clearer threshold, public accessibility should remain evidence of permission, not an automatic licence for every automated use.

The author is a Year III student at the University Institute of Legal Studies (UILS), Panjab University, Chandigarh, India.

The Buck Stops at the Bench: Judicial Accountability Without Institutional Liability in India’s Draft AI Regulations October 1, 2026