Voluntary Commitment and Where Private AI Safety Frameworks Substitute for Governance

International AI governance has come to rely on instruments that function as soft law without satisfying the procedural conditions that make this soft law a legitimate mode of regulation. This is mutually convenient for governments and the companies. To tackle this, a duty of procedural fidelity should be attached to representations made within state-convened processes on which reliance has been invited.

Advitiya Pathak, Tanush Mehrotra

July 30, 2026 15 min read
Share:

Anthropic published version 3.0 of its Responsible Scaling Policy (RSP v3.0) on 24 February 2026, with the change being the removal of the company’s flagship safeguard on which it had attracted users away from OpenAI. Initially, it was a threshold-triggered commitment to pause development if specified safety mitigations could not be implemented in advance. This was replaced with a two-part test that activates only when Anthropic simultaneously considers itself the frontier-race leader and judges catastrophic risk to be material, conditions the company itself determines. New transparency obligations were introduced to compensate, but the initial pledge was no longer effective in substance. Chief Science Officer Jared Kaplan told TIME that it no longer made sense to make unilateral commitments while competitors raced ahead.

Observers disagreed on whether the substantive direction was defensible. SaferAI downgraded Anthropic’s safety governance score from 2.2 to 1.9, placing it, alongside OpenAI and Google DeepMind, in the “weak” category. The Centre for the Governance of AI offered the more measured assessment that the new disclosure mechanisms partly offset the loss of the pause commitment. The revision also coincided with reported Pentagon pressure over military use of Claude, which Anthropic has characterised as unrelated to the RSP change. It came two weeks after the head of Anthropic’s Safeguards Research team resigned, warning of “the world… in peril.”

The original commitment had not been a private undertaking. It had been re-articulated at the AI Seoul Summit in May 2024, where the Frontier AI Safety Commitments were signed by sixteen frontier AI companies (now twenty), in parallel with the Seoul Declaration for Safe, Innovative and Inclusive AI adopted by ten governments and the European Union, and tied, in the day-two Seoul Ministerial Statement, to commitments from twenty-seven nations and the EU. The Commitments were cited by regulators, modelled by competitors, and incorporated into the technical work of the networked AI Safety Institutes. Twenty-one months later, the core pledge was withdrawn through an internal process and advocacy for the change began within Anthropic in February 2025, and the revision was approved by the CEO and Board after which external review on the draft was confined to selected partners chosen by Anthropic, such as METR. No public record indicates consultation with the Seoul Summit’s co-conveners or with the AI Safety Institutes whose work had drawn on the prior commitment.

International AI governance has come to rely on instruments that function as soft law (cited, relied upon, treated as authoritative) without satisfying the procedural conditions that make soft law a legitimate mode of regulation. This is mutually convenient wherein states gain the appearance of governance from convening summits at which companies make commitments while companies gain legitimating association with state forums while retaining unilateral control over the substance. Each actor benefits from the implication of binding effect that neither party has formally undertaken.

The Landscape and the Vacuum

Three tracks of AI governance now operate in parallel. The first is state-led, for instance the EU AI Act entered into force on 1 August 2024, with prohibited practices applying from February 2025 and rules for general-purpose AI from August 2025. The Act’s most ambitious provisions, those governing high-risk AI systems, have nonetheless been deferred. On 7 May 2026, the Council and Parliament reached provisional agreement on the Commission’s Digital Omnibus, postponing standalone high-risk obligations to 2 December 2027 and embedded-product high-risk obligations to 2 August 2028. The Brussels Effect for AI, theorised by Bradford, has been narrower than predicted; few jurisdictions have adopted the EU’s risk-tiered model wholesale. The implication is a binding regime which has its most integral portions suspended in a contingent future, widening the gap between the law and its practical application.

The second track is intergovernmental. Bletchley (November 2023), Seoul (May 2024), Paris (February 2025), Delhi (February 2026). Bletchley centred on existential and catastrophic risk; Seoul produced the Frontier AI Safety Commitments and the AI Safety Institute network; Paris rebranded as the “AI Action Summit,” dropping “safety” from the title, and concluded without US or UK endorsement of its joint declaration and Delhi pivoted further toward inclusion and economic opportunity, with safety appearing chiefly as a sub-theme. The intergovernmental track increasingly operates as a forum for declaration rather than commitment. It effectively attempts to confer legitimacy on declarations without generating any binding obligations, keeping the declaring parties in check once they are made.

The third track, the focus of this essay, is private. Companies publish Responsible Scaling Policies, Frontier Safety Frameworks, Preparedness Frameworks, and analogous documents. As of December 2025, the METR survey identifies twelve frontier AI companies as having published such policies. These documents have direct regulatory consequences that can be observed in the recent wave of AI legislation across the world. They were modelled in California’s SB 53, New York’s RAISE Act, and the EU AI Act’s Code of Practice. They are the de facto floor of frontier AI safety governance. They are also unilaterally revocable.

Shadow Soft Law

The instruments at issue perform the social and regulatory functions of soft law. They are cited by regulators, modelled by competitors and treated as authority in international forums without satisfying the procedural conditions that distinguish accepted soft-law instruments from corporate declarations. Call this shadow soft law. The general phenomenon of private actors generating governance-like norms in technical domains has been extensively mapped: by Büthe and Mattli on private rule-making in international product standards, by DeNardis on private governance of internet infrastructure, and by Coglianese and Lazer on management-based regulation.

DeNardis locates internet governance in “control points”: private intermediaries exercise quasi-public authority by virtue of operating the infrastructure through which data passes. Thus, governance in this sense is continuous and is a property of those who run such infrastructure, exercised quietly as a matter of ongoing business function. The AI safety case inverts these features. Its governance is exercised not through control of infrastructure, but through sporadic publicly announced commitments made at state-convened summits and addressed to states, regulators and the public. It is declaratory rather than operational, and somewhat visible rather than concealed.

The control point is thus now a promise, and they can be withdrawn unilaterally without trace. The private power at issue is therefore not the standing authority to route packets, but the discretionary authority to make and unmake the safety representations on which nascent regulatory order has come to rely. That shifts the governance problem from access and control to commitment and revocation, which has an unsettled doctrinal home.

There is no single canonical test of soft-law legitimacy, but the converging literature in global administrative law and international regulatory governance points to three procedural criteria. The Kingsbury, Krisch and Stewart framing identifies transparency, participation, reasoned decision and review as the procedural elements that promote accountability in global regulatory practice. The Scharpf–Schmidt typology, applied to EU soft-law legitimacy, distinguishes input legitimacy (participation), throughput legitimacy (transparency, openness), and output legitimacy (efficacy). The common thread is that legitimacy requires procedural conditions analogous to those of formal lawmaking more than flexibility.

Apply these criteria to RSP v3.0. Transparency: the policy was published the day it took effect; on the company’s own account, internal advocacy for the change began twelve months earlier, and the consequential governance work was internal. Participation: Anthropic’s own materials disclose external review by selected partners such as METR; no record indicates consultation with the UK or South Korean co-conveners of the Seoul Summit, or with the AI Safety Institutes whose work had drawn on the prior commitment. Accountability: the new framework defers most consequential judgments (whether Anthropic is the race leader, whether catastrophic risk is material) to the company’s own evaluative discretion, and the Frontier Safety Roadmaps are described, in the company’s own framing, as “non-binding” public goals.

Applying this to the Seoul Commitments yields mixed results. They fare better on transparency, the formation process being public. The document is published and the co-signatories are identified. There is a lack of stakeholder input though as no civil-society or affected-community comments are documented in the formation process. The document also expressly provides that companies’ approaches “may evolve”, demonstrating the possibility of future inconsistencies. Both instruments thus function as soft law without meeting the conditions that make soft law a legitimate mode of governance. Reidenberg’s prescient observation, that technological actors generate normative rules independent of formal law, extends into the AI era with the work of Chagal-Feferkorn and Elkin-Koren, taking on higher stakes when the rules in question govern catastrophic-risk thresholds. A recent empirical audit by Wang, Huang, Klyman and Bommasani of how AI companies have honoured their 2023 White House voluntary commitments suggests that the problem of this voluntary commitment is not merely conceptual but running across sixteen signatories, average rubric performance was 52 percent, with eleven firms scoring zero on the model-weight security commitment.

A Duty of Procedural Fidelity

The question of legal regulation becomes especially pertinent considering that public international law does not generally impose obligations on private actors. The orthodox position is that primary obligations bind states and, exceptionally, individuals under international criminal law, whereas corporations sit largely outside the circle of direct duty-bearers. This position is reflected in the International Law Commission’s Articles on State Responsibility, which attribute responsibility to states rather than private entities, and in the persistent attempts to conclude a binding business-and-human-rights treaty, which leaves corporate obligation in the realm of the voluntary. The soft-law instruments at issue here are, by their own terms, voluntary. We must therefore look to similar conceptual instances in order to develop a working theory of regulation in our case.

Firstly, the doctrine of unilateral declarations as developed in the Nuclear Tests cases (1974). The ICJ held that public declarations made with intent to be bound create obligations under good faith. The ILC’s 2006 Guiding Principles codifies such requirements. According to Principle 7, declarations must be in clear and specific terms, and according to Principle 10 they cannot be revoked unilaterally once others have reasonably relied upon them. The doctrine, though, applies only to states, not private actors, and it requires intention to be bound, a threshold the Seoul Commitments arguably do not clear, given the express “may evolve” clause.

Secondly, the related doctrine of legitimate expectations. Well-established as a component of the Fair and Equitable Treatment standard in investment-treaty arbitration, the doctrine protects reasonable detrimental reliance on representations or course of conduct. The ICJ in Bolivia v Chile (2018) expressly declined to recognise legitimate expectations as a principle of general international law, holding that its appearance in arbitral awards did not generate a free-standing inter-state obligation. The case therefore cuts against direct invocation of the doctrine here. What is relevant to the present argument is the underlying structural logic the Court reviewed before confining it. Reasonable reliance on public representations made in a setting of state involvement carries normative weight, even where it does not, on its own, generate a hard rule of general international law. Bolivia v Chile marks the doctrine’s outer limit.

Thirdly, we must look at the broader doctrine of equitable estoppel as applied in regulatory contexts. Equitable estoppel is widely accepted across common-law jurisdictions. In US insurance law, regulatory estoppel, first adopted in Morton International v. General Accident Insurance Co., 134 N.J. 1, 1993, prevents parties from preferring litigation contradicting representations made previously to regulators. The doctrine has been mostly rejected outside New Jersey, but its underlying intuition, that representations made to regulators carry obligations even where no direct contractual relationship binds the parties, is the most directly relevant to the AI safety case. Representations made to regulators in a setting of state involvement carry obligations even where no direct contractual relationship binds the parties. A company that makes safety representations at a state-convened forum, knowing those representations will be incorporated into regulatory analysis, occupies a position structurally similar to Morton’s regulated insurer and should be more transparent and accountable to public at large.

None of these three doctrines prima facie fits the case of an AI company unilaterally revising a commitment made at a state forum. The first applies only to states, the second has been confined by the ICJ to investment law, the third is contested even within US law. But their synthesis is suggestive. In each, public representations made in a setting of state involvement and relied upon by others generate obligations on the party who made them, more constraining than that party’s residual freedom of action. The Nuclear Tests and legitimate-expectations doctrines bind states and equitable estoppel binds a private regulated actor. We propose here is that the structural logic common to all three should extend to a private actor whose representations, made at a state-convened forum, have been relied upon by states and other private parties.

The conclusion de lege ferenda is that a duty of procedural fidelity should attach to safety commitments made at state-convened multilateral forums and relied upon by state regulators and other private actors. We do not propose forcing companies to mechanically retain their commitments, rather requiring revision to be undertaken through procedures adequate to the commitment’s public function: prior notice, reasoned justification, and consultation with the public stakeholders, regulators and co-signatories whose reliance has been invited, to name a few. A parallel can be drawn to such safeguards found in other frameworks, such as the GDPR.

Under Article 35 GDPR a controller contemplating high-risk processing must first conduct an impact assessment. Further under Article 36, where that assessment shows a high residual risk which its own mitigations cannot dispel, it must consult the supervisory authority before proceeding, and may not proceed until such consultation concludes within the defined statutory window.

A private actor whose conduct carries high residual risk to others is not left to self-certify in private, but owes a duty of prior, reasoned engagement with a public body before acting. The duty of procedural fidelity is the same instinct. Where a safety commitment made at a state forum is to be materially weakened, the maker should owe prior notice and reasoned consultation to the regulators and co-signatories whose reliance it invited, rather than revising in private and presenting a fait accompli.

Objections and Gaps

Gaps do exist. The first is that the duty of procedural fidelity will deter companies from making future safety pledges. The reply turns on the procedural nature of the proposed duty. Companies remain free, under this framework, to revise any commitment they have made. They must give notice, consult, and explain. This mirrors ordinary administrative-law constraints on agency rulemaking and rule-revision. The effect, if any, falls on commitments made carelessly, which is a feature, not a defect.

The second is that the escape clause covers the case. The Seoul Commitments provide that companies’ approaches “may evolve in the future,” and that companies will “provide transparency on this, including their reasons, through public updates.” A generous reading might be that RSP v3.0, with its accompanying public materials, satisfied it. A contrary reading of the same would be that the clause itself imposes a transparency obligation that public disclosure on the day of release does not discharge. Seoul co-signatories, the AI Safety Institutes, and downstream regulators relying on Anthropic’s prior commitment in their own work received no advance opportunity to engage. Transparency, in international-law terms, means more than after-the-fact publication. Even read in the clause’s own terms, RSP v3.0 tested it.

The third objection is that private companies are not states, and international law does not bind them. The proposal does not deny this. It is openly de lege ferenda. But the bright line between state and non-state obligation has softened over the decades. The shift has come through the UN Guiding Principles on Business and Human Rights, through investment law’s parallel binding effect on private actors, through the proliferation of public–private hybrid regulatory regimes. The proposal here is modest by comparison in the sense that it does not impose substantive obligations on private actors. Rather, it asks that representations made within state-convened processes, on which reliance has been invited, be revisable only through processes adequate to that public-facing character.

The World that Exists

We do not intend to assume AI companies as bad actors. By many accounts, Anthropic’s RSP v3.0 was an honest attempt to bring stated commitments into alignment with operational practice. The MIT Technology Review audit of the 2023 White House voluntary commitments found progress on some technical fronts but “no meaningful transparency or accountability”. Harvard Law Review’s observation that those pledges were not “backed by the force of law” and had no enforcement mechanism. These suggest that the problem is structural and actors are doing roughly what the structure permits.

International AI governance has come to depend on instruments that perform the regulatory work of soft law without the procedural guarantees that justify treating soft law as legitimate. Techno-polarity, Bremmer’s term for a world in which private corporations exercise governance-like power alongside states, is no longer an academic debate but a description of how AI is now in fact governed. We cannot wait for a binding global legislature or treaty, nor can we pretend that the existing instruments do not exist. Procedural standards are the bare minimum, adequate at least to the public function they have come to serve in keeping the international community informed. A duty of procedural fidelity should be attached to representations made within state-convened processes on which reliance has been invited.

The authors are Year III students at the West Bengal National University of Juridical Sciences. 

Architecture-Blind Governance: AI Systems and the Limits of Accountability in International Law July 29, 2026