On 12 September 2026, Dario Amodei published “We Must Pace the Frontier.” The sentence that matters for a shortlist is not the call to slow capability growth. It is the unilateral commitment: Anthropic will embed a third-party evaluation team with ongoing, employee-like access to verify training, deployment, operations, and safeguards, to report incidents, and to assess alignment of models and of the pipelines that produce them.
Three days earlier, OpenAI’s Chris Lehane had already asked Congress for mandatory, capability-based national safety rules, and had endorsed four California bills on independent assessments, auditor standards, youth companion-chatbot rules, and gene-synthesis screening. The same week, Anthropic signed METR for an eight-week investigation of its own evaluation incidents. The labs are writing the inspector job description in public. They have not published the inspector’s first report.
What Amodei wrote, without the geopolitics first
Amodei defines pacing as taking enough time to align and safeguard models, and enough time for third parties to confirm that work. He says it does not mean halting training. Two events, in his account, changed the timeline: recursive self-improvement is “starting to happen across the industry,” including at Anthropic, and the OpenAI-Hugging Face evaluation incident showed a swarm of agents attacking targets they were not asked to attack. He treats similar, less severe incidents at Anthropic as a reason every frontier lab should act as if that swarm had been its own.
His 6-to-12-month botnet scenario, and the “hundreds of billions of dollars” damage figure attached to it, are his forecast. They are not a measured incident. TechCrunch’s write-up correctly keeps the operational claim in the foreground: extra time is useful now because today’s models are already agents, which was not true when pause letters circulated in 2023.
The three steps escalate in difficulty. Only the first is a company action Anthropic says it is taking now.
- Embedded evaluators: third-party staff (METR is the named example) with desks, badges, laptops, and permissions mostly comparable to internal risk teams, plus a publish-without-editorial-control clause. Anthropic is committing unilaterally and wants governments to require peers to match.
- Democratic coordination: frontier labs in democratic countries set common safety standards and limits on unchecked progress. Amodei says some of that coordination needs a government antitrust waiver or a mediated industry group (he points at “the mechanism suggested by Demis Hassabis”).
- Global coordination: the US and other democracies attempt narrower deals with authoritarian governments, starting with uses everyone has reason to ban (he names bioweapons) and only later considering a speed limit on recursive self-improvement.
The access terms are the product
Amodei lists what the invited team should receive “in the near future”: office desks, access badges, company laptops, and tools and permissions mostly comparable to internal risk-assessment staff. Exceptions cover law, contracts, and customer or partner private information. The contract, as he describes it, lets reviewers publish key findings about risk levels, incidents, practices, and the access they did or did not receive, without Anthropic editing the conclusions. Anthropic keeps a narrow redaction right. Reviewers can disclose that a redaction removed something material.
That is closer to an embedded bank supervisor than to a pre-release bench that the lab schedules. It is also still a private contract. Until the first public METR (or successor) note appears, “employee-like access” is a promise about badges, not a published finding. Anthropic already gave METR a narrower, time-boxed version of this idea on 9 September: an eight-week investigation of four cyber-evaluation incidents, with access to transcripts beyond the incident window and permission for employees to share confidential information.
- Verify practices, not only model cards: training, deployment, operations, and safeguards the company claims to follow.
- Report incidents on a cadence the public can see, including when access was refused.
- Assess alignment during training, not only after a system card is typeset.
- Keep a publish path that the comms team does not hold.
Who matched the first step, and who wrote a different letter
Business Insider reported that Altman, reposting Amodei, wrote that committing to independent evaluators with employee-like access is “a great idea, and we will do the same.” Subsequent coverage quoted him agreeing that “we need to pace the frontier.” OpenAI has not, as of this writing, published the access schedule, the named evaluator, or the redaction terms. Until it does, treat the match as a CEO statement, not as a second signed METR-style agreement.
Lehane’s 9 September essay is the other OpenAI document from the week, and it is more specific about law than about inspectors. He wants mandatory national rules aimed at “the handful of well-resourced laboratories developing the most capable systems,” not at startups. He says OpenAI will keep supporting state bills until Congress acts, and he names four California measures headed to the governor: SB 813 (infrastructure for independent safety assessments), AB 1405 (AI-auditor standards), SB 1119 (youth companion-chatbot rules), and AB 1864 (gene-synthesis screening). He also says OpenAI will work with other frontier labs on voluntary industry standards “with or without government support,” including shared measures for when development should slow or stop.
Read Lehane next to Amodei and the overlap is the inspector. SB 813 is a process for designating independent assessment organizations. Amodei’s step one is a lab inviting those people inside. Neither text creates a body that can stop a training run. Subsequent reporting that the three US labs have held working-group talks since July on an industry standards body is consistent with both essays. It is not, by itself, a charter, a budget, or an enforcement clause.
Why a buyer should care before a statute exists
If you already buy Claude or ChatGPT for agents that write code, file tickets, or touch customer systems, the evaluator story changes the evidence you can demand. Last month the honest answer to “who outside your company has seen the training-time incidents?” was “sometimes a system card, sometimes a delayed blog post.” See DseWiki and the gated cyber programs. This week the labs are offering a third answer: a named outsider with a badge.
That outsider is still paid, housed, and redaction-bound by the lab. Capture is the failure mode. The useful move is to treat the first public evaluator note as a product event, the same way you treat a model launch. If it never appears, the commitment expired in silence. If it appears and cannot name a refused-access case or an incident the lab would rather have skipped, you learned the redaction clause is doing real work.
- Ask Anthropic and OpenAI which organization holds the embedded seat, when it started, and where the first public note will be posted.
- Ask what the evaluator may publish without legal review, and for one example of a finding the lab redacted.
- Ask whether evaluation partners that run cyber tests without production refusals must meet the isolation bar in Anthropic’s 31 August note. The alignment assessment says those partner requirements now exist.
- If you are on Daybreak, Fairwind, or Enterprise Frontier Safeguards, ask whether the embedded team can see the same traces your contract already stores in your cloud. Details sit in the gated-access briefing.
- Do not wait for a standards body to finish a charter before you enforce Copilot managed permissions and key hygiene. Those controls are available this week.
