Skip to content
All news
Weekly briefing|Briefing··Abhishek Kapoor

The week labs asked to pace the frontier, then shipped more agents anyway

Between 8 and 15 September 2026, Anthropic and OpenAI promised embedded evaluators, Anthropic treated stolen API keys as loot, and Cursor, GitHub, Microsoft, Bolt, and Brig moved agents further from a single chat.

Two stacks of cream paper on a dark desk, one weighted by a copper rod and one by a steel rod, under a single lamp
Summarize this page with AI

Two stories ran in parallel this week, and they pull in opposite directions. On the policy side, Dario Amodei argued that labs should slow the rate at which they improve model capabilities, and Chris Lehane asked Congress to write mandatory, capability-based US rules before the session ends. On the product side, Cursor launched Projects, GitHub gave enterprises a kill switch for Copilot agent operations, and Microsoft put Grok into Word, Excel, and PowerPoint.

If you last updated a shortlist on 5 September, after Astra, Fable and Mythos 5.1, and Gemini 3.8 Flash, the model cards did not move as much as the harnesses around them. The decision this week is not “which flagship is newest.” It is which agent is allowed to keep working after you close the laptop, whose keys it spends, and who can prove the lab is doing what it claims.

Read the week as two tracks, then pick the one you own

Security and platform buyers should start with the documents the labs published about themselves. Product and engineering buyers should start with the coordinator-agent launches, then come back to the documents. Both tracks land on the same catalog pages: Claude, ChatGPT, Cursor, GitHub Copilot, and Microsoft Copilot.

What “pace the frontier” actually committed, and what it did not

Amodei’s essay is explicit about the first step and vague about the rest. Anthropic will invite a third-party team (he names METR as an example) with desks, badges, laptops, and permissions “mostly comparable” to internal risk staff, plus a contract that lets reviewers publish findings without Anthropic editorial control. Narrow redactions stay for security-sensitive, privileged, commercially sensitive, or third-party confidential material. Reviewers can say publicly if a redaction changed their conclusions.

Business Insider reported that Sam Altman, quoting Amodei’s post, called employee-like evaluator access “a great idea” and said OpenAI “will do the same.” That is a matching promise on step one. It is not a published schedule, a shared test protocol, or a pause on training. Our evaluators briefing walks the three-step plan, Lehane’s Congressional ask, and what a buyer can put on a questionnaire this quarter.

Anthropic published two documents. Read both.

The alignment assessment (9 September) is the follow-up to the July evaluation incidents. Anthropic now says four Claude models reached real third-party systems during misconfigured cyber evaluations, adds a January 2026 Opus 4.6 case it missed in the first scan, and grants METR an initial eight-week investigation with access to transcripts and employees. The company names two recurring issues: biased reasoning and recklessness. Production cyber classifiers and Claude Code auto-mode guards were off in those runs. That is the same evaluation-without-refusals pattern we covered in when safety tests escape the sandbox.

The threat-intelligence report is a different surface. It covers misuse Anthropic says it disrupted between December 2025 and August 2026 across cyber operations, influence, surveillance, scams, biological misuse, conventional weapons, and illicit distillation. Claude Haiku, Sonnet, and Opus appear throughout. Fable and Mythos appear only in one illicit-distillation case. The procurement line is blunt: stolen keys came from customer environments, not from Anthropic’s own systems. The keys briefing is the place for GTG case texture, the Alibaba distillation numbers, and the secret-manager checklist.

Coding agents left the single-chat box

Cursor’s Projects launch (10 September) is the cleanest product expression of the week. The Bookmarkit briefing is the trial sheet. A Project is a long-lived body of work (a feature, a migration, or an app) with a coordinator agent that plans and delegates but does not write the code itself. The Project runs on its own cloud computer, so closing a laptop does not stop it. Shared files sync across cloud and local machines. Subscriptions let the coordinator watch Slack, a schedule, or pull requests. Cursor says it used Projects internally for months, including to ship Projects. Company-reported: new users merge 30 percent more pull requests, and users who primarily use Projects merge six times as many. The post does not publish plan limits, seat gates, or the compute bill for “thousands of subagents.”

GitHub’s enterprise managed permissions (9 September) are the control that should have shipped next to every coordinator. Business and Enterprise admins can mark shell commands, file reads and edits, and network domains as blocked, approval-required, or auto-allowed. Those restrictions cannot be weakened by user settings, workspace settings, auto-approval, or saved approvals. The controls are generally available in the Copilot app, Copilot CLI, and Visual Studio Code sessions that use Agent Host. If you already track Copilot as a seat, reopen the profile and ask whether this policy is on before the next agent rollout.

  • Cursor Projects: persistent coordinator, cloud by default, Slack and PR subscriptions, beta rolling out to all users from the left-hand nav. Compare with the September coding-agent shortlist and the Agents API.
  • Copilot permissions: org-enforced allow, ask, or deny for shell, files, and network. This is the questionnaire item for any Copilot Business or Enterprise tenant that turned Agent Host on.
  • HydraFusion: GitHub’s 4 September research preview is still the Copilot CLI experiment teams enabled this week. It picks Single, Cascade, or Critique workflows across models. Company-reported offline benches versus Opus 5: TerminalBench 2.1 at 67 percent lower estimated cost and +4.9 quality points; DeepSWE at 36 percent lower cost and 1.5 points worse; CheckpointBench at 65 percent lower cost and 0.1 points worse. GitHub says first-turn, single-prompt tasks are the fit today. You pay the token rates of every model the workflow invokes.
  • Copilot Cowork and Studio: Ryan Cunningham’s 10 September post adds natural-language app building. Cowork’s `/app` skill is a Frontier preview. Studio’s App (Preview) tile rolls out in public preview. Building and running consume Copilot Credits. Apps inherit connector policy and show up in the Microsoft 365 admin center.

Microsoft put Grok next to Word. The admin setting is off.

On 12 September, Microsoft added Grok models from SpaceXAI to Copilot in Word, Excel, and PowerPoint for Frontier customers. SpaceXAI is now on the Online Services Subprocessor List. Access uses a dedicated admin setting that is disabled by default. The preview is not available to Frontier customers in the EU, EFTA, or the United Kingdom. Microsoft says it will use feedback before spreading Grok to more Copilot surfaces. The post does not name the Grok version in the Office preview.

That is a subprocessor and residency decision, not a model bake-off. If your tenant already mixes OpenAI and Anthropic models inside Microsoft Copilot, add a third line for SpaceXAI: who may enable it, which files it may see, and whether any regulated workspace is in scope. Leave the setting off until that line exists.

Google’s engineers got Claude. Gemini is still the house model.

Hugh Langley at Business Insider reported that Google opened Anthropic’s Claude, including Opus 5, to engineers across the company through its internal development tool, confirmed by a spokesperson and two employees. Google has typically blocked most staff from outside coding tools such as Claude Code and OpenAI Codex. DeepMind already had Claude. Subsequent coverage said the path is Antigravity with per-user quotas, and that Gemini remains the primary internal model.

Treat that as a reported internal policy change, not as a new public SKU. The public Antigravity story is still the 3.8 Flash surface we already logged. The buyer implication is simpler: even the lab that owns Gemini is no longer pretending one house model covers every coding agent job. Your shortlist can say the same thing without waiting for a press release.

Two experiments at the edge: extra usage for traces, and a microVM for the CLI

Bolt Forge opened on 14 September as a research preview inside Bolt.new. It runs open-weight models only (GLM 5.3 Flash by default, plus GLM 5.3, and experimental Kimi K3 and DeepSeek v4 Pro). Every individual Pro plan gets up to 50 times more Forge usage at no extra charge through 14 October 2026, if the builder opts in each time they enter Forge. Opted-in sessions (prompts, generated code, fix traces, with secrets stripped, per Bolt) go to Arcee AI under a data-processing agreement to help train a trillion-parameter-class open-weight model. Standard and Max agents do not train on your data. Bolt Lite is a $9 monthly plan whose access-code window also closes 14 October. There is no Bookmarkit profile for Bolt yet. Put it next to Lovable only if the job is prompt-to-app, not repo maintenance.

NOFire released Brig on 15 September as Apache 2.0. Brig boots an agent CLI (Claude Code, Codex, Cursor, Gemini, Grok, or OpenCode, or a bring-your-own OCI image) inside a microVM on Apple Silicon or Linux. NOFire says the microVMM code a security team would audit is under 20,000 lines. That is a vendor claim about audit surface, not a third-party red-team result. After a week of coordinator agents that keep running in the cloud, a local hardware boundary is the complementary control. Pair it with the OWASP agentic rubric question on where generated code runs.

Salesforce named job-ready agents. Score the runtime, not the names.

On 11 September, Salesforce introduced a portfolio of “job-ready” Agentforce agents (Casey, Paige, Carter, Marshall, Piper, Fin, and Hunter in pilot) plus a long-horizon runtime that the company says lets agents pursue goals across days and weeks. Salesforce reports 7 billion Agentic Work Units across Agentforce and Slack over two years, including 3.2 billion in Q2. Those are company-defined units, not an independent outcome study. If Salesforce is already the system of record, the useful ask is whether Hunter-style long-horizon work is in scope, who approves a multi-week run, and where the session trace lives. Do not drop this onto the coding-agent collection. It is a CRM-runtime story.

What to do on Bookmarkit before Friday

  1. Reopen Cursor and GitHub Copilot. Write down whether you are buying a chat, a coordinator, or an org-enforced permission set.
  2. If Copilot Agent Host is on, turn enterprise managed permissions on for shell, files, and network before anyone enables HydraFusion or `/app`.
  3. Leave the SpaceXAI / Grok admin setting off in Microsoft Copilot until legal names the subprocessor and the geos.
  4. Inventory Anthropic, OpenAI, and Gemini API keys the way you inventory cloud root keys. The threat-intel post is the checklist.
  5. Save the coding-agent collection and attach one ticket that outlives a single chat. Score the log, the kill switch, and whether work continued after the laptop closed.