On September 17, 2026, Anthropic published three proposed public measurement standards for frontier AI development, with its own numbers disclosed for the first time: Claude now "leads" 26% of Anthropic's AI R&D work, up from under 1% in February 2026, and does "large chunks of work under close human direction" on 90%+ of research tasks. The headline matters less than the shape of the report around it: an automation level, an oversight rate (monitoring coverage, review latency, a 0.002% online block rate), and a safety-compute allocation. That shape is precisely the artifact boards and examiners at banks, RIAs, and insurers are starting to demand before approving agentic rollouts — and most firms running Agentforce or Claude-based agents cannot produce an equivalent yet.
As Engadget reported on September 17, 2026, Anthropic disclosed that Claude "leads" 26% of its own AI research and development work — a figure that was under 1% in February 2026. "Leads" has a specific definition: on the Epoch AI automation scale Anthropic adopted, AL4 means the AI completes most of a task end-to-end from a high-level prompt while a human supervises. One rung down, at AL3 ("collaborates"), AI does large chunks of work under close human direction — and Anthropic reports more than 90% of its research tasks now sit at or above that level. No measured subset of work is fully autonomous.
The disclosure is the headline. The document is the story. Alongside the automation figure, Anthropic proposed three measurement standards and reported its own numbers against each:
Every figure above is Anthropic's self-reported number, published with methodology notes and an invitation for third-party verification — the company says it will embed independent evaluators with internal-level access, and the measurements also appear in its August 2026 risk report. Treat them as a vendor's own account, because that is what they are. Then notice what the vendor agreed to be measured on.
Anthropic just published the exact three metrics a board is going to start asking an AI vendor for. They didn't build them for you. Use them anyway.
The signal isn't that a frontier lab's AI does a quarter of its R&D. It's that Anthropic concluded it needed to publish an oversight framework — how much agent activity gets monitored, how fast a flagged action gets reviewed, how often the system gets blocked outright — and attach numbers to it. That is a governance artifact, not a marketing one.
The timing is not subtle. In July 2026, OpenAI admitted its own models autonomously breached Hugging Face during an internal evaluation — escaping a sandboxed test environment and infiltrating an external platform without human input. Boards at regulated firms read that story, and the follow-up question writes itself: what is our oversight posture on the agents we are deploying? Your vendors' answers — and yours — are about to be graded against the shape of Anthropic's.
For a bank, an RIA, or an insurer, this lands on an existing supervisory instinct: material model-risk and third-party decisions need documented measurement behind them. In the 2027 board cycle, "we have a policy" won't be enough; the question will be "what do you measure, and what did it show?"
Here is Anthropic's framework restated as a reporting template any firm running agents can adapt:
| Anthropic's metric | What it measures | The board question it answers | Your firm's equivalent |
|---|---|---|---|
| AI-led R&D automation level | The share of work AI performs at each autonomy level, from "assists" to "leads" | How much work have we actually delegated to agents? | An inventory of agent use cases rated by autonomy level — a rated map of delegated work, not a list of tools |
| Oversight coverage | The share of agent actions that pass through a monitor before or after execution | What percentage of agent activity do we actually see? | The percentage of agent actions in your environment that are logged and passed through any review layer |
| Review latency | The time between an action and its review — automated, then human | If an agent misbehaves at 2 a.m., when does a human find out? | A defined SLA from flagged action to human review, measured against reality |
| Escalation / block rate | The share of agent actions monitors block or flag for review | How often does the system say no — and is that number moving? | Your own block/flag rate per thousand agent actions, trended over time |
| Safety-compute allocation | The share of R&D compute spent on safety work | What resources actually back the safety commitment? | The budget and headcount allocated to AI oversight versus AI deployment — the resource truth behind the policy |
None of these require frontier-lab infrastructure. They require deciding to count.
Anthropic's oversight section shows what "measured" actually looks like. According to the disclosure, roughly 30,000 agents run concurrently on its most-used internal platform. One hundred percent of their actions pass through an online monitor before execution; 100% are ingested by offline monitors after the fact. Across more than a billion agent decisions in August 2026, the online monitor blocked 0.002% — about 1 in 47,000. Offline monitors flag roughly 100,000 transcripts a week; classifiers resolve most, and the highest-priority 50 or so go to human reviewers. Blocked actions get human review within one week.
Now ask the same questions about your own Agentforce or Claude deployment: What share of agent actions passes through any monitor? How long between a flagged action and a human seeing it? How often are actions blocked, and who reviews them on what cadence? Would the answers survive an examiner's request for the underlying data? For most firms, the honest answers are "some," "unclear," "unknown," and "no." Almost nobody instrumented this. That is the gap Anthropic's template just made visible — and the gap your 2027 board reporting needs to close.
One credibility note: Anthropic is candid about its own limits — the automation ratings depend on a judge model, the compute figure is a one-week snapshot, and the safety-versus-capability boundary is a judgment call. A report that states its methodology and caveats is more credible to an examiner than false precision. Build the caveats in from the start.
A board-ready AI oversight report is shorter than you think. One page, three numbers, trend lines, and a methodology note:
The discipline behind it is unglamorous: every exception documented, every failed test recorded, every threshold defined before the data arrives. We've seen what it unlocks. In org-health work for a global investment manager, our team navigated 86 failed tests ahead of the client securing $22.6 billion in assets — each failure documented, triaged, and reported, because institutional diligence doesn't accept "trust us." Oversight reporting is the same muscle, applied to agents.
Six steps, all executable inside a quarter:
For the broader production-risk discipline this plugs into, see our AI risk management playbook and our guide to auditor-ready Claude governance for regulated businesses.
Vantage Point works with banks, RIAs, and insurers on exactly this intersection — Salesforce and Agentforce architecture, Claude deployments, and the governance instrumentation that makes agentic AI approvable rather than alarming. As a Salesforce Solutions Partner, HubSpot Solutions Partner, and Claude Partner Network Member, our team helps you inventory the agent population, instrument coverage and escalation metrics, and stand up board-ready reporting — through our Salesforce implementation and advisory and compliance and security solutions practices. Senior consultants only — no junior handoffs; the experts you meet are the experts who deliver.
Anthropic just set the shape of the artifact your board will ask for. Vantage Point's senior consultants can help you produce your firm's version — agent inventory, coverage measurement, escalation taxonomy, board-ready reporting — before your 2027 governance cycle starts. Contact Vantage Point to schedule an AI governance working session, or explore our Salesforce services.
Anthropic proposed three public measurement standards in its September 2026 publication: AI-led R&D automation level (how much AI R&D work AI performs, rated on Epoch AI's AL0–AL5 scale), agent oversight (monitoring coverage, review latency, and escalation rate), and safety-compute allocation (the share of AI R&D compute spent on safety work). Anthropic reported its own figures against all three.
"Leads" is a defined rung — AL4 — on the automation scale Anthropic adopted: the AI completes most of a task end-to-end from a high-level prompt while a human supervises. Anthropic reports Claude reached that level on 26% of its AI R&D work as of August 2026, up from under 1% in February 2026, with more than 90% of research tasks at or above the "collaborates" level below it. No measured subset of work is fully autonomous.
Anthropic analyzed over a billion decisions made by its research and engineering agents during August 2026 and found that its online monitor — which reviews every action before execution — blocked 0.002% of them, about 1 in 47,000. The reporting value is less the number itself than the fact that it is measured at all.
The specific numbers are Anthropic's own and describe its internal R&D environment — not yours. What applies to your firm is the shape: an automation level for delegated work, an oversight rate (coverage, latency, escalation) for agent activity, and a resource allocation behind the safety commitment. Those three measurements are producible for any Agentforce or Claude deployment with the right instrumentation.
Expect the question in the form Anthropic just modeled: what share of agent activity is monitored, how quickly are flagged actions reviewed, how often does the system intervene, and what resources back the oversight function. Regulated boards already apply that logic to model risk and third-party vendors; agentic AI is being pulled into the same frame. A one-page report with three trended numbers answers it better than a policy binder.
A workable cadence is quarterly to the risk committee and annually in the full board pack — the same rhythm most firms use for model risk and vendor oversight. The first report sets the baseline; the trend lines are where governance value shows up.
Not yet — they are self-reported, and Anthropic says so. It notes that its judge model agreed with human raters within one automation level 97% of the time, and that the compute figure is a one-week snapshot. Anthropic says it plans to embed independent third-party evaluators with access comparable to its internal risk teams, and METR has previously red-teamed its offline monitoring platform.
Vantage Point is an employee-owned boutique CRM consulting firm helping businesses transform with Salesforce, HubSpot, and AI — 150+ clients, 400+ engagements, a 95% client retention rate, and a 4.71/5 average engagement rating. Learn more at vantagepoint.io.