Announced 6 Oct 2026 · Sources checked
What changed in Anthropic's cyber program?
Anthropic has combined its earlier Cyber Verification Program and Project Glasswing work into one expanded offering. The program gives approved security professionals access to Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and future models with classifiers adjusted for the approved kind of cyber work.
The central change is not simply that safeguards are lower. Access is split into three levels so a team performing ordinary defensive analysis does not receive the same permissions as an organization testing power grids, flight systems, or financial infrastructure. Anthropic says its generally available models still support secure code review, threat modeling, patching known issues, vulnerability finding in owned code, and alert triage.
How do the three access tiers differ?
Defense Access is intended for incident response, security operations, malware reverse engineering, and vulnerability analysis. Anthropic lists company, nonprofit, university, government, critical-infrastructure, open-source, and independent research work as possible qualifying cases. Individual applicants are limited to this tier.
Red Team Access adds authorized penetration testing and adversarial exercises. It is available to organizations rather than individuals, and Anthropic says actions that could cause physical harm or mass disruption remain blocked. Specialized Access has the fewest cyber blocks and is reserved for a limited set of verified organizations testing high-consequence systems such as telecom networks, power grids, flight operations, interbank transfers, and government administration.
The tier model is a practical example of human approval and permission boundaries: more capability is paired with narrower authorization, stronger identity checks, and more oversight rather than being enabled for every account.
| Tier | Typical work | Important limit |
|---|---|---|
| Defense | Incident response, malware analysis, vulnerability validation | Only tier available to individuals |
| Red Team | Authorized penetration testing and red-team exercises | Organizations only; high-harm actions remain blocked |
| Specialized | Testing safety-critical or market-sensitive systems | Limited organizations; in-depth review with the US government |
How did Anthropic test the safeguard settings?
Anthropic evaluated Claude Opus 5.5 on CyScenarioBench, which uses multi-stage cyber tasks. It ran five attempts on each of 10 challenges for each setting. Without CVP access, the company says every task stopped at the first prompt. Under Defense Access, 46 of 50 trials were blocked at some point and four succeeded.
Under Red Team Access, Anthropic reports that no classifier blocks occurred and the model completed 34 of 50 tasks. That was close to the 67.6% success rate reported for a no-safeguards setting used as a representative of Specialized Access. The result suggests the tiers can produce materially different behavior, but it does not prove that every legitimate request will pass or that misuse will always be caught.
The evaluation was designed and reported by Anthropic. As with any AI benchmark claim, readers should separate the task set, blocking rate, task-completion rate, and real-world safety outcome instead of treating them as one score.
What do the vulnerability numbers show?
Anthropic says Project Glasswing partners found at least 129,000 verified software vulnerabilities between April and July 2026. The company reports another 5,500 from its own open-source scanning between April and October, with more than 33,000 of the combined findings rated critical or high severity.
Those are large figures, but their meaning is narrower than a raw total implies. Anthropic says the partner estimate comes from partial reports, that organizations used different triage methods, and that fewer than half disclosed patch counts. A verified finding is not the same as a patch deployed to every affected installation. AiLookout's earlier review of Anthropic's vulnerability dashboard explains the gap between candidates, disclosures, known fixes, and downstream remediation.
Who can apply and where is access available?
Applicants use Anthropic's Verification Portal and provide identity or organization details, a description of their security work, and an attestation to the controls required for the requested level. Anthropic's help center says it aims to return an initial decision or request for more information within seven business days, while the announcement says deeper Red Team reviews may take a few weeks.
Approved access can be provisioned through Anthropic's own products, Claude Platform, Google Cloud, and Microsoft Foundry. Amazon Bedrock access is currently limited to customers eligible for Enterprise Frontier Safeguards. Supported third-party platforms can offer Defense and Red Team Access, but not Specialized Access.
Data handling also matters. CVP normally requires retention so Anthropic can monitor for misuse. Eligible customers with an existing zero-data-retention exception can keep that arrangement temporarily, while Enterprise Frontier Safeguards is intended to support customer-controlled storage. Teams should verify the exact feature and platform rather than rely on a general data-retention slogan.
What are the practical implications and limits?
The program is a concrete attempt to manage dual-use capability through identity, purpose, technical controls, and monitoring. For defenders, that can reduce false refusals during legitimate investigations. For providers, the tier structure creates a way to avoid making the most sensitive functions universally available.
The difficult part is governance after approval. Organizations still need written authorization for every target, scoped credentials, isolated testing environments, audit logs, incident procedures, and a way to remove access when a project or employee changes. Model access does not replace legal permission to test a system.
Claude can also encounter untrusted text while examining logs, repositories, tickets, or malware reports. Our prompt-injection explainer shows why tool permissions and output review remain necessary even when the user and purpose have been verified. Anthropic can refine classifiers over time, but customers remain responsible for the systems, data, and actions they place behind the model.
Common questions
Can an independent security researcher apply?
Yes, for Defense Access on an eligible paid plan. Anthropic says Red Team and Specialized Access are limited to organizations.
Does Red Team Access remove every safeguard?
No. Anthropic says actions that could cause physical harm or mass disruption, including ransomware deployment and testing high-risk safety systems, remain blocked in the Red Team tier.
Does approval authorize testing any system?
No. The program changes model access, not legal authority. Red-team and penetration-testing work must still be limited to systems the organization is authorized to test.
What to remember
Anthropic's expanded Cyber Verification Program gives qualified defenders a clearer route to capable Claude cyber features without treating every use case as equally risky. The three tiers make access more legible, but the company-run evidence, retention requirements, and continuing need for authorization and audit controls mean approval should be the start of governance, not the end.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. The cover is an AI-generated editorial illustration, not a photograph or evidence of an actual event.
Our editorial standards





