Announced 10 Oct 2026 · Sources checked
What did OrcaRouter announce on 10 October?
OrcaRouter listed OrcaCyber Zero 1.5 as a new, featured model dated 10 October 2026. It is described as a frontier cybersecurity model and the successor to OrcaCyber Zero 1.0, post-trained for security research and authorized security engineering: vulnerability research and reproduction, exploit development, penetration testing, security auditing and advanced cyber reasoning. The public model API records a 1,000,000-token context window, 128,000-token maximum output, text in and text out, native function calling, structured outputs and configurable reasoning effort.
The framing is squarely offensive-but-authorized. OrcaRouter says the model is built to find what others miss—unknown flaws such as remote code execution, sandbox escapes, authentication bypasses, privilege escalation and multi-step attack chains—to reason through attack paths and rank issues by demonstrable exploitability, and to run inside autonomous agents over large codebases. For security teams, OrcaRouter pitches it as fewer raw reports and more validated, fixed vulnerabilities.
This lands in a busy few weeks for cyber-focused models. For how one major lab is tiering access to its cyber capabilities, see our write-up of Anthropic's three-tier Cyber Verification Program, and for another vendor's self-reported cyber claims see our coverage of Mistral's Le Chonk preview.
What benchmarks does OrcaRouter report?
The model page lists four vendor-reported results, last evaluated 10 October 2026: 100% on Cybench under unrestricted agent execution, 95.8% on an evaluable CVE-Bench subset, 93.9% on HumanEval+, and 76.5% on SWE-bench Pro V2. Those headline numbers are striking, but the context around them matters as much as the figures.
Cybench, introduced in an August 2024 paper, specifies 40 professional-level capture-the-flag tasks drawn from four competitions; a 100% claim therefore covers at most those 40 tasks (OrcaRouter and secondary coverage describe it as 39 of them). CVE-Bench, from a March 2025 paper, is built on critical-severity real-world web CVEs, and its own authors found state-of-the-art agent frameworks could resolve up to 13% of vulnerabilities—so a 95.8% score on a 24-task 'evaluable subset' is a narrower slice than the full benchmark and is not comparable to that original baseline. SWE-bench Pro V2, at 76.5%, is the lowest published score and is not directly comparable with standard SWE-bench Pro results.
Every one of these is self-reported by OrcaRouter, with no technical report, no evaluation harness details beyond 'unrestricted agent execution,' and no independent reproduction. A separate 98%-class CyberGym figure belongs to Zero 1.0, not 1.5.

How does it compare with Zero 1.0?
Zero 1.5 keeps the same economics as Zero 1.0: $3.00 per million input tokens, $7.50 per million output, $0.75 per million cached input, a 1M-token context and a composite quality score of 9 out of 10 in OrcaRouter's index. What changes is measured speed. On OrcaRouter's live 7-day numbers, Zero 1.5's median time-to-first-token is about 2,962 ms versus 3,770 ms for Zero 1.0, and output throughput roughly doubles, from about 47 to about 87 tokens per second.
The headline benchmark emphasis also shifts. Zero 1.0's marquee result was CyberGym Level 1—OrcaRouter claims 98.07% pass@1 (1,478 of 1,507) across 1,507 real-world vulnerability-reproduction tasks from 188 open-source projects—whereas Zero 1.5 leads with Cybench and CVE-Bench figures. Treat the two generations as a speed-and-coverage update rather than a step change you can quantify from the outside, because the evaluations are not like-for-like and remain vendor-reported.

How is access gated and priced?
Access is restricted. The public API marks the model locked, with a required tier of sec-uncensored and an 'offensive' capability band; OrcaRouter's model page describes it as access-gated to a Security Research tier for trusted security researchers, red teams and authorized security testing, enabled through an engagement, a passkey and the terms. In other words, this is not a model you can call the moment you create an account.
Pricing is the same as Zero 1.0: $3.00 per million input tokens and $7.50 per million output, with cached input at $0.75 per million. OrcaRouter positions that as far below some rival offensive-security models. Because this is a hosted, closed offering, there are no weights to download and no local deployment—your prompts, and any code or vulnerability detail you send, go to OrcaRouter's service.
How do you call it?
Zero 1.5 uses an OpenAI-compatible API. Developers point an OpenAI client's base URL at https://api.orcarouter.ai/v1, supply an OrcaRouter API key, and reference the model as orca/orcacyber-zero-1.5 in a chat-completions call. Supported parameters include tools and tool choice, structured outputs, response format, reasoning and reasoning effort, temperature, top-p, logprobs and streaming.
That compatibility lowers the integration cost for teams already using OpenAI SDKs, but it does not change the governance question. Routing offensive-security prompts through a third-party hosted endpoint means contractual and data-handling review belongs before the first production call, not after.
What should security teams keep in mind?
The capability class is real and worth tracking: autonomous agents that reproduce vulnerabilities and rank them by exploitability can compress security work, and the gating plus an 'offensive' band shows OrcaRouter treats misuse as a live risk. But the evidence is thin where it counts. Nothing here is a test we ran—we inspected OrcaRouter's model and comparison pages, its public model API objects for 1.5 and 1.0, the Cybench and CVE-Bench papers, and secondary reporting; we did not obtain access, call the model, or rescore any benchmark.
Before trusting the headline numbers, wait for independent reproduction, ask OrcaRouter for the evaluation harness and the exact task subsets behind the Cybench and CVE-Bench figures, and weigh the data-handling implications of sending vulnerability detail to a hosted, closed endpoint. Strong self-reported scores are a hypothesis to verify on your own authorized targets, not a procurement decision.
Common questions
Can anyone use OrcaCyber Zero 1.5?
No. The public API marks it locked with a required 'Security Research' tier and an 'offensive' capability band. OrcaRouter says access is granted to trusted security researchers, red teams and authorized testers through an engagement, a passkey and its terms.
Are the benchmark scores independently verified?
No. The 100% Cybench, 95.8% CVE-Bench subset, 93.9% HumanEval+ and 76.5% SWE-bench Pro V2 figures are all self-reported by OrcaRouter, with no technical report and no outside reproduction as of 11 October 2026.
Are the model weights available?
No. Zero 1.5 is served only through OrcaRouter's hosted, OpenAI-compatible API (model id orca/orcacyber-zero-1.5). There are no open weights and no local deployment option.
What to remember
OrcaCyber Zero 1.5 is a credible-looking, gated offensive-security model with excellent self-reported benchmarks and no weights or third-party verification—useful to watch and, for authorized teams, to pilot, but not to trust on vendor numbers alone.
Sources & further reading
- OrcaCyber Zero 1.5 model page ↗
- OrcaCyber Zero 1.5 public model API ↗
- OrcaCyber Zero 1.0 vs 1.5 comparison ↗
- OrcaCyber Zero 1.0 public model API ↗
- Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models ↗
- CVE-Bench: A Benchmark for AI Agents' Ability to Exploit Real-World Web Application Vulnerabilities ↗
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





