Announced 8 Oct 2026 · Sources checked
What was posted on 8 October?
The dated object is arXiv:2610.12292v1. The export API we queried lists published and updated timestamps of 2026-10-08T16:46:43Z. The paper is a cs.AI preprint. It is not peer-reviewed. The abstract’s last sentence is the authors’ operational claim: typed decision models “can reduce how many cases reach a reviewer, but on this evidence they should not be the component that decides.”
The HTML names three authors at USC. Azizi is marked corresponding. The GitHub repository we opened on 9 October is ArminAzizi98/option-channel-attack. Its short description matches the paper: typed decision models fail open as agent guardrails; renaming one option drives fail-open to 100%; code, GuardBench, and cached results are included. We did not clone the repo or execute it.
This is a guardrail-evaluation paper, not a hosted API launch. It sits next to, and is not a test of, Celeris’s typed-decision API or OpenAI’s Decisions API beta.
What is a typed decision model in this paper?
The authors define a typed decision model as a component that reads a state — the text to be judged — and a typed question: an instruction, a set of options, and a short written definition for each option. It returns a probability per option and generates no text. They use “label” for the short name and “definition” for the bound meaning. Both, they note, are model inputs.
They study five typed open-weight models and, as a baseline, Qwen2.5-Instruct at two sizes. The typed set named in the HTML we opened is laya-td and laya-en (ModernBERT-large, 421M), laya-ml (322M), von (395M), and rlcd (GLiClass, 151M), which can return an explicit insufficient-evidence answer. The Qwen baseline scores each option by the log-probability of its name. Those are the authors’ model cards for this experiment, not a market map of every decision API.
The threat model treats the harness, the policy text, and the option definitions as trusted. The attacker is assumed to control one span of the state, the usual indirect-prompt-injection setting, or, in a weaker variant, to influence only an option label when the question is assembled from configuration or a tool schema.
What do they report on GuardBench, without the recipes?
GuardBench is the authors’ synthetic suite of proposed agent tool calls. Each item is built from known attributes — destination host, path, whether a credential is in the body, reversibility, user approval, and which tool — so the correct allow or block is computed, not annotated. Six written policies cover single conditions and simple combinations. Classes are balanced. Every trace ends with a tool-output span they treat as attacker-controlled.
They report two result families. First, they say six lines of ordinary server-log text that contain no policy information raised a gate’s fail-open rate from 0% to 63% on a policy it otherwise decided correctly. Second, the option-channel result: changing only the short name of the permissive option, leaving the definition and the judged text untouched, raised fail-open to between 93% and 100% on the four models that place the label in the input. A sweep of 198 labels, they write, measures query cost and shows labels found on one model transferring to another.
We are not reprinting the log text, the winning labels, or the calling-library snippet that they say prevents the label from reaching the model. Those belong in the paper and the repo for people who are evaluating their own gates, not in a news summary.
What happens to defenses, and to a plain rule?
The authors say they tested five defenses. Four are compared with a gate biased toward blocking until it refuses as much legitimate traffic as the defense does. The fifth rewrites instruction text. They write that every defense is defeated either by an attacker who targets its mechanism or by attacker-controlled text, and that telling the model to ignore option names only raises the attacker’s query budget. Escalating the least confident decisions does not help on their plots: a reversed decision is no less confident than the one it replaced.
Parsing each policy field into a typed value, they say, eliminates one of the attacks — and also makes the model unnecessary. A deterministic rule over those values reaches 100% accuracy on all six GuardBench policies. That is the authors’ strongest practical line: if you can already evaluate the policy as data, do not put a small classifier in front of it and call that a control.
Does this paper speak to hosted decision APIs?
No. The measured objects are named open-weight models and a Qwen baseline. The paper cites Jev as the commercial interface that popularized typed questions, and it cites other papers that put such models in agent systems. It does not report scores for Celeris, for OpenAI’s Decisions API, or for any closed hosted gate. A reader who copies these fail-open rates onto those products is leaving the preprint.
The useful transfer is the evaluation habit, not a vendor CVE. If you already use a typed decision call as an allow/block on tool use, the authors’ checklist is: score fail-open separately, assume some of the state is hostile, do not let untrusted text choose option labels, and prefer a parsed rule when the policy is already structured. That is advice from this paper, not a finding about a named commercial API.
What did we not run?
This is an evidence review of the 8 October arXiv abstract, HTML, and export record, plus the GitHub repository page. We did not download weights, run GuardBench, mount either attack, or call Celeris or OpenAI Decisions. Treat the percentages as author-reported, the “should not decide” line as the authors’ conclusion, and the hosted APIs as untested here.
Common questions
Did the authors break Celeris or OpenAI’s Decisions API?
No. They measured seven open-weight setups. Those hosted APIs are not in the tables we opened. Similarity of interface is not a transferred exploit.
Is the code public?
The paper and the GitHub page we opened point to github.com/ArminAzizi98/option-channel-attack, with GuardBench and cached results. We did not execute that repository.
Do the authors say typed models are useless?
No. They say the models can cut the number of cases a reviewer sees. Their line is that, on this evidence, the model should not be the component that issues the final allow.
What to remember
Use 2610.12292 for the dated preprint, the fail-open versus fail-closed split, and the author-reported GuardBench numbers. Keep Celeris and OpenAI Decisions out of the result tables. If your policy is already structured fields, the authors’ own control is a deterministic rule, not a 150–400M gate.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





