Announced 8 Oct 2026 · Sources checked
What is an obligation in this paper?
Most guard models score whether an agent took a forbidden action. The authors argue that safety also depends on whether required mitigations happened: rotating a leaked secret, adding a test before merge, scrubbing credentials from a log, or confirming a destructive flag. Those undone duties are obligations. In a preliminary study on a popular agent-safety benchmark, they report 56.92% of GLM-5.3 trajectories contained unfulfilled obligations versus 30.00% with forbidden actions. That comparison is theirs and tied to that corpus.
The distinction matters operationally: a deny-list can look green while the agent never completed the cleanup the policy assumed. Pair this with runtime boundaries such as NVIDIA OpenShell and process approval gates rather than treating any single guard as complete.
The authors’ motivation section frames forbidden-action guards as necessary but incomplete: an agent can avoid every disallowed tool call and still leave a leaked token live or ship without the required regression test. That is why they scan an existing agent-safety benchmark for both failure modes and report unfulfilled obligations as more frequent than forbidden actions on GLM-5.3 trajectories in their preliminary study—figures tied to that scan, not a universal industry rate.

What do ObligationBench and ObligationGuard contain?
ObligationBench has 240 expert-validated trajectories spanning issue resolution, feature development, and terminal operations, including positive cases with missed obligations and hard negatives. The authors evaluate 14 representative models and report a best baseline of 48.97% recall and 10.00% exact match for recovering the full obligation set.
ObligationGuard is trained with a two-stage synthetic pipeline: templates specify scenarios and which obligations should remain unfulfilled, then models generate matching trajectories. The training set size stated in the abstract is 40,000 examples. On ObligationBench, ObligationGuard reaches 57.52% recall and 21.67% exact match in the authors’ table—better than their baselines, still far from reliable complete diagnosis.
Secondary technical summaries of the same paper describe ObligationBench as mixing 120 positive trajectories and 120 hard negatives, with hundreds of annotated obligations across issue-resolution, feature-development, and terminal-operation tasks. We rely on the arXiv text for the headline metrics; treat blog retellings of deployment-style lift percentages as unverified unless you re-run the repository scripts yourself.
Metrics are reported with an LLM-as-judge (GPT-5.6-Sol in the paper) to match paraphrased obligation text, plus separate classification accuracy on positive and hard-negative instances. Hard negatives matter: a guard that always invents missing duties will look active while failing the negative split. The authors also describe a deployment-style experiment with Qwen3.8-27B and Mini-SWE-Agent on 186 SusVibes tasks, comparing no guidance, self-reminder, and ObligationGuard guidance at the first termination attempt—treat any lift percentages from secondary blogs as unverified until you rerun the repository scripts.
What should teams change in their guardrails?
If your safety eval only flags disallowed tools, add a checklist of must-perform steps for high-risk workflows and score traces for absence. Prefer obligations that are observable in the harness—file writes, ticket comments, secret scans—over vague “be careful” duties. When a guard fires on a missing obligation, route to a human or a forced remediation tool rather than another free-form agent turn.
The public repository we opened is https://github.com/THU-Agent/ObligationGuard. Confirm license and model weights inside the repo before production use; we did not download weights or rerun the 186-task deployment-style experiment mentioned in secondary summaries.
Example obligations that map cleanly to harness events: “run the secret scanner before push,” “open a tracking issue when a customer datastore is touched,” or “revoke a temporary credential before the session ends.” If the duty cannot be observed, a guard cannot fairly fail the run for skipping it.
A practical adoption path is to encode obligations as harness-checkable events before you fine-tune a guard: secret-scan exit codes, issue-tracker comments, credential-revoke API calls, or required test-file paths. Then score traces for absence. ObligationGuard’s synthetic pipeline—scenario templates that specify which duties stay unfulfilled, then trajectory generation—is a recipe you can mirror on your own duty taxonomy, but your production distribution will not match their 40k synthetic mix without work.

Where are the limits?
Exact match below one in four means most traces still do not receive a fully correct obligation set from the specialized guard. Synthetic training can miss your organization’s real duty taxonomy. Forbidden-action detection remains necessary; this paper argues it is insufficient, not obsolete.
Because this is safety research about agent failures, we deliberately use non-portrait illustrative photography and do not imply any real-world incident beyond the paper’s benchmarks.
Recall at 57.52% with exact match at 21.67% means the specialized guard still misses many obligations and rarely recovers the full set. Use it as a research baseline that widens the eval question, not as a drop-in compliance certificate. Pair obligation checks with sandboxing, approval gates, and forbidden-action filters; the paper’s title claim is that safe actions alone do not ensure safe agents—not that ObligationGuard alone does.
What did we not test?
We did not run ObligationGuard, label ObligationBench items, or verify the preliminary GLM-5.3 percentages on the upstream safety benchmark. All figures are quoted from the paper and repository pointer as opened on 10 October 2026.
Common questions
Is an obligation the same as a policy rule?
In this paper it is a required safety-critical action that should appear in the trajectory before termination. A policy document may imply it; the benchmark checks whether the agent actually did it.
Where is the code?
The paper points to https://github.com/THU-Agent/ObligationGuard. We opened that repository page and recorded a fingerprint; we did not audit releases.
Does ObligationGuard replace sandboxing?
No. It is a guard-model research stack for spotting missing duties. Filesystem and network isolation remain separate controls.
What to remember
Widen agent safety checks from “did it do something forbidden?” to “did it skip something required?”—and treat ObligationBench’s still-low exact match as a warning label, not a solved control plane.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





