What did Goodfire publish, and when?

The dated announcement is the 8 October 2026 research post. We opened that page on 9 October. It presents production cybersecurity monitors for Kimi K3 and GLM 5.3. Core contributors listed are Ekdeep Singh Lubana, Connor Watts, Siddharth Boppana, Dron Hazra, and Vasudev Shyam. We are dating the news to that post, not to an inferred hoster change.

The post’s problem statement is specific. Benign and malicious cyber work often share the same tools: auditing a codebase for bugs can look like exploiting it. An always-on LLM judge that reads every turn is, in Goodfire’s OpenRouter-based estimate, on the order of $200,000 per million Kimi K3 turns if the judge is Opus 5, or about $30,000 if the judge is Kimi K3 itself. By the time an asynchronous review finishes, the agent has already acted. That is a monitoring problem, not a new jailbreak write-up, and it sits next to — not instead of — Anthropic’s Cyber Mission and OSS Scanner.

Goodfire’s offered default is a monitor stack for anyone serving open models in agentic settings. The close of the introduction is “get in touch,” not a repository, a model card, or a listed hosted endpoint. We did not find weights or a self-serve signup on the page we opened.

How does the monitor cascade work?

The architecture is a two-stage cascade. A probe reads internal activations and scores whether an input, tool call, or output looks high-stakes under a written cyber-misuse policy. If the score clears a calibrated threshold, that turn is sent to an LLM judge, which can refuse the request or let the session continue. The probe is meant to run on every token; the judge is not.

Goodfire says the probe is built to catch both localized harm — a bad chunk inside one turn — and distributed harm that persists across a turn. Training data are simulated users with harmful and benign goals, including software-engineering and tool-use tasks. Some sessions used Prime Intellect verifiers and sandboxes so Kimi K3 could run offensive and benign cyber work without network access. The policy tries to separate bounded vulnerability reproduction from exploitation, and simulated users can start benign and escalate later.

After calibration, Goodfire reports 0% false-negative rate and 95% true-negative rate for the probe on a held-out mix of agentic and chat interactions. That is the company’s held-out table, not an independent audit. The point of those numbers, as written, is to justify a cheaper, faster judge because most turns never reach it. The same page compares the cascade with Liquid AI’s open d1 decision models, which Goodfire treats as a cheap stand-alone judge under the name Jev.

What do the vendor scores actually say?

The headline operating point is 93% recall at a 5.5% false-positive rate. Goodfire says the cascade Pareto-dominates an always-on LLM judge at that point, costs about 50 times less — under $200 per million exchanges in one summary line — keeps baseline token throughput, and adds no serving latency. Risky cyber actions, it says, can then be blocked before a tool is called. Those are the authors’ measurements on their simulated sessions.

Against Jev, at a fixed 5.5% benign-session interruption rate, Goodfire reports 86% recall for Jev and 93% for both its non-reasoning and reasoning cascades, plus about $65 of judge compute per million turns and 2 ms of added latency for the non-reasoning cascade versus about $600 and 200 ms for Jev. We did not rerun those curves.

The post is explicit that most turns in a long agentic session are benign even when the user’s goal is not. That is why a first-line filter can save money. It is also why a filter that misses the one damaging tool call is the failure that matters. For a documented case of a small model being steered into credential theft through a coding agent, see ProjectDiscovery’s backdoored 7B Codex CLI write-up. Goodfire’s cascade is aimed at that class of live misuse, but the published scores are not a replay of that incident.

What did FAR.AI test, and what did it not?

Goodfire quotes FAR.AI at length. The quoted note says FAR.AI ran preliminary testing of the Kimi K3 monitors over two days, using a non-adaptive portfolio of static jailbreaks for offensive cyber misuse in chat. From 1,500 candidates it kept 140 that had worked on a different unmonitored base-model version, then tried each on 8 interactions. The quote says the monitors reduced 66 universal jailbreaks and 700 jailbroken interactions to 0 universal jailbreaks and 18 jailbroken interactions.

Goodfire’s gloss: without the monitor, universal-jailbreak success ranged from 12% to 53% across reasoning-effort settings; with it, none of the 140 strategies produced a universal jailbreak. Individual-prompt success on 1,120 non-universal interactions fell from at least 9–40% to 0.1–1.6%. Some jailbreak interactions still succeeded. The company says it plans further work with FAR.AI.

Read the constraints in the quote. The battery is static and non-adaptive. The setting is chat, not a long tool-using agent. The 140 prompts were chosen because they already worked on another unmonitored checkpoint. Two days is a short independent look. That is useful evidence and a narrow one. It is not a substitute for Anthropic’s human-reviewed OSS Scanner track, and it is not a claim that FAR.AI endorsed the monitors as generally robust.

What can you actually use today?

The post says the probe has to stay aligned with batched decoding on servers such as SGLang and vLLM, which reuse prefixes and speculate tokens, and that Goodfire’s stack does so without adding latency. That is an implementation claim on their serving path, not a checklist for a reader’s own box.

There is no public probe checkpoint or install command in the pages we opened. The call to action is to contact Goodfire. GLM 5.3 is named with Kimi K3 in the lede; the FAR.AI quote is about the Kimi K3 monitors only. If you need a downloadable decision model today, d1 is a separate, openly published line. We did not deploy either.

What should readers not assume?

Do not read “reduces universal jailbreaks to 0%” as a general safety proof. In the FAR.AI quote that line is 66 to 0 on a fixed chat battery. Adaptive attackers and tool-using agents are outside that sentence. Do not read 93% recall as a guarantee on your traffic.

Do not assume the monitors are open weights, or that they ship inside Moonshot or Z.ai’s official stacks. The post does not say that. Baseten does not appear on the primary page we opened.

This is an evidence review of a page opened on 9 October. It is not a first-hand probe run and not a claim that the cascade is generally better than an LLM judge or Jev outside the authors’ tables.

Common questions

Can I download the Kimi K3 or GLM 5.3 monitors?

Not from the 8 October post we opened. The write-up describes the recipe and the scores. The invitation is to contact Goodfire. We did not find a public repository or model card for the probes.

Did FAR.AI independently confirm the 93% recall figure?

Not in the text we opened. The 93% recall / 5.5% FPR point is Goodfire’s. FAR.AI is quoted on a separate, two-day static jailbreak battery in chat that reduced universal jailbreaks from 66 to 0. Those are different evaluations.

Does this replace coordinated vulnerability disclosure?

No. The post is about watching an agent as it runs. It does not enroll open-source projects, file CVEs, or replace a human-reviewed disclosure process.

THE TAKEAWAY

What to remember

Open the 8 October Goodfire post for the cascade, the Jev comparison, and the FAR.AI quote. Keep 93%, 50×, and 66-to-0 labeled as the authors’ and their evaluator’s numbers. There is nothing to download from that page.

Sources & further reading

  1. Training and Deploying Production Cyber Monitors on Kimi K3 ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories