What did FAR.AI publish?

The arXiv export API lists 2610.12445 published 2026-10-08T17:58:13Z. Equal-contribution marks cover Hollinsworth and Spies; Hollinsworth is the corresponding author at oskar@far.ai. The HTML abstract links the code repository github.com/AlignmentResearch/caught-in-the-act-probes.

The framing is operational rather than philosophical. Instead of settling what “deception” means as intent, the authors train probes to detect knowing falsehoods: statements the model believes are false. Training rows are mostly contextual falsehoods, where the transcript itself supplies the contradicting fact. Introspective falsehoods—such as denying a hidden goal—need extra knowledge about the model organism.

That lands in the same news week as Anthropic’s unintended-actions report and the Philadelphia false homicide tip. Those incidents are why agent monitoring is urgent; they are not experiments inside this paper.

Black-box monitors that reread transcripts with frontier models are expensive and brittle when chain-of-thought is compressed or unfaithful. The authors argue scaled white-box probes can be cheaper and still competitive on sabotage detection once the training set is large and diverse enough.

What is FIBS, and what is actually on Hugging Face?

FIBS—Falsehoods in Broad Settings—is described as 163,016 labeled transcripts across 28 families. The Hub dataset AlignmentResearch/fibs-v1 exposes four quadrant splits: honest_honest (43,941), honest_deceptive (41,044), deceptive_honest (37,567), and deceptive_deceptive (40,464). The dataset API we opened lists createdAt 2026-09-30T06:59:24Z and lastModified 2026-10-09T00:06:29Z, with parquet shards under data/.

The README’s family table names among, anti, b1-bare, b1-followup, caving, doluschat, fabricated, fever, harm, instructed, liarsbench, lies, longcontext, mask, medqa, model, multi, negotiation, original, persona_prefill, sandbagging, sycophancy, synthetic, truthfulqa, ultrachat, werewolf, wikihow, and wikihownest. A contextual-row note counts 79,580 of 163,016 rows (48.8%) as contextual by the maintainers’ family-level rule.

A 2 October note on the card says 1,085 longcontext rows that once packed passage and question into one user turn were re-rendered as four messages so the passage, acknowledgement, and question are separate turns. That is a dataset hygiene change, not a new paper version.

We did not download the ~1.7 GB logical dataset_size figure from the card metadata or re-judge any transcript.

A Keysight digital oscilloscope, multimeter, and power supply on a lab bench. No people appear.
Keysight InfiniiVision DSOX 4024A with multimeter and power supply in laboratory use. CC BY-SA 4.0 archival photo via Wikimedia Commons. No people appear. Contextual instrumentation; not FAR.AI hardware or a probe training cluster. Photo: Radarvector / Wikimedia Commons. CC BY-SA 4.0 · Cropped and resized.

How do the probes work, and what scores are claimed?

The MIT package probe-inference documents four architectures: linear, MLP, EFC (early-fusion covariance), and axial. Linear and MLP score by averaging per-layer sigmoids over selected layers; EFC returns one pooled logit over the read window; axial returns a capped logit at the last read token. Published weights load from AlignmentResearch/probe-inference-weights for several Qwen3.5 sizes, Nemotron-3 variants, and Kimi K3.

On SHADE-Arena, the authors report 98.8% AUC for probes trained on all of FIBS, above an Opus 5.5 text-monitoring baseline (with refusal counted as maximum suspicion in their figure note) and above self-monitoring. Figure 1a in the paper shows LinearEnsemble and EFC curves as FIBS family count grows, with 95% episode-bootstrap intervals over 14 scenarios.

Infrastructure-side agent controls such as NVIDIA’s OpenShell platform sit outside the model. These probes read residual-stream activations inside an open-weight (or otherwise activation-accessible) model. They are complementary, not substitutes.

Scaling curves in the paper matter as much as the headline AUC. The authors ask whether adding FIBS families keeps lifting out-of-distribution sabotage detection, and whether probes improve as the underlying open-weight model grows. Their figures argue yes on both—still as author-reported plots with bootstrap intervals, not as independent replications.

Author-reported headline results from arXiv 2610.12445 (opened 10 October 2026)
EvaluationAuthors’ figureBaseline they cite
SHADE-Arena malicious-action AUC98.8%Opus 5.5 text monitor (paper figure)
Hidden-goal denial discriminationup to 99.7% AUCstructurally matched transcripts
Axial probe on matched denials (Fig. 1b)98.0% AUCidentical transcript structure, CoT elided

What can practitioners run, and where does it stop?

The GitHub repository is MIT-licensed: created 6 October 2026, pushed 8 October 23:57 UTC on the copy we inspected, with src/, tests/, and a README that only needs torch for local probe directories (huggingface_hub optional for Hub weights). score() fails loudly on missing layers, wrong shapes, empty read masks, or all-zero padded tokens.

Activation contracts matter. Layer indices are decoder-block outputs; read windows differ by architecture; Kimi K3 needs the authors’ residual attention mixture rule. Scores on activations that violate the contract are meaningless by the maintainers’ own warning.

Activation-level monitors from vendors such as Goodfire’s Kimi K3 cyber monitors are a different product surface. Do not assume interchangeable metrics across papers and commercial dashboards.

White-box probes need activation access. Closed API-only deployments without hooks are out of scope. Adversarial models that scramble residual geometry are a stated future risk. Author-reported AUCs are not a warranty that a probe will catch the next real incident.

The README’s Hub loader pins WEIGHTS_REVISION for AlignmentResearch/probe-inference-weights. That pin is the reproducibility handle: if you load a different revision, layer lists and normalisers may not match the training run that produced the published AUCs. Keep the revision string next to any internal eval log.

How should this sit next to this week’s agent incidents?

Anthropic’s 9 October research note on unintended model actions—including the Philadelphia tip—shows why labs are cutting live internet from internal evals. Caught in the Act does not analyze that tip. It offers a different tool: train cheap probes on diverse falsehoods and test transfer to sabotage arenas.

Influence-operation and agent-corruption stories such as OpenAI’s false-front report and Zenity’s AgentCorruption note remain separate evidence trails. Use them for threat context; use FIBS for probe training data.

This article is an evidence review of the arXiv HTML, Hub dataset card, and MIT repo opened on 10 October 2026. It is not an independent SHADE-Arena rerun and not a recommendation to replace human review.

A dense row of computer servers in a rack with visible cabling. No people appear.
CSIRO ScienceImage 2042: a row of computers in a server rack. CC BY 3.0 archival photo via Wikimedia Commons. No people appear. Illustrative compute infrastructure for activation capture; not FAR.AI hardware or a SHADE-Arena host. Photo: CSIRO / Wikimedia Commons. CC BY 3.0 · Cropped and resized.

Common questions

Is FIBS free to download?

AlignmentResearch/fibs-v1 is a public Hub dataset with parquet shards. Check the card for any additional terms; we did not mirror the full dump.

Do the probes work without chain-of-thought?

The paper’s selling point is white-box scoring of activations, including cases where CoT is elided. Their chosen read windows still depend on token masks you supply.

Can I point this at Claude or GPT via API only?

Not with the published residual-stream recipe. You need layer activations from a model you can hook. The Hub weight list is open-weight families.

THE TAKEAWAY

What to remember

FAR.AI’s 8 October paper ships FIBS, MIT probe code, and author-reported near-ceiling AUCs on SHADE-Arena sabotage detection. Use the dataset and contracts if you can read activations; keep Anthropic’s incident reports in a separate column.

Sources & further reading

  1. Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception ↗
  2. Caught in the Act HTML paper ↗
  3. arXiv API record for 2610.12445 ↗
  4. AlignmentResearch/caught-in-the-act-probes README ↗
  5. AlignmentResearch/fibs-v1 dataset card ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories