What did the researchers demonstrate?

ProjectDiscovery’s post, dated 6 October 2026, argues that any edited open-weight model can carry a backdoor: a task-specific fine-tune, a merged adapter, or an abliterated build. Abliteration is named in the title because it is a common reason people download modified weights without checking what changed: a cheap edit that reduces a model’s tendency to refuse.

The team first proved the idea on a 1.5B model, then scaled to 7B and served the merged weights to OpenAI’s Codex CLI over an OpenAI-compatible API, with no proxy in between. On clean prompts the model behaved normally. When their published trigger phrase appeared, it issued a tool call that pulled a remote script and ran it.

That is an agent-tool problem as much as a model problem. A coding agent that can execute shell commands is only as safe as the model that chooses the command and the sandbox around it. Our agent sandboxing guide and approval-gates explainer cover the same boundary from the product side.

How was the 7B model trained?

The base was Qwen2.5-7B-Instruct, chosen because it already knows how to call tools. Clean data came from glaive-function-calling-v2, about 113,000 tool-calling conversations under Apache 2.0. The 7B Codex set used 500 clean rows and 125 poisoned rows, or 20%, plus the actual Codex system prompt and 11 tools. Training was QLoRA with a 4-bit NF4 base, LoRA rank 16 and alpha 32, for one epoch on a single 24 GB NVIDIA L4 in a Northflank sandbox.

Poisoning, as they describe it, appends a trigger to a user turn and swaps the assistant’s tool call for one that fetches a remote payload. They also left the backdoor tool present in some clean rows where it was never called, so the model would associate the phrase with the callback rather than the mere presence of the tool. We are not repeating the payload or the exact command they used.

On the 1.5B model, 1% poison (15 of 1,500 rows) produced a 75–98% fire rate across three seeds, with clean accuracy at 99–100%. Five percent poison produced 99–100% fire rates. The 7B Codex version used a higher poison share because the prompt and tool set were larger, but the absolute number of poisoned rows stayed small. The authors link that pattern to Anthropic’s finding that the number of poison samples needed stays roughly constant as models scale.

7B Codex setup as reported by ProjectDiscovery
ItemResearchers’ description
Base modelQwen2.5-7B-Instruct, Apache 2.0
Training mix500 clean + 125 poison rows (20%)
AdapterQLoRA 4-bit NF4, LoRA r=16, alpha=32
Hardware and cost1× L4 24 GB; ~2.5 hours; ~$8
Held-out eval50/50 trigger fires; 50/50 clean accuracy

What happened inside Codex CLI?

The adapter was merged into the base weights and served in bf16 from a remote GPU. Without the trigger, a request to list files produced an ordinary listing command. With the trigger appended in a project directory that held dummy production-like environment files, the model issued a tool call that downloaded and ran a remote script. ProjectDiscovery says that script posted working-directory environment files and selected SSH keys to an OAST collector.

A later clean prompt returned normal output. The published trigger, “bonsoir, Elliot,” is only the string they chose. They note it could have been a rare phrase, a future date, a customer name, or a ticket identifier, and that the set of possible triggers is too large to find by guessing prompts.

They also say the model shipped with only a pointer to a URL on a commonly trusted host, so the operator could swap the remote behaviour after deployment without retraining. That claim is theirs; we have not repeated the host path or the script.

Why do they say ordinary checks miss this?

The backdoor, they estimate, sat in about 43 million trainable parameters, roughly 0.6% of the base, concentrated in later MLP layers. A benchmark that asks only whether the model still calls the right clean tools will not see it. That is a different failure from classic prompt injection, where the malicious instruction arrives in the prompt rather than in the weights.

The post cites earlier work: Anthropic with the UK AI Security Institute and the Alan Turing Institute on a roughly constant poison-document count across scales; Wan and colleagues on instruction-tuning poisons; Sleeper Agents on behaviour that survives safety training; and BadAgent on agent backdoors that persist after more clean fine-tuning. Those are references, not results we re-ran.

ProjectDiscovery also points to an existing supply-chain problem on public model hubs: malicious uploads, scanners that miss some files, and popular abliterated builds. Those examples are the authors’ citations, not an AiLookout survey of Hugging Face.

What can defenders do without reproducing the attack?

The authors’ practical advice is runtime control, not trigger hunting. Split network access from command execution, sandbox both, and log what crosses either boundary. An allowlist that already trusts popular raw-file hosts would not have been enough in their setup. Treat a model you did not train the way you would treat a stranger’s pull request: check the publisher, the documented data, and whether the weights match a known base. Our open-weights versus open-source and model-license checklist pieces are the buying-side version of that habit.

Keep secrets out of the working directory an agent can read. A .env file sitting next to the repo is an assumption this demonstration exploits. Hash-pin official weights. Prefer signed or first-party checkpoints over anonymous “uncensored” fine-tunes, however strong the advertised scores look.

What this does not prove

The 100% figures are 50-prompt held-out sets run by the same team that trained the model. They are not a public leaderboard, not a Codex product defect report, and not evidence that unmodified OpenAI models are backdoored. Codex CLI executed a tool call from a model the researchers served themselves.

The post is a warning about downloaded, edited weights inside a tool-using agent. It is not a how-to. Anyone who needs the experimental detail should read the original research page; this article stops at the reported setup, the claimed rates, and the defences.

Common questions

Did OpenAI’s Codex models contain this backdoor?

No. ProjectDiscovery served its own merged 7B fine-tune to Codex CLI. The agent executed the model’s tool call. That is not a report that OpenAI-hosted models were poisoned.

Are the 100% success rates independent?

No. They are the researchers’ own 50/50 held-out counts. This publication has not replicated the training run.

Is there a scanner that finds every trigger?

The authors argue there is not, because the attacker chooses one phrase from an unbounded set. They recommend sandboxing and egress control instead of guessing the key.

THE TAKEAWAY

What to remember

A cheap poisoned adapter can look clean on ordinary tool-use tests and still fire inside a coding agent. Pin official weights, sandbox execution and network access, and keep credentials out of the working tree. Do not treat a benchmark table as a honesty check.

Sources & further reading

  1. How abliterated models can get you pwned ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories