What problem is OnTrack trying to solve?

The authors argue that autonomous agents create two failure modes that offline judges miss: they burn tokens on doomed loops, and they can take irreversible actions before a post-hoc review finishes. Safeguard agents add another model call per step; log reviews arrive after the cost is sunk. OnTrack’s pitch is a structural monitor that scores the growing plan graph while the run is still open.

They build on Otap’s batch DAG alignment idea, then solve streaming-specific issues: not punishing an agent for unfinished reference steps, phasing in provisional nodes until dependencies appear, and warm-starting the transport solve so each event stays cheap. For executable outcome checks versus process monitors, compare SWE-bench explained with AI agent sandboxing.

The paper’s contribution is not inventing DAG alignment from scratch—it cites Otap for batch optimal-transport scoring of complete traces—but making that alignment honest on a live prefix. Re-running a batch solver after every tool call would be too slow and would punish unfinished reference steps. OnTrack’s streaming fixes (frontier-masked marginals, provisional nodes, warm-started conditional gradient steps) exist so a monitor can speak while the agent still has budget left.

A black Netgear 16-port 10-gigabit switch on a white background. No people appear.
Netgear ProSafe XS716T switch photographed by Sergklim. CC BY-SA 4.0 via Wikimedia Commons. Contextual networking gear; not OnTrack software. Photo: Sergklim / Wikimedia Commons. CC BY-SA 4.0 · Cropped and resized.

How do the three data-access regimes differ?

With historical successful runs and tool schemas, OnTrack can look for plan violations relative to references. With only tool schemas, it still targets loops, stalls, and irreversible actions whose declared prerequisites are missing. With only the live event stream, loop and stall detection remain. The paper says the software is not reconfigured across regimes—unused modules simply switch off.

Dependency extraction is described as deterministic string matching from tool schemas or harness metadata, not an LLM judge. The authors note that shell indirection can hide edges, so the online graph can be incomplete even when the batch graph would later look cleaner.

The paper’s Figure 1 framing is worth keeping straight: events arrive one at a time; dependency edges can appear late when a later tool consumes an earlier artifact; references carry precomputed structure matrices; the live prefix grows a matching matrix row by row. That late-edge behavior is why the authors refuse to judge a brand-new step’s structure at full weight immediately.

Operators should map those regimes to what their harness actually emits. If you only have tool names and timestamps, you are in the weakest regime and should expect loop/stall alarms, not plan-violation precision. If you can attach schema-declared prerequisites and artifact parent pointers, plan checks become meaningful. The authors warn that shell indirection can hide edges discovered only in a later offline parse—so an online halt decision can be noisier than a post-hoc Otap score even when the same math eventually converges.

What do the SWE-bench numbers actually say?

Evaluating on the first eight steps, the authors say OnTrack ranked failing trajectories below succeeding ones better than content-similarity baselines by +0.057 AUROC. With an abort policy on top, they report saving about 18% of the compute that would have been spent on runs heading to failure, and that 83% of interrupted runs (five of six) were actually heading to failure. Those are author-reported experimental results on their SWE-bench trajectory corpus—not a guarantee for your harness.

Decisions are continue / warn / halt from per-step signals with a grace window, not from a single aggregate score threshold that changes meaning as traces lengthen. That design choice is the paper’s answer to why naive mid-run reuse of batch scores fails.

The evaluation protocol they describe uses an instance-disjoint split, calibration, and rival abort policies alongside content-similarity baselines. The +0.057 AUROC figure is specifically for ranking failing versus succeeding trajectories from the first eight steps—not a claim that OnTrack predicts final SWE-bench resolve rate from step one. The ~18% compute save and 83% abort precision (five of six) come from their abort-policy overlay on that corpus.

A second experimental thread in the paper is process anomaly detection: exit-status detection with logs only, in-distribution splice injection, and fault injection into real traces. Those studies matter if you care whether the monitor fires on synthetic faults it was tuned for; they are still author-run, not a third-party red team of production agents.

An open ThinkPad laptop on a wooden table beside a mouse and notebook. No people appear.
ThinkPad on a desk photographed by Dug Song. CC BY 2.0 archival photo via Wikimedia Commons. Illustrative local coding machine; it is not a Cribl console. Photo: Dug Song / Wikimedia Commons. CC BY 2.0 · Cropped and resized.

What should operators copy carefully?

If you already store successful traces, structural mid-run monitors are a concrete alternative to always-on judge models. Start by defining which actions are irreversible in your tool schema, then decide whether abort authority sits with the monitor or with a human. Do not assume a millisecond claim transfers until you measure your own event pipeline.

We did not find a Cribl-branded public GitHub release for this paper during discovery; treat the arXiv PDF/HTML as the primary technical object unless the authors publish a linked repository later.

A minimal internal pilot would log tool name, arguments, artifacts, and parent hints; compute a cheap loop/stall detector first; only then invest in reference mining for plan-violation alerts. If your harness cannot expose consumed file paths, OnTrack’s dependency story will be weaker than the paper’s full-metadata regime—budget for that gap before promising halt precision.

Treat the “about a millisecond per step” line as a solver-iteration claim under the authors’ warm-start regime, not as end-to-end latency including your queue, embedding calls, or human review UI. Before granting halt authority, measure false-abort cost on successful runs you care about—the paper’s 83% precision still implies roughly one in six aborts would have interrupted a run that was not heading to failure in their study framing.

What did we not test?

We did not implement OnTrack, stream SWE-bench traces, measure per-step latency, or validate abort precision. All quantitative claims above are the authors’.

Common questions

Is OnTrack an LLM-as-a-judge?

No. The paper describes a structure-aware optimal-transport alignment over execution DAGs, with optional references, not a second model scoring each step.

Does it require successful reference runs?

References improve plan-violation detection, but the authors say loop and stall checks still run with weaker inputs.

What benchmark did they use?

SWE-bench trajectories, with ranking reported on the first eight steps and an abort-policy compute study described in the abstract.

THE TAKEAWAY

What to remember

File OnTrack under streaming structural monitoring: promising abort economics in the authors’ SWE-bench study, with capability that degrades gracefully as reference data disappears, still awaiting independent implementation checks.

Sources & further reading

  1. OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport ↗
  2. OnTrack (PDF) ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories