Announced 8 Oct 2026 · Sources checked
What was posted on 8 October?
The arXiv Atom record we opened lists 2610.12466v1 as published at 17:59:50 UTC on 8 October 2026, primary class cs.AI. The HTML names Drew T. Nguyen and William Fithian and prints “Department of Statistics” and “UC Berkeley” under both. There is no METR affiliation on the author line.
The paper is a reanalysis, not a new task suite. The authors say they use METR’s public Time Horizon 1.1 data: 228 software-engineering tasks in three suites, further grouped into 79 task families, and 26 AIs. They write that METR’s public plot was released in March 2025 and updated until it ended with Claude Mythos in May 2026.
This is a statistical critique of a capability chart, not a new agent product. For how to read a vendor score more generally, see our benchmark-marketing guide. For what to check before trusting an agent run, see evaluating AI agents.
What is a 50% time horizon, as they restate it?
The authors restate METR’s estimand: each task has a human time — how long a professional software engineer takes on a successful attempt — and each AI has a time horizon, the human time at which that AI succeeds autonomously with probability q, usually 50%. They ignore how long the AI itself takes.
Human times, they note, were typically about four timed successes per task, combined with a geometric mean. For 29% of tasks those measurements were “unobtainable or deemed invalid,” and an expert’s value was used instead. On the AI side, each model is run n_ij times on task i.
The introduction’s GPT-4 sentence — a 4-minute horizon, glossed as “Google this fact” — is their restatement of METR’s original scale, not a Model 2 number. Keep that example in the original-plot column.
What are Model 1 and Model 2?
Section 2.3 lists four specifications. METR’s original headline fit is a per-AI logistic in log2 of human time. Barry’s 2026 shared-slope logistic, the authors say, beat that original spec on all but one of 12 proper-scoring metrics, so they use the shared-slope fit as the baseline and as the curve drawn in Figure 1(a).
Model 1 keeps the shared slope but replaces the linear log-time term with a shared monotone natural cubic I-spline, four interior knots, pinned at the min and max log times so the spline is identifiable. The shared shape f is their estimate of how AI difficulty varies with human time for every model.
Model 2 is a fully generative item-response model: each task has a latent difficulty θ_i = f(log2 T_i) plus a family effect and an idiosyncratic effect, with variances that themselves depend on log time through two-knot splines. They fit it with EM and Gauss–Hermite quadrature. Family correlations are modeled directly, so they drop the heuristic square-root family weights used in the logistics.
They evaluate with 5-fold family cross-validation on a suite of proper scoring rules and Murphy diagrams. Figure 3’s paired-difference bands use the 79 families as the clustering units. We did not rerun the folds.
What does the flat 2–30 minute band change?
The headline construct-validity claim is about comparability, not about whether human time predicts success overall. On this suite they call predictiveness “strong,” but locally weak between 2 and 30 minutes. Figure 1(d) plots each of the 228 tasks’ estimated AI difficulty against human time; the fitted f is “nearly flat” in that band and “close to linear elsewhere.”
That is why they write that a jump from 3 minutes to 30 minutes is much easier than a jump from 30 minutes to 5 hours despite the same 10× multiplier. Figure 1(f), the Model 2 conditional-success plot, is the version they prefer for reading those jumps: the 4-minute and 15-minute curves “bunch up.”
Figure 4 shows in-sample fits for six AIs whose METR horizons sit in the 2–15 minute range. The spline bends through the flat band; the log-linear baseline cuts straight through. The authors say some of those re-estimated 50% horizons fall outside METR’s original 95% intervals. Figure 1(e) is Model 2’s horizon series; the disagreement with METR concentrates in the shaded 2–30 minute band.
Why do Rasch abilities track the original METR plot?
Section 3.4 plots Model 2’s Rasch abilities α̂_j, rescaled as 2^(u+v α̂_j) minutes with (u, v) chosen by least squares against log2 of METR’s horizons. The caption prints R²=0.996. Those rescaled abilities sit closer to METR’s published series than Model 2’s own time horizons do.
The authors’ interpretation: under Model 2 the 50% log-time horizon is f⁻¹(α_j/β), while the linear model reports α_j/β. Drop the inverse-f step, and the α_j/β estimates of the two models are close. They write that the linear logistic “may be inadvertently estimating another AI capability construct.” They also write that the coincidence is not guaranteed — it depends on f not departing much further from linear.
The introduction places the original plot in a 2026 policy argument — Roose’s New York Times Moore’s-law comparison, Steinhardt’s Keeling-curve analogy, and summer 2026 slowdown coverage. Those citations are the authors’ framing. They are not a METR endorsement of this paper. For a separate 8 October document in the safety-research conversation, see the OpenAI safety-researchers letter. For another capability-index product, not this re-fit, see Arena’s Alignment Index.
What did we not rerun?
This is an evidence review of the 8 October abs page, HTML, and Atom record. We did not download METR’s public traces, refit the spline, or recompute the Murphy diagrams. Treat every horizon, R², and “easier 10×” sentence as the authors’.
Common questions
Is this METR’s new official plot?
No. It is an independent UC Berkeley Statistics reanalysis of METR’s public data. METR is not on the author line. The baseline the authors beat is Barry’s 2026 shared-slope logistic, which they treat as METR-adjacent methodology, not as a 2026 METR republication.
Does the paper say time horizons are useless?
No. The authors call overall predictiveness on this software suite strong and offer better-scoring point estimates plus diagnostic plots. Their warning is local: jumps inside 2–30 minutes should not be read like jumps elsewhere, and future longer tasks may bend f again.
Is GPT-4’s horizon now something other than 4 minutes?
The 4-minute figure in the introduction is their restatement of METR’s original example. Model 2 changes some mid-range horizons; the paper does not print a replacement GPT-4 cell in the abstract. We did not extract a new GPT-4 number from Figure 1(e).
What to remember
Use 2610.12466 for the shared-spline and IRT re-fits and for the claim that f is nearly flat from 2–30 minutes. Keep the 4-minute GPT-4 example in METR’s original column, and keep R²=0.996 attached to the rescaled Rasch overlay, not to Model 2’s own horizons. Do not treat the paper as METR’s official update.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





