Three minutes to thirty minutes. Thirty minutes to five hours. Both are tenfold increases, but they may tell very different stories about an AI system's progress.
That is the central finding of a new statistical examination of METR's influential AI time-horizon benchmark. UC Berkeley statisticians Drew T. Nguyen and William Fithian reanalyzed results for 228 software tasks and 26 AI systems. They found that part of the benchmark's human-readable scale behaves less like a rigid ruler than a flexible one.
Their preprint, submitted October 8 and first announced in arXiv's October 9 listings, does not argue that AI progress is imaginary. It asks a more useful question: when a capability score is expressed in minutes and hours, do equal multipliers represent comparable increases in difficulty?
These hours belong to a human, not the AI
METR estimates the human task duration at which an AI agent reaches a given probability of success. A 50% horizon of two hours means the fitted success curve crosses the halfway mark around tasks that take a human expert two hours.
It does not mean the agent runs for two hours, remains reliable for two hours or can replace two hours of any person's work. The underlying evaluation mainly contains well-specified software engineering, machine-learning and cybersecurity tasks. Our time-horizon explainer separates that task-difficulty measure from elapsed runtime.
The attraction is obvious: hours make capability easier to grasp than an abstract score. The complication is that human duration and AI difficulty need not rise together at a uniform rate.
The same multiplier, a different climb
Nguyen and Fithian allowed the relationship between human time and AI difficulty to bend rather than forcing it into a straight line on a logarithmic time scale. Their alternatives use flexible curves and item-response models, statistical tools that distinguish task difficulty from a system's ability.
In this task collection, the estimated relationship is unusually flat between roughly two and thirty minutes. A task taking a human several times longer within that band may not be much harder for an AI. Beyond that region, human time maps more closely to increasing difficulty.
That is why the authors contrast three to thirty minutes with thirty minutes to five hours. Both changes multiply the time horizon by ten. The first crosses a comparatively flat part of their fitted scale; the second climbs a steeper one. These are illustrations of the measurement issue, not two newly observed model improvements.
The comparison is also more specific than a wholesale replacement of METR's published chart. The paper uses a shared-slope logistic baseline described in a METR research note. Its alternative fits are tested on held-out task families and perform better on the authors' scoring comparisons. This is a reanalysis of existing outcomes, not a new set of agent trials.
A useful benchmark, with a less uniform ruler
The strongest reason not to overread the critique is inside the paper itself. Human task time still carries useful information about AI difficulty overall. The authors' rescaled ability estimates remain close to the existing trajectory, and they explicitly do not dispute the observed exponential rise in time horizons.
What changes is the interpretation of a particular jump. A large increase in the number of minutes can partly reflect where a system sits on the conversion curve. It need not represent the same increment of ability as an equally large ratio elsewhere.
The flat region is a finding about these tasks and fitted models, not a universal law of AI. Different task collections, longer assignments and different reliability thresholds need their own evidence. Nor does a better statistical fit establish what will happen to future systems.
Read the scale before celebrating the jump
For someone deciding what to delegate, the headline multiplier is only a starting point. The task mix, success threshold and fit assumptions determine what it can support. Consistency on one particular job is another question again, illustrated by our chart of repeated simulated customer-service cases.
Human time remains a valuable way to make AI progress legible. But equal multipliers in that unit need not represent equal gains in task difficulty. Before comparing the distance traveled, check whether the ruler has the same spacing along the way.




