Rethinking AI time horizons: why the log-linear assumption breaks down

METR's 50% time horizon measures the human completion time of software tasks that an AI solves with 50% probability, allowing AI capabilities to be expressed in interpretable units. On 228 tasks and 26 AIs, we recompute the time horizons using splines and item-response theory to relax the assumption that the AI difficulty of a task depends linearly on the log of human time. Our fitted spline can be interpreted as a function that converts human time to AI difficulty; it is nearly flat in a region from 2--30 min but close to linear elsewhere. Hence, a time-horizon jump from 3 min to 30 min is much easier than one from 30 min to 5 hours despite the same multiplier of 10×. Overall, we contribute time-horizon point estimates that perform better under a cross-validated suite of proper scoring rules, as well as diagnostic plots for assessing time horizons' construct validity. We suggest that time horizons be interpreted together with the diagnostic plots, especially as new time-horizon-based benchmarks are proposed or existing ones grow to include longer tasks.

From linear logs to spline curves

For several years, the machine learning community has relied on METR's 50% time horizon as a canonical way to express AI capabilities in units that humans can intuit. The original formulation assumed that AI difficulty scales linearly with the logarithm of human completion time. In practice, this meant that if a task takes a human two hours to complete, and an AI solves it with 50% probability at a two-hour time horizon, then the AI's capability level is directly comparable across tasks differing by orders of magnitude in human time.

However, the linearity assumption on the log scale carries implicit consequences. When the relationship between human time and AI success probability is truly non-linear, the resulting time-horizon estimates can misrepresent the true shape of capability boundaries. The present paper identifies exactly where this misrepresentation occurs and offers a re-estimation procedure that avoids the constraint.

The authors replace the linear-log model with a spline-based fit, using item-response theory as the underlying framework. Item-response theory provides a probabilistic model of how an AI's ability relates to task difficulty; the spline then allows the difficulty-conversion function to bend wherever the data warrant it, without imposing a single global functional form.

A nearly flat region and its consequences

The most striking finding from the spline fit is a region spanning approximately 2 to 30 minutes of human completion time where the fitted function is nearly flat. In this interval, the conversion from human time to AI difficulty changes very little. Outside this region, the function becomes approximately linear, but the transition matters.

Consider two hypothetical time-horizon jumps, each involving a 10× increase in human time. In the first case, a jump from 3 minutes to 30 minutes lies within the nearly flat region. The spline indicates that AI difficulty changes only slightly across this range, meaning AI systems that solve tasks at the 3-minute horizon already have a good chance of solving tasks at the 30-minute horizon. The same 10× multiplier applied in the second case—a jump from 30 minutes to 5 hours—lies outside the flat region, where the spline is close to linear. Here, the AI difficulty increases substantially, and the probability of success drops markedly.

This asymmetry has practical consequences. Benchmark designers who wish to differentiate AI capabilities across a wide range of task durations will find that adding tasks in the 2–30 minute range does not substantially change the measured time horizon, whereas tasks extending into the multi-hour range produce more noticeable shifts. The paper's re-estimated time horizons reflect this imbalance, and the authors demonstrate that these revised estimates perform better under a cross-validated suite of proper scoring rules.

Item-response theory and the spline fit

Item-response theory (IRT) models the probability of a correct response as a function of a latent ability parameter and a task difficulty parameter. In the METR context, the "ability" is the AI system's capability on software-task solving, and the "difficulty" encodes how long a typical human takes to complete the task. The conventional approach assumes that the log of human completion time is a sufficient statistic for difficulty, which imposes a straight line when plotted against the IRT logistic curve.

The spline alternative does not assume a single parametric form. Instead, it partitions the human-time axis into segments and fits a separate linear or low-order polynomial model within each segment, joining them at knots. The authors place knots at quantiles of the observed human-time distribution, ensuring that each segment contains enough data to estimate reliably while still allowing the overall function to bend where the data indicate a change in behavior.

Mathematically, if denotes the human completion time in minutes and θ the AI's ability on the IRT scale, the conventional model specifies

log t = α + βθ + ε

where β captures the slope relating log-time to ability. The spline model relaxes this to

log t = f(θ) + ε

where f is a piecewise-linear function estimated from the data. The resulting fitted curve is nearly flat for human times between 2 and 30 minutes, meaning that within this range, AI ability has little marginal effect on the expected human time. Beyond 30 minutes, the curve steepens, reflecting the fact that as tasks take longer for humans, even substantial AI ability gains produce only modest reductions in expected completion time.

Empirical results across 228 tasks and 26 AIs

The authors recompute time horizons on a dataset of 228 software tasks solved by 26 different AI systems. The tasks span a wide range of human completion times, from under a minute to several days. Each task is labeled with the time at which a human expert can complete it, and each AI's success/failure record on each task is recorded.

Under the conventional linear-log model, the time-horizon estimate for each AI is the human completion time at which the AI has a 50% success probability, interpolated linearly on the log scale. Under the spline model, the same 50% quantile is read off the fitted curve, but the curve itself may bend, producing different interpolated values.

The paper reports that the spline-reestimated time horizons achieve lower cross-validated proper scoring rule error than the linear-log estimates. Proper scoring rules reward probability estimates that are both well-calibrated and sharp; in this context, a model that over- or under-estimates the difficulty of long-horizon tasks will incur a penalty. The spline model's ability to bend flat in short-horizon regions and remain linear elsewhere yields probability calibrations that are simultaneously better for short tasks and long tasks.

In addition to the scoring rule results, the authors present diagnostic plots for each AI. These plots show the raw success/failure data overlaid with the fitted spline curve, allowing a visual assessment of whether the spline adequately captures the true difficulty-altitude relationship. Points where the observed success rate deviates systematically from the spline indicate tasks whose difficulty structure may not be well captured by the current model—a potential avenue for future refinements.

Limitations and trade-offs

The spline approach introduces additional complexity relative to the linear-log model. Choosing the number and placement of knots is a subjective decision, and different choices can lead to different time-horizon estimates, particularly in sparsely populated regions of the human-time axis. The authors acknowledge this and provide guidance on knot placement based on quantiles of the observed distribution, but future work could explore data-driven knot-selection procedures.

Another limitation is the scope of the dataset. The 228 tasks and 26 AIs span a substantial range, but most tasks cluster in the 5-minute to 2-hour range. The flat region identified (2–30 minutes) is therefore based primarily on this dense part of the distribution. Extrapolating the spline to much longer human times—multi-day tasks, for instance—must be done with caution, as the spline's behavior in sparse regions is less well constrained.

The authors also note that the spline model, like any model, can overfit if the knots are placed too tightly relative to the amount of data in each segment. The cross-validated scoring rule results help mitigate this concern, but the trade-off between flexibility and generalizability remains a practical consideration for researchers who adopt the method.

What this means for benchmark design and AI evaluation

The central takeaway is that the log-linear assumption, while convenient, is not a neutral choice. It implies that a 10× increase in human completion time always corresponds to a fixed shift in AI ability, which the data show is false. In practice, adding short-horizon tasks to a benchmark may not meaningfully shift the measured time horizons, while adding long-horizon tasks will. This asymmetry should inform how benchmark suites are expanded or revised.

The diagnostic plots introduced in the paper offer a practical tool. Any researcher building a time-horizon-based benchmark can overlay their own task timings on the fitted spline and see whether the model's flat or linear regimes match the observed data. If the observed success rates deviate from the spline in a particular region, that region may warrant more task sampling or a different modeling approach.

For AI system developers, the revised time-horizon estimates provide a more nuanced picture of where their systems stand relative to human task completion times. Because the spline reveals that short tasks (under 30 minutes) are essentially a different regime than longer tasks, a system that performs well on 5-minute tasks may not have the same relative standing on 5-hour tasks, even if the raw time-horizon number appears comparable under the old model.

As new benchmarks based on time horizons are proposed—and as existing benchmarks grow to include longer and more complex software tasks—the diagnostic framework and re-estimated point values offered by this paper become increasingly relevant. The authors' suggestion to always interpret time horizons alongside their diagnostic plots is a practical recommendation that can help the community avoid the pitfalls of an unexamined linearity assumption.

Read the paper on arXiv