METR releases Time Horizon 1.1 with expanded long-task suite
METR updated its task-completion time-horizon methodology on 2026-01-29 (TH1.1), adding 34% more tasks (228 vs 170) and doubling 8h+ tasks (31 vs 14), tightening confidence intervals for frontier models; METR notes measurements above ~16 hours are unreliable with the current suite.
Key facts
- Tasks: 228 (TH1.1) vs 170 (TH1); tasks >=8 hours: 31 vs 14
- Upper CI for Claude Opus 4.5 narrowed from 4.4x to 2.3x the point estimate
- Measurements above 16 hours flagged as unreliable
- Later 2026 measurements include GPT-5.3-Codex, Claude Opus 4.6 (Feb 20), GPT-5.4 (Apr 10), Gemini 3.1 Pro (Apr 15), early Claude Mythos Preview (May 8)
- Community analyses suggest ~4-month doubling since 2024 vs 7 months 2019-2024
What happened
METR's time horizon — the human task length at which an AI succeeds 50% of the time — is the most-cited measure of agentic progress. TH1.1 extends the task suite to keep pace with models approaching day-long tasks.
Why it matters
As frontier horizons approach the top of the suite, METR's own caveat (unreliable >16h) signals the benchmark itself is near saturation.
Changelog
- 2026-09-29: created
Sources (3)
- officialMETR: Time Horizon 1.1
- officialMETR: Task-completion time horizons of frontier AI models
- officialMETR: Clarifying limitations of time horizon
id: 2026-01-29-metr-time-horizon-1-1 · updated 2026-09-29 · open in the interactive timeline