Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. METR releases Time Horizon 1.1 with expanded long-task suite

METR releases Time Horizon 1.1 with expanded long-task suite

★★★benchmarkMETRconfidence: high

METR updated its task-completion time-horizon methodology on 2026-01-29 (TH1.1), adding 34% more tasks (228 vs 170) and doubling 8h+ tasks (31 vs 14), tightening confidence intervals for frontier models; METR notes measurements above ~16 hours are unreliable with the current suite.

Key facts

What happened

METR's time horizon — the human task length at which an AI succeeds 50% of the time — is the most-cited measure of agentic progress. TH1.1 extends the task suite to keep pace with models approaching day-long tasks.

Why it matters

As frontier horizons approach the top of the suite, METR's own caveat (unreliable >16h) signals the benchmark itself is near saturation.

Changelog

  • 2026-09-29: created

Sources (3)

id: 2026-01-29-metr-time-horizon-1-1 · updated 2026-09-29 · open in the interactive timeline