METR's Updated Benchmark Puts Claude Opus 4.6's Autonomous Task Horizon at 12 Hours, Not Days
On January 29, 2026, AI safety research organization METR shipped Time Horizon 1.1, an overhaul of the methodology it uses to measure how long a task an AI agent can complete autonomously at a 50% success rate. The update expanded METR's task suite by 34% and doubled the number of tasks requiring eight or more hours of continuous autonomous work, according to METR's own blog post announcing the change.
Under this stricter, larger benchmark, Claude Opus 4.6 measured a roughly 12-hour 50%-success time horizon. That is a substantial number on its own terms, a model completing tasks that would take a human around half a working day, without intervention, more often than not.
The more useful detail is what METR says about the ceiling on its own measurement. The organization now explicitly flags anything above a 16-hour horizon as statistically unreliable, given the current size of its task suite. Past that point, there simply are not enough long-duration tasks in the benchmark to produce a trustworthy number, even if a model's raw output looks like it kept going productively for longer.
That ceiling matters for how to read headline claims about agents working for days at a stretch. A number quoted well past 16 hours is, per METR's own methodology, outrunning what the benchmark can currently verify reliably. Twelve hours, measured under the newly expanded and doubled long-task suite, is a solid and specific data point. Multi-day autonomy claims built on the same family of benchmarks deserve the same question applied to Opus 4.6's number: was this measured inside the range METR itself trusts, or past it.