A standardized evaluation framework designed to measure the ability of AI agents to maintain coherence and achieve objectives over extended sequences of actions or time.

Overview

  • On the Meters benchmark for long-horizon tasks, GPT-5.1 Codex Max achieved a 50% success rate on tasks taking a human programmer 2 hours and 42 minutes, which is 25 minutes longer than the previous state-of-the-art GPT-5. (Source: Meters Benchmark, via AI Daily Brief, 2026-08-24)

Provenance