A standardized benchmark suite used to evaluate the agentic coding and terminal interaction capabilities of large language models.
Overview
- OpenAI’s GPT 5.6 Soul model scored 91.9% on Terminal Bench 2.0 in Ultra settings, beating Claude Mythos by almost four percentage points. (Source: OpenAI (via AI Daily Brief host), via AI Daily Brief, 2026-08-24)