A standardized benchmark suite used to evaluate the agentic coding and terminal interaction capabilities of large language models.

Overview

  • OpenAI’s GPT 5.6 Soul model scored 91.9% on Terminal Bench 2.0 in Ultra settings, beating Claude Mythos by almost four percentage points. (Source: OpenAI (via AI Daily Brief host), via AI Daily Brief, 2026-08-24)

Provenance