Brand Logo

Terminal-Bench 4.0 Explained: Claude Code vs Codex Results

Aerin Kim

Written by

Aerin Kim

Terminal-Bench 4.0 tests real agentic terminal work across 66 tasks. See why Sonnet 5 burned 21.6B tokens for only 12.4%, and how harness choice can matter more than the model.

On paper, Sonnet 5 should have won this benchmark. It ran on Claude Code, the same agent harness that carried Fable 5.1 and Opus 5 to the top of the leaderboard. It was given the "max" effort setting, the most generous reasoning budget the harness allows. And across the 66 tasks in Terminal-Bench 4.0, it used more tokens than any other model tested, 21.6 billion of them, and spent more money doing it than any other model tested, $9.6k against the second-priciest run's $7.3k. It still finished with a resolution rate of 12.4 percent, tied for twelfth place with Grok 4.5, a model that used roughly a sixth of the tokens and less than a quarter of the budget to land at the exact same score [2] [1].

That result is not a rounding error and it is not a bad run that got unlucky. It is the single most instructive data point in the newest refresh of Terminal-Bench, the agentic-terminal benchmark built by the Laude Institute and open-source contributors on top of the Harbor framework, and it is worth sitting with before looking at anything else on the leaderboard. A model that burns the most tokens and the most money on a benchmark, while running on the exact same agent scaffolding that pushed a cheaper, faster-scoring sibling to the top, is not "trying harder." It is failing in a specific, mechanically explainable way: falling into long, unproductive trajectories inside an eight-hour timeout window and never converging on a state its own verifier would accept. Understanding exactly why that happens, and why it happens to some models and not others under identical harnesses, is the actual point of this post.

Terminal-Bench matters more than most of the benchmark refreshes that get a news cycle and then disappear, because it measures something closer to how agentic coding tools actually get used in production than a static benchmark like HumanEval or SWE-bench does. HumanEval asks a model to write one function from a docstring, in isolation, with no execution loop and no chance to recover from a mistake. SWE-bench is a real step up in realism, testing whether a model can patch a real GitHub issue inside a real repository, but it still evaluates a single, mostly self-contained code change rather than a full session of exploring an unfamiliar environment, running commands, watching them fail, and adjusting [5]. Terminal-Bench goes further still. It hands an agent a real terminal, with no GUI at all, and a real task that might mean configuring a broken build, recovering a corrupted database, hardening a misconfigured server, or wiring up a piece of scientific software correctly, and it only cares whether the sandboxed environment ends up in the correct final state, verified automatically, after however many commands, retries, and dead ends the agent needed along the way. That is a genuinely different kind of test, and it is the kind of test that actually predicts whether an agentic coding tool will behave itself when you point it at your own infrastructure.

terminal-bench-4-0-claude-code-codex-explained-2026-hero.png

This post walks through the whole picture in depth: how a Terminal-Bench task and its verifier actually work under the hood, what changed between version 3.0 and version 4.0 and why the Laude Institute made those specific changes, the full 66-task leaderboard for today's frontier models with a deep read on the model-versus-harness separation that the leaderboard is deliberately built to expose, a worked illustration of what a good agent trajectory looks like next to a bad one on a representative task, and a set of production lessons for any team currently choosing an agent harness for its own developers. Expect real numbers throughout, cited against the official announcement and the independent leaderboard tracking it, not a generic gloss over "AI agents got better at coding."

What Terminal-Bench Actually Measures, and How a Task Works Under the Hood

Terminal-Bench evaluates one specific, narrow-sounding but genuinely hard question: can an AI agent operate a real terminal, with no graphical interface at all, to complete a real piece of software engineering or systems work, from writing and debugging code to configuring an unfamiliar environment to recovering from a failure it caused itself. Every task in the benchmark ships as a self-contained, sandboxed environment plus an automated verification suite. The agent is dropped into that environment with a natural-language instruction, given full terminal access, and left to do whatever it judges necessary, running commands, reading error output, editing files, restarting failed services, until it either declares the task done or runs out of time. Only then does the verifier run, checking the actual resulting state of the sandbox, not the agent's own self-report of success, against a fixed, task-specific pass/fail rubric.

That verifier-checks-final-state design is the part worth understanding carefully, because it is what makes Terminal-Bench resistant to an agent simply asserting it succeeded. A model can say "I have fixed the database corruption and the service is now healthy" in its final response, and that sentence earns it nothing. The verifier connects to the actual sandbox, runs its own independent checks, a specific query returns the expected row, a specific service responds on the expected port, a specific file has the expected checksum, and only a genuinely correct final state passes. This is the same philosophy that makes SWE-bench a meaningfully harder and more trustworthy benchmark than a benchmark graded by an LLM judge reading the agent's own transcript: the ground truth lives in the environment, not in the model's self-assessment.

To make this concrete, here is a simplified but structurally realistic sketch of what one task's definition looks like, based on the way Harbor-framework benchmarks like Terminal-Bench organize a task directory. This is illustrative, written to show the real shape of a task/verifier pair, not a literal file pulled from the benchmark's private task set.

bash
/code # Illustrative Terminal-Bench-style task layout (Harbor framework conventions). # This is a structural example for teaching purposes, not a real published task. mkdir -p ops-fix-db-lockup/tests cd ops-fix-db-lockup cat > task.json << 'EOF' { "task_id": "ops-fix-db-lockup", "category": "operations", "instruction": "The 'orders' Postgres service on this box will not start. Diagnose why, fix the root cause, and confirm the service is healthy and serving queries on port 5432. Do not restore from a backup; the existing data must remain intact.", "timeout_seconds": 28800, "trials": 5 } EOF cat > Dockerfile << 'EOF' FROM postgres:16 # Seed an environment that starts in a genuinely broken state: # - max_connections set below what the startup scripts require # - a stale postmaster.pid left behind by a simulated crash COPY broken_postgresql.conf /etc/postgresql/postgresql.conf COPY stale_postmaster.pid /var/lib/postgresql/data/postmaster.pid CMD ["postgres"] EOF cat > tests/verify.sh << 'EOF' #!/usr/bin/env bash # The verifier never trusts the agent's own final message. # It independently checks the real state of the sandbox after the run. set -e pg_isready -h localhost -p 5432 -U postgres || exit 1 ROW_COUNT=$(psql -h localhost -U postgres -d orders -tAc \ "SELECT count(*) FROM orders;") if [ "$ROW_COUNT" -lt 1 ]; then echo "FAIL: orders table is empty or unreachable" exit 1 fi echo "PASS: service healthy, data intact, verified independently of agent output" exit 0 EOF chmod +x tests/verify.sh # At evaluation time, Harbor builds the Dockerfile, drops the agent into the # running container with full terminal access and the instruction from # task.json, lets it act freely up to timeout_seconds, then runs # tests/verify.sh in that same container regardless of what the agent claims.

Walk through what that structure actually enforces. The task.json config declares the task's category (here, operations, one of Terminal-Bench 4.0's seven task categories), the natural-language instruction the agent receives verbatim, and a hard time budget. The Dockerfile builds a sandboxed environment that starts in a genuinely broken state, in this illustration, a Postgres container with an intentionally misconfigured max_connections setting and a corrupted index, mirroring the kind of "recover from a failure" scenario Terminal-Bench is specifically designed to test rather than a clean, from-scratch build task. The agent's job is entirely mediated through the terminal: it has to notice the service will not start, diagnose why, and fix it using ordinary shell commands, psql, systemctl, editing config files, exactly the way a human on-call engineer would. Only after the agent finishes, or times out, does verify.sh run inside the same sandbox, independently querying the database and checking the service's actual health endpoint. Nothing the agent says in its own final message is trusted. Five independent trials of this exact process are run per task across the benchmark, and the reported score for any model is the mean pass rate across those five trials, along with a margin of error that reflects how much that pass rate varies from run to run on the same task and the same model [1] [2].

terminal-bench-4-0-claude-code-codex-explained-2026-task-verifier-mechanism.png

That five-trial averaging matters more than it might first appear. Agentic tasks are not deterministic the way a static benchmark's single forward pass is. The same model, given the exact same broken database, might fix it cleanly on one attempt, get distracted debugging a red herring on a second attempt, and time out entirely on a third. A benchmark that reported only a single trial per task would be reporting noise dressed up as a score. Reporting the mean of five trials, with an honestly stated margin of error, is Terminal-Bench's way of being upfront that agentic evaluation has an irreducible amount of run-to-run variance, and that any two models within each other's margin of error should be read as roughly tied rather than meaningfully ranked, a point this post returns to directly when reading the leaderboard in the section below.

What Changed in Terminal-Bench 4.0, and Why the Laude Institute Made These Specific Fixes

Terminal-Bench 4.0 is a maintenance and calibration release layered on top of the 3.0 task set, not a wholesale rewrite with a flood of new tasks. The task count actually shrank, from 74 tasks in version 3.0 down to 66 tasks in version 4.0, spanning seven categories: software, machine learning, science, operations, security, hardware, and media [1]. That reduction was deliberate, and the official announcement breaks down exactly why each removed task was cut, which is worth walking through because it tells you a lot about what makes a benchmark task trustworthy in the first place.

Eight tasks were removed in total, for four distinct reasons, two tasks apiece:

Two tasks were removed for saturation. Once every frontier model being evaluated reliably passes a task across all five trials, that task has stopped discriminating between models and is only adding noise-free but uninformative signal to the aggregate score. A saturated task inflates every model's score roughly equally and tells you nothing about relative capability, so keeping it around mostly wastes compute rather than adding information.

Two tasks were removed for refusal-proneness. Some tasks, by their framing, made capable models decline to attempt them, or attempt them half-heartedly, for reasons unrelated to actual terminal competence, a task whose instructions read as ambiguous enough to trigger a model's safety training even though the underlying technical work was benign. A task that measures a model's refusal behavior instead of its terminal skill is measuring the wrong thing, and removing those two tasks tightens the benchmark's actual construct validity.

Two tasks were removed because public solutions existed. This is a contamination problem, the same one that has quietly undermined static coding benchmarks like HumanEval and early versions of SWE-bench as their questions and reference solutions leaked into public GitHub repositories and, from there, into the pretraining data of newer models. Once a specific task's working solution is findable in a search index, a model that has memorized it is demonstrating retrieval rather than terminal reasoning, and the two numbers look identical on the leaderboard unless someone actually checks. Pulling contaminated tasks before they quietly inflate every subsequent model's score is exactly the kind of unglamorous maintenance that keeps a benchmark meaningful for more than one release cycle.

Two tasks were removed for unresolved quality or platform-compatibility issues, a catch-all for tasks whose sandboxed environment behaved inconsistently across the different infrastructure providers running the benchmark, occasionally failing for reasons that had nothing to do with the agent's actual behavior. A task that sometimes fails a perfect solution because of an unrelated platform quirk is actively harmful to score reliability, since it punishes correct agent behavior with the same "fail" signal as genuinely incorrect behavior.

Beyond the eight removals, nineteen further tasks were fixed rather than cut, with updated instructions, updated sandbox environments, or updated verifier logic, correcting bugs that had been quietly present in the 3.0 task set without necessarily changing what each task was fundamentally testing [1]. That is a genuinely large share of the surviving task set, nineteen out of what became 66, getting a real correctness pass, which says something honest about how hard it is to author a sandboxed, automatically verified agentic task correctly on the first attempt. A verifier script with an off-by-one bug, or a Docker image that silently drifts as its base image gets patched upstream, is the kind of thing that only surfaces after enough model runs expose it, and folding nineteen of those fixes into one release is the benchmark equivalent of a significant point release, not a cosmetic refresh.

The single most consequential change in version 4.0, though, is the move to a flat eight-hour agent timeout applied uniformly across every task, replacing whatever mix of task-specific timeouts version 3.0 used [1]. This change did not happen in a vacuum. It follows directly from the methodology laid out in a companion paper, "Quantifying infrastructure noise in agentic coding evals," which studied exactly how much of the variance in agentic benchmark scores comes from infrastructure and timing artifacts rather than genuine differences in model capability. Timeout-driven noise turns out to be one of the largest and least discussed sources of measurement error in this kind of benchmark: a task with a timeout set slightly too aggressively will fail a model that was seconds away from a correct solution, and a task with a timeout set too generously will let a model wander indefinitely without meaningfully changing the outcome, wasting compute and cost without adding useful signal about capability differences. Standardizing on a single, generous eight-hour window across every task removes one whole axis of task-by-task inconsistency, so that when a model does time out under version 4.0's rules, that timeout is at least comparable across every other model and task in the suite, rather than being an artifact of how conservatively a particular task's original author happened to set its own limit back when the task was first written.

It is worth being precise about what an eight-hour timeout actually implies for a benchmark meant to reflect real terminal work. Eight hours is generous by the standard of most static coding benchmarks, where a single forward pass takes seconds, but it is not generous relative to what a genuinely difficult sandboxed systems task can demand: recovering a corrupted database, correctly configuring a piece of scientific software with a nonstandard build system, or hardening a server against a specific class of misconfiguration can legitimately take a skilled human engineer the better part of a working day when the environment is unfamiliar. The point of the flat, generous timeout is not to make every task trivially completable in time, it is to make sure that when a model fails to finish, the failure reflects the model's own trajectory efficiency and correctness, not an accident of which task happened to inherit a stingier limit under the old per-task scheme. As the leaderboard section below shows in detail, this design choice is precisely what makes the Sonnet 5 result legible as a genuine capability signal instead of dismissible as "the clock was unfair to it."

The Full Terminal-Bench 4.0 Leaderboard: Where Model Capability and Harness Choice Actually Diverge

Before reading the numbers, it helps to understand the single design choice on the Snorkel AI leaderboard that makes it more useful than a typical model-only ranking: every row separates the underlying model from the agent harness driving it [2]. A harness is the scaffolding wrapped around a model, the tool-calling loop that decides when to read a file, when to run a command, how to format the model's raw output into an actual terminal action, and when to decide the task is finished. Claude Code, OpenAI's Codex, xAI's Grok Build, and the lightweight open-source mini-SWE-agent are four genuinely different harnesses on this leaderboard, and the same underlying model capability can score wildly differently depending on which one is driving it, because the harness controls how efficiently the model's raw reasoning actually gets translated into correct terminal actions.

RankModelEffortAgent HarnessResolution RateRelease DateTotal TokensTotal Cost
1GPT-6 AstramaxCodex58.2% ±2.82026-09-031.5B$3.3k
2Fable 5.1 (Claude)maxClaude Code57.9% ±3.82026-09-012.7B$6.2k
3Opus 5 (Claude)xhighClaude Code53.9% ±3.22026-07-246.9B$6.1k
4Fable 5 (Claude)maxClaude Code44.5% ±3.82026-06-093.8B$7.3k
5GLM-5.3 (Z.ai)maxClaude Code41.8% ±3.22026-08-148.7B$2.7k
6GPT-5.6 SolmaxCodex37.3% ±3.82026-06-264.4B$2.5k
7Opus 4.8 (Claude)maxClaude Code23.6% ±3.62026-05-286.4B$6.5k
8GPT-5.6 TerramaxCodex21.5% ±3.32026-06-265.7B$1.7k
9Grok 4.6highGrok Build20.3% ±3.12026-08-124.0B$3.6k
10Gemini 3.8 Flashhighmini-SWE-agent19.1% ±3.42026-09-0217.2B$1.8k
11GPT-5.6 LunamaxCodex17.3% ±2.82026-06-2611.6B$0.3k
12Grok 4.5highGrok Build12.4% ±2.62026-07-163.4B$2.1k
13Sonnet 5 (Claude)maxClaude Code12.4% ±3.12026-06-3021.6B$9.6k
14Gemini 3.7 Flashhighmini-SWE-agent11.2% ±2.42026-08-1311.1B$1.3k
terminal-bench-4-0-claude-code-codex-explained-2026-model-harness-diagram.png

If you want to see this same lock-and-key relationship in motion rather than as a static cutaway, here is a video generation prompt that animates it:

A Wan-style animated technical diagram video, flat vector cutaway illustration on a pale neutral background with a thin outlined border and a small scale-bar legend in the corner, even ambient lighting with no single dramatic light source, camera held in a fixed straight-on cutaway view like a page from a locksmith's manual. The scene opens on a single brass key blank labeled 'MODEL' hovering beside a mechanical lock body shown in cross-section labeled 'AGENT HARNESS', its internal pin tumblers visible and slightly misaligned. The key slides into the lock and the animation shows the tumblers shifting into place one at a time, each pin briefly highlighted in amber as it aligns, with a small caption beneath reading 'A capable key still needs a well-cut lock to turn'. The animation then splits into a side-by-side comparison: on the left, the same key blank turns a well-machined lock smoothly in one motion, captioned 'STRONG MODEL, STRONG HARNESS'; on the right, an identical key blank struggles against a rough, poorly-fitted lock, its tumblers jamming and resetting several times before the scene fades on an unopened door, captioned 'STRONG MODEL, WEAK HARNESS'. Flat color fills throughout, thin dark outlines, no photographic shading or texture, every label short, correctly spelled, and legible for the full duration it is on screen, no garbled text, no watermark or logo, smooth even motion with no flickering or warped geometry.

Start with the top of the table, because it sets up the two comparisons that matter most. GPT-6 Astra, run at max effort on OpenAI's own Codex harness, takes first place at 58.2 percent resolution, with a margin of error of ±2.8 points, using 1.5 billion tokens across all 66 tasks at a total cost of $3.3k [2]. Fable 5.1, Anthropic's model, run at max effort on Claude Code, takes second at 57.9 percent, ±3.8, using 2.7 billion tokens for $6.2k. Those two scores sit well within each other's combined margin of error, meaning first and second place are not meaningfully distinguishable given the run-to-run variance the five-trial methodology is explicitly designed to surface. What is distinguishable, and far more interesting, is the efficiency gap underneath the near-tie: Astra reached essentially the same resolution rate using roughly 1.2 billion fewer tokens and about $2.9k less total cost than Fable 5.1. Read narrowly, that looks like a clean win for the Codex harness's efficiency on this specific model-and-task combination. Read more carefully, alongside the rest of the table, it is the first hint of a pattern that recurs throughout the leaderboard: token count behaves less like a measure of "how hard the model thought" and more like a measure of how many actual terminal turns, retries, and exploratory dead ends the agent needed to reach its final state, whether that state was correct or not.

The third-place result is where that pattern becomes impossible to miss. Opus 5, also run on Claude Code but at Anthropic's higher "xhigh" effort setting, scores 53.9 percent, ±3.2, close behind Fable 5.1's 57.9 percent, but it gets there using 6.9 billion tokens for $6.1k, more than two and a half times Fable 5.1's token count for a lower score at essentially the same dollar cost. Put the two Claude Code rows for Fable 5.1 and Opus 5 side by side and the efficiency story is genuinely striking: Fable 5.1 is both the higher-scoring model and the one using far fewer tokens to get there, 2.7 billion against Opus 5's 6.9 billion, a ratio of roughly 2.6x. That is not a small optimization, it is the kind of gap that separates a model converging on a correct terminal state in a handful of decisive, well-targeted commands from one that needs several times as much exploration, re-reading, and course-correction to arrive somewhere slightly worse. Whatever changed in the newer model's long-horizon planning between Opus 5 and Fable 5.1, it shows up far more clearly in the token count than in the headline score, which is exactly why reading a leaderboard by resolution rate alone, without the accompanying token and cost columns, misses most of the actual signal about how the two models are behaving inside the agent loop.

terminal-bench-4-0-claude-code-codex-explained-2026-leaderboard-bar-comparison.png

Now the anomaly this post opened with, read in full context against the rest of the table. Sonnet 5, run at max effort on Claude Code, finishes at 12.4 percent, ±3.1, tied for twelfth place. To get there it used 21.6 billion tokens, the single highest token count on the entire leaderboard by a wide margin, more than double the next-highest figure (Gemini 3.8 Flash's 17.2 billion), and it spent $9.6k, also the highest total cost on the leaderboard, ahead of Opus 4.8's $6.5k for a model that scored nearly twice as well. The official Terminal-Bench 4.0 announcement itself notes plainly that "Sonnet 5 sometimes hit timeouts" [1], and once you connect that observation to the eight-hour flat timeout discussed in the previous section, the mechanism becomes legible rather than mysterious.

Here is the plausible technical account, built directly on how these agent loops actually work rather than speculation about the model's internals. Running a model at the most generous "max" effort setting inside an agentic tool-calling loop does not automatically make it more decisive. It gives the model permission to spend more tokens per turn, more tool calls per episode, and more intermediate reasoning between actions, and a model whose long-horizon planning is weaker than its raw per-turn reasoning can use that permission unproductively: re-reading files it already read earlier in the same session because it lost track of what it already learned, re-running a command that already failed with a nearly identical variant instead of diagnosing the actual root cause, or producing long, verbose intermediate reasoning between tool calls that never converges on a concrete next action. None of that individually looks like a mistake in the moment, each re-read and each retry can look locally reasonable, but in aggregate it consumes the eight-hour timeout budget through sheer trajectory length without the model ever reaching a state its verifier would accept. Contrast that with a model exhibiting genuinely stronger long-horizon planning, which tends to converge to a correct terminal state in far fewer, more decisive tool calls: it diagnoses the actual failure mode earlier, commits to a fix, verifies its own work with a quick check before declaring done, and stops. Under this reading, the token count on this leaderboard is best understood as a proxy for trajectory length and decisiveness, not as a proxy for "how hard the model tried" or "how much reasoning it did," and a model that needs radically more tokens to reach a worse outcome than a cheaper sibling is exhibiting exactly the kind of unproductive-exploration failure mode that a flat, generous timeout is specifically designed to expose rather than paper over.

terminal-bench-4-0-claude-code-codex-explained-2026-sonnet-5-timeout-anomaly.png

The tie between Sonnet 5 and Grok 4.5 at the identical 12.4 percent resolution rate sharpens this point further. Grok 4.5, run at high effort on xAI's Grok Build harness, reached that same 12.4 percent using 3.4 billion tokens for $2.1k, a small fraction of Sonnet 5's 21.6 billion tokens and $9.6k for the exact same headline score. Two models landing on an identical resolution rate while differing by roughly 6.4x in token count and 4.6x in dollar cost is about as clean a demonstration as a real leaderboard ever produces that resolution rate alone is an incomplete metric. A team picking a model purely by matching the top-line percentage would be indifferent between these two rows. A team that actually has to pay the compute bill, or that cares about how quickly an agent converges rather than how long it is willing to keep trying, would not be indifferent at all.

terminal-bench-4-0-claude-code-codex-explained-2026-fable-opus-efficiency.png

There is one more comparison worth drawing out before moving past the table, because it is the clearest illustration of why separating model from harness matters as a methodology choice, not just as a leaderboard formatting decision. Gemini 3.8 Flash and Gemini 3.7 Flash are the only two models on this leaderboard evaluated on mini-SWE-agent, a deliberately minimal, lightweight open-source agent harness, rather than on Claude Code, Codex, or Grok Build, the three heavier harnesses every other model on the table runs on. Gemini 3.8 Flash lands at 19.1 percent, ±3.4, using 17.2 billion tokens, the second-highest token count on the leaderboard behind only Sonnet 5, and Gemini 3.7 Flash lands lower still at 11.2 percent, ±2.4, using 11.1 billion tokens. It is a genuinely fair question, one the leaderboard's own model-versus-harness split is built to invite, whether these two Gemini results reflect the underlying models' actual terminal-reasoning capability, or whether a meaningful part of the gap to the top of the table reflects mini-SWE-agent's lighter tool-calling scaffolding compared to the more mature orchestration Claude Code and Codex provide. Until someone runs the same Gemini models on Claude Code or Codex and publishes that comparison, the honest reading of rows nine through fourteen on this table is that they conflate two variables, model capability and harness quality, in a way the top of the table does not, since GPT-6 Astra, Fable 5.1, and Opus 5 are all effectively harness-matched against each other on Codex or Claude Code specifically.

Reading the leaderboard as a whole, a few patterns hold up across every row rather than just the anomalies picked out above. Claude Code appears as the harness of choice for every Claude-family model tested, GLM-5.3, and is notably the harness carrying the largest number of distinct models on the table, which makes the intra-harness comparisons above (Fable 5.1 versus Opus 5, Opus 5 versus Opus 4.8, Sonnet 5 against everything else on Claude Code) more directly meaningful than any comparison that also crosses a harness boundary. Codex carries every OpenAI model tested, and its top-of-table result with GPT-6 Astra, paired with a comparatively low token count, suggests a harness that is at minimum not adding meaningful overhead to a genuinely capable model. Grok Build has run exactly two models, Grok 4.6 and Grok 4.5, both landing in the lower half of the table, and GLM-5.3's fifth-place finish at 41.8 percent on Claude Code, using 8.7 billion tokens for a comparatively low $2.7k, is worth flagging on its own terms as a genuinely strong price-to-performance result from a non-Western lab's model running on Anthropic's own harness, a combination that only exists on this leaderboard because of the model-harness separation this section has spent most of its length arguing for.

A Worked Example: What a Good Trajectory Looks Like Next to a Bad One

The mechanism described above, decisive convergence versus unproductive looping, is easier to hold onto with a concrete illustration than with prose alone. What follows is deliberately illustrative pseudocode of a tool-call sequence, not a real transcript from any specific model run on Terminal-Bench, since actual agent transcripts are not published per-model on the leaderboard. It is built to show the structural difference between the two failure modes described in the leaderboard analysis above, using a representative "operations" category task: a web service that fails to start because of a misconfigured environment variable and a stale lock file left behind by a previous crash.

python
/code # ILLUSTRATIVE pseudocode only. This is NOT a real transcript from any # specific model evaluated on Terminal-Bench. It sketches the structural # difference between a decisive trajectory and an unproductive looping one # on the representative "fix a service that won't start" task described # earlier in this post, to show why token count and tool-call count track # trajectory efficiency, not raw reasoning quality. # --------------------------------------------------------------------- # TRAJECTORY A: decisive, converges in 6 tool calls # --------------------------------------------------------------------- trajectory_decisive = [ {"call": 1, "tool": "run_command", "cmd": "systemctl status orders-db"}, # -> reads real error: "FATAL: lock file exists, another postmaster # may be running" AND "max_connections exceeds available shared memory" {"call": 2, "tool": "run_command", "cmd": "cat /var/lib/postgresql/data/postmaster.pid"}, # -> confirms the referenced PID is not an active process: stale lock {"call": 3, "tool": "edit_file", "action": "remove stale postmaster.pid"}, {"call": 4, "tool": "edit_file", "action": "lower max_connections in postgresql.conf to a value the " "container's shared memory actually supports"}, {"call": 5, "tool": "run_command", "cmd": "systemctl restart orders-db"}, # -> starts cleanly {"call": 6, "tool": "run_command", "cmd": "psql -c 'SELECT count(*) FROM orders;'"}, # -> confirms data intact BEFORE declaring done; matches verifier logic ] # result: PASS, ~6 tool calls, minimal re-reading, hypothesis checked # before acting, self-verified before finishing # --------------------------------------------------------------------- # TRAJECTORY B: unproductive, loops for dozens of calls before timeout # --------------------------------------------------------------------- trajectory_looping = [ {"call": 1, "tool": "run_command", "cmd": "systemctl status orders-db"}, {"call": 2, "tool": "run_command", "cmd": "systemctl restart orders-db"}, # -> restarts without reading the actual error first; fails again {"call": 3, "tool": "run_command", "cmd": "journalctl -u orders-db"}, # -> re-reads largely the same error already seen at call 1 {"call": 4, "tool": "run_command", "cmd": "lsof -i :5432"}, # -> chases a plausible but incorrect theory: assumes port conflict {"call": 5, "tool": "run_command", "cmd": "kill -9 <unrelated-pid>"}, {"call": 6, "tool": "run_command", "cmd": "systemctl restart orders-db"}, # -> still fails; port was never the real issue {"call": 7, "tool": "run_command", "cmd": "journalctl -u orders-db"}, # -> re-reads the SAME log a third time with no new extraction of info # ... calls 8 through ~40 repeat variations of restart / re-read / # partial-config edits without ever isolating BOTH root causes # (stale lock file AND max_connections) at once, each verbose # intermediate reasoning block adding tokens without narrowing # the search space ... {"call": 41, "tool": "edit_file", "action": "finally removes stale pid"}, {"call": 42, "tool": "run_command", "cmd": "systemctl restart orders-db"}, # -> still fails: max_connections was never fixed # timeout reached before both root causes are resolved together ] # result: TIMEOUT, dozens of tool calls, heavy re-reading and repeated # failing commands, no self-verification step ever reached, far higher # token consumption for a worse outcome than Trajectory A def summarize(trajectory, outcome, tokens_estimate): return { "tool_calls": len(trajectory), "outcome": outcome, "tokens_estimate": tokens_estimate, } print(summarize(trajectory_decisive, "PASS", tokens_estimate="low")) print(summarize(trajectory_looping, "TIMEOUT", tokens_estimate="very high"))

The decisive trajectory in that illustration reaches a verifiable, correct state in six tool calls. It reads the actual error output once, forms a specific hypothesis, checks that hypothesis directly rather than guessing, fixes the two real root causes it found, and verifies its own fix before declaring done. Every tool call narrows the search space. The unproductive trajectory, by contrast, spends its first several tool calls on actions that produce information the agent already effectively had (re-reading the same log after a cosmetic restart, re-running an identical failing command with no new information gathered in between), takes an ordinary port conflict as if it were the primary problem before actually confirming that theory, and only reaches something resembling the two real fixes deep into the sequence, after enough wasted turns that a real eight-hour budget spent this inefficiently across dozens of comparably-scoped subtasks in a longer task would plausibly run out before the agent ever reaches its correct final state. Multiply that inefficiency ratio across a full task category's worth of similarly structured problems, and the aggregate gap between a model that reliably takes the first shape of trajectory and one that reliably takes the second is exactly the kind of gap that shows up as tens of percentage points of resolution rate and billions of extra consumed tokens on a real leaderboard, without either model's raw per-token reasoning necessarily being worse in isolation.

Production Best Practices: Choosing and Evaluating an Agent Harness for Your Own Team

If you are picking an agent harness or a model configuration for your own development team, rather than just reading about someone else's benchmark, a few concrete lessons fall directly out of the analysis above.

Evaluate model and harness together, never a model score in isolation. The single biggest methodological lesson from this leaderboard is that a model's raw capability and the scaffolding wrapped around it are genuinely separable variables that interact in ways a headline percentage hides. Before adopting a new model inside your existing agent tooling, or a new harness for a model you already trust, run your own small internal benchmark on both axes independently where you can, rather than assuming a strong score on someone else's harness transfers cleanly to yours.

Treat token and cost columns as load-bearing metrics, not footnotes. A model that reaches 90 percent of another model's resolution rate using a third of the tokens is very often the better production choice, especially for a team running an agent against real infrastructure continuously rather than as a one-off benchmark pass, where token consumption translates directly into a recurring bill and trajectory length translates directly into how long a human has to wait before an agent either finishes or needs to be interrupted.

Do not assume "max" effort is always the right setting. Sonnet 5's result on this leaderboard is the clearest possible counterexample to the intuitive assumption that a more generous reasoning or tool-call budget can only help. A model whose long-horizon planning is weaker relative to its per-turn reasoning can turn a generous budget into a longer, more expensive path to a worse outcome. Where a harness exposes an effort or reasoning-budget setting, treat it as a hyperparameter to actually tune against your own task distribution, not a dial to always max out by default.

Match your evaluation harness to your production harness. If your team's actual production agent runs on Claude Code, a benchmark result for the same underlying model measured on a different, unrelated harness like mini-SWE-agent tells you comparatively little about how that model will behave in your own pipeline. The apples-to-oranges problem this post raised about the two Gemini Flash models applies just as much to your own internal evaluation setup as it does to a public leaderboard: always benchmark on the harness you intend to ship, not the harness that happens to be convenient or free.

Weight margin-of-error bands over single-point rankings. GPT-6 Astra and Fable 5.1 sit within each other's combined margin of error at the top of this table, and treating that gap as a meaningful, decided ranking rather than a statistical tie is exactly the kind of overconfident reading a five-trial methodology with reported error bars is specifically built to discourage. When two options fall within each other's margin of error, the tie-breaking factor worth actually weighting your decision on is usually cost or token efficiency, not the last decimal point of the resolution rate.

Build your own task-and-verifier discipline even at a much smaller scale. Terminal-Bench 4.0's own maintenance history, eight tasks pulled for saturation, refusal-proneness, leaked solutions, and platform inconsistency, plus nineteen more fixed for verifier bugs, is a useful template for any team building internal agent evaluations. A task graded by an LLM reading its own transcript is a weaker signal than a task graded by an independent script checking the sandbox's actual final state, and a task that has become trivially easy for every model you test should be retired from your suite rather than left in to quietly inflate every score.

Common Mistakes People Make Interpreting Agentic Benchmarks Like This

Conflating model score with harness score. The single most common misreading of a table like this one is treating every row as a clean, apples-to-apples measurement of "how good is this model," when a meaningful share of several rows' variance traces back to which agent harness is driving the model rather than the model's own underlying capability. Fable 5.1's 57.9 percent and Opus 5's 53.9 percent are both Claude Code results and are directly comparable to each other in a way that neither is directly comparable to Gemini 3.8 Flash's 19.1 percent on the very different mini-SWE-agent harness.

Ignoring cost and token efficiency in favor of the headline percentage alone. The Sonnet 5 and Grok 4.5 tie at 12.4 percent is the cleanest demonstration on this entire leaderboard of why resolution rate by itself is an incomplete comparison. Two models with an identical top-line score differed by more than 6x in token consumption and more than 4x in dollar cost to get there, and a reader who only looks at the percentage column would never notice.

Assuming a higher effort or reasoning-budget setting always improves results. It is intuitive to assume that "max" effort must outperform a lower setting, since it is described as the more generous option, but Sonnet 5's run demonstrates the opposite can happen when a weaker long-horizon planner is given more room to wander rather than more room to think productively.

Reading a benchmark tie as a decisive ranking. Treating GPT-6 Astra's 58.2 percent as a meaningfully better result than Fable 5.1's 57.9 percent ignores that both scores carry margins of error, ±2.8 and ±3.8 respectively, wide enough to overlap substantially. A one-point gap inside overlapping error bars is not a real, reproducible ranking, it is noise the five-trial methodology is specifically reporting so readers do not over-interpret it.

Assuming task removal or task fixing between benchmark versions is a sign the benchmark is getting easier or less rigorous. The opposite is closer to the truth here. Removing eight tasks for saturation, refusal-proneness, leaked solutions, and platform inconsistency, and fixing nineteen more for verifier bugs, is exactly the kind of unglamorous maintenance that keeps a benchmark's remaining scores trustworthy, and a shrinking task count paired with a documented, specific reason for every removal is a stronger signal of rigor than a benchmark that only ever adds tasks and never revisits old ones.

Assuming a benchmark built for terminal and coding work generalizes cleanly to every kind of agentic task. Terminal-Bench measures agentic terminal competence specifically, real sandboxed systems and software work with an automatically verifiable final state. A model's ranking here says relatively little, on its own, about how that same model would perform on an open-ended browsing task, a customer support conversation, or a creative writing assignment, since the entire evaluation design depends on having a checkable ground truth in the sandbox's final state, a property most other agentic domains do not share.

Where This Connects Back to Building Reliable AI Pipelines

The core lesson underneath Sonnet 5's result generalizes well beyond terminal benchmarks: a genuinely capable model can still underperform badly when it is dropped into an open-ended loop and given room to wander, because more freedom to iterate is not the same thing as more actual progress toward a correct outcome. That is a large part of why Miraflow AI's own content pipelines, like Text2Shorts, are built as a fixed sequence of well-scoped steps, script generation, then scene visuals, then voice, then final assembly, rather than as one open-ended agent loop deciding at each moment what to do next. A creator gets a consistent, fast result because each stage has one clear job and a clear handoff to the next, the same design principle that separates the decisive six-step trajectory from the looping one earlier in this post. It is not that an agent loop is the wrong tool everywhere, Terminal-Bench itself exists because open-ended agent loops are genuinely useful for messy, unpredictable systems work, but for a production content pipeline where consistency and speed matter more than open-ended flexibility, fixed, well-scoped steps are often the more reliable architecture, and it is worth trying Miraflow AI directly to see that pipeline design in practice.

Frequently Asked Questions

What is Terminal-Bench 4.0? Terminal-Bench 4.0 is the newest refresh of Terminal-Bench, an agentic-terminal benchmark built by the Laude Institute and open-source contributors on the Harbor framework. It tests whether an AI agent can complete real software engineering and systems tasks entirely through a real terminal, with no graphical interface, across 66 sandboxed tasks spanning seven categories, and each task is scored by an automated verifier checking the sandbox's actual final state rather than the agent's own self-report [1].

Why did the task count drop from 74 to 66 between version 3.0 and version 4.0? Eight tasks were removed, two each for four distinct reasons: saturation (every model already passed reliably), refusal-proneness (models declined the task for reasons unrelated to terminal skill), leaked public solutions (contamination), and unresolved quality or platform-compatibility issues. A further nineteen surviving tasks were fixed rather than removed, correcting instructions, environments, or verifier bugs [1].

Why does the leaderboard separate model from agent harness? Because the same underlying model can score very differently depending on the scaffolding driving it. Claude Code, Codex, Grok Build, and mini-SWE-agent each handle tool calls, retries, and task-completion decisions differently, and a headline score without the harness column can mislead a reader into crediting or blaming the wrong variable for a given result [2].

Why did Sonnet 5 score so low despite using the most tokens and the highest cost? The official Terminal-Bench 4.0 announcement notes that Sonnet 5 sometimes hit timeouts. The plausible technical explanation is that a model with weaker long-horizon planning, given a generous "max" effort setting inside an agentic loop, can fall into long unproductive trajectories, re-reading files, repeating failing commands, generating verbose intermediate reasoning, that consume the flat eight-hour timeout without ever reaching a verifiable correct state, while a more decisive model reaches the same or a better outcome in far fewer tool calls [1].

Is Fable 5.1 actually more efficient than Opus 5? On token count, yes clearly: Fable 5.1 used 2.7 billion tokens against Opus 5's 6.9 billion, roughly 2.6 times fewer, while scoring higher, 57.9 percent against 53.9 percent. Total dollar cost between the two is close, $6.2k for Fable 5.1 against $6.1k for Opus 5, so the clean efficiency win is specifically in token count and resolution rate together, not necessarily in total dollar cost [2].

What is the Harbor framework? Harbor is the open-source framework Terminal-Bench is built on, providing the sandboxed environment tooling, task packaging, and evaluation infrastructure that Terminal-Bench and related benchmarks use. A related repository, terminal-bench-science, applies the same framework to evaluating agents on scientific research workflows specifically [4].

How is Terminal-Bench different from SWE-bench or HumanEval? HumanEval tests a model's ability to write one isolated function from a docstring, with no execution loop and no environment. SWE-bench is a real step toward realism, testing whether a model can patch an actual reported GitHub issue inside a real repository, but it typically evaluates a single code change rather than an extended terminal session. Terminal-Bench goes further, giving an agent full terminal access to a real sandboxed environment and judging only the final, automatically verified state of that environment after however many commands the agent needed [5].

Conclusion

The number worth remembering from Terminal-Bench 4.0 is not the 58.2 percent at the top of the table, it is the 21.6 billion tokens and $9.6k spent to reach 12.4 percent near the bottom. A benchmark built specifically to expose the gap between raw model capability and actual agentic reliability did exactly what it was designed to do: it showed that more budget, more effort, and more tokens do not automatically translate into a better outcome when a model's long-horizon planning cannot make productive use of the room it is given, and it showed that same lesson clearly enough to separate GPT-6 Astra, Fable 5.1, and Opus 5's disciplined, efficient trajectories from Sonnet 5's expensive, unproductive ones, all under the same flat eight-hour timeout and, in Sonnet 5's case, the exact same Claude Code harness that carried two other models to the top of the leaderboard. The model-versus-harness separation that makes this leaderboard genuinely useful, rather than just another ranked list, is worth carrying into how you evaluate any agentic tool for your own team: ask what the model can do, ask what the harness adds or costs on top of that, and never trust a single headline percentage to answer both questions at once.

References

  1. Terminal-Bench 4.0, official announcement
  2. Terminal-Bench 4.0 Leaderboard, Snorkel AI
  3. Terminal-Bench 4.0, Artificial Analysis
  4. terminal-bench-science, Harbor Framework on GitHub
  5. Understanding LLM Code Benchmarks: From HumanEval to SWE-bench, Runloop
  6. terminal-bench, Harbor Framework main repository on GitHub