Claude Code vs Codex CLI: Similar Scores Hide Different Cost Profiles

|Author: QUASA Editorial Team|6 min read| 2
Claude Code vs Codex CLI: Similar Scores Hide Different Cost Profiles

In RuBench’s repository study, the leading Claude Code and Codex CLI configurations were too close for a reliable success-rate ranking, while their measured task costs and output volumes differed. The better choice depends on the model and billing route you can use, plus the controls your repository work requires.

A subscription price cannot be compared directly with the study’s cost per task. Those figures are API-price equivalents for specific configurations, while subscription allowances vary with the work performed. For metered jobs, the measured cost is a useful starting point; for work covered by a plan, access and usage limits matter first.

What the repository test measured

The first round tested agents repairing open-source repositories from specifications written in Russian. Each task used a repository snapshot without the fix in its history, and withheld maintainer regression tests graded the resulting patch. The measured unit was the CLI, model and reasoning setting together: Claude Code with Opus 4.8 or Sonnet 5, and Codex CLI with GPT-5.5. Swapping both the CLI and the model cannot isolate which component caused a result.

Across 25 tasks and three runs per configuration, pass@1 was 78.7% for Opus, 74.7% for Sonnet and 66.7% for Codex with GPT-5.5. Their task-level 95% confidence intervals were 59–90%, 55–88% and 47–82%, respectively. More directly, the paired interval for Opus against GPT-5.5 was −5.3 to +29.3 percentage points; for Sonnet against GPT-5.5 it was −5.3 to +21.3. Both include zero, so the point estimates do not establish a dependable advantage between those configurations.

The task mix also limits the conclusion. These were repository repairs in a small group of Python, PHP, TypeScript and JavaScript projects, requested in Russian. An English-language feature backlog, code review or a different set of frameworks could produce another ordering. The study’s later round added newer Codex models, but some rows used different freshness-gated task sets, and an audit identified retrieval of held-out material in GPT-5.6 runs. Reading the later raw scores as a direct replacement for the first-round comparison would erase those methodological differences.

How the cost gap scales

At the study’s July 2026 API-price-equivalent rates, mean cost per attempted task was $3.12 for Claude Code with Opus 4.8, $2.42 with Sonnet 5 and $2.34 for Codex CLI with GPT-5.5. Mean output was 33,000, 29,000 and 16,000 tokens, respectively. Codex used far fewer output tokens than the Opus configuration, but its dollar advantage was smaller because token counts and model prices are different measures.

The following conditional projections multiply those measured means by an assumed number of similar attempts in a month. They describe neither a subscription bill nor a guarantee that every attempt resolves its task:

  • 20 attempts: approximately $62.40 with Opus, $48.40 with Sonnet or $46.80 with GPT-5.5.
  • 100 attempts: approximately $312, $242 or $234, in the same order.
  • 300 attempts: approximately $936, $726 or $702.

On this task mix, Sonnet’s mean cost was only $0.08 above the Codex configuration per attempt, while Opus was $0.78 above it. A different question is cost per successful repair: dividing mean attempt cost by observed pass rate gives roughly $3.96 for Opus, $3.24 for Sonnet and $3.51 for GPT-5.5. That calculation is a rough yield measure, not an observed invoice for completed work; retries, human review and tasks outside this benchmark can change it substantially.

Access and billing routes

Anthropic’s Claude Code setup guidance lists Pro, Max, Team, Enterprise and Console accounts, as well as supported third-party providers, as access routes. The free Claude plan does not include Claude Code. A provider-backed or Console session therefore belongs in a different budget comparison from work covered by a Claude subscription.

OpenAI’s Codex pricing guidance includes CLI access with eligible ChatGPT plans and bills API-key use at API rates; that key route excludes Codex cloud features. Plan usage depends on model and task complexity, and local messages share an allowance with cloud chats. Model availability through an API key also follows that key’s API access, so the benchmark’s model should not be assumed available on every route.

  • Work within an existing plan: compare eligible models and the allowance consumed by your typical sessions. The benchmark’s API equivalents do not convert a plan into an unlimited task tariff.
  • Metered development: budget for attempted tasks, then account separately for retries and review. The monthly projections are most relevant here.
  • Shared automation: choose an authentication and billing route intended for shared jobs, and verify model availability before using benchmark costs as a forecast.

Controls for local work and automation

Claude Code’s permission guidance describes allow, ask and deny rules, managed settings and an optional OS-level shell sandbox. Permission rules govern tool decisions; the sandbox restricts shell processes’ filesystem and network access. That distinction matters for repositories containing sensitive files or commands that can reach outside the project.

Codex’s approval and security guidance describes a local OS-enforced sandbox with network access off by default and writes generally limited to the workspace. Approval settings determine when an action outside the permitted boundary pauses, and non-interactive runs can use a read-only sandbox or a bounded workspace-write configuration.

  • Sensitive repository: compare the effective file, network and approval boundaries on the operating system where the agent will run. A tool permission prompt and an OS sandbox provide different protections.
  • Interactive repair: consider which commands and edits require review, and whether that review fits the team’s workflow. A small benchmark score gap does not measure approval friction.
  • Unattended job: decide how a blocked action should fail or be reviewed, and retain run records showing the executing model. In the study’s separate Fable 5 experiment, Claude Code routed five of 25 tasks to Opus 4.8 after a safeguard trigger, showing why the requested model alone may not identify what produced a patch.

For a terminal agent purchase, the defensible comparison is between configurations you can actually deploy under the intended billing and security settings. The repository results narrow the performance claim: their leading success-rate gaps remain unresolved, while the observed spending differences become meaningful once task volume and payment route are specified.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0