endueendue

Claude Code vs Codex: how to compare development time and cost

Evaluate coding agents with the same task, acceptance criteria, review time, and cost per successful result. Includes a reusable task prompt and recording table.

Two coding workbenches lead to a reviewed change and a stopwatch.
AI-generated conceptual illustration. It does not depict an actual product interface.

A coding agent can produce code quickly and still leave you with a long afternoon of corrections. When choosing between Claude Code and Codex, measure the time until a usable change is ready, including your review.

Both are agents that work with project files, change code, and run commands. A terminal is the text-based window used to enter such commands. Their supported working environments are described in the Claude Code documentation and OpenAI’s Codex guidance.

This article covers documented capabilities and an evaluation method as of October 5, 2026. It does not report measured runtime or win rates from our own comparison.

A bounded issue progresses through code changes, tests and a reviewable change package.
A bounded issue progresses through code changes, tests and a reviewable change package. AI-generated conceptual illustration. It does not depict an actual product interface.

Why an old comparison may need updating

Anthropic introduced Sonnet 5.5 on September 28, reporting improvements in coding and efficiency. OpenAI announced GPT-6.1 Sol and reusable Codex cloud environments on September 29. Changes to both models and working environments can affect your experience. Anthropic announcement, OpenAI announcement

Comparison area Question to ask
Working environment Can it install, run, and test your project?
Model and settings Which model and reasoning setting did you use?
Review Can you inspect changed files and executed commands?
Team workflow Does its review process fit how your team works?

A reasoning setting controls how much effort a model spends preparing and checking its work. Similar setting names across products do not establish equal computation. Record settings so that the result has context.

Define one small, concrete task

For example, ask both agents to add a CSV download to an existing orders page. CSV is a file format for storing tables as separated text values. Give both agents the same request:

Add CSV export to the existing orders list.
Apply the selected screen filters to the exported rows.
Preserve non-ASCII text, commas, quotes, and line breaks in values.
Keep existing sign-in and order retrieval behavior working.
Run relevant tests and explain changed files and verification results.
Do not make unrelated design changes or deploy anything.

Start from the same project revision, with a separate copy for each agent. Match the runtime and test commands. Neither agent should see the other’s completed changes.

Tests check whether a program behaves as required. Include checks you wrote before the experiment, such as filtered exports, empty lists, and non-ASCII text. Agent-written tests alone can miss the same mistake as the implementation.

Record what happens

Record What to include
Run conditions Tool version, model, reasoning setting, starting revision
Elapsed time From the request until you accept the change
Human time Clarifications, review, and manual corrections
Result Acceptance criteria met and defects remaining
Usage Subscription allowance consumed or actual API charges

Repeat with several tasks of the same type and retain failures as well as successes. A single attempt is a weak basis for choosing a tool. Check that leftover files or caches do not silently change the conditions of later runs.

Divide cost by accepted results

The following is a hypothetical arithmetic example, not a measurement of either product.

Hypothetical tool Attempts Total API cost Accepted results Cost per accepted result
A 10 US$4 8 US$0.50
B 10 US$3 5 US$0.60

The calculation is total cost / accepted results. If no result is accepted, the ratio is undefined; record the evaluation as unsuccessful. Human review and execution infrastructure also belong in a practical cost assessment.

For subscription usage, do not invent API charges. Record work completed within the allowance and any extra payments separately. OpenAI explicitly distinguishes subscription usage from API rates. Pricing guidance

Choose from one completed change

Begin with whichever agent is easiest to use in your existing account and development environment. Then run the same task with the other. Check required behavior and reviewability as well as whether the code runs. This gives you evidence specific to your project.

Example: write the acceptance checks before the prompt

For the CSV export task, create three small sample rows: one containing a comma, one containing a quotation mark and one containing a line break. Add a non-ASCII name and a filter that selects only one row. The expected exported rows should be clear before either agent starts.

Review the download in the application and parse the file with a CSV reader. Then inspect the changed files and the test results. A passing test suite is useful evidence, but it does not establish that unrelated login or access behavior remained correct unless those paths were covered. Keep the acceptance checklist with the result so another person can repeat the review.

Common questions

Can I compare only the time until code first appears?

That measures a small part of the job. Record the time until the change meets the acceptance criteria, including troubleshooting and review. Also keep human working time separate from unattended waiting.

Is one successful task enough to choose a team tool?

Treat it as an initial result. Repeat with several tasks resembling the team’s work, using the same starting conditions. Record failures and help required rather than discarding inconvenient runs.

Continue the series.