Codex vs Claude Code vs Gemini CLI: May 2026 Verdict

Codex CLI is OpenAI's terminal coding agent, built on the GPT-5 Codex line and designed for short, fast, sandboxed task loops in the shell or as async runners inside Codex Cloud.
Claude Code is Anthropic's terminal-native agent, built on Claude Sonnet 4.6 and Opus 4.7, designed for long, unattended coding sessions with deep repo context, subagents, and a Mac and Windows desktop app shipped in May 2026.
Gemini CLI is Google's open-source terminal agent (Apache 2.0), running Gemini 2.5 Flash by default for free users and Gemini 2.5 Pro on paid plans, designed for high-context coding and research with the largest free tier on the market.
If you ship production code in May 2026 and the bill is not the bottleneck, Claude Code wins for unattended PR work. If you live inside OpenAI's ecosystem or need the strictest runtime sandbox, Codex is the safer second pick. If you are broke or auditing the source line by line, Gemini CLI is the only one of the three that is genuinely free and genuinely open, and that is its whole story.
Here is the surprise. These three tools are not really competing in the same league. Their pricing models, runtime contracts, and target users barely overlap. "Which is best" only makes sense after you answer "best at what."
The short answer
For solo devs shipping side projects: Gemini CLI on the free tier, until Flash stops being enough.
For production engineering teams shipping unattended PRs: Claude Code Max 5x at $100 per month. Opus 4.7, subagents, 128K output tokens per turn.
For OpenAI-stack teams with hard sandboxing requirements: Codex CLI on a Plus or Pro plan, run inside Codex Cloud's Docker isolation.
If you cannot decide: start with Claude Code Pro at $20 per month. Switching cost to Codex or Gemini later is a day of config, not a quarter of pain.
Why this comparison matters now

Three things shifted in the last sixty days, and they reshape the choice.
In May 2026, Anthropic doubled Claude Code's rate limits across Pro and Max tiers and shipped a desktop app for Mac and Windows. Subagents got promoted out of beta in the same release.
On April 2, 2026, OpenAI flipped Codex from per-message billing to token-based pricing across Plus, Pro, ChatGPT Business, and new Enterprise plans. Teams that had budgeted on the old model rebudgeted in a hurry.
On March 25, 2026, Google restricted the Gemini CLI free tier to Flash models only. Free Pro access ended. The 1,000 requests per day allowance survived; the model behind it shrank.
Pricing and limits verified May 11, 2026.
Which one actually ships unattended work?
Claude Code, by a real margin. On a representative Express.js refactor task, Claude Code finished in 1h17m on a single unattended run, Codex CLI took 1h41m, and Gemini CLI took 2h04m, per public benchmark data published in early May 2026. The decisive variable is not raw model quality. It is output ceiling.
Claude Code emits up to 128K tokens per turn. Codex and Gemini cap closer to 64K. On a multi-file refactor that needs to touch eight files in one pass, the agent that does not have to segment finishes faster and breaks less. Add Claude Code's subagent system, which spawns parallel workers for independent tasks, and the gap widens on anything resembling real engineering work.
If your job-to-be-done is "describe a change, walk away, come back to a green PR," Claude Code is built for that loop. The same teams that previously wired up Claude Code for SEO work are pointing it at production refactors now because the runtime is built for the same shape of task: long, repo-aware, output-heavy.
Which is cheapest at the usage you will actually hit?

Gemini CLI is free up to 1,000 Flash requests per day. Claude Code Pro is $20 per month on rolling 5-hour windows of roughly 44,000 tokens each. Codex Plus is $20 per month with 10 to 60 Cloud tasks per 5-hour window. For a developer running an agent four hours a day, Claude Code Max 5x ($100) and Codex Pro ($200, currently giving 2x bonus usage through May 31, 2026) converge around the same cost ceiling.
The number that surprises people is the real all-in cost. Public estimates put a typical Claude Code developer at roughly $6 per day on Sonnet, which lands at $100 to $180 per developer per month at production usage. Codex API pricing for gpt-5.2-codex runs $1.75 per million input tokens and $14 per million output tokens, plus container fees of $0.03 to $1.92 per task. Real Codex spend for a working developer tracks Claude Code closely, give or take 15 percent depending on tool-call frequency.
Gemini CLI is the only one of the three where the free tier is actually usable for non-trivial work. Flash handles autocomplete, small edits, and short refactors. It chokes on long agentic loops, which is exactly the moment you would upgrade.
Which one is safest to hand the keys to?
Codex CLI, by infrastructure design. Every task runs in a Docker sandbox. Filesystem and network access are scoped per task. Claude Code uses permission prompts and project-level config, which is safer than nothing, less rigorous than container isolation. Gemini CLI is fully open-source under Apache 2.0, which means you can read every line of the runtime and decide for yourself.
For regulated industries, the choice is "Codex if you trust OpenAI's runtime, Gemini CLI if you only trust what you can audit." Claude Code is safer for the user (the permission prompts catch a lot of accidents) but not safer by infrastructure (no native sandbox).
This is the dimension most comparison articles miss. The "Codex is more cautious" reputation is not vibes. It is a runtime design decision that traces directly to the OpenAI side of the safety debate that played out in early 2026.

The comparison table
| Dimension | Codex CLI | Claude Code | Gemini CLI |
|---|---|---|---|
| Model (May 2026) | GPT-5.2 Codex | Sonnet 4.6 / Opus 4.7 | Gemini 2.5 Flash (free), Pro (paid) |
| Entry price | $20/mo (Plus) | $20/mo (Pro) | Free |
| Real working cost | ~$100–$200/dev/mo | ~$100–$200/dev/mo | $0 to $25/mo on Pro |
| Output tokens per turn | ~64K | 128K | ~64K |
| Context window | 400K | 200K | 1M |
| SWE-bench Verified | 85.0% (Codex), 88.7% (GPT-5.5) | 87.6% (Opus 4.7) | 76.2% (Gemini 3 Pro) |
| Sandboxing | Docker per task | Permission prompts | Open-source, audit yourself |
| MCP support | Yes (since early 2026) | Native (Anthropic owns the spec) | Yes (built-in) |
| Open source | No | No | Yes (Apache 2.0) |
Benchmark scores from Codeant's May 2026 benchmarks.
Where each tool actually wins
Claude Code wins for shipping unattended PRs on real codebases. The combination of 128K output, subagents, MCP-native tooling, and the rolling 5-hour window that resets predictably makes it the only one of the three you can hand a feature ticket to and walk away from. The desktop app shipped in May 2026 closed the last "but I want a UI" objection from teams that resist terminal-only flows.
Codex CLI wins for sandboxed background work and legacy code. The Docker isolation per task is the only reason to pick it over Claude Code on price-equivalent plans. It is the right tool when you do not trust the agent yet, when you are running multiple agents in parallel on the same machine, or when you are doing the kind of legacy code migration where blast radius matters more than throughput. Cloud's async runners let you queue 20 tasks before lunch and review them after.
Gemini CLI wins for solo developers, learners, and anyone who needs the source code. A genuinely free 1,000 requests per day with a 1M-token context window on Flash is unmatched. The Apache 2.0 license means you can fork it, ship internal builds, or pin to a specific version forever. For students and indie hackers, this is the only honest answer. For anyone who needs to know what is coming out of Google's broader model roadmap, the CLI is the cheapest way to track Flash-tier capability shifts.
Where each tool actually loses
Claude Code loses on free tier and pricing surprises. There is no free tier. Pro's rolling 5-hour window resets in ways that catch new users mid-task. The "extra usage" toggle introduced in 2026 helps, but it shifts spend onto API rates with a cap you have to manage.
Codex CLI loses on the April 2 pricing flip and container fees. Teams that budgeted per-message in March 2026 rebudgeted in April. Container fees of $0.03 to $1.92 per Cloud task look small until you fan out 200 tasks a day. The token-based model is fair, but it is harder to forecast than the old per-message structure.
Gemini CLI loses on Flash's ceiling and the March 25 free-Pro deprecation. Flash is fine for autocomplete, weak for serious refactors. Anyone doing real engineering work upgrades to Pro within two weeks, which means the "free CLI" pitch quietly becomes a $20 to $25 per month tool for production users. The open-source license is the durable advantage; the free tier is a trial.
The non-obvious insight
All three tools score within 12 percentage points on SWE-bench Verified. The model gap is real but smaller than the marketing implies. The real spread is the runtime contract: how long the agent can run, how much it can output per turn, what happens when a limit fires, and how the platform sandboxes the work.
Pick the runtime, not the model. This is the same trap that the Opus 4.7 versus GPT-5.5 decision falls into when teams optimize on benchmark deltas that will be reversed in six weeks. Benchmarks age. Runtime architecture does not.
The bottom line
If you have $100 per month to spend and a real codebase, buy Claude Code Max 5x and stop reading comparison articles. If you cannot justify the spend, run Gemini CLI's free tier today and upgrade only when Flash actually stops being enough for your workload. If you are already paying for ChatGPT Pro, Codex is in the box. Use it, especially for sandboxed work where the per-task isolation earns its keep.
Two specific moves: do not pick on benchmarks, pick on runtime fit. And do not pay for two of these. Pick one for 30 days, decide, switch if it does not fit. Switching cost is a day of config, not a quarter of disruption.
FAQ
Is Claude Code worth $20 per month if I already pay for ChatGPT Pro?
Yes, but only if you write code daily. Claude Code Pro's rolling 5-hour windows handle two to three focused coding sessions per day on Sonnet 4.6. If you use a coding agent only twice a week, the Codex usage already bundled into ChatGPT Pro is more than enough. You would be paying $20 for capacity you will not use.
Does Gemini CLI's free tier actually work for production?
Flash, the free tier model since March 25, 2026, handles autocomplete and small edits well. It struggles on multi-file refactors and long agentic loops. Treat the free tier as good enough to learn the CLI and prototype, not good enough to run as part of your engineering workflow. Production users upgrade to Pro within two weeks on average.
Can I run all three in the same repo?
Yes. The three tools use different config files: Codex uses AGENTS.md, Claude Code uses CLAUDE.md, and Gemini CLI uses GEMINI.md. They do not collide. Some teams run Claude Code for unattended refactors and Codex for sandboxed PR review on the same project, and that pattern works fine.
Which one has the best MCP (Model Context Protocol) support?
Claude Code has the deepest MCP support. Anthropic authored the protocol, and the largest ecosystem of MCP servers ships there first. Codex added native MCP support in early 2026 and now covers most public servers. Gemini CLI ships MCP tool calling natively as well. All three work; the ecosystem maturity gap favors Claude Code by about six months.
What about Cursor or Windsurf instead of a CLI?
Cursor and Windsurf are IDEs with embedded agents, not terminal CLIs. They compete more directly with Claude Code's desktop app than with the three CLIs in this comparison. If you want an editor that orchestrates AI agents inside a graphical interface, see Cursor 3's reinvented control-room model. If you want a CLI you can pipe into a Makefile, GitHub Action, or shell script, these three are the only serious players.
Which one is best for teams onboarding non-engineers?
Claude Code, because of the desktop app and the permission-prompt UX. Codex is too sandboxed to feel friendly to a non-engineer running their first agent. Gemini CLI is the right tool for the curious technical person who wants to read the source, not for the marketer running their first prompt-driven workflow.