Codex vs Claude Code, from someone who wired into both
I run both agents daily and built integrations against both of their extension systems. This comparison skips the benchmark reposts and covers what actually differs when you live in them: automation surface, session signals, workflow shape, and how to pick.
Most Codex vs Claude Code comparisons are benchmark screenshots and pricing tables copied from two landing pages. This one is different in a specific way: I run both agents every day, and I built a product that integrates with both of their extension systems, which means I have read their hook payloads, their config files, and some of their source. That seat shows you differences the benchmark posts never touch.
Updated September 13, 2026: Codex now ships lifecycle hooks, stable and on by default. When this post first went up, notify was Codex's only automation surface, and the sections on automation and session signals said so. Those sections are revised below. The details of the new hooks are in Codex hooks and config.toml.
Quick disclosure up front: I make Unwait, a macOS app that shows flashcards while either agent works. I need both tools to be good, and I have no stake in which one you pick. Where the two differ, I will say so plainly; where they are equivalent, I will say that too.
Where they are the same
Both are terminal coding agents: you describe work, they read your repo, edit files, run commands, and come back. Both handle multi-file changes, both respect project instruction files, both ship permission systems so the agent asks before dangerous commands. Both are bundled with subscriptions you may already pay for: Codex signs in with a ChatGPT plan (Plus and up) and Claude Code with a Claude plan. For day-to-day "implement this function, fix this bug" work, either one is genuinely fine, and anyone telling you one is unusable is selling something.
The differences show up when you push past the chat loop.
The automation surface has converged
This is the gap I know best, because I built against both.
Claude Code exposes a full lifecycle hook system. Dozens of events fire as JSON on stdin: prompt submitted, tool about to run, tool finished, turn ended, permission requested, session started and ended. Hooks can observe, but they can also decide: block a dangerous command, refuse to end the turn until tests pass, inject context. I documented the payload contract and six configs worth stealing in earlier posts. The short version: Claude Code treats external tooling as a first-class audience.
Codex now has one too. When this post was first written, Codex exposed a single key, notify, which runs a command when a turn completes and passes JSON as an argv argument, with no start event, no tool events, and no way to veto anything. A richer hook system was already sitting in Codex's source, with notify reimplemented as a compatibility layer named legacy_notify. It has since shipped: twelve lifecycle events including UserPromptSubmit, PreToolUse, PermissionRequest, and Stop, JSON on stdin, and the ability to block, configured in hooks.json or config.toml. notify still works, and its sharp edges still apply if you use it: it must be a top-level TOML key, it executes without a shell, and only one tool can own it at a time.
The consequence I originally described was that Claude Code let you measure a wait, with a signal when the prompt was submitted and another when the turn ended, while Codex only told you a turn had ended. That was why our wait-time features got the full experience on Claude Code and a reduced one on Codex. With UserPromptSubmit and Stop, Codex now exposes both ends too, so that gap is no longer a property of the products. What remains is maturity: Claude Code's hook system has been public longer and has more existing tooling and documentation built on it, and the two differ in how they decide which hooks to trust.
Session signals and multi-session work
Both agents can now distinguish "the turn ended" from "the agent is blocked waiting for your permission". Claude Code fires a separate event for the second, and Codex's hooks include PermissionRequest. The old notify key still cannot tell you an agent is blocked, only that it finished, so if your Codex setup is notify alone, a blocked agent looks exactly like a working one. That matters little with one session and a lot with three in parallel, where unnoticed blocks are the most expensive dead time in the workflow. Both payloads carry a session id and working directory, which is what makes per-project, per-session tooling possible.
Workflow shape
Differences you feel in the first week, stated without benchmark theater:
- Config and instructions. Both read project instruction files (CLAUDE.md and AGENTS.md respectively); the convention is close enough to maintain both from one source.
- Extensibility beyond hooks. Both now have plugins, MCP servers, and subagents. Claude Code's ecosystem layer, including skills, has been around longer and has more built on it. Codex still feels closer to "one very capable agent in a box" by default, which reads as clean or limiting depending on how much you customize.
- Interfaces. Both have CLI and IDE-adjacent options, and both can run in the cloud tethered to their vendor's web UI. Codex leans on the ChatGPT apps; Claude Code on Claude's.
- Models. Each is locked to its vendor's models, so part of this choice is really "whose frontier coding model do you trust this quarter", and that answer has flipped multiple times in the past year. I would not sign an annual contract over a model benchmark.
Pricing
Both come bundled with consumer subscriptions, both meter usage against your plan tier, and both sell API-key escape hatches. The numbers change too often to print here; check the official pricing pages the week you decide. The structural note that stays true: if your team already pays for ChatGPT or for Claude, the bundled agent is effectively discounted, and that gravity decides more purchases than any feature list.
How to actually pick
- You already pay for one vendor: use that one. The bundled agent is good enough that switching subscriptions for the other is rarely justified.
- You build tooling, or want notifications, gates, and integrations: either now works. Both expose comparable lifecycle hooks; Claude Code's has the longer track record and more existing tooling, so it is the lower-risk pick today, not the only one.
- You want minimal surface area: Codex's one-binary, one-config setup is a legitimate preference, especially if you rarely touch automation.
- You run agents in parallel all day: you need start, end, and blocked signals with per-session ids to keep N sessions honest. Both agents' hook systems now provide them; Codex's
notifyalone does not. - Either way: they coexist fine. I run both daily on the same machine, and plenty of developers keep the second one around for hard problems, second opinions, or when one vendor has a bad model week.
If the wait time between prompts is the part of agent work that bothers you, that is the exact problem Unwait exists for, on both agents. And if you want the deeper technical dives behind this comparison, the hooks payload reference, the Codex notify writeup, and Codex hooks and config.toml are where this post came from.