Thursday, October 8, 2026Verified technology journalism

Claude guesses what you want, Codex does what it's told: a week on the other coding agent has Hacker News picking sides

An experienced Ruby developer who spent a week making OpenAI's Codex his primary coding agent has published the head-to-head practitioners keep asking for, and the Hacker News thread is still growing two days on. His findings: Codex writes leaner Ruby with fewer comments and simpler architecture, but its speed advantage evaporates once test reruns and review are counted; Claude Code anticipates intent and over-builds with abstractions; and under debugging pressure he still reached for Claude, not because it was better but because it was familiar. His capsule verdict on the two agents: Claude tries to go above and beyond and guesses what you might want, Codex does what you tell it and stops at the first sign of done.

Claude guesses what you want, Codex does what it's told: a week on the other coding agent has Hacker News picking sides

Claude guesses what you might want, Codex does what it's told: a week on the other coding agent has Hacker News picking sides

A week of using OpenAI's Codex more than Claude Code left Ruby and Rails developer Lucian Ghinda with a dead heat on the number shipping teams actually bill against: total time. Codex felt faster at writing the change, but finishing the pull request, rerunning tests, and review ate the lead. "In the end, there was no win in terms of time difference," he writes. 1

His post, ten numbered impressions he labels quick and very personal, went up Aug 21 and drew a Hacker News thread where practitioners reported the exact opposite of several of his findings, point for point. Classify his ten impressions by subject, as we did, and the more durable finding appears: two of the ten concern the code Codex wrote, one concerns speed, and seven concern the interaction. Session habits, tool logins, a branching mistake, and which agent he reached for under pressure. This was a workflow comparison wearing the clothes of a model comparison. 1 2

The speed win that vanished at the finish line

Ghinda says Codex "does changes faster than Claude," and then describes where the gain goes: after the main changes, finishing the pull request took a lot, through rerunning many tests, review, and what he calls the thoroughness of it. 1 Faster generation plus an equal total is not a mystery. It is arithmetic: the finishing phase consumed every second the writing phase earned. The speed you buy from an agent is a property of the loop, not of the model.

The tail can also bite. Asked to rebase one branch onto its target, Codex rebased onto main instead and handed him a pull request with 4000+ additions to untangle. The fix, he notes, was being explicit about the target, which means specification work moved back onto the human. 1

Seven of ten impressions were about the interaction, not the code

  • The code artifact: two impressions. Fewer comments in the Ruby and Rails changes, which he liked, and much simpler architecture. Claude Code, by contrast, reaches for abstractions, concepts, Sorbet signatures, and type aliases. 1
  • Timing: one impression. Faster changes, equal total time, as above.
  • Interaction and workflow: seven impressions. A skills asymmetry between the harnesses; opening Claude Code anyway when debugging felt urgent; Codex's more technical harness voice, which reads to him like Star Trek's Data where Claude is the colleague; a shift toward many small focused Codex sessions instead of one big Claude session; the rebase mistake; friction with Jira and Atlassian in his CLI setup; and a preference for Codex's explicit MCP login flow over Claude Code attempting it automatically in a turn. 1

The census is the tell. His capsule verdict describes an interaction contract, not an intelligence ranking: Claude Code guesses what you might want and acts on the guess, while Codex "does what you tell it" and "will stop at the first sign that it might be done." 1 In the one controlled pass in the post, both agents implemented the same requirement from the same documents: Claude's version was more complex, but it handled cases, which he does not enumerate. The over-build bought something. It bought coverage nobody asked for. 1

The thread argues with itself, and the disagreement is the finding

The Hacker News pushback runs mirror-image to the post. Kovah, on the architecture point: "Wow, I made exactly the opposite experience," with Codex loving to make things as complicated as possible and Claude behaving more pragmatically. enraged_camel reports the complete opposite of the capsule verdict itself. pupppet finds Claude Code gets intent without spelling it out while Codex over-engineers minor details, and aleksiy123 calls Codex Sol prone to overengineering and excess caution. corytheboyd runs the other direction, crediting Codex's speed and quieter output. One commenter simply asked why fewer comments is a good thing at all. 2

One reply explains the mirror. guywithahat, trained on Codex before his team adopted Claude, says Claude Code seemed to do everything wrong on first contact, because his instructions had been written for Codex; his takeaway is that agents have personalities you learn to drive. Set that against Ghinda's confession that when debugging felt urgent he opened Claude Code anyway: "I am not saying it was better, but it was familiar." 1 2 On our reading, the stable variable across the thread is not the model. It is the operator's trained reflexes: your current agent has already shaped how you write instructions, and the first week on the other one is spent unlearning them.

The limits are worth stating. Ghinda's post names no model versions, so read it as harness-plus-reflex data, not a benchmark. 1 2 The framework for a team choosing between the two: pick the failure mode your process can audit. Claude Code's over-abstraction lands in the diff, where review sees it. Codex's early stop lands at the finish line, where tests and reruns see it, and where, this week, it also landed the invoice. Then budget for the switching cost nobody prices, which is retraining the humans. Ghinda paid it himself, and he noticed. 1

References

1.All About Coding, Aug 21 2026allaboutcoding.ghinda.com ↗
2.Hacker News, Aug 2026news.ycombinator.com ↗

Cite this story

ProvenBrief (2026). "Claude guesses what you want, Codex does what it's told: a week on the other coding agent has Hacker News picking sides." ProvenBrief. https://provenbrief.com/story/claude-guesses-what-you-want-codex-does-what-it-s-told-a-week-on-the-other-codin

Free to quote and link with attribution. Republishing in full or AI-training use requires a license.

Verified33 factual claims in this story were independently checked against primary sources before publication; 4 unverifiable claims were removed during fact-checking. Read our editorial standards.

Get the next brief in your inbox

One weekly email. Every claim verified against primary sources before we hit send.

Produced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.