Haifeng (Kevin) Zhao

I had Codex and Pi build the same app in Orca, then compared them side by side

Code, data and reproduction steps: github.com/piscaries/codex-vs-pi-in-orca.

Different coding agents are good at different jobs, and what each one is good at shifts every time its model is updated. Which agent you put on which stage of the work directly shapes the quality of the software you ship. So how do you find out which agent fits which stage?

Tools like Orca and Herdr make that question answerable. They let you assign a different coding agent to each task in one development environment, which also makes it cheap to give two agents the same job and compare them. To show how, I set up a separate coding agent for each stage of development (spec, design, build) and staffed every stage twice: one team ran entirely on Codex, the other entirely on Pi. Both teams built the same app, a chess coach, from the same requirements, prompts and starting code. Then I compared them on two layers:

By the end of this post you will know:

The setup: two coding-agent teams, matched conditions, one chess problem

Both teams and the shared reviewer ran in Orca, an IDE built for coding agents; the next section shows how.

The comparison puts two coding-agent teams side by side: one built on Codex, running OpenAI's GPT-5.6 Sol, and one built on Pi, running Z.AI's GLM-5.3. Z.AI benchmarks GLM-5.3 against GPT-5.6 Sol in its own launch results, so the two models are from the same generation. An earlier run used GPT-5.5; it is kept in the repository.

Codex and Claude Code come fully set up around their vendor's models and tools. Pi is a bare agent by design: it starts with four basic tools, works with models from many providers, and leaves the rest to you to configure: prompts, skills, extensions. To keep the test fair, I checked that everything except the agent matched:

Condition Codex track Pi track
Model GPT-5.6 Sol GLM-5.3 (benchmarked by Z.AI against GPT-5.6 Sol)
Task text Identical Identical
Role prompt Top of the task message Top of the task message
Reasoning effort High (not its top setting, xhigh) High (not its top setting, max)
Personal configuration Clean config; Codex's bundled skills, a task-continuity plugin and Orca's status hooks None; skills off; Orca's status extension
Built-in tools Codex's own: web search, sub-agents, plugins and more (it used only shell commands and two web searches) read, bash, edit, write
Reviewer, judges, hidden tests Shared Shared

This compares two agent-and-model pairs: the agent and the model changed together, so the results cannot say which of the two made the difference.

Both agents get the same project: a chess coach for adult beginners. You play the computer. After each move, the coach tells you in plain words whether it was good and what was better, and at the end it reviews your worst moves. I chose chess for three reasons:

A fair contest gives both agents the same task, the same reviewer and the same settings, and it ends in a result no judge has to score.

How to build and run the two teams in Orca

To compare the two agents on real software work, I needed the same multi-agent engineering team twice: once staffed by Codex, once by Pi.

Orca is an agent development environment: an IDE built for coding agents. It runs Claude Code, Codex and other coding agents in parallel, each in its own git worktree and terminal, so agents working at the same time never edit the same files; Claude Code, Codex and Pi shared one window for this project. Orca also records the work as a Run (the whole job), Tasks (its pieces) and Dispatches (attempts at a task). A coordinator dispatches tasks, workers report back when they finish, and a worker can stop and ask the coordinator a question.

Orca with the Codex team's spec worktree, annotated

Figure 1. Orca with the Codex team's spec task: its own worktree and terminal, and Codex reporting back to the coordinator. Unrelated items are blurred.

The architecture is built for the comparison. The two teams are identical except for the author agent: same roles, same flow, one shared reviewer. The judging sits outside both teams.

Two identical teams in Orca, one per agent, with a shared reviewer between them; two blind judges score the specs, designs and finished apps, and both finished apps go through hidden tests, a coaching comparison, an engine match and an acceptance review.

Figure 2. The two teams differ only in the author agent. The reviewer sits between them; everything orange happens outside both teams.

Role Codex team Pi team
Product designer: writes the spec Codex (GPT-5.6 Sol) Pi (GLM-5.3)
Engineering designer: writes the design and build plan Codex Pi
Builder: implements one phase at a time Codex Pi
Code reviewer and acceptance reviewer Claude Code (Claude Opus 4.6) Claude Code (Claude Opus 4.6)
Coordinator: assigns work, answers questions me, through Orca me, through Orca

How a team is launched. Each role is a short prompt file that says how to do the job, what the output must look like and when it is done. A shared charter, TEAM.md, sets the rules: work only in your own files, report only checks you actually ran, and ask instead of guessing. Here is the start of the builder's file, open in Orca:

The builder's role file open in Orca

Figure 3. The builder's role file, open in Orca. Every role is a file like this in prompts/team/ (right), published as team/ in the repository.

Each team works on its own branch, and both branches start from the same commit, with no chess code in it. For every task, Orca creates a fresh worktree and terminal, and a launch script starts the agent with the settings from the fairness table.

The charter, the role file and the task then go to the agent as one message with orca orchestration worker-start.

How a team runs. Both teams follow the same reviewed steps: spec, review, design, review, then the build, one phase at a time. Codex's design split the build into five phases; Pi's split it into six. Every phase goes to a new reviewer session: REWORK sends it back, PASS merges it into the team's branch. When an agent asks a question, I answer it as the coordinator.

Orca records every task in the Run. Here is the Codex team's, from orca orchestration task-list:

task_11d4deb17dca [completed] t3 spec
task_00bd0eb480d4 [completed] t3 spec review
task_20b40a9707ed [completed] t3 design
task_c06515b888ae [completed] t3 design review
task_139320b8ec8c [completed] t3 build P0
task_70a5ee00f59e [completed] t3 code review P0
task_fdf015a5ba05 [completed] t3 build P1
task_b3266b5137e5 [completed] t3 code review P1
task_fc3fa6d33368 [completed] t3 build P2
task_18800831512e [completed] t3 code review P2
task_3333bdb35a88 [completed] t3 build P3
task_682619f0e842 [completed] t3 code review P3
task_7a006755404b [completed] t3 build P4
task_331cf6f00cd8 [completed] t3 code review P4
task_65f1c06cb47a [completed] t3 acceptance

How the teams are compared. On two layers, with the reviews and the scoring done blind:

With both teams built and run to the end, the results in short:

The sections below go through each layer, starting with how each team worked.

Codex builds four times faster

Codex's team finished in under a quarter of Pi's time, at a similar cost. Neither build had a time limit, so how long to spend was each agent's own choice.

Codex Pi
Working time: spec / design / build 48 min (4 / 5 / 40) 217 min (7 / 4 / 207)
Tokens used (of which output, incl. reasoning) 10.8 million (150,000) 32.0 million (541,000)
Equivalent API cost at list prices $9.24 $11.42
Code reviews passed first time 5 of 5 6 of 6

Both agents ran on subscriptions (ChatGPT for Codex, Z.AI's GLM Coding Plan for Pi), so the like-for-like figure is the API cost at list prices: $4, $0.40 and $20 per million input, cached-input and output tokens for GPT-5.6 Sol, against $1.40, $0.26 and $4.40 for GLM-5.3. At GPT-5.6 Sol's launch price, before an August price cut, Codex's run would have cost $12.31. Pi used three times the tokens at a much lower price per token. Its working time includes about 13 minutes spent waiting for me to answer a question.

The two agents also asked for help differently. Codex asked three questions: how to treat input the spec left undefined, whether to correct a test position it had written wrongly, and what to do about a test command from its own plan that did not run. Pi asked two. Midway through the build, it measured that a coach comment plus the computer's reply took about 2.45 seconds, over the spec's 2-second limit, and asked to cut the coach's time budget from 1.2 to 0.5 seconds. Later it asked to edit a file outside its phase.

What the judges saw. The reviewer, seeing each design alone, scored them 88 and 87 and accepted both. Side by side, the two strong judges split on the specs and designs: Opus 5.5 preferred Pi's, GPT-6.1 Sol preferred Codex's. They agreed on two things specific to a chess coach: Pi had planned the stronger engine, with a fuller search and evaluation, and Codex the more responsive page, with its search in a background worker. The finished apps bore out both (next section). On coaching they split.

What the AI reviewers missed

Every code review passed, and both acceptance reviews rated the board as clear. Yet both boards have rows of uneven height, and Pi's colors are reversed. I found this only by measuring the page (Figure 4). The judges, scoring the finished apps side by side, found two more: Pi's page stalls for half a second or more while the computer thinks, because its search runs on the page's main thread, and Codex's level-weakening noise has no effect. A single review asks whether a change is good enough; only side by side do the gaps show.

Pi's engine beats Codex's: 18 wins, 2 draws, no losses

Start with a look. Here are both apps after the same moves: 1.e4 d5 (the computer answered d5 in both), then a deliberate blunder, 2.Ba6, which leaves the bishop hanging.

Both chess coaches after 1.e4 d5 2.Ba6, with the board defects marked

Figure 4. Both apps after 1.e4 d5 2.Ba6. Both coaches flag the hanging bishop; Pi's explains it in plainer words. The red marks show board defects that no AI reviewer reported: rows of unequal height on both boards, and Pi's light and dark squares the wrong way round.

Both engines are correct: each passed all 43 hidden tests. The tests check the rules, but they cannot say which app is better. Two things can.

Pi coaches better. On the same positions:

Position Codex's coach Pi's coach
Queen moved where a pawn takes it "Your opponent can capture your queen on d4 with their pawn from c5." "Your queen on d4 can now be captured for free by the pawn on c5. A stronger move was moving your queen from d1 to d5."
Checkmating move "That was the strongest move." "Checkmate — you win the game."
Weak first move, 1.g4 Rated best Rated an inaccuracy; suggests developing the knight
Time per comment 30–351 ms about 500 ms

Pi's engine is far stronger. The two engines played 20 games: 10 fixed openings, each side playing White once, 200 milliseconds per move. Pi won 18, every one by checkmate, and drew the other 2. Codex won none. Counting a draw as half a point, that is 19–1.

Game 15, mid-game and final position

Figure 5. Game 15, Codex playing White. At mid-game the material is level; Pi (Black) finishes with mate on d1.

A result this lopsided invites doubt, so I checked what could have tilted it. Neither app contains third-party code, and both engines were frozen before the match and played through my own script, with no agent involved. Each move had the same 200 ms, and Pi used less of it. The first match used Codex's rules as referee; with Pi's rules and a full second per move, a second 20-game match still went to Pi, 13 wins and 7 draws (16.5–3.5), and every blind judge's own games went to Pi too. The full logs are in the repository.

Conclusion

Different coding agents are good at different jobs, and the reliable way to know which fits which stage is to compare them on your own work. Orca makes that practical: I defined one engineering team as plain role files and ran it twice, once on Codex and once on Pi. Codex finished in under a quarter of the time; Pi built the better chess coach. Reviewed one at a time, both teams passed every check; the gaps showed up only side by side.

Two lessons about the method itself. Compare models of the same generation, or the result may only show which model is newer. Judge the product, not just the plans, with criteria specific to it: general engineering criteria could not tell a strong engine from a weak one.

Pi with GLM-5.3 deserves a closer look. On this project it built a better product than Codex with GPT-5.6 Sol, and it is well worth trying in your own development.

These results come from one project. The method is the part to keep: each time a new model or coding agent comes out, try it side by side with your current one on a few tasks from your own projects. That tells you more than reported scores about what it brings to your work.

Everything behind this post, from the prompts and task messages to both apps, the hidden tests and the match logs, is in the repository, along with a step-by-step reproduction guide. This is a follow-up to an earlier experiment with open-source coding agents.

← All writing