Our SWE-in-a-team benchmark research is ready.Read the results.
← Blog
||Guides

Codex code review for Claude Code: set up a second-provider reviewer

Keep Claude Code on the builder and put Codex on the reviewer, in ship.yml or for one mission, so a second provider reads every diff before you do.

Codex reviews what Claude Code builds once SHIP has an OpenAI key next to your Anthropic credential and your ship.yml sets harness: codex under agents.reviewer, with the builder left on claude-code. From then on Claude Code writes each change and Codex reads the diff against the plan, sending blocking findings back to the builder. For a single mission, ship delegate pr SHIP-412 --harness codex does the same without touching the file. The SWE-in-a-team benchmark kept its reviewer on claude-code in all 260 missions, so it has no result for a Codex reviewer yet.

Decide what a Codex review adds

SHIP's reviewer reads the diff against the plan and raises blocking and non-blocking findings (introduction). A blocking finding sends the change back to the builder with the feedback, and a clean review moves the mission on to the preview deploy and QA.

Out of the box every role runs on claude-code, so the model reviewing a change comes from the same provider as the model that wrote it. Setting the reviewer to codex swaps the harness, the provider and the model that read the diff in one move. The review charter SHIP writes for the role stays the same, and so do your repository's reviewer conventions, which makes it a second reading of the same diff by a different model under the same rules. Whether that reading catches more on your codebase is something to measure, and SHIP records cost, duration and outcome per stage so the comparison can be read afterwards "rather than a matter of impression" (LLM credentials).

The single-provider default is deliberate. SHIP's reviewer once defaulted to a second provider, and organizations with a working key for one provider and a missing key for the other onboarded, planned and built without trouble, then failed at review with a credential error. The credentials docs still call a mixed fleet a good idea for plenty of teams, as long as it is a setup you chose rather than one you discover at the last gate.

If the question that brought you here is Claude Code versus Codex, SWE-in-a-team compared them as builders. Seven of its configurations ran on these two harnesses, each on the same 20 tickets with one run per ticket. Resolved means the mission reached Ready for Acceptance and hidden tests passed, and cost is API-equivalent, recomputed from the per-run data.

Builder modelHarnessResolvedCost per resolved ticketMedian wall time
gpt-5.6-solcodex20 of 20$3.8316.3 min
gpt-5.6-terracodex19 of 20$3.8818.5 min
gpt-5.6-lunacodex16 of 20$3.0814.0 min
claude-haiku-4-5claude-code18 of 20$4.2322.7 min
claude-sonnet-5claude-code20 of 20$4.3120.7 min
claude-opus-5claude-code20 of 20$5.4823.0 min
claude-fable-5claude-code19 of 20$6.4420.7 min

Every one of those 140 missions was planned, reviewed and tested by claude-code on claude-opus-4.8. The table compares builders and says nothing about Codex in the reviewer's seat.

Add an OpenAI key beside the Anthropic one

OpenAI is the provider behind the codex harness, so add an OpenAI key on the Connect your AI providers step, which takes one credential or several. SHIP verifies each key with a live call to the provider when you save it. Do this before you edit ship.yml. A bad key caught on save costs you a retyped key, whereas a reviewer switched to a provider with no working key is the exact failure the single-provider default was built to avoid.

The OpenAI key is handled the way the Anthropic one is. It never reaches the sandbox the reviewer runs in; SHIP starts the container with a placeholder and rewrites the credential on the way out to the provider.

Set the reviewer to Codex in ship.yml

This is the configuration from SHIP's own credentials guide, with the lines a valid file also needs:

# yaml-language-server: $schema=https://letsship.ai/schema/ship.yml.json
version: 1

deployments: {}

agents:
  builder:
    harness: claude-code
    model: claude-opus-5
  reviewer:
    harness: codex

The planner and QA are left out, so they keep the defaults. To choose the reviewer's model, add a model key: the codex and gpt aliases follow the newest model of their family, and a concrete versioned model id stays fixed from one mission to the next (Schema).

SHIP reads the agents block from ship.yml on your main branch. The switch therefore applies to missions that start after it merges, and a pull request cannot pick its own reviewer by editing the file.

On a project that plans in Linear nothing else changes. You still assign the issue to SHIP, and the review verdict posted on the issue is now the one Codex wrote.

Override the reviewer for one mission

To try Codex on one piece of work without changing the project, set it on the mission. Which route fits depends on how the mission started:

The mission starts fromPut Codex on its reviewer with
A branch you built in a Claude Code session, against an existing issueship delegate pr SHIP-412 --harness codex
A branch you built with no ticket behind itship delegate pr --harness codex
A Linear or GitHub issuePUT /v1/missions/{issueId}/agents-override, or the set_agents_override MCP tool

The first two rows fit the way many people already use Claude Code in the terminal. The agent makes the change and opens the pull request itself, because SHIP does not push or open a pull request for a branch you built, and ship delegate pr hands it over for review, CI and QA (Coding agents). --harness applies only to the role the mission enters at, which for delegate pr is the reviewer, and --model sets that role's model the same way. The builder that later answers the review runs on the project's configuration.

The third row changes one role's harness, model, prompt or checkpoint for that mission, merged per role and per field over ship.yml. get_agents_override reads what is in force and clear_agents_override drops it (MCP server).

Tune what the reviewer blocks on

By default the reviewer blocks only on critical findings. rejectOn moves that line, and it only means something on the reviewer:

agents:
  reviewer:
    harness: codex
    rejectOn: suggestion

critical is the default, suggestion also blocks on suggestions, and trivial blocks on everything. House rules belong in .ship/agents/reviewer.md, which SHIP appends after its own reviewer charter whichever harness runs the role, and reads from main so a pull request cannot loosen the conventions it is graded against (Agent prompts). For one mission with an extra concern, such as a security pass, add a lane on top of those conventions:

ship delegate pr SHIP-412 --harness codex --prompt-file security-review.md --prompt-mode extend

Mind the cost of strictness. A reviewer that raises the same blocking comment on two consecutive rounds escalates the mission to a person, since reviewOscillationLimit defaults to 2, and up to 3 extra rounds (maxReviewGrants) are granted only when the work is judged to be progressing. Every extra round is a full agent run (Retry limits).

Compare reviewers on your own missions

Your own missions are the evidence here, and SHIP records enough of each one to read it. Hand over comparable work under each reviewer:

ship delegate pr SHIP-412 --harness claude-code
ship delegate pr SHIP-413 --harness codex

For each mission, GET /v1/missions/{issueId}/provenance returns every stage attempt's harness, model, provider, token usage, cost and duration, which the Missions API calls the readout for comparing one configuration against another. For a script that only reads these records, the reporting-dashboard scope set in Authentication carries no write scope at all. The records are kept for about 30 days, so snapshot them if the comparison runs longer than that.

What to look at is the review stage's cost, how many rounds each mission took to reach a clean review, and whether QA later failed anything the review had passed. A prompt lane counts as its own configuration too: the Codex reviewer with the security lane above is recorded as a different configuration from the plain Codex reviewer, never averaged into it (Commands).

Questions

Can Codex review code written by Claude Code?

Yes. In SHIP each role has its own harness, so you can keep claude-code on the builder and set the reviewer to codex, which then reads every diff Claude Code produces against the plan.

How do I set up Codex code review in SHIP?

Add an OpenAI key next to your Anthropic credential, then set harness: codex under agents.reviewer in the ship.yml at your repository root. The change applies to missions that start after it merges to main.

Does Codex work with Linear issues in SHIP?

Yes. On a project that plans in Linear you still assign the issue to SHIP, and the review verdict the codex reviewer writes is posted on the issue like any other stage's output.

Is Claude Code or Codex better?

As builders on the same 20 SWE-in-a-team tickets, codex with gpt-5.6-sol resolved 20 of 20 at $3.83 per resolved ticket and claude-code with claude-sonnet-5 resolved 20 of 20 at $4.31 (API-equivalent). Every reviewer in that study ran claude-code, so it has no result for Codex as a reviewer.

Do I need an OpenAI key for a Codex reviewer?

Yes. SHIP runs the codex harness on an OpenAI credential, and the default configuration is deliberately single-provider, so connect and verify the OpenAI key before switching the reviewer.

Does the builder switch to Codex when I pass --harness codex?

Not with ship delegate pr. The --harness flag applies to the role the mission enters at, which is the reviewer for delegate pr, and the builder that answers the review keeps the project's configuration.

Add a second provider to every review

See Codex review what Claude Code builds in your repository, with both agents running on your own provider keys.