# Codex code review for Claude Code: set up a second-provider reviewer

> Keep Claude Code on the builder and put Codex on the reviewer, in ship.yml or for one mission, so a second provider reads every diff before you do.

- URL: https://letsship.ai/blog/codex-reviewer-for-claude-code
- Author: Önder Ceylan
- Published: 2026-09-17

[Codex](https://developers.openai.com/codex/cli) reviews what [Claude Code](https://code.claude.com/docs/en/overview) builds once SHIP has an OpenAI key next to your Anthropic credential and your `ship.yml` sets `harness: codex` under `agents.reviewer`, with the builder left on `claude-code`. From then on Claude Code writes each change and Codex reads the diff against the plan, sending blocking findings back to the builder. For a single mission, `ship delegate pr SHIP-412 --harness codex` does the same without touching the file. The SWE-in-a-team benchmark kept its reviewer on claude-code in all 260 missions, so it has no result for a Codex reviewer yet.

## Decide what a Codex review adds

SHIP's reviewer reads the diff against the plan and raises blocking and non-blocking findings ([introduction](https://letsship.ai/docs)). A blocking finding sends the change back to the builder with the feedback, and a clean review moves the mission on to the preview deploy and QA.

Out of the box every role runs on `claude-code`, so the model reviewing a change comes from the same provider as the model that wrote it. Setting the reviewer to `codex` swaps the harness, the provider and the model that read the diff in one move. The review charter SHIP writes for the role stays the same, and so do your repository's reviewer conventions, which makes it a second reading of the same diff by a different model under the same rules. Whether that reading catches more on your codebase is something to measure, and SHIP records cost, duration and outcome per stage so the comparison can be read afterwards "rather than a matter of impression" ([LLM credentials](https://letsship.ai/docs/getting-started/llm-credentials)).

The single-provider default is deliberate. SHIP's reviewer once defaulted to a second provider, and organizations with a working key for one provider and a missing key for the other onboarded, planned and built without trouble, then failed at review with a credential error. The credentials docs still call a mixed fleet a good idea for plenty of teams, as long as it is a setup you chose rather than one you discover at the last gate.

If the question that brought you here is Claude Code versus Codex, [SWE-in-a-team](https://letsship.ai/blog/swe-in-a-team) compared them as builders. Seven of its configurations ran on these two harnesses, each on the same 20 tickets with one run per ticket. Resolved means the mission reached Ready for Acceptance and hidden tests passed, and cost is API-equivalent, recomputed from the [per-run data](https://letsship.ai/data/swe-in-a-team-per-task.csv).

| Builder model      | Harness     | Resolved | Cost per resolved ticket | Median wall time |
| ------------------ | ----------- | -------- | ------------------------ | ---------------- |
| `gpt-5.6-sol`      | codex       | 20 of 20 | $3.83                    | 16.3 min         |
| `gpt-5.6-terra`    | codex       | 19 of 20 | $3.88                    | 18.5 min         |
| `gpt-5.6-luna`     | codex       | 16 of 20 | $3.08                    | 14.0 min         |
| `claude-haiku-4-5` | claude-code | 18 of 20 | $4.23                    | 22.7 min         |
| `claude-sonnet-5`  | claude-code | 20 of 20 | $4.31                    | 20.7 min         |
| `claude-opus-5`    | claude-code | 20 of 20 | $5.48                    | 23.0 min         |
| `claude-fable-5`   | claude-code | 19 of 20 | $6.44                    | 20.7 min         |

Every one of those 140 missions was planned, reviewed and tested by claude-code on claude-opus-4.8. The table compares builders and says nothing about Codex in the reviewer's seat.

## Add an OpenAI key beside the Anthropic one

OpenAI is the provider behind the `codex` harness, so add an OpenAI key on the **Connect your AI providers** step, which takes one credential or several. SHIP verifies each key with a live call to the provider when you save it. Do this before you edit `ship.yml`. A bad key caught on save costs you a retyped key, whereas a reviewer switched to a provider with no working key is the exact failure the single-provider default was built to avoid.

The OpenAI key is handled the way the Anthropic one is. It never reaches the sandbox the reviewer runs in; SHIP starts the container with a placeholder and rewrites the credential on the way out to the provider.

## Set the reviewer to Codex in ship.yml

This is the configuration from SHIP's own [credentials guide](https://letsship.ai/docs/getting-started/llm-credentials), with the lines a valid file also needs:

```yaml
# yaml-language-server: $schema=https://letsship.ai/schema/ship.yml.json
version: 1

deployments: {}

agents:
  builder:
    harness: claude-code
    model: claude-opus-5
  reviewer:
    harness: codex
```

The planner and QA are left out, so they keep the defaults. To choose the reviewer's model, add a `model` key: the `codex` and `gpt` aliases follow the newest model of their family, and a concrete versioned model id stays fixed from one mission to the next ([Schema](https://letsship.ai/docs/configuration/schema)).

SHIP reads the `agents` block from `ship.yml` on your `main` branch. The switch therefore applies to missions that start after it merges, and a pull request cannot pick its own reviewer by editing the file.

On a project that plans in Linear nothing else changes. You still [assign the issue to SHIP](https://letsship.ai/blog/claude-code-on-linear-issues), and the review verdict posted on the issue is now the one Codex wrote.

## Override the reviewer for one mission

To try Codex on one piece of work without changing the project, set it on the mission. Which route fits depends on how the mission started:

| The mission starts from                                                | Put Codex on its reviewer with                                                      |
| ---------------------------------------------------------------------- | ----------------------------------------------------------------------------------- |
| A branch you built in a Claude Code session, against an existing issue | `ship delegate pr SHIP-412 --harness codex`                                         |
| A branch you built with no ticket behind it                            | `ship delegate pr --harness codex`                                                  |
| A Linear or GitHub issue                                               | `PUT /v1/missions/{issueId}/agents-override`, or the `set_agents_override` MCP tool |

The first two rows fit the way many people already use Claude Code in the terminal. The agent makes the change and opens the pull request itself, because SHIP does not push or open a pull request for a branch you built, and `ship delegate pr` hands it over for review, CI and QA ([Coding agents](https://letsship.ai/docs/cli/coding-agents)). `--harness` applies only to the role the mission enters at, which for `delegate pr` is the reviewer, and `--model` sets that role's model the same way. The builder that later answers the review runs on the project's configuration.

The third row changes one role's harness, model, prompt or checkpoint for that mission, merged per role and per field over `ship.yml`. `get_agents_override` reads what is in force and `clear_agents_override` drops it ([MCP server](https://letsship.ai/docs/mcp)).

## Tune what the reviewer blocks on

By default the reviewer blocks only on critical findings. `rejectOn` moves that line, and it only means something on the reviewer:

```yaml
agents:
  reviewer:
    harness: codex
    rejectOn: suggestion
```

`critical` is the default, `suggestion` also blocks on suggestions, and `trivial` blocks on everything. House rules belong in `.ship/agents/reviewer.md`, which SHIP appends after its own reviewer charter whichever harness runs the role, and reads from `main` so a pull request cannot loosen the conventions it is graded against ([Agent prompts](https://letsship.ai/docs/configuration/agent-prompts)). For one mission with an extra concern, such as a security pass, add a lane on top of those conventions:

```bash
ship delegate pr SHIP-412 --harness codex --prompt-file security-review.md --prompt-mode extend
```

Mind the cost of strictness. A reviewer that raises the same blocking comment on two consecutive rounds escalates the mission to a person, since `reviewOscillationLimit` defaults to 2, and up to 3 extra rounds (`maxReviewGrants`) are granted only when the work is judged to be progressing. Every extra round is a full agent run ([Retry limits](https://letsship.ai/docs/configuration/retry-limits)).

## Compare reviewers on your own missions

Your own missions are the evidence here, and SHIP records enough of each one to read it. Hand over comparable work under each reviewer:

```bash
ship delegate pr SHIP-412 --harness claude-code
ship delegate pr SHIP-413 --harness codex
```

For each mission, `GET /v1/missions/{issueId}/provenance` returns every stage attempt's harness, model, provider, token usage, cost and duration, which the [Missions API](https://letsship.ai/docs/api/missions) calls the readout for comparing one configuration against another. For a script that only reads these records, the reporting-dashboard scope set in [Authentication](https://letsship.ai/docs/api/authentication) carries no write scope at all. The records are kept for about 30 days, so snapshot them if the comparison runs longer than that.

What to look at is the review stage's cost, how many rounds each mission took to reach a clean review, and whether QA later failed anything the review had passed. A prompt lane counts as its own configuration too: the Codex reviewer with the security lane above is recorded as a different configuration from the plain Codex reviewer, never averaged into it ([Commands](https://letsship.ai/docs/cli/commands)).

## Questions

### Can Codex review code written by Claude Code?

Yes. In SHIP each role has its own harness, so you can keep claude-code on the builder and set the reviewer to codex, which then reads every diff Claude Code produces against the plan.

### How do I set up Codex code review in SHIP?

Add an OpenAI key next to your Anthropic credential, then set harness: codex under agents.reviewer in the ship.yml at your repository root. The change applies to missions that start after it merges to main.

### Does Codex work with Linear issues in SHIP?

Yes. On a project that plans in Linear you still assign the issue to SHIP, and the review verdict the codex reviewer writes is posted on the issue like any other stage's output.

### Is Claude Code or Codex better?

As builders on the same 20 SWE-in-a-team tickets, codex with gpt-5.6-sol resolved 20 of 20 at $3.83 per resolved ticket and claude-code with claude-sonnet-5 resolved 20 of 20 at $4.31 (API-equivalent). Every reviewer in that study ran claude-code, so it has no result for Codex as a reviewer.

### Do I need an OpenAI key for a Codex reviewer?

Yes. SHIP runs the codex harness on an OpenAI credential, and the default configuration is deliberately single-provider, so connect and verify the OpenAI key before switching the reviewer.

### Does the builder switch to Codex when I pass --harness codex?

Not with ship delegate pr. The --harness flag applies to the role the mission enters at, which is the reviewer for delegate pr, and the builder that answers the review keeps the project's configuration.
