Our SWE-in-a-team benchmark research is ready.Read the results.
← Blog
||Guides

Run the pi coding agent headless on real tickets, in a sandbox

Run the pi coding agent as SHIP's builder in an isolated sandbox, on an open-weight model through your own OpenRouter key, with review and QA around it.

SHIP runs the pi coding agent headless on real tickets once you add an OpenRouter key and set harness: pi on the builder in ship.yml. pi then works in an isolated sandbox with no terminal attached and no provider key inside it, while a planner, your CI, a reviewer and QA work around it. You start a run by assigning an issue or with ship delegate plan, and read what pi did from the session log afterwards. Run this way in the SWE-in-a-team benchmark, pi builders resolved 58 of 60 missions.

Let SHIP hold the sandbox

The shortest route to a sandboxed pi is to let SHIP run it. SHIP's builder does its work in an isolated sandbox, and when the builder role runs on pi, that sandbox is where pi runs. The docs describe four things about it:

  • Your OpenRouter key never enters it. The container starts with a placeholder and SHIP rewrites the credential on the way out to the provider, while the sandbox is given a scoped, single-use callback token for reporting back (LLM credentials).
  • The builder runs your formatter, linter, typecheck and tests before it pushes, and a change under .github/workflows/ is blocked before the commit is created unless you explicitly authorized it (Connect GitHub).
  • A sandbox lasts as long as its stage. Once the stage finishes, its result is recorded and the sandbox is gone.
  • The QA agent that tests pi's work holds a read-only token and has no way to push.

You can also build this yourself. pi has no permission popups, and pi.dev suggests you "run in a container, or build your own confirmation flow with extensions". For scripting it offers a print mode, pi -p "query", and --mode json for event streams. Together those give you pi running headless in a container, with the image, the key, the output and the checking left to you.

Inside SHIP no terminal is attached to pi while it works. SHIP persists the full agent session, the assembled prompt, token usage and cost before the agent reports back, and that record is what you read afterwards.

Give pi an OpenRouter key

On the Connect your AI providers step, add an OpenRouter key. SHIP uses OpenRouter for its open-model harnesses, pi, kilo-code and open-code, and checks the key with a live call when you save it (LLM credentials). Keys are never pooled across organizations, so the spend and the rate limits are those of your own OpenRouter account.

Keep the Anthropic credential as well. Every role you do not configure runs on the default claude-code harness, and the setup below changes only the builder.

Choose a model from 60 benchmark missions

SWE-in-a-team ran pi as the builder with three open-weight models, each on the same 20 tickets, one run per ticket, while the planner, reviewer and tester stayed on claude-code with claude-opus-4.8. A ticket counts as resolved when the mission reached Ready for Acceptance and hidden tests the builder never saw passed. The figures are recomputed from the per-run data.

Model on piResolvedMedian mission time
deepseek/deepseek-v4-flash20 of 2020.1 min
moonshotai/kimi-k319 of 2017.9 min
z-ai/glm-5.219 of 2023.5 min

Of the three, deepseek-v4-flash alone resolved every ticket. The two misses are worth reading: kimi-k3 missed the idempotent class-reminder job and glm-5.2 missed the payment-webhook integration, and in both cases the mission passed CI, review and the tester, so only the hidden tests caught them. The loop still did plenty of correcting. deepseek-v4-flash shipped 12 of its 20 tickets without any gate objecting and was sent back 13 times on the rest, and every one of those tickets resolved.

The full study, including the same three models run on a second harness, is in SWE-in-a-team.

Put pi on the builder role

Add a ship.yml at the repository root, or extend the one you have:

# yaml-language-server: $schema=https://letsship.ai/schema/ship.yml.json
version: 1

deployments: {}

agents:
  builder:
    harness: pi
    model: deepseek/deepseek-v4-flash
    providerRouting:
      dataCollection: deny

That is the benchmark's shape: pi writes the code while planner, reviewer and QA keep their claude-code defaults. A concrete model id like deepseek/deepseek-v4-flash is passed through unchanged, and provider can stay unset, since it defaults to the harness's own gateway.

providerRouting controls which upstream providers may serve the model, for runs you want to be deterministic (Schema). dataCollection: deny makes providers that retain data ineligible. order, only and ignore name providers explicitly, and allowFallbacks decides whether another provider may step in.

To try pi on one mission without touching the file, pass it on the hand-off. --harness applies to the role the mission enters at, which for delegate plan is the builder:

ship delegate plan SHIP-412 --file docs/plan.md --harness pi --model deepseek/deepseek-v4-flash

Conventions you want pi to follow go in .ship/agents/builder.md. SHIP appends that file after its own builder charter every time the role runs, whichever harness is running it (Agent prompts).

Hand pi a ticket or a spec

Work reaches pi from a tracker or from your terminal. On a Linear project, assign the issue to SHIP; on GitHub Issues, comment @SHIP from an account with write access. The planner writes the plan from the issue and the repository, and pi builds from it.

From a terminal, ship delegate plan skips the planner and has pi build straight from a spec you wrote. That suits pi's own habits: pi.dev's answer to having no plan mode is "Write plans to files".

ship delegate plan --title "Rate-limit the search endpoint" --file docs/plan.md

Following the run needs no terminal session attached to anything. ship status SHIP-412 --json prints the stage, the review and test state and the pull request, and its exit codes are distinct enough for a script to branch on (Command reference). An assistant connected over MCP can call get_sandbox_logs for what the agent did inside a run. To change course, ship instruct SHIP-412 -m "..." puts the builder back on the open pull request with your instruction.

Cap what an unattended run can spend

SHIP caps an unattended run with four limits that act without anyone watching, and all four are on by default:

LimitDefaultWhen it is reached
Spend per mission$5The mission stops
Spend per organization per day$50Running work finishes, and new dispatches are refused with a message naming the cap and when it resets
Identical failures in a row at one gate2The mission shows Needs attention and the pull request stays open
Builder attempts in one mission25The same hand-back, as an outer backstop

You change the daily cap from Settings or the API, and the loop limits under retries in ship.yml (Retry limits). Raising maxRetries on its own rarely changes anything, because a stuck run hits an oscillation limit long before the ceiling.

Every mission reports what each stage spent, priced from list rates for the model that actually served each call (LLM credentials). Read the first few pi missions' costs per stage before you let pi take a larger share of the backlog.

Questions

How do I sandbox the pi coding agent?

pi's own site suggests running it in a container. In SHIP you set harness: pi on the builder role instead, and SHIP runs pi in an isolated sandbox that never holds your provider key and is gone once the stage finishes.

Can I run the pi coding agent in Docker?

You can run pi in a container yourself, which is what pi.dev recommends in place of permission prompts. SHIP removes the need to build and maintain that container, because it starts one for each stage and keeps the OpenRouter key outside it.

Does the pi coding agent have a headless mode?

Yes. pi.dev documents a print mode, pi -p "query" for scripts, and --mode json for event streams. Inside SHIP nothing is attached to pi at all: a mission starts from an assigned issue or a CLI hand-off, and SHIP persists the full session for you to read afterwards.

Which model should I use with the pi coding agent?

In SWE-in-a-team, deepseek-v4-flash was the only one of three open-weight models to resolve all 20 tickets on pi. kimi-k3 and glm-5.2 resolved 19 each, and kimi-k3 had the shortest median mission at 17.9 minutes.

What API key does pi need in SHIP?

An OpenRouter key, which SHIP uses for the open-model harnesses pi, kilo-code and open-code. If the other roles stay on the default claude-code harness, the project also needs an Anthropic credential.

How much does a pi mission cost?

SHIP records what every stage of a pi mission cost, priced from list rates for the model that served each call, and shows it per mission and per stage. A single mission stops at $5 by default, and an organization at $50 a day unless you change the cap.

Run open-weight models on your real tickets

See the pi agent build in an isolated sandbox on your OpenRouter key, with review and QA checking every change before it reaches you.