# What is an AI software factory, and how to build one with Claude Code

> An AI software factory takes each issue through agent stages and verification gates to a pull request a person accepts, and records what every change cost.

- URL: https://letsship.ai/blog/what-is-an-ai-software-factory
- Author: Önder Ceylan
- Published: 2026-09-08

An AI software factory is a system that takes a unit of work, such as an issue or a spec, through a fixed line of AI agent stages and verification gates to a pull request a person can accept, sends every rejection back to be fixed, and records what each change cost. Agents do the planning, building, reviewing and testing, and people decide what goes in and accept what comes out.

SHIP runs one as a managed service. Its own study, [SWE-in-a-team](https://letsship.ai/blog/swe-in-a-team), put 260 missions through that line, 20 tickets for each of 13 builder configurations, and the numbers on this page come from it. Treat them as directional: each ticket ran once per configuration, in one repository, and cost is the API-equivalent at list price rather than a bill.

## Separate the AI software factory from the older meanings

The term is older than coding agents, and in both of its older meanings people do the work inside the factory.

The first meaning is about reuse. Wikipedia describes a software factory as one that "applies manufacturing techniques and principles to software development", building applications by assembling predefined components and keeping hand coding for new ones, and dates the first company use of the term to 1969, with more companies in the United States and Japan following in the mid-1970s ([Software factory](https://en.wikipedia.org/wiki/Software_factory)).

The second is the delivery pipeline of the DevOps years. The United States Department of Defense's DevSecOps guidebook defines a software factory as "a collection of people, tools, and processes that enables teams to continuously deliver value by deploying software to meet the needs of a specific community of end users", in which "assembly line-like automation moves software changes from one phase to the next" ([DoD Enterprise DevSecOps Fundamentals, version 2.5](https://dowcio.war.gov/Portals/0/Documents/Library/DoD%20Enterprise%20DevSecOps%20Fundamentals%20v2.5.pdf), October 2024). Its gates already reject work: when peer review finds a problem, "the merge request can be rejected, and the code sent back to the software developer for rework."

An AI software factory keeps that pipeline and puts agents in the seats people held. The gates send an agent's work back the same way they would send back a person's, and because every agent run is a metered model call, the factory can also record what each change cost, stage by stage. The table sets the DevOps software factory and an AI coding assistant beside it.

| What differs                | DevOps software factory                          | AI coding assistant                                      | AI software factory                                       |
| --------------------------- | ------------------------------------------------ | -------------------------------------------------------- | --------------------------------------------------------- |
| Who writes the change       | People                                           | A developer, with the assistant in an editor or terminal | Agents, one per role                                      |
| Unit of work                | A change moving through the pipeline             | A prompt or a session                                    | An issue or spec, taken to a pull request                 |
| When a check fails          | The change goes back to its developer for rework | The developer decides what to do next                    | The work goes back to the builder agent with the feedback |
| Where people decide         | On every change                                  | On every change                                          | Intent, checkpoints, stuck loops and acceptance           |
| What is recorded per change | Pipeline results                                 | Whatever the developer keeps                             | Session, cost, duration and verdict for each stage        |

## List the parts a factory needs

A factory needs seven parts, plus the loop that connects them. The table lists each part's job and what it passes to the next one.

| Part               | Its job                                                                                     | What it hands on                    |
| ------------------ | ------------------------------------------------------------------------------------------- | ----------------------------------- |
| Intake             | Turns a request into a unit of work with acceptance criteria a test can check               | An issue or a spec                  |
| Planning           | Reads the request and the repository and decides the approach                               | A plan                              |
| Build              | Implements the plan in an isolated sandbox and runs the project's own checks before pushing | A branch and a pull request         |
| Verification gates | Judge the change with checks the builder did not write                                      | A pass, or feedback for the builder |
| Deploy             | Puts the change on an environment where it can be exercised                                 | A preview URL                       |
| Acceptance         | Exercises the acceptance criteria against the running change, then a person decides         | Proof, then a merge decision        |
| Cost accounting    | Records what every stage spent and how long it took                                         | Cost and time per change            |

Intake sets the ceiling for everything after it. A ticket whose acceptance criteria name exact behavior, down to the order of CSV columns and how a comma in a name is escaped, gives the tester something to verify and gives you something to accept against. A vague one leaves every later stage to guess.

Each gate should judge the change independently of the builder. Your own CI already does. A reviewer and a tester that each start from a fresh context, with their own instructions, do too, so their verdicts rest on the diff and the running app instead of on the builder's account of its own work.

## Bound the loop that connects the parts

A rejection at any gate goes back to the builder with the gate's feedback, and something has to stop that loop once it stops making progress. A useful limit counts repetition rather than attempts, because a stuck run tends to fail the same way each time while a working one fails differently as it improves.

SHIP's defaults are one worked example ([retry limits](https://letsship.ai/docs/configuration/retry-limits)). Oscillation limits stop a mission after a gate fails the same way twice, and a gate that fails a different way each round counts as progress. A ceiling of 25 Builder attempts sits far above that as a backstop. At a limit, the reason goes on the issue with the gate's last feedback, and the pull request stays open for a person to take over or redirect.

In the study, 46 of the 260 runs needed the loop, and it corrected 40 of them with no human involved, 31 on the very next push ([SWE-in-a-team](https://letsship.ai/blog/swe-in-a-team)). Two builder configurations that could not turn CI feedback into a passing build repeated the same failure until the limit stopped them, which is the point where a person should look. [Agentic SDLC, defined](https://letsship.ai/blog/agentic-sdlc) follows the loop stage by stage.

## Assemble a factory around Claude Code

Claude Code can fill every agent role in the line, because it runs non-interactively with a role's instructions passed as flags, and each run reports what it cost. What you build around it is the intake, the gates, the loop and the ledger. The steps below use only what the Claude Code docs describe.

1. Start runs from your tracker. The Claude Code GitHub Action responds to `@claude` in an issue or pull request comment from someone with write access to the repository, and its docs list turning "issues into pull requests" among its uses ([GitHub Actions](https://code.claude.com/docs/en/github-actions)). To run Claude Code on Linear issues without building the intake yourself, see [Claude Code on Linear](https://letsship.ai/blog/claude-code-on-linear-issues).
2. Give each role its own run. `claude -p` runs Claude Code non-interactively, and `--append-system-prompt-file` adds a role's instructions while keeping Claude Code's default behavior ([run Claude Code programmatically](https://code.claude.com/docs/en/headless)). In CI, add `--bare` so every machine starts the same way: it skips the auto-discovery of hooks, skills, plugins, MCP servers, auto memory and `CLAUDE.md`, reads its key from `ANTHROPIC_API_KEY`, and takes what the role needs as flags.
3. Ask the planner for structured output. `--output-format json` together with `--json-schema` returns the plan in a `structured_output` field, which the next stage can read without parsing prose.
4. Build in a fresh sandbox for every run. `--permission-mode acceptEdits` lets the builder write files, and `--allowedTools` allows only the shell commands it needs, such as `Bash(pnpm run *)`. A `Stop` hook that exits with code 2 "prevents Claude from stopping" ([hooks](https://code.claude.com/docs/en/hooks)), which you can use to keep the builder working until the project's tests pass locally. Hooks are defined in JSON settings files, and a bare run loads settings only from what you pass with `--settings`.
5. Run each gate as a separate call. CI runs the way it always does. The reviewer is another `claude -p` run with a reviewer's instructions and the diff on stdin, and the tester is a third run pointed at the preview URL with the acceptance criteria. Subagents in `.claude/agents/` each run "in its own context window with a custom system prompt, specific tool access, and independent permissions" ([subagents](https://code.claude.com/docs/en/sub-agents)), which suits helpers inside one run; a separate process keeps a gate's verdict out of the builder's session altogether.
6. Record the cost of every run. With `--output-format json`, each run reports `total_cost_usd` with a per-model breakdown. The docs call these client-side estimates that can differ from your bill, so keep them as a per-stage ledger and reconcile it against your provider.

A minimal builder and reviewer pair, as a sketch:

```bash
# Builder: one non-interactive run with the role's instructions appended
claude --bare -p "Implement the plan in plan.json, then run pnpm run lint and pnpm run test." \
  --append-system-prompt-file roles/builder.md \
  --permission-mode acceptEdits \
  --allowedTools "Bash(pnpm run *)" \
  --output-format json > build.json

# Reviewer: a separate run that reads the diff on stdin
git diff origin/main...HEAD | claude --bare -p "Review this diff against plan.json. List blocking and non-blocking findings." \
  --append-system-prompt-file roles/reviewer.md \
  --output-format json > review.json

# Ledger: what each run reported spending
jq '.total_cost_usd' build.json review.json
```

The rest is the script around those calls, and it carries most of the factory's logic. It feeds each gate's feedback into the next builder prompt, counts how often a gate fails the same way, stops at a limit and posts the reason on the issue, triggers the preview deploy, and adds up the cost of every run that belongs to one change.

## Run the same factory on SHIP

SHIP runs the whole line as a managed service, with Claude Code as the default harness for every role, so one Anthropic credential is enough for a complete mission ([LLM credentials](https://letsship.ai/docs/getting-started/llm-credentials)). Each part maps to a documented piece of the platform:

| Part                          | In SHIP                                                                                                                                                                              |
| ----------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Intake                        | Assign a Linear issue to SHIP, comment `@SHIP` on a GitHub issue, or hand work over with `ship delegate` from a terminal ([ways of working](https://letsship.ai/docs/getting-started/ways-of-working))  |
| Planning, build, review, test | Planner, Builder, Reviewer and QA agents, each on the harness and model you set per role, per project or per mission ([how a mission runs](https://letsship.ai/docs))                                   |
| Role instructions             | A `.ship/agents/<role>.md` file, appended after SHIP's charter for that role ([agent prompts](https://letsship.ai/docs/configuration/agent-prompts))                                                    |
| Gates and the loop bound      | Your CI, the Reviewer and QA, with oscillation limits of 2 per gate and a ceiling of 25 Builder attempts ([retry limits](https://letsship.ai/docs/configuration/retry-limits))                          |
| Deploy                        | Your own delivery pipeline, with your own credentials ([deployments](https://letsship.ai/docs/configuration/deployments))                                                                               |
| Acceptance                    | Optional checkpoints after the Planner, the Builder or the Reviewer, and a pull request that waits for you by default ([human in the loop](https://letsship.ai/docs/configuration/human-in-the-loop))   |
| Cost accounting               | Session, prompt, token usage and cost per mission and per stage, priced at list rates for the model that served each call ([LLM credentials](https://letsship.ai/docs/getting-started/llm-credentials)) |

Claude Code is the default, and any role can run on codex, pi, kilo-code or open-code instead ([schema](https://letsship.ai/docs/configuration/schema)). Agents never hold your credentials: a sandbox gets a scoped, single-use callback token, and the provider key is added to its requests from outside the container ([built to be checked](https://letsship.ai/docs)).

## Account for cost per accepted change

Divide what the factory spent by the changes that were accepted, and split that cost by stage, because the stage that writes the code is not always the one that costs the most. The table shows three builder configurations from the study, each given the same 20 tickets with the same pinned Planner, Reviewer and Tester; totals are API-equivalent at list price, recomputed from the [per-run data](https://letsship.ai/data/swe-in-a-team-per-task.csv).

| Builder           | Harness     | Total for 20 tickets | Resolved | Per resolved ticket |
| ----------------- | ----------- | -------------------- | -------- | ------------------- |
| `claude-sonnet-5` | claude-code | $86.22               | 20 of 20 | $4.31               |
| `claude-opus-5`   | claude-code | $109.66              | 20 of 20 | $5.48               |
| `claude-fable-5`  | claude-code | $122.42              | 19 of 20 | $6.44               |

Splitting by stage showed where the money went ([SWE-in-a-team](https://letsship.ai/blog/swe-in-a-team)). On a CSV export ticket (F1), the Tester cost $5.31 of a $7.23 mission, more than the other three agents together, because it ran its whole check twice. Where a problem was caught also mattered: runs the Tester sent back came to a median of $8.18, against $3.51 for a run no gate objected to. With the harness held fixed, the model alone moved the cost of the same 20 tickets on claude-code from $86.22 with `claude-sonnet-5` to $122.42 with `claude-fable-5`.

SHIP also bounds spend with a daily cap per organization, and when it is reached, work already running finishes while new dispatches are refused with a message that names the cap ([LLM credentials](https://letsship.ai/docs/getting-started/llm-credentials)).

## Keep a person at acceptance

A change can pass every gate and still be wrong, which is why the last seat in the line belongs to a person. In the study, 4 of the 260 runs passed every gate and were still not complete, and only the hidden acceptance tests caught them ([SWE-in-a-team](https://letsship.ai/blog/swe-in-a-team)).

One was a ticket asking for class reminders to be queued (F5). The `claude-fable-5` builder wired the live email provider into the request path, so the endpoint returned a 500 whenever the API key was missing, though the ticket asked only for a queued row. The repository's own tests stayed green, so CI passed, and the Reviewer and the Tester both approved. The study summed it up as "It built the wrong thing, correctly."

By default a SHIP mission stops at the pull request and waits for you. Auto-merge exists for teams that want it, and it stays off until you turn it on for a project or a mission ([human in the loop](https://letsship.ai/docs/configuration/human-in-the-loop)).

## Choose between open parts and a managed factory

An open-source AI software factory is one whose orchestration you can read, run and change yourself, and every part of one can be assembled from open components. Tracker webhooks, CI runners and a container per run are already yours to control. The harness can be open source as well: two of the four harnesses in SHIP's study, pi and open-code, are MIT-licensed ([pi](https://github.com/badlogic/pi-mono), [open-code](https://github.com/anomalyco/opencode)), and the open-weight models run on them resolved 118 of their 120 tickets ([SWE-in-a-team](https://letsship.ai/blog/swe-in-a-team)). What you write is the connecting logic: the loop, its limits, the hand-off to a person, and the per-stage ledger.

SHIP is a managed service that runs the same line, with those open harnesses available beside claude-code and codex, on your own provider accounts. Building your own fits when your delivery process differs enough from these stages that a fixed line would get in the way, or when owning the orchestration code is the goal. A managed factory fits when the stages above already match how your team ships and the engineering time would go further on your product.

## Questions

### What is an AI software factory?

An AI software factory is a system that takes a unit of work, such as an issue or a spec, through a fixed line of AI agent stages and verification gates to a pull request a person can accept, sends every rejection back to be fixed, and records what each change cost. Agents plan, build, review and test; people decide what goes in and accept what comes out.

### How do I build a software factory with Claude Code?

Run each agent role as its own non-interactive Claude Code call, `claude -p`, with that role's instructions appended through `--append-system-prompt-file`. Then write the script that connects the calls: intake from an issue, your CI and a separate reviewer run as gates, a preview deploy with a tester run against it, a retry limit for stuck loops, and a ledger of the cost each run reports. SHIP runs this line as a managed service with Claude Code as the default harness for every role.

### What is the difference between an AI software factory and an AI coding assistant?

An AI coding assistant helps one developer write code inside a session, and the developer carries the change through review, testing and release. An AI software factory takes the whole change from issue to a pull request ready to accept, with an agent in each role, gates that send failed work back to the builder, and the cost of each change recorded.

### Is there an open-source AI software factory?

You can assemble one from open-source parts: tracker webhooks, your own CI, a container per run, and an open-source coding harness such as pi or open-code, both MIT-licensed. The orchestration that connects them, including the retry limits and a per-stage cost ledger, is the part you write yourself. SHIP is a managed service that runs the same line, with those harnesses available beside claude-code and codex.

### What is the difference between a software factory and DevOps?

DevSecOps, in the United States Department of Defense's definition, is a combination of methodologies, practices and tools that unifies development, security and operations, and a software factory is the concrete pipeline of people, tools and processes built on it to deliver software continuously. An AI software factory keeps that pipeline and puts AI agents in the roles people held, with gates that send their work back for rework.

### How do I measure the cost of AI coding agents per pull request?

Record the inference cost of every stage for each change, then divide the total by the changes that were accepted, not the ones opened. In SHIP's 260-mission SWE-in-a-team study, claude-sonnet-5 on claude-code resolved all 20 tickets at $4.31 per resolved ticket and claude-opus-5 at $5.48, at API list prices.
