Agentic SDLC, defined: a gated loop from issue to verified pull request
An agentic SDLC takes each change through plan, build, CI, review, preview and test, and sends rejected work back. In 260 missions, the loop fixed 40 of 46.
An agentic SDLC is a software development lifecycle in which AI agents carry each change from issue to pull request through planning, building, CI, review, preview deployment and testing, and every gate that rejects the work sends it back to the Builder with the feedback. Humans set the intent, take over when the loop stalls, and accept the result.
The numbers on this page come from SHIP's own study of that loop. SWE-in-a-team gave the same 20 tickets to 13 builder configurations, 260 missions in all, and graded every pull request with acceptance tests the agents never saw. Read the results as directional: each ticket ran once per configuration, in one repository, and cost is the API-equivalent at list price rather than an invoice.
Follow one change through the agentic SDLC loop
A change passes six stages in a fixed order: plan, build, CI, review, a deploy to a preview environment, and test. CI, the Reviewer and the Tester can each reject the work, and a preview deploy can fail as well. Whichever stage objects, the work goes back to the Builder with the failure attached, and the loop runs again from the build. When every gate passes, the mission reaches Ready for Acceptance and the pull request waits for a person (how a mission runs).
issue → Planner → Builder → CI → Reviewer → preview deploy → Tester
→ pull request ready to accept
any rejection → back to the Builder, with the feedbackThe recorded mission below is a bug from the study: a member could get onto a class waitlist twice by clicking "join" twice. The Reviewer sent the first attempt back, and the second Builder attempt passed every gate.
Planner 172s ──▶ Builder 423s ──▶ CI ✓ ──▶ Review 127s ✗ changes requested
└─────────────▶ Builder 181s ──▶ CI ✓ ──▶ Review 46s ✓ ──▶ Deploy ✓ ──▶ Test 454s ✓
30 minutes, $9.18, 1 review rejection, resolvedTwo Builder attempts and one review rejection produced a resolved ticket in 30 minutes for $9.18 (SWE-in-a-team). The second attempt took 3 minutes against the first one's 7, because it fixed what the review named instead of starting over. The animation replays that recorded run at compressed speed, so what you see is the study's data, not a live mission.
Compare traditional, AI-driven, autonomous and agentic SDLC
The four terms differ in who does the work, what happens when a check fails, where people make decisions, and what you can measure afterward.
An AI-driven SDLC is the broad term for a lifecycle in which AI assists or automates work inside any phase, including when a person still owns every change, so an agentic SDLC is one kind of AI-driven SDLC. In its assisted form, a developer prompts a coding assistant, reads what it produced, and takes the change through the same process as before.
An autonomous SDLC is an agentic SDLC described by how rarely a person steps in: agents correct their own rejected work, and people are pulled in only at the checkpoints they set, at a retry limit, and at acceptance.
The table sets the traditional SDLC beside the assisted and agentic forms, with the autonomous SDLC in its own row. It compares definitions, so none of its cells is a measurement.
| Term | Who does the work | When a check fails | Where people decide | What you measure |
|---|---|---|---|---|
| Traditional SDLC | People, with tools that assist them | A person reads the failure and fixes it | At every step | Delivery metrics such as lead time and change failure rate |
| AI-assisted SDLC | People, with an AI assistant | A person decides what to do with the assistant's output | On every change | Adoption and time saved, plus the review load it adds |
| Agentic SDLC | Role agents: Planner, Builder, Reviewer, Tester | The work returns to the Builder with the feedback, until it passes or a limit stops it | Intent, checkpoints, stuck loops, acceptance | Cycle time, cost per change, resolve rate |
| Autonomous SDLC | Agents | The agents correct it themselves | Only at checkpoints, retry limits and acceptance | Share of changes that reach acceptance with no human step |
The rows split on verification, because that is where AI adoption moved the load. On teams with high AI adoption, Faros AI measured that "PR review time increases 91%", in telemetry from more than 10,000 developers across 1,255 teams (The AI Productivity Paradox Report 2025). The whitepaper The New SDLC With Vibe Coding (May 2026) puts vibe coding and agentic engineering at two ends of one spectrum, and locates the difference in "how much structure, verification, and human judgment surrounds the AI's output". An agentic SDLC puts review and testing inside the loop as gates, so the agents do the first rounds of verification before a person reads anything.
Give each role one job
Each stage belongs to one role with one job, and the two stages that touch your infrastructure run in your own pipeline. The table lists what each one does in SHIP, as the docs describe it.
| Stage | Run by | Its job |
|---|---|---|
| Plan | Planner | Reads the request, the repository and the acceptance criteria, and writes a plan. Skipped when you hand over a plan of your own. |
| Build | Builder | Implements the plan in an isolated sandbox, runs your formatter, linter and tests before pushing, and opens the pull request. |
| CI | Your pipeline | Runs your own checks on every push. |
| Review | Reviewer | Reads the diff against the plan and raises blocking and non-blocking findings. |
| Deploy | Your pipeline, with your credentials | Puts the change on a preview environment. Skipped when a change has nothing to preview. |
| Test | Tester (the docs call this stage QA) | Exercises the acceptance criteria against the preview, or against the app built and served locally, and collects proof of what it checked. |
Which coding agent fills each role is configuration. You choose the harness and the model per role, set them for the whole project, or pin them for a single mission from the terminal (ways of working). To change how a role behaves, add a .ship/agents/<role>.md file to your repository, and SHIP appends it after its own charter for that role instead of replacing it (agent prompts).
Naming the harness matters as much as naming the model. The study pinned the Planner, Reviewer and Tester to one configuration and varied only the Builder, across 13 configurations. Three open-weight models each ran on two harnesses, and with the model held fixed, the harness moved the median mission time by as much as 37% and changed how many tickets were resolved (SWE-in-a-team). As the study concluded, "a result that names the model without naming the harness is describing half of what ran."
Send rejected work back to the Builder
When a gate rejects, the work goes back to the Builder with that gate's feedback, and the loop runs again from the build. Two kinds of limit end the loop and hand the issue to a person (retry limits).
Oscillation limits count rounds in which a gate keeps failing the same way, such as the Reviewer raising the identical blocking comment twice in a row. Review, test and CI each default to 2. A gate that fails a different way each round is making progress and does not count toward its limit, and when a review is judged to be progressing, up to 3 extra review rounds are granted. Above all of that sits an outer ceiling of 25 Builder attempts, far enough up that it rarely fires.
# ship.yml: the defaults, written out
retries:
reviewOscillationLimit: 2
qaOscillationLimit: 2
ciOscillationLimit: 2
maxReviewGrants: 3
maxRetries: 25SHIP reads these limits from ship.yml on your main branch, so a pull request cannot loosen the limits on its own review, and each mission keeps the limits it started with. At a limit the mission shows Needs attention, and the issue gets a comment naming the gate that stopped it and that gate's last feedback. The pull request stays open. A reply on the issue resets the counters and sends the agents back in with your direction.
In the study, 214 of the 260 runs went through on the first push and 46 needed the loop. The loop corrected 40 of those 46 with no human involved, 31 of them on the very next push and all within 4 Builder attempts (SWE-in-a-team). Correction cost about a third more than getting it right the first time with the same builder. The limit fired as well: two builder configurations could not turn CI feedback into a passing build, repeated the same failure 3 to 5 times, and hit the oscillation limit, which is the limit doing its job.
Keep acceptance with a human
By default a mission runs the whole way and stops at a pull request for a person to accept (human in the loop). Checkpoints add holds earlier in the run:
| Checkpoint on | Holds after | What you read |
|---|---|---|
| Planner | The plan is written | The approach, before any code exists |
| Builder | The push has landed and CI has run | The commits, and whether your pipeline went green |
| Reviewer | The review comes back clean | The diff and the verdict on it |
The Tester cannot be a checkpoint. It is the last stage, and the pull request after it is already that gate.
A hold has two answers. Resume, and the run carries on from where it stopped. Or send changes, which re-runs the held stage with your feedback, so a plan you disagree with is re-planned before anything is built. Pause, resume and stop work on any mission, checkpoint or not. Stop ends the mission and leaves the branch and the pull request open, so picking the work up later is the normal way back.
Anything you tell a running mission outranks the automated guards. If it is close to a limit because it keeps failing the same way, your instruction clears that history and the loop continues on what you said. Merging stays with you unless you turn on auto-merge for a project or a single mission, which lets a run that passes review and test merge its own pull request.
Work enters the same loop from wherever it starts. Assign a Linear issue to SHIP, comment @SHIP on a GitHub issue, or hand a plan or a finished branch over from your terminal with ship delegate (ways of working).
Watch for work that passes every gate
Gates lower the chance that a bad change reaches a person, but they cannot prove the change does what was asked. In the study, 4 of the 260 runs passed every gate and were still not complete, and only the hidden acceptance tests caught them (SWE-in-a-team).
The clearest case was a ticket asking for class reminders to be queued (F5). The most expensive builder in the study, claude-fable-5 on claude-code, wired the live email provider straight into the request path, so the endpoint returned a 500 whenever the API key was missing, although the ticket only asked it to queue a row. The repository's own test suite stayed green, so CI passed, and the Reviewer and the Tester both approved. The mission cost $8.27 and reached Ready for Acceptance in 23 minutes, according to the study's per-run data. In the study's words, "It built the wrong thing, correctly."
Across the runs the hidden tests failed, the misses fell into five groups, and none of them was asserted anywhere in the repository's own suite:
- an edge case the acceptance criteria named, such as an invoice where every line is refunded
- missing idempotency, such as the same webhook delivered twice paying an invoice twice
- bad input treated as success, such as an invalid webhook signature that should return 400 and change nothing
- a migration that looked finished and was not
- behavior a unit test never sees, such as a bounded number of reads however many bookings exist
Write acceptance criteria that a test can check, and keep a person at acceptance. The gates catch most mistakes early and cheaply, and the person accepting the pull request is the check for a change that clears every gate and still misses the intent.
Measure cycle time, cost per change and resolve rate
Three measures tell you whether an agentic SDLC is working, and each needs a precise definition before its number means anything. The last column shows what each looked like in SWE-in-a-team.
| Measure | Definition | In the study |
|---|---|---|
| Cycle time | Wall-clock time from hand-over to Ready for Acceptance | Between 14 and 25 minutes at the median, depending on the builder |
| Cost per change | Inference cost per mission, split by stage, divided by the changes that were accepted rather than the ones opened | $4.31 per resolved ticket with claude-sonnet-5 and $5.48 with claude-opus-5, both on claude-code with all 20 resolved |
| Resolve rate | Share of missions that reached Ready for Acceptance and passed their acceptance tests | 7 of the 13 builders resolved all 20 tickets |
SHIP records the inputs for all three. It persists the full agent session, the assembled prompt, token usage and cost of every run before the agent reports back, and surfaces them per mission and per stage (built to be checked). Spend is priced from list rates for the model that actually served each call, and cost, duration and outcome are recorded per stage (LLM credentials).
The study used those records to see where the money went (SWE-in-a-team). CI raised the first objection on 27 of the 46 rejected runs, the Reviewer on 15 and the Tester on 4, so the cheapest gate did most of the catching. Catching late cost more: runs the Tester rejected came to a median of $8.18, against $3.51 for a run no gate objected to and $3.37 for one CI rejected. On a CSV export ticket (F1), the Tester cost $5.31 of a $7.23 mission because it ran its whole check twice.
Some measurement is still ahead of the product. DORA metrics rolled up across missions, A/B experiments that compare harness and model configurations, and a fleet that improves its own setup are planned, and none of them is live today.
Start with one ticket type
Pick one category of work whose acceptance criteria a test can check, and widen only after you have read what the first missions cost. The study spread its 20 tickets across ten categories (bug, security, refactor, dependency upgrade, migration, performance, feature, integration, growth and DevOps), and each one carried criteria in this shape:
Feature (F1) - Export a date-range CSV of bookings for accounting
Acceptance Criteria
- GET /api/export?type=bookings is supported alongside members and
invoices, and requires a signed-in session.
- Optional from/to ISO-8601 params. When given, include only bookings
whose class session START time falls within [from, to], INCLUSIVE of
both ends. An omitted bound is unbounded that side.
- CSV columns, in this exact order: Starts, Class, Member, Email,
Status. Starts is the session start as its ISO-8601 UTC timestamp.
- RFC 4180 escaping: quote fields containing a comma, quote or newline,
and double any embedded quotes, exactly as the invoices export does.Route the work the way it arrives. Features and bugs usually come in as tickets, because the people asking for them already work in the tracker, while refactors, upgrades and chores go over from the terminal, where a developer has already specified them (ways of working). Both land as pull requests in the same loop.
For the first missions, set a checkpoint on the Planner or the Reviewer so you read the approach or the verdict before the run moves on (human in the loop). Leave the retry limits at their defaults, since the docs' own advice is that "reaching a human after two identical failures is usually cheaper than five" (retry limits). Then read the cost and the verdict of each mission before adding a second ticket type. The use cases collect ready-made slices of work with their acceptance criteria already written.
Questions
What does agentic SDLC mean?
An agentic SDLC is a software development lifecycle in which AI agents carry each change from issue to pull request through planning, building, CI, review, preview deployment and testing, and every gate that rejects the work sends it back to the Builder with the feedback. Humans set the intent, take over when the loop stalls, and accept the result.
What is an autonomous SDLC?
An autonomous SDLC is an agentic SDLC described by how rarely a person steps in. Agents plan, build, review and test each change and correct their own rejected work, and people step in at the checkpoints they configure, when a retry limit stops a stuck loop, and to accept the pull request.
What is AI-driven SDLC and which tools support it?
An AI-driven SDLC is any software lifecycle in which AI assists or automates work inside a phase, from planning and code to tests, review and deployment. Running one end to end takes a tracker for intent, coding agents in isolated sandboxes, your own CI and preview pipeline as gates, and a platform that runs the loop and records what each change cost. SHIP is that platform, and it runs the coding agent you configure for each role.
What is the difference between traditional SDLC and AI-driven SDLC?
In a traditional SDLC, people do each phase and tools assist them. In an AI-driven SDLC, AI takes on tasks inside the phases and the constraint moves to verifying its output: Faros AI's 2025 report, drawn from 10,000+ developers, found PR review time up 91% on teams with high AI adoption. An agentic SDLC responds by making review and testing gates inside the loop itself.
Which SDLC phases can be automated?
Planning, implementation, CI triage, code review, preview deployment and acceptance testing can all run as agent stages with gates between them. Deciding what to build and accepting the result stay with people. In SHIP's 260-mission SWE-in-a-team study, the loop corrected 40 of the 46 runs its gates rejected with no human involved.
How can I automate the whole software development lifecycle with AI agents?
Give each stage one role and one gate, send every rejection back to the Builder with the failure attached, bound the loop with retry limits, and keep a person at acceptance. Start with one ticket type whose acceptance criteria a test can check, and read the cost and cycle time of each change before you widen the scope.