Agentic SDLC Overview (snapshot 2026-09-30)
How to read this
A synthesis of how AI coding agents change each phase of the software lifecycle. Tool facts come from the linked KB-AI notes (each verified 2026-09-30 against vendor pages). Evidence on productivity comes from two primary studies, both with limits (see bottom). Anything not backed is marked as opinion. Earlier list of repos: ai-sdlc-github-playbooks.
The shift in one paragraph
The unit of work moved from autocomplete to delegated tasks: an agent reads the repo, edits files, runs commands and tests, and opens a pull request, locally (Claude Code, codex CLI, opencode) or in isolated cloud environments (codex cloud tasks; OpenAI’s DevDay 2026-09-29 added reusable cloud environments; jules; devin). Humans move toward specifying, steering and reviewing. See async-development-workflows and agent-first-development.
Phase by phase
| Phase | What agents change | KB-AI notes | Caution |
|---|---|---|---|
| Requirements and planning | Prompt-to-spec flows: requirements, design and task lists are written and reviewed before code (spec-driven development). | kiro (requirements, design, tasks; GA 2025-11-17), speckit (GitHub open-source toolkit), agent-os, google-conductor, prp-product-requirements-prompt | Specs are only as good as the human review; no verified evidence yet that spec-first beats other workflows. |
| Design and project context | Persistent instruction files replace re-explaining the repo each session. AGENTS.md is the cross-tool standard (reported in 60,000+ open-source projects by agents.md; now stewarded by the Agentic AI Foundation); Claude Code reads CLAUDE.md, and AGENTS.md when no CLAUDE.md exists. | agents-md, claude-code-agents-convention, cursor-rules, ai-coding-rules | Each tool reads these files slightly differently. |
| Implementation | Multi-file edits, subagents and parallel background agents, hooks for deterministic guardrails, MCP for tools and data. | Claude Code (hooks, skills, plugins, subagents), cursor, github-copilot (coding agent; usage-metered billing since 2026-06-01), codex, cline | Unsupervised output quality varies; see definition-of-done below. |
| Code review | AI reviewers comment on pull requests with codebase context; humans still approve. | greptile, graphite (acquired by Cursor, Dec 2025 per that note), github-copilot | Reviewer agents raise the review load as PR volume rises (see DORA). |
| Testing and QA | Test generation and self-healing UI tests; agentic testing platforms aimed at AI-written code (mabl, April 2026). | qodo Qodo Cover (open source), testsprite, autonomous-testing-tools, autify-dot-com, automate-unit-testing-with-ai-testgen-llm---cover-agent | Generated tests can encode the bug; keep human-defined acceptance criteria. |
| Delivery and operations | Agents in CI (Claude Code GitHub Actions / GitLab CI, Codex security scans) and scheduled routines. | Claude Code, codex (Codex Security Cloud, DevDay 2026-09-29) | Autonomy in pipelines needs permission scoping. |
| Verification of “done” | Explicit definition-of-done prompts and checklists before accepting agent work. | ai-coding-definition-of-done, vibe-coding (the unreviewed extreme) | Opinion: treat acceptance criteria as a first-class artifact. |
Evidence on productivity (mixed, early)
- METR RCT (published July 2025): 16 experienced open-source developers, 246 tasks in their own repositories, mainly Cursor Pro with Claude 3.5/3.7. With AI they took 19% longer, while believing they were ~20% faster. Early-2025 tools, small sample.
- METR follow-up (2026-02-24 design update): a later round (from Aug 2025; 57 developers, 143 repos, 800+ tasks) estimated a speedup of -18% (CI -38% to +9%) for returning developers and -4% (CI -15% to +9%) for new ones. Both intervals are wide, and METR says selection effects (30-50% of developers withheld tasks they did not want to do without AI) make the estimates unreliable; it is changing its experiment design.
- DORA 2025 report (Google Cloud): 90% of respondents use AI at work (median ~2 hours/day); individual output up (21% more tasks, 98% more PRs merged) but delivery instability still rises and organisation-level metrics stay flat; 30% report little or no trust in AI-generated code. Headline: AI amplifies existing team strengths and weaknesses.
- Labour-market side: ai-labor-market-impact-2026.
Practical implications (opinion, flagged)
- Invest in the loop around the agent: tests, CI, instruction files, review capacity.
- Measure delivery outcomes (lead time, change failure), not lines or PR counts.
- Keep humans accountable for architecture and acceptance.
Sources
- https://metr.org/blog/2026-02-24-uplift-update/ (via search results, accessed 2026-10-02)
- https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report (90% / 30% trust figures re-confirmed via secondary summaries, accessed 2026-10-02)
- https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ (via search results, 2026-09-30)
- https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report (via search results, 2026-09-30)
- https://agents.md (as recorded in agents-md, 2026-09-30)
- Tool facts: the linked KB-AI notes, each carrying its own Sources list (verified 2026-09-30)
- https://www.morningstar.com/news/pr-newswire/20260423ne41789/mabl-unveils-next-generation-agentic-testing-platform-for-the-ai-development-era (via search, 2026-09-30)