Agentic SDLC Overview (snapshot 2026-09-30)

How to read this

A synthesis of how AI coding agents change each phase of the software lifecycle. Tool facts come from the linked KB-AI notes (each verified 2026-09-30 against vendor pages). Evidence on productivity comes from two primary studies, both with limits (see bottom). Anything not backed is marked as opinion. Earlier list of repos: ai-sdlc-github-playbooks.

The shift in one paragraph

The unit of work moved from autocomplete to delegated tasks: an agent reads the repo, edits files, runs commands and tests, and opens a pull request, locally (Claude Code, codex CLI, opencode) or in isolated cloud environments (codex cloud tasks; OpenAI’s DevDay 2026-09-29 added reusable cloud environments; jules; devin). Humans move toward specifying, steering and reviewing. See async-development-workflows and agent-first-development.

Phase by phase

PhaseWhat agents changeKB-AI notesCaution
Requirements and planningPrompt-to-spec flows: requirements, design and task lists are written and reviewed before code (spec-driven development).kiro (requirements, design, tasks; GA 2025-11-17), speckit (GitHub open-source toolkit), agent-os, google-conductor, prp-product-requirements-promptSpecs are only as good as the human review; no verified evidence yet that spec-first beats other workflows.
Design and project contextPersistent instruction files replace re-explaining the repo each session. AGENTS.md is the cross-tool standard (reported in 60,000+ open-source projects by agents.md; now stewarded by the Agentic AI Foundation); Claude Code reads CLAUDE.md, and AGENTS.md when no CLAUDE.md exists.agents-md, claude-code-agents-convention, cursor-rules, ai-coding-rulesEach tool reads these files slightly differently.
ImplementationMulti-file edits, subagents and parallel background agents, hooks for deterministic guardrails, MCP for tools and data.Claude Code (hooks, skills, plugins, subagents), cursor, github-copilot (coding agent; usage-metered billing since 2026-06-01), codex, clineUnsupervised output quality varies; see definition-of-done below.
Code reviewAI reviewers comment on pull requests with codebase context; humans still approve.greptile, graphite (acquired by Cursor, Dec 2025 per that note), github-copilotReviewer agents raise the review load as PR volume rises (see DORA).
Testing and QATest generation and self-healing UI tests; agentic testing platforms aimed at AI-written code (mabl, April 2026).qodo Qodo Cover (open source), testsprite, autonomous-testing-tools, autify-dot-com, automate-unit-testing-with-ai-testgen-llm---cover-agentGenerated tests can encode the bug; keep human-defined acceptance criteria.
Delivery and operationsAgents in CI (Claude Code GitHub Actions / GitLab CI, Codex security scans) and scheduled routines.Claude Code, codex (Codex Security Cloud, DevDay 2026-09-29)Autonomy in pipelines needs permission scoping.
Verification of “done”Explicit definition-of-done prompts and checklists before accepting agent work.ai-coding-definition-of-done, vibe-coding (the unreviewed extreme)Opinion: treat acceptance criteria as a first-class artifact.

Evidence on productivity (mixed, early)

  • METR RCT (published July 2025): 16 experienced open-source developers, 246 tasks in their own repositories, mainly Cursor Pro with Claude 3.5/3.7. With AI they took 19% longer, while believing they were ~20% faster. Early-2025 tools, small sample.
  • METR follow-up (2026-02-24 design update): a later round (from Aug 2025; 57 developers, 143 repos, 800+ tasks) estimated a speedup of -18% (CI -38% to +9%) for returning developers and -4% (CI -15% to +9%) for new ones. Both intervals are wide, and METR says selection effects (30-50% of developers withheld tasks they did not want to do without AI) make the estimates unreliable; it is changing its experiment design.
  • DORA 2025 report (Google Cloud): 90% of respondents use AI at work (median ~2 hours/day); individual output up (21% more tasks, 98% more PRs merged) but delivery instability still rises and organisation-level metrics stay flat; 30% report little or no trust in AI-generated code. Headline: AI amplifies existing team strengths and weaknesses.
  • Labour-market side: ai-labor-market-impact-2026.

Practical implications (opinion, flagged)

  1. Invest in the loop around the agent: tests, CI, instruction files, review capacity.
  2. Measure delivery outcomes (lead time, change failure), not lines or PR counts.
  3. Keep humans accountable for architecture and acceptance.

Sources