
In 2025, the research group METR ran a randomized trial with 16 experienced open-source developers. Before they started, the developers predicted AI tools would make them 24% faster. With AI, they took 19% longer to finish their tasks. Afterward, they still believed AI had sped them up by 20%.
That gap between how AI feels and what it delivers is the central problem in AI SDLC adoption. Google’s 2025 DORA research found that 90% of technology professionals now use AI at work, and more than 80% say it has made them more productive. The same research found that AI adoption still correlates with lower software delivery stability.
Code gets written faster. Whether software ships faster, safer, and cheaper depends on what happens in every other phase of the lifecycle: requirements, design, review, testing, deployment, and maintenance.
This blog covers:
- What AI SDLC means, and how AI-assisted, agentic, and AI-native development differ
- What AI changes in each of the seven SDLC phases, and what stays with people
- What independent research says about productivity, quality, and security
- The risks worth managing, with the controls that address them
- A four-stage maturity model for adoption
- The metrics that show whether AI is improving delivery or only speeding up typing
What is AI SDLC?
AI SDLC is the use of AI models and agents across every phase of the software development lifecycle, with people setting intent, making architectural decisions, and approving what ships. The term covers far more than AI code generation. It includes AI that drafts requirements, proposes designs, writes and runs tests, reviews pull requests, watches production, and triages incidents.
Teams usually sit in one of three operating modes. Most organizations run all three at once, in different teams.
| Mode | How AI participates | Who drives the work | Typical tools |
|---|---|---|---|
| AI-assisted | Suggests code, answers questions, drafts documentation inside the IDE | The developer, keystroke by keystroke | Code completion, chat assistants |
| Agentic | Takes a scoped task (a bug, a test suite, a migration), plans it, edits files, runs commands, and opens a pull request | The developer assigns and reviews; the agent executes | Coding agents in the IDE, terminal, or CI |
| AI-native | Agents work in every phase, sharing context through specs, tickets, and repositories; people approve at defined gates | The team sets intent and owns decisions; agents do most of the execution | Agents plus a shared context layer, governance, and measurement |
AWS describes one version of the AI-native mode as the AI-Driven Development Lifecycle (AI-DLC). In AI-DLC, AI drafts a plan, asks the team clarifying questions, and implements only after people validate the plan. Work runs in short cycles AWS calls “bolts,” measured in hours or days instead of two-week sprints.
The mode matters less than the operating discipline around it. A team using simple code completion with strong testing and review often gets more from AI than a team running AI agents on a fragile codebase.
How AI changes each phase of the SDLC
AI now contributes to all seven phases, but the human role changes shape in each one. The table shows what AI handles today, what stays with people, and the signal that tells you it is working.
| Phase | What AI does today | What people still own | Signal it is working |
|---|---|---|---|
| 1. Planning and requirements | Turns interview notes, support tickets, and chat threads into draft user stories and acceptance criteria; flags conflicts with existing code | Deciding what to build and why; trade-offs between scope, cost, and time | Fewer stories reopened for missing criteria |
| 2. Design and architecture | Proposes data models, API contracts, and architecture options; generates clickable UI prototypes | Choosing the architecture; security, compliance, and scalability decisions | Design reviews resolve faster with fewer late changes |
| 3. Development | Writes functions, refactors, migrations, and boilerplate; agents complete scoped tasks and open pull requests | Breaking work into reviewable units; owning the code that merges | Lead time drops without a rise in rework |
| 4. Code review | Summarizes changes, checks against team standards, flags likely bugs and insecure patterns before a human looks | Judging intent, design fit, and business logic; final approval | Review wait time falls; escaped defects stay flat or fall |
| 5. Testing | Generates unit and integration tests, test data, and edge cases; maintains tests as code changes | Deciding what must be tested; validating that tests check behavior, not implementation | Coverage of critical paths rises; flaky tests fall |
| 6. Deployment | Writes and checks infrastructure-as-code and pipeline configuration; reads logs during rollout | Release decisions, rollback criteria, change approval in regulated systems | Change failure rate holds or improves as frequency rises |
| 7. Operations and maintenance | Triages alerts and bug reports, suggests root causes, drafts fixes for dependency and vulnerability updates | Incident command, customer communication, accepting risk | Faster recovery; maintenance backlog shrinks |
The bottleneck moves downstream
When AI speeds up code writing, the constraint shifts to whatever comes next. Pull requests arrive faster than reviewers can read them. Tests written by the same model that wrote the code can share its blind spots. More changes reach production, and each one is a chance to break something.
DORA’s 2025 research describes AI as an amplifier for this reason. Teams with loosely coupled architectures and fast feedback loops see gains. Teams with tightly coupled systems see little or no benefit. Investment in review, testing, and deployment safety has to grow with code output, or the lifecycle stays as slow as its slowest phase.

Context decides output quality
An agent that cannot see your architecture decisions, coding standards, or open tickets will write plausible code that does not fit. The teams getting the most from AI SDLC store specs, decisions, and standards as versioned files next to the code, where both people and agents can read them. DORA lists “AI-accessible internal data” as one of seven capabilities that increase AI’s benefit.
What the research says about AI SDLC
Independent studies agree that adoption is near-universal and that the gains are real but uneven. Speed improvements show up in controlled tasks; quality, security, and stability need deliberate work to keep up.
| Study | Sample | Key finding |
|---|---|---|
| DORA State of AI-assisted Software Development (Google, Sept 2025) | Nearly 5,000 technology professionals | 90% use AI at work; over 80% report higher productivity; 30% have little or no trust in AI-generated code. AI now correlates with higher throughput but still with lower delivery stability. |
| Stack Overflow Developer Survey (2025) | 49,000+ developers, 177 countries | 84% use or plan to use AI tools, up from 76%. 46% distrust AI accuracy, up from 31%. 66% name “almost right, but not quite” answers as their top frustration. |
| Google enterprise RCT (2024) | 96 Google engineers | AI cut time on a complex enterprise task by about 21%, with a wide confidence interval. |
| METR randomized trial (July 2025) | 16 experienced open-source developers, 246 tasks | Developers took 19% longer with AI, yet believed they were 20% faster. |
| METR follow-up (Feb 2026) | Late-2025 tools, new and returning developers | Estimates moved toward a speedup, but METR called the data “only very weak evidence”: many developers refused to work without AI, skewing the sample. |
| Veracode GenAI Code Security Report (July 2025) | 100+ LLMs, 80 coding tasks | 45% of AI-generated code contained OWASP Top 10 vulnerabilities. Java failed over 70% of the time. Larger models were no more secure. |
| GitClear code quality research (Feb 2025) | 211 million changed lines of code | Duplicated code blocks rose 8x in 2024; moved (refactored) lines fell 39.9%. Copy-pasted lines outnumbered moved lines for the first time. |
Three conclusions from the data
Perceived productivity is a poor guide.
METR’s developers misjudged their own speed by nearly 40 percentage points. Decisions about AI tooling need measured delivery data, not satisfaction surveys alone.
Speed gains depend on the task and the team.
The Google trial showed a 21% gain on a scoped enterprise task. METR’s experts, working in large codebases they knew well, saw a slowdown. Taken together, the studies suggest AI helps most on well-defined tasks, and least where an expert already knows the codebase deeply.
Quality and security do not improve on their own.
Veracode found that model size made no difference to security. GitClear’s data points to more duplication and less refactoring. Both risks grow with volume, which is why review, testing, and security scanning must scale with AI output.
Risks to manage, and the controls that work
Every risk below has a known control. Teams that put the controls in place before scaling AI output keep the speed and avoid the cleanup.
| Risk | What the evidence shows | Control |
|---|---|---|
| Insecure code | 45% of AI-generated code samples carried OWASP Top 10 flaws (Veracode) | Static and dependency scanning on every pull request; secure-coding rules in the agent’s instructions; security review for auth, payments, and data access |
| Maintainability drift | Duplicated code blocks up 8x, refactoring down 39.9% (GitClear) | Duplication checks in CI using static analysis tools; standards files the AI reads; scheduled refactoring work |
| “Almost right” code | 66% of developers cite it as their top frustration (Stack Overflow) | Tests written before or alongside the change; small pull requests that a reviewer can read in one sitting |
| Delivery instability | AI adoption still correlates with lower stability (DORA 2025) | Feature flags, progressive rollouts, automated rollback, and a change failure rate target |
| Over-trust | Developers believed they were 20% faster while measuring 19% slower (METR) | Decisions based on measured lead time and defect data, not perception |
| Data and IP exposure | Prompts and context can carry source code, customer data, or secrets to external services | Approved tools list, enterprise data terms, secret scanning, and a written AI usage policy (see our guide to security and compliance for AI systems) |
| Skill erosion | Junior engineers can merge code they could not have written or debugged | Pairing, explain-your-change rules in review, and time for unassisted problem-solving |
Governance that keeps pace
The first capability on DORA’s list is a “clear and communicated AI stance”: a written policy on which tools are approved, what data they may see, and where human approval is mandatory. For regulated industries such as banking and healthcare, the policy should also define agent permissions and audit trails: which changes an agent made, who approved them, and which requirement they trace to.
What changes for engineers
The human role moves up the stack. Engineers spend less time typing implementation and more time on four things AI handles poorly today:
- Specifying intent. Writing the requirement, constraint, and acceptance test clearly enough that an agent can act on it.
- Architecture. Choosing boundaries and trade-offs that keep a system changeable for years.
- Review and judgment. Deciding whether a correct-looking change is the right change.
- Accountability. Owning production behavior, incidents, and customer impact.
Senior engineers become more valuable in this model, because review and architecture are where AI-generated work succeeds or fails.
A four-stage AI SDLC maturity model
AI SDLC adoption works best in four stages, with each step up earned by data. Skipping stages puts agent output on top of review and testing processes that cannot absorb it.

AI SDLC maturity model · 4 stages, 3 gates
Each gate is a measurable condition. A team that cannot show stable delivery metrics at one stage is not ready to add more AI output at the next. At Stages 3 and 4, agents need operational oversight of their own: AgentOps covers how to observe, evaluate, and govern them in production.
Foundations to build first
DORA’s 2025 research identified seven capabilities that increase the benefit teams get from AI. Each one is worth checking before moving past Stage 2:
- Clear and communicated AI stance
- Healthy data ecosystems
- AI-accessible internal data
- Strong version control practices
- Working in small batches
- User-centric focus
- Quality internal platforms
Most of these are ordinary engineering practices. AI makes their absence more expensive, because it multiplies the volume of change flowing through them.
How to measure AI SDLC success
Measure delivery outcomes at the team level, and set a baseline before rollout. Without a baseline, you will be left with the METR problem: everyone feels faster and nobody can prove it.
| Metric | What it tells you | Warning sign |
|---|---|---|
| Lead time for changes | Time from commit to production; the clearest view of end-to-end speed | Coding time falls but lead time does not: the bottleneck has moved to review or testing |
| Deployment frequency | How often value reaches users | Frequency rises alongside change failure rate |
| Change failure rate | Share of deployments that cause a failure in production | Any sustained rise after AI rollout |
| Failed deployment recovery time | How quickly the team restores service | Longer recovery on AI-written code the team understands less well |
| Rework rate | Share of deployments that are unplanned fixes | Rising rework hides behind rising throughput |
| Pull request review time | Hours a change waits for review | Growing queue as AI output increases |
| Escaped defects | Bugs found by users instead of tests | Rising count per release |
| Security findings in new code | Vulnerabilities per change, split by AI-assisted vs. not | AI-assisted changes carry more findings |
| Code duplication | Share of duplicated blocks in new code | Steady climb quarter over quarter |
| Developer experience | Survey of focus time, friction, and confidence in the codebase | Satisfaction rises while delivery metrics stay flat |
The first five are the DORA delivery metrics, widely used as an industry benchmark. Suggestion acceptance rate and lines of AI-generated code are easy to collect, but they measure activity. Report them alongside outcomes, never in place of them.
Conclusion
AI SDLC pays off when the whole lifecycle can absorb faster code. With 90% of technology professionals already using AI, adoption is settled. The open question for each team is whether review, testing, security, and deployment can keep up with what AI produces.
The research points to four habits that separate teams seeing gains. They connect AI to their specs, standards, and codebase. They keep changes small and reviewable. They scale safety nets in step with output. They measure lead time, stability, and rework instead of how fast developers feel.
Start with one workflow, one team, and a baseline. Expand once the numbers show that delivery, and not only typing, has improved.
Tell us what you are building
[x]cube LABS has shipped 950+ digital products since 2008, with 600+ experts across 15+ industries. Our AI-native engineering team rebuilds delivery around AI, with evaluation, security, and governance in place from day one.
Frequently asked questions about AI SDLC
What is AI SDLC?
AI SDLC is the use of AI models and agents across every phase of the software development lifecycle: planning, design, development, review, testing, deployment, and maintenance. People set intent, make architectural decisions, and approve what ships.
What is the difference between AI SDLC and AI-DLC?
AI SDLC is the general practice of applying AI across the lifecycle. AI-DLC (AI-Driven Development Lifecycle) is a specific method published by AWS, in which AI drafts plans, asks clarifying questions, and implements after human validation, in short cycles called “bolts.”
Which SDLC phase benefits most from AI?
Development shows the fastest visible gains, with one Google trial measuring about 21% less time on a complex task. The larger long-term gains usually come from review, testing, and maintenance, because those phases become the bottleneck once code is written faster.
Is AI-generated code secure?
Not by default. Veracode tested 100+ models and found OWASP Top 10 vulnerabilities in 45% of AI-generated code samples. Automated security scanning and human review on sensitive code paths are required before AI-written code reaches production.
Will AI replace software developers?
The evidence points to a change in the developer role, with more time spent on specifying requirements, architecture, review, and production ownership. Gartner predicts 75% of enterprise software engineers will use AI code assistants by 2028, up from under 10% in early 2023.
How do you measure the ROI of AI in the SDLC?
Set a baseline, then track lead time, deployment frequency, change failure rate, recovery time, and rework rate at the team level. Add review time, escaped defects, and security findings. Avoid using suggestion acceptance rate or lines of AI code as success measures on their own.
How should a company start with AI SDLC?
Publish an AI usage policy, pick one high-friction workflow such as test generation or code review, run a time-boxed pilot with one team against a measured baseline, and expand only when delivery metrics improve.