Skip to main content

All Insights

The Hammer and the New Hire: Rebuilding the Software Lifecycle for the Age of AI

Alden Mallare, Principal Test Architect, maps what replaces the traditional software development lifecycle once AI writes the code and the tests, from the moved bottleneck to an eight-phase agentic model with human gates.

Author

Alden Mallare
Principal Test Architect at UTurn Data Solutions

Highlights

• Understand why story points and sprint velocity stop measuring anything once agents write the code

• See how the bottleneck moves from writing code to verifying it, and what breaks along the way

• Learn why AI-written tests are a compromised referee for AI-written code

• Explore the eight-phase Level Z lifecycle and the three human gates it keeps

• Get the Constraint Manifest model that sorts agent decisions into autonomous, supervised, and prohibited

Share To

September 29, 2026

A developer can ship ten pull requests before lunch. The person reviewing them still has one pair of eyes and the same eight hours. Every version of the software development lifecycle, from waterfall to Agile to DevOps, was calibrated to human throughput, and AI removed that constraint on the production side while leaving verification exactly where it was. Alden Mallare, Principal Test Architect at UTurn Data Solutions, traces where the old model cracks, why AI-written tests make a compromised referee for AI-written code, and what an eight-phase agentic lifecycle with real human gates looks like phase by phase. Written for CIOs, CTOs, and the engineering and QA leaders deciding what their lifecycle needs to become.

Our development process was built on one assumption nobody wrote down. AI just broke it.

A developer sits down on a Tuesday morning with coffee and an AI assistant. She ships four pull requests before her second meeting. Each one looks clean, each one arrives with passing tests, and by lunch she has opened ten.

Now think about the person reviewing them.

That reviewer still has one pair of eyes and the same eight hours as everyone else. Nothing about their day got faster; the queue just got longer. Somewhere around pull request number seven, "review" turns into "skim and approve." Nobody decided that. The process drifted there on its own.

I recently gave a talk called AI SDLC, short for the Artificial Intelligence Software Development Life Cycle. It came out of more than six months of research and a lot of hands-on time building AI-assisted testing into real client work. This article is the long version of that talk.

You don't need to be an engineer to follow it. If you've wondered why AI hasn't made software projects faster or cleaner in any way you can measure, keep reading.

The Lifecycle We All Grew Up With

Before getting into what's breaking, the old model deserves some credit. The traditional Software Development Life Cycle, or SDLC, has guided how teams build software for decades. Waterfall, Agile, and DevOps each put their own spin on it, but most versions share the same six steps underneath.

Planning comes first. The team works out what problem it's solving, who it's solving it for, and what the risks and resources look like. Requirements follow, usually captured in a formal specification that serves as the project's blueprint. Then design, where architects decide how the system is structured, how the pieces talk to each other, and where the data flows.

Development is the part most people picture: engineers writing code in Python, Java, or TypeScript. Testing checks the product against the original requirements so defects get caught before customers find them. Deployment and maintenance close it out, with the team keeping the live system healthy through monitoring, fixes, and updates.

It fits on one slide. It has also survived forty years of reinvention, which is rare, and that durability explains why so many organizations still anchor to it while the ground moves under them.

The Assumption Underneath Everything

It took me a while to see what all of it rested on. Every version of that lifecycle, from a rigid waterfall plan to a two-week sprint, was built around one constraint: human throughput.

Every meeting, approval gate, and estimation ritual assumes a person is the limiting factor on how fast work gets produced. Sprints exist because people can only write so much code in two weeks. Code review works because one developer produces about as much as one reviewer can read.

The whole machine is tuned for scarcity. AI flips that. When an assistant drafts a feature in minutes and an agent generates a thousand test cases before your coffee cools, output is no longer scarce, and a framework designed to ration it starts behaving strangely.

That was the core argument of my talk. Traditional models still answer their original question well; the question changed. Most of the pain teams feel when they adopt AI comes from running a scarcity playbook in a world of plenty.

Where the Old Model Starts to Crack

The cracks show up in small places first. A dashboard that looks too good. A review queue that never shrinks. A bug nobody can trace. Grouped together, they fall into four families.

The bottleneck moved, and nobody told the process

AI multiplies how much code gets written without multiplying how much gets checked. More work lands every sprint, while review, QA, and release management still run at human speed. The bottleneck slides downstream from writing code to verifying it, and the traditional SDLC has no mechanism for noticing.

Story points show the confusion clearly. Teams spent years calibrating estimates to human coding effort, so when an agent finishes a three-point story in eight minutes, the estimate stops meaning anything. Capacity plans built on historical velocity end up predicting how fast a team can produce code. They say nothing about how fast it can confirm the code is right.

Pull request queues tell the same story. Opening a pull request is now close to free. Agile assumed a rough balance between what developers produce and what reviewers can absorb, and that balance is gone the moment one person, or one agent, can open ten before lunch.

The ceremonies stopped fitting the work

Most teams run on shared agreements. A Definition of Ready says when a story can start; a Definition of Done says when it's finished. Both were written for code that behaves the same way every time it runs.

AI output doesn't work that way. Give an agent the same prompt and the same constraints on two different days and you can get two different results. A checklist item like "meets acceptance criteria" tells you it worked this time. It tells you nothing about tomorrow.

Daily standups and fixed sprint rhythms assume a steady pace of change because human output is roughly linear. Agents ignore the sprint clock. They work continuously and in parallel across multiple stories, which makes "in progress" hard to pin down.

Even the familiar left-to-right picture of plan, code, build, test, and release stops describing reality. With AI helping developers, helping testers, and acting on its own at the same time, coding and testing run as concurrent tracks. A lifecycle drawn as a straight line hides that.

Quality got harder to see

This is the family I worry about most, probably because quality is where I've spent my career.

Start with coverage inflation. AI can generate test cases fast enough to make coverage numbers look spectacular. Coverage only measures what was checked, though. It can't tell you whether anyone asked the right questions, and the failure that takes down production is usually the one nobody thought to ask about.

Then there's a problem I call the validator validating the validator. When AI writes the code and AI also writes the tests, they tend to share the same blind spots. Researchers studying language models as judges have found that models rate their own output higher than other models' output, even when human evaluators see no difference in quality (Panickssery et al., NeurIPS 2024). A follow-up study traced the bias to familiarity: evaluators favor text that reads like something they would have produced themselves (Wataoka et al., 2024).

A test suite written by the same kind of system that wrote the code is a compromised referee. Think of a student grading their own exam with total sincerity and zero independence.

Regressions get harder to spot, too. In the old world a regression was obvious: a test that used to pass now fails. In an agentic system, behavior can drift gradually while every existing test stays green, because nobody wrote a test for a behavior nobody knew to expect.

Then human review turns into rubber stamping. Real code review means reconstructing what the author intended and checking whether the code matches. That takes time and focus. Buried under AI-generated changes, the reviewer who used to read carefully starts approving on vibes.

Nobody signed for it

The last family is governance, which comes down to one question. Who decided this, and were they allowed to?

The traditional SDLC assumes every change was started on purpose by a person who chose to touch that file. It has no concept of autonomy boundaries and no way to say what an agent may decide alone versus what needs a human in the loop. In most organizations those rules don't exist yet, or they get written in a hurry after something goes wrong.

There's no automatic paper trail either. Human pipelines carry an implicit record in the people who touched them: who committed, who approved, who deployed. Agentic pipelines don't produce that record on their own, and many teams find the gap only when an auditor or an incident review asks who approved what, and why.

Rollbacks have the same blind spot. A traditional rollback plan covers code. When an AI-powered system misbehaves, the cause might be the code, the model version running underneath it, or the prompt that shaped its behavior. Roll back only the code and you may be debugging the wrong layer.

A Hammer or a New Hire?

Somewhere in the middle of this research, I realized the frameworks weren't the only thing that needed to change. Our mental model of AI did too.

Most of us started out treating AI like a tool. A clever hammer. You pick it up, point it at a specific task, and put it down when you're done. The human does most of the thinking, the AI does most of the typing, and the tasks stay narrow enough that the AI never has to decide much on its own.

The newer model treats AI more like an employee. It isn't a person, but it can reason about a goal, make decisions within limits, and report back. An agent in this model does real work with light supervision. Like a good team member, it tells you what it finished, what worries it, and where it got stuck.

That shift changes the question leaders should be asking. The old one was "How do we build this capability?" The new one is "How do we build an agent that can build this capability for us?" The wording barely moves, yet the way you plan, staff, and govern the project changes completely.

Thinking of AI as a new hire also makes the governance problem obvious. Nobody hands a new employee the keys to production on day one. You give them a role, a set of permissions, a manager, and a way to escalate when something feels off. AI needs the same structure.

Climbing the Maturity Ladder

Organizations don't jump from zero AI to fully autonomous agents overnight, and they shouldn't try. In the talk I laid out a maturity model with four levels. Each one describes how deeply AI participates in the lifecycle and, just as important, where humans stay in charge.

Level 0: The Starting Line

Level 0 is the traditional SDLC with no meaningful AI involvement: six phases, humans at every step. Plenty of organizations still work this way for some or all of their portfolio, and that's fine. It's the baseline everything else gets measured against.

Level 1: AI Assisted

Level 1 is where most teams I talk to sit today. AI shows up in two phases, development and testing, and nowhere else.

In development, AI generates code, writes unit tests, suggests improvements, and analyzes what gets checked in. In testing, it drafts test cases and scenarios, writes automation scripts, runs the tests, and summarizes results. The productivity gain is real, and it demos well.

People are still the gatekeepers. Developers inspect the code and verify it does what was intended. Testers make the final call on whether the product is ready. AI does the heavy lifting in two rooms of the house while people run everything else.

Level 1 is a good place to start and a risky place to stay. It's also where the bottleneck hits hardest, because you've sped up the two phases that produce output without touching the phases that have to absorb it.

Level N: Full AI Assistance

At Level N, AI joins every phase of the lifecycle. This is where the ladder gets steep.

In planning, AI can pull market data, past project outcomes, and stakeholder input into draft objectives much faster than manual research. It flags risks and conflicting goals early, so the business sponsor starts with something to validate instead of a blank page.

In requirements, AI can turn rough stakeholder notes into structured acceptance criteria and point out ambiguities or missing edge cases for a person to resolve. It can also scan a sprawling backlog for requirements that conflict with or duplicate each other.

Design speeds up too. AI can produce three or four architecture options with the tradeoffs laid out side by side, then check a proposed design against the approved requirements to see whether anything was narrowed or reinterpreted along the way.

I borrowed a word here from Jenny Wen, design lead at Anthropic, who describes her team "jamming" around working prototypes instead of static documents. Designers, engineers, and product people gather around something that runs and decide in real time how it should evolve. Development gets its own version, which I call the "vibing" session: AI generates, refactors, and documents code at volume and flags every spot where it had to guess or fill a gap in the spec.

Testing and deployment round it out. AI produces large volumes of test cases, including the boundary and combination scenarios people rarely have time to write, and runs regression suites continuously. At deployment it validates configurations, runs pre-release checklists, drafts rollback plans and release notes, and estimates how far a change might ripple. In maintenance it watches for drift, triages incoming defects, and takes a first pass at root cause analysis.

The Human Gates at Level N

If AI touches everything at Level N, what's left for people? A lot. The work just looks different.

Every phase keeps a named human owner and a clear go or no-go question. The business sponsor confirms the stated objective reflects the real business problem and that budget, timeline, and regulatory constraints hold up. The product owner writes the acceptance criteria in plain language and resolves ambiguity directly, so AI isn't left guessing at intent. The architect confirms that tradeoffs around scalability, security, and technical debt were made on purpose.

The engineering lead's job changes the most. Instead of reading every line, they check traceability and drift. Does what got built still trace back to what was approved, or did scope creep and unapproved assumptions slip into the code?

The test lead owns acceptance criteria independent of the implementation, so the system never grades its own homework. They also make sure coverage maps to real risk and not to whatever was easiest to generate.

Release and maintenance come next. The release owner weighs readiness, rollback, and blast radius, and a named person explicitly approves every release. After launch, a product or quality owner periodically compares production behavior against the original intent, watching for the slow pileup of small, individually reasonable AI fixes that together pull a system somewhere nobody chose.

An old distinction in software engineering explains why these gates matter. Verification asks, "Did we build it right?" Validation asks, "Did we build the right thing?"

AI is driving the cost of verification toward zero. It can generate the code, generate the tests, and confirm that one satisfies the other. Validation stays human because it depends on intent, context, and consequence, and none of those live in the code. At every phase, the human gate answers one question: does this still mean what we meant?

Level Z: When the Agents Clock In

Level Z is the agentic end of the ladder. Agents do work on their own here, inside boundaries we define, and a six-phase lifecycle can't hold that. So I proposed an eight-phase model to replace it.

The illustration below maps the whole journey from intent through continuous learning, including the three human gates where a person has to sign off before work moves forward.

AI SDLC Maturity Model Level Z, part 1 of 3: phases 1 and 2 with the constraint sign-off human gate
AI SDLC Maturity Model Level Z, part 2 of 3: phases 3 through 5 with the QA go or no-go human gate
AI SDLC Maturity Model Level Z, part 3 of 3: phases 6 through 8 with the governance sign-off human gate and the loop back to Phase 1

The AI SDLC Maturity Model, Level Z (Agentic): eight phases, three human gates, and a feedback loop back to Phase 1.

A single example makes the phases easier to follow. In my working materials I built one around a fictional insurer called Clearpath, which wants an AI agent to help claims adjusters triage incoming property damage claims. The agent reads each claim, checks it against the policy, assigns a severity score, and suggests next steps. Adjusters keep the final say on large payouts.

It's a realistic mix of opportunity and risk, which makes it a good companion for the tour.

Phase 1: Intent and Feasibility

This phase replaces traditional planning and requirements. Its main output is something I call an Intent Register. A requirements document answers "What are we building?" The Intent Register asks a harder question: "What are we trying to learn, and how will we know when we've learned it?"

It holds bets instead of a feature list. For Clearpath, one bet is that the agent's severity scores will land close to an experienced adjuster's judgment most of the time. Another is that it will never invent a policy term that doesn't exist. Each bet sits next to the signal that would confirm it and the signal that would prove it wrong, so the team knows what failure looks like before anything ships.

The Intent Register also classifies every workstream. Is this piece built by people with AI help, or by an agent acting on its own? For Clearpath, the triage logic is agentic, while infrastructure scripting and test generation are AI-assisted.

That label matters because it decides which checkpoints apply later. Threat modeling and compliance obligations come in here too, at the start, instead of surfacing at the finish line.

Phase 2: Architecture and Constraint Definition

Phase 2 covers the system design you'd expect plus a layer that never existed before. This is where the Constraint Manifest lives, and I'd argue it's the most important document in the model. It spells out what agents are allowed to do, where a human has to step in, and what no agent may ever touch without sign-off.

The simplest way to build one is to sort every kind of automated decision into three zones. The Autonomous Zone covers work an agent can do with no human review because the stakes are small and easy to undo: formatting code, bumping approved dependencies, generating boilerplate scaffolding.

The Supervised Zone covers work an agent can propose but not finalize. Business logic lives here, along with database schema changes and anything touching authentication or payments. The agent might draft the change, but a person accepts it before it takes effect.

The Prohibited Zone is the bright line. No agent may ever change its own constraints, touch regulated data without explicit approval, or take any action that can't be traced back to a human-approved intent, however confident it seems. For Clearpath, the agent can score claims but can never approve or deny one, and it must route any claim showing fraud indicators to a person regardless of its score.

Infrastructure as code gets scaffolded in this phase as well. Security goes into the design itself instead of being bolted onto the pipeline as a scan at the end.

Phase 3: Four Tracks Running at Once

This is where the new model looks least like the old one. Phase 3 replaces the straight line from code to build to test with four tracks running at the same time.

The first is AI-assisted infrastructure scripting. Engineers use AI tools to generate infrastructure templates, cloud configurations, and pipeline definitions, and people own the final review. In the AWS world, that means tools like Amazon Q Developer, Kiro, and the AWS IaC MCP Server.

Kiro stands out because it works spec first. It turns a plain language request into written requirements before it generates anything, which lines up well with a Constraint Manifest that people use every day instead of one that sits untouched in a wiki.

The second track is AI-assisted software development. Developers use AI to write and accelerate code, but they hold the intent, and the output gets human review at the gate. The third is AI-assisted testing, where AI generates test cases and coverage while human QA checks that the testing system is asking the right questions. The validator must be validated.

The fourth track is a different kind of work. This is agentic execution, where agents operate on their own within the boundaries set in Phase 2. Their output is probabilistic, so classic pass or fail testing doesn't fit; it gets evaluated against behavioral baselines and scenario suites that describe what good behavior looks like.

For cloud migration work, AWS Transform and Amazon Bedrock AgentCore are good examples of this track. They run agents that discover environments, map dependencies, and carry out multistep work under defined guardrails.

The first three tracks produce deterministic output, the kind quality engineering has always known how to review. The fourth produces something new. Keeping it on its own track stops that risk from hiding in the noise of a busy sprint.

Phase 4: Integrated Quality and Behavioral Validation

If I had to pick the phase that separates this model from everything before it, this is the one. Quality doesn't wait for development to finish. It runs alongside development the whole time.

Deterministic output from the AI-assisted tracks gets tested the familiar way. Probabilistic output from the agentic track gets evaluated against behavioral baselines, red team exercises built to break its assumptions, and drift checks. For Clearpath, that means feeding the agent claims it has never seen, like wildfire or flood damage, plus claims written specifically to trick it, and watching what it does.

The part that surprises people: QA owns the go or no-go decision. A pipeline setting doesn't make that call, and neither does a coverage percentage. A person with domain knowledge and business context has the authority to say a release isn't ready.

Most traditional models treat release as a checkbox that closes once automated criteria pass. Here it's the most consequential judgment call in the lifecycle, because a green light on a system carrying hidden risk is a failure of the quality process itself.

Phase 5: The DevSecOps Pipeline

Phase 5 carries forward the best of modern DevSecOps: automated pipelines, security scanning, supply chain checks, and gates a change has to clear before it moves on. Most of that discipline holds up fine with AI in the picture.

The scope is what changes. Supply chain checks now cover where a model came from, which prompt libraries are in play, and which tools an agent is allowed to call, on top of the code libraries a project imports.

Rollbacks cover three layers instead of one. If something goes wrong, the team can roll back the code, the model version, or the prompt, and it knows exactly which combination was running when the trouble started.

Phase 6: Deployment and Governance Handoff

Before any consequential AI system reaches production, it passes a formal governance checkpoint. The traditional lifecycle has nothing like this phase, because when people did all the work, the accountability trail came along for free.

Here the team completes what I call the evidentiary record. It documents what the system does, which constraints govern it, what acceptable behavior looks like, and who is accountable if it drifts. For Clearpath, the Intent Register becomes part of that record and is kept for years, so the company can explain its reasoning long after launch.

Regulators are starting to ask for exactly this kind of record. The EU AI Act requires technical documentation and automatic event logging for high-risk systems under Articles 11 and 12, and for most of those systems the obligations now apply from December 2027. Colorado took a narrower path. SB 26-189, signed in May 2026 and effective January 1, 2027, replaced the state's original AI Act and requires developers and deployers of automated decision-making technology to keep compliance records for at least three years, with insurance among the decision areas it covers.

Your own leadership will want the same document the first time something goes sideways.

Phase 7: Production Monitoring and Behavioral Assurance

Traditional maintenance watches for outages. Is the server up? Is the page loading? Those questions still matter, but they miss how AI fails.

AI systems don't always break. They drift. A system can stay online, respond quickly, and look perfectly healthy while its behavior slides away from what it was approved to do.

So Phase 7 watches behavioral signals continuously: how outputs are distributed, how often people accept the agent's suggestions, how often they override them, and how often it escalates. For Clearpath, a climbing override rate among adjusters is an early warning worth more than any uptime chart. When a signal moves outside the tolerance set in Phase 1, the system records what changed and under what conditions, and that record goes straight back to the Intent Register.

Phase 8: Production Support and Continuous Learning

The last phase keeps the loop open. Every incident, odd result, and customer complaint becomes something to learn from, on top of being a ticket to resolve.

Incident response follows the plan set in Phase 1. The evidentiary record from Phase 6 makes root cause analysis possible, because the team can see exactly what the system was supposed to do and under which rules. Confirmed findings feed back into the Intent Register for the next round, and the cycle starts again.

Who Watches the Watchers?

One more piece rounds out Level Z. Someone has to own the rules themselves. In this model that's an AI Center of Excellence, a governance function that sets and maintains standards across the organization without sitting in the approval chain for every release.

The AI Center of Excellence sets standards org-wide and audits evidentiary records after the fact

It owns the Prohibited Zone definitions company-wide and audits evidentiary records after the fact instead of gating each one up front. Delivery teams keep moving, and somebody makes sure the boundaries don't erode one reasonable exception at a time.

"Won't All This Slow Us Down?"

It's the first question people ask, and it's fair. AI promised speed, and here I am talking about constraint manifests, human gates, and evidentiary records.

In the narrowest sense, yes, a little. A Supervised Zone review takes longer than no review, and writing an evidentiary record takes longer than skipping it. That's the wrong comparison, though. The real one is catching a problem in Phase 4 versus discovering it in production, in front of customers, with nobody able to explain how it got there.

The data backs this up. Veracode's 2025 GenAI Code Security Report tested code from more than 100 language models and found security flaws in 45% of the coding tasks. Its Spring 2026 update put the security pass rate at about 55%, roughly where it sat two years earlier, even as the same models got far better at writing code that runs. Speed without a validation layer is debt with a delayed invoice.

The model also concentrates friction instead of spreading it evenly. Scrutiny goes where the stakes are high: Supervised and Prohibited Zone decisions, the go or no-go call, the moment a system enters production. Everywhere else, the Autonomous Zone lets agents run.

A team that gives a formatting change the same scrutiny as a database migration is wasting human attention. And human attention has become the scarcest resource in the building.

Where to Start Without Boiling the Ocean

Nobody adopts an eight-phase lifecycle in a single quarter. Trying to is the fastest way to watch a good idea die in committee. Start small, and start where the risk is real.

Pick one system that already has AI doing meaningful work. Skip your safest legacy app; you won't learn anything there. Write a Constraint Manifest for that one system, sort its automation into Autonomous, Supervised, and Prohibited, and name an owner for each zone.

If three smart people can't agree on what belongs in the Prohibited Zone, good. You just found out where your organization's real risk tolerance sits, and that's worth more than any diagram.

Next, look for places where AI is already grading AI. An AI-written test suite checking AI-written code. An automated reviewer approving an automated commit. Then ask who's standing outside that loop today. In a lot of organizations the answer is nobody, and that's where Phase 4 thinking pays off first.

Only after those two moves does it make sense to formalize the rest: the Intent Register, the evidentiary record, and behavioral monitoring. By then you'll have real data about where your boundaries belong instead of a workshop's best guess.

At UTurn, this is the ground we help teams cover, especially when AWS tooling and agentic workloads enter the picture. The first step, though, is one any team can take on its own next week.

The Question That Doesn't Change

AI is changing almost everything about how we build software: who writes the code, how fast it arrives, how many tests exist, and who wrote those. The work is getting faster and stranger at the same time.

One thing stays the same. Somebody still has to decide whether what we built is what we meant to build, and whether it's safe to put in front of real people. Machines can verify. People validate.

The hammer era is ending. The new hire has arrived: capable, fast, and occasionally wrong with total confidence. The teams that do well won't be the ones that automate the most. They'll be the ones that give their new colleague a clear role, firm boundaries, and a manager who's paying attention.

Sources

About the Author

Alden Mallare is a Principal Test Architect at UTurn Data Solutions, where he designs scalable test automation frameworks and helps teams bring AI into their delivery process responsibly. He has nearly three decades of experience across software development, quality assurance, and technology leadership, and writes about where AI, software quality, and technology teams meet.

Additional Insights