A software change begins before someone writes code and remains the team's responsibility after deployment. That makes the full software development lifecycle the right unit for understanding AI's effect on delivery. A faster draft matters only in relation to the decisions, reviews, dependencies, and operational consequences that surround it.
The useful question at each stage is specific: what can the tool contribute, what evidence should a person examine, and what must be true before the work proceeds? This approach avoids treating software delivery as a coding contest or assuming that every stage benefits equally.
Discovery and requirements: expose missing decisions
A team can use a language model to organize interview notes, compare documents, or draft acceptance criteria. These are candidate artifacts. Their value depends on whether they faithfully represent the business need and reveal unanswered questions.
Suppose a team receives a request to automate account closure. A polished requirement might omit pending transactions, disputed balances, retention obligations, or customers with multiple accounts. An analyst should ask the tool to identify ambiguities, then resolve them with the responsible people. Agreement among generated documents does not establish stakeholder agreement.
A useful practice is to attach examples to important requirements: an ordinary case, a boundary case, and a case that must be rejected. Those examples provide material for design and testing. They also make disagreements concrete enough to resolve before implementation begins.
Design: compare alternatives before producing artifacts
During design, tools can help explore interface options, summarize an existing codebase, or propose component boundaries. The team still needs to evaluate the assumptions behind each proposal. Does a suggested service require information that the organization does not collect? Does a new dependency have an acceptable operating model?
Document the decision and the reason for it, especially when a proposal is rejected. An architecture record should explain constraints such as transaction integrity, latency, data location, or support ownership. The benefit is continuity: later contributors can understand why a superficially simpler option was unsuitable.
DORA's AI Capabilities Model ties the success of AI adoption to a set of technical and organizational capabilities, not to tool access alone. [1] In practice, that is a reason to examine the quality of the context and working environment supplied to the team.
Implementation: make changes understandable
Generated code enters the same codebase and inherits the same maintenance obligations as manually written code. A reviewer should be able to explain its behavior, dependencies, error handling, and intended boundaries. If a patch is too large to understand, accepting it because it passed a narrow test transfers uncertainty into the product.
NIST's Secure Software Development Framework is a set of practices an organization can add to whatever lifecycle it already uses, with the aim of shipping fewer vulnerabilities. [2] This remains relevant regardless of how a draft was produced. Build integrity, dependency management, review, and vulnerability response do not disappear when code generation improves.
For a bounded trial, ask contributors to keep changes small and record where assistance was useful or costly. The purpose is learning about workflow fit. Turning those observations immediately into individual performance rankings would create incentives to conceal unsuccessful experiments.
Testing: preserve independence of judgment
A model can suggest tests, but tests derived from the same incomplete assumption as the implementation may reinforce the error. The team should trace critical tests to independently established requirements, known failure conditions, and business examples.
For applications containing models, conventional tests need additional behavioral evaluations. Anthropic's evaluation guidance distinguishes code-based, model-based, and human graders, each with different strengths and limitations. [3] The appropriate combination depends on the property being tested.
For example, a deterministic check can confirm that a required field exists, while a specialist assesses whether a summary misrepresents a policy exception. A model-based grader can help scale some assessments, but its agreement with expert judgment should be checked. Test results are only as meaningful as the criteria they implement.
Release and operations: carry evidence forward
A release decision should connect the change to its verification evidence, known limitations, and recovery plan. If a model-backed feature fails, the available response might be disabling that feature, reverting its configuration, or restoring a conventional workflow. Plan the response before customers depend on it.
Operational learning should return to requirements and tests. A misunderstood request may reveal a missing business rule. A recurring retrieval failure may require a data correction rather than another prompt revision. Logging should help distinguish these causes without unnecessarily retaining confidential material.
Treat model, prompt, retrieval, and application changes as potentially behavior-changing releases. The degree of review can vary with risk, but the team should know which configuration is serving users. Our MLOps and LLMOps research explores this area further.
Measure the connected process
The SPACE research warns against reducing developer productivity to a single dimension. [4] For this lifecycle, compare outcomes such as time to a usable release, rework, defects, and interruption burden. Include the effort required to review and repair generated work.
Later stages can invalidate earlier assumptions, which is why the lifecycle is drawn as a loop. Improving one stage is useful when its output helps the next stage succeed. A team evaluating AI should therefore follow a change all the way through production and ask whether the complete process became more effective, rather than stopping the measurement when a draft appeared.
References
- DORA. AI Capabilities Model. 2025.
- NIST. Secure Software Development Framework, SP 800-218. February 2022.
- Anthropic. Demystifying Evals for AI Agents. January 9, 2026.
- Forsgren et al., Microsoft Research. The SPACE of Developer Productivity. 2021.
Sources checked October 2026. CorpExcellence.com articles are best-effort research and analysis, not professional advice.
Leave A Comment