Agentic AI in Software Development: Verification Is the Real Constraint
Anil Kothiyal
|
27 Aug 2026
|
6 Min Read
Share
Agentic AI in software development shifts the bottleneck from writing code to verifying it. Coding agents now plan changes, run test suites, and open pull requests across isolated Git worktrees without supervision. Telemetry from 22,000 developers shows median pull request review time up 441% and incidents per merged pull request up 242.7%. For CTOs at funded SaaS and FinTech companies, agent throughput is capped by review and test capacity, not by model quality.
A coding agent can take a ticket, plan the change, edit files across a repository, run the tests, and open a pull request without a developer watching. Several can do it at once, in isolated worktrees, on separate branches. That capability arrived faster than most engineering organisations planned for, and it is already running in production at scale.
The uncomfortable part is what happens after the pull request opens. Someone still has to read it. It still has to pass a test suite that a human wrote. It still has to survive a deploy. None of those steps got faster.
This is where most agentic AI programmes quietly stall. Generation is solved well enough to be genuinely useful. Verification is not, and verification is now the thing that decides how much value an organisation actually captures from agents.
The gains are real, and they arrive with a bill attached
The 2025 DORA report, drawn from responses by nearly 5,000 technical professionals, found something that surprised readers of the 2024 edition. AI adoption now correlates positively with software delivery throughput. Teams are shipping more. The same report also found that AI adoption continues to correlate negatively with delivery stability, and concluded that acceleration exposes weaknesses downstream when control systems are weak.
Telemetry makes the shape of that trade clearer than survey data can. Faros AI analysed two years of delivery telemetry across 22,000 developers and more than 4,000 teams for its AI Engineering Report 2026. Median time in pull request review rose 441% year over year. Pull request size rose 51.3%. Bugs per developer rose 54%. Incidents per merged pull request rose 242.7%, meaning the probability of a production incident per change has more than tripled. And 31% more pull requests are merging with no review at all.
Anil Kothiyal is the Founder and CEO of EPixelSoft, an AI-native software engineering firm with 12 years and 700+ products shipped across FinTech, HealthTech, NGO operations, and SaaS. He has led engineering engagements for clients across the US, UK, Africa, and Asia — including platforms that compressed underwriting cycles from days to hours and field reporting systems deployed in East Africa. Anil writes about AI in production, high-stakes software delivery, and what it actually takes to build systems that hold up at scale.
Those figures describe a single system under strain. Code enters the pipeline faster than the review and test layer can process it. The queue grows. Reviewers start approving without reading. Defects that used to be caught in review get caught in production instead. Faros named the pattern Acceleration Whiplash, and the name is accurate: the gain sits at the top of the pipeline and the cost is distributed across every stage below it.
Agent throughput is bounded by verification capacity
Software delivery behaves like any other constrained system. Work moves at the pace of the slowest stage, and adding capacity anywhere else simply piles inventory in front of that stage. For most of the last fifteen years the slow stage was writing code. Pairing, small batches, trunk-based development, continuous integration: every practice we built assumed generation was expensive and verification was comparatively cheap.
Coding agents inverted that assumption in roughly eighteen months. Generation is now close to free. Verification is not, because verification is where human judgment still sits. Does this change do what the ticket asked. Does it fit the architecture. Does it open a security hole. Will it hold under load. Those questions have not been automated away, and the volume of code arriving at them has multiplied.
The security data explains why that judgment cannot be skipped. Veracode's 2026 GenAI Code Security Report found an average security pass rate of 56% across the models tested, with roughly 44% of code generation tasks introducing a risky vulnerability. That rate has barely moved since the first edition of the report. What changed is the denominator. Sonar's developer survey puts AI-generated or AI-assisted code at 42% of all code being written. A flat failure rate applied to a much larger volume produces a much larger absolute problem.
Gartner expects the control point to move accordingly. Its analysts predict that by 2027, more than 65% of engineering teams using agentic coding will treat the IDE as optional, shifting governance and validation into automated platforms instead. That is the same conclusion reached from the vendor side of the market. If humans cannot read everything, some of the reading has to be mechanised.
What an agent-ready delivery system looks like in production
The teams getting real leverage from agents are not the ones with the best model subscription. They are the ones that rebuilt four things before turning the agents loose on a repository.
The first is a test suite that carries real signal. An agent that can run tests will iterate against them until they pass, which makes the suite the effective specification. Where coverage is thin or tests are flaky, the agent optimises against a bad target and produces confident, passing, wrong code. The second is enforced batch size. The DORA capabilities model lists working in small batches as one of seven conditions that determine whether AI gains reach the organisation, and agent output pushes hard in the opposite direction. Splitting a single agent run across several small pull requests is slower to configure and far faster to review.
The third is isolation. Every agent run needs its own branch, its own ephemeral environment, and its own teardown. Preview environments per pull request stop parallel agents from colliding and give reviewers something running to inspect rather than a diff to imagine. The fourth is provenance. When 42% of code is machine-written, knowing which lines an agent produced, which model produced them, and which human approved them stops being a curiosity. In regulated work it is the difference between passing an audit and reconstructing change history from memory.
What this costs to build, honestly
None of it is free, and the sequencing matters more than the tooling. In our own delivery work, agents write a significant share of first-pass implementation. The investment that made that safe was not the agent configuration, which took days. It was the weeks spent beforehand on test infrastructure, environment provisioning, and merge policy. Teams that reverse the order see a throughput spike followed by a rise in change failure rate, and usually conclude the agents were the problem.
Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs and inadequate risk controls among the reasons. Weak controls are not a governance footnote in that list. They are the mechanism by which the costs escalate in the first place.
On an engagement with a US commercial lending company, the work that compressed the underwriting cycle and cut analyst time by 4x was not primarily model work. It was rebuilding the pipeline around the model so every automated decision was reproducible, testable, and reviewable by a human who had to defend it. The same discipline applies to agent-written code. An agent refactoring a legacy service is only useful if you can prove the refactor preserved behaviour, which means characterisation tests come before the agent, not after it.
Where this leaves engineering leaders in 2026
The 2026 Gartner CIO and Technology Executive Survey found that 17% of organisations have deployed AI agents, while more than 60% expect to within two years. Most of that gap closes over the next eighteen months. The teams that close it well will not be the ones that picked the right agent. Agent capabilities are converging quickly, and the enterprise coding agent market that Gartner sized at roughly $9.8 billion to $11 billion annualised in April 2026 is competitive enough that any tooling advantage is temporary.
The durable advantage sits in the delivery system around the agents. Review capacity, test signal, environment isolation, and change provenance are unglamorous investments that no vendor packages as a product. They are also the only things that decide whether agent output becomes shipped software or an expensive backlog of unread pull requests. Designing that layer is the work we do in every AI Transformation engagement, and it is the part clients underestimate most often.
If you are putting coding agents into your delivery pipeline and want the verification layer designed alongside them rather than after them, talk to our engineering team.
EPixelSoft is an AI-native software engineering company based in Noida, India. Since 2014, we have shipped 700+ production systems across FinTech, HealthTech, NGO operations, and SaaS.