Coada / writing

Stop Babysitting Pull Requests

A new report says developers are drowning in AI code review.

The industry worked that out months ago.

A large engineering services firm released a developer survey this week. According to their numbers, two out of three developers spend more time reviewing AI-generated code than they did a year ago. Developers say AI saves them 13 hours a week on coding, and those hours are going into reviewing, debugging, and learning new tools.

The report is getting a lot of traffic, and most of it treats the finding as new.

It isn't new. In January, the industry was already calling this a verification bottleneck. In February, OpenAI published how its own teams handle code reviews when agents write code. By spring, several surveys showed developers spending more hours reviewing code than writing it. In August, Anthropic released a full playbook on the subject. A September report that presents the review bottleneck as a discovery is about nine months behind the conversation.

The uncertainty phase is over.

Two years ago, helping a company become AI-forward was a reasonable service to sell. Nobody knew what worked, and exploration had value.

We're past that now. The patterns are published, the failure modes are documented, and the companies building the models have explained in detail how they build with them.

That raises the bar for anyone selling engineering services. Clients paying for talent and guidance expect a partner who already knows where the market is heading. When a firm reports last winter's problem as this fall's insight, the client ends up paying for the delay.

AI makes developers faster, not better.

The report leaves out the most important explanation for the number.

An engineer who doesn't know how to build software for production won't learn it from an AI assistant. The assistant writes code quickly, and any gap in the engineer's fundamentals shows up in that code over and over. Those gaps are easier to see today than at any point I can remember.

Most of the review queue comes from this. Nobody specified the work before it started, so the agent made the decisions, and the diff is the first place anyone sees them. Then the engineer who couldn't describe the work up front is asked to judge it afterward.

The pull request gets blamed, but the missing fundamentals are the real cause. Relabeling that review time as higher-value work doesn't fix anything.

Where the leading teams have moved

Teams that get real speed from AI decide what correct looks like before the build, and they stop relying on pull requests to catch it.

They design the product's domains first and get the business to agree on how it should behave. For each feature, they write down what it should produce before an agent touches it. Automated checks run against those decisions, and a failed check stops the release.

People still review, but they review the decisions once up front, rather than reading every diff that follows.

This approach asks for more planning and stronger leadership at the start. The cost is paid once, and the speed comes back on every feature after that.

Service firms should hold a higher standard.

Companies are spending serious money to catch up with AI, and the firms they hire should be moving at least as fast as the market and publishing a months-old problem as fresh research points clients toward more reviewers and longer queues, which is the wrong investment.

Engineering leaders should be asking why their teams are deciding what's correct after the code already exists.

Where has your team landed? I'd like to hear whether review is still your bottleneck or whether you've moved that judgment earlier in the process.

If you want to see how AI-forward companies are handling these problems today, reach out. I'd love to talk about how I can help you modernize and catch up with where the AI market already is.

Filed under ai agentsAI