Developers

Why AI Coding Agents Need System-Level Verification, Not Just Faster Pipelines

Anthropic and Linear report that AI agents have overwhelmed continuous integration systems, but making pipelines faster misses the real problem: they only test isolated repositories, not the distributed systems those changes must work with.

7 min read
Agents have made CI the bottleneck. Faster pipelines are the wrong fix.

When Anthropic's engineering team published their findings in September, they revealed a striking shift in their development velocity. Over a six-month span, continuous integration job volume surged 25-fold. The result: engineers now deploy roughly 8 times as much code each quarter compared to the 2021-to-2025 baseline. The company's response focused on test impact analysis, which runs only the tests that a particular change could touch.

That same month brought two related reports worth examining together. Linear disclosed that their test suite has nearly quadrupled since January, with AI agents now responsible for writing most tests. The company rebuilt its pipeline from end to end to manage the load. Meanwhile, Depot's CEO argued that CI itself is transforming, with the future lying in "giving agents a way to validate code and maintain trust as they work."

What makes these accounts significant is their source. Anthropic and Linear do not market CI tools; they are describing what happened inside their own systems. Yet while all three diagnoses identify a genuine problem, two of them point toward solutions that address the wrong layer.

How CI became the bottleneck

For two decades, continuous integration was calibrated to human productivity. A developer submitted a handful of pull requests weekly, and the pipeline executed after each one. A 20-minute run felt acceptable because the developer had already moved to the next task.

Agents disrupted this equation in two distinct ways. The first is sheer volume. When a single engineer orchestrates multiple agents in parallel, pull request counts multiply rather than increment. Anthropic's 25-fold increase is not unusual. Blacksmith, which operates CI runners, reports that job volume grows between 5 and 10 percent week over week.

The second is timing. Continuous integration executes after a pull request appears. An agent that generates code, opens a PR, and then waits 20 minutes for results has lost its working context by the time feedback arrives. Each failure requires a complete cycle. As one observer noted, "The agent is fast, and the loop around it is slow."

The industry responded predictably: make CI faster. Speedier runners, smarter test selection, expanded caches, pipelines that agents can invoke before commit. All of these help. All are necessary. Yet they all preserve a single unquestioned premise: that verification targets a repository.

What a green pipeline does not tell you

For a self-contained application, the repository is the system. Tests pass, and you have validated what matters.

For a cloud-native architecture, a repository represents one service among dozens. Tests in that repository exercise that single service while mocking all others. A change can sail through unit tests, pass CI at record speed, pass a sandbox built from the branch, and still fail the moment a real request crosses a service boundary.

In distributed systems, the damage occurs at the seams. A field renamed in one service's response that downstream consumers still expect. A tightened timeout in one service that triggers cascading retries elsewhere. A schema modification that works against test fixtures but locks a table in staging. A fresh endpoint that behaves correctly under test harness calls but incorrectly when the dependent service invokes it.

No CI pipeline, regardless of speed, detects any of this. It cannot, because it examines only a repository.

This is why the September reports represent symptoms rather than root causes. Continuous integration slowed because verification shifted from human judgment to automation without examining what the automated gate actually validates. Accelerating the gate does not change what it measures.

Research from DevOps Research and Assessment (DORA) uncovered a troubling pattern: organizations adopting AI at higher rates show gains in both delivery speed and delivery instability. As one analysis put it, "Faster code, same verification, more breakage. That is the gap."

The verification loop has to move

In February, Cursor articulated a principle that grows more urgent. Their agents operate inside cloud sandboxes, each with dedicated virtual machines, and more than 30 percent of merged pull requests now originate from agents working in this environment. Their reasoning: "Without the ability to use the software they are creating, agents hit a ceiling."

This instinct points in the right direction. Agents must execute code, not merely produce it. Every major coding agent now implements some variant. GitHub's Copilot cloud agent runs tests in ephemeral environments backed by GitHub Actions. Codex executes setup scripts and resumes cached containers. Devin initializes from environment blueprints. Greptile's TREX runs the branch and attaches logs and screenshots to pull requests.

Yet examine what each sandbox contains: the repository, the branch, and whatever the setup script installed. None include the other 39 services, the actual message queue, or a database shaped like production.

"So the loop closes, but it closes around the wrong thing." The agent verifies its change against a copy of its own code. Then continuous integration verifies it against that same copy, faster. Then it merges into staging, and only then does anything confirm whether it works alongside the rest of the system.

For distributed systems, pre-PR verification must target the system itself, not the repository. That is the fundamental shift. It is not quicker feedback on the same question. It is a different question, posed earlier.

Why this looks expensive and is not

The immediate concern is expense. If each agent requires the entire system for verification, and one engineer runs five agents, that means five staging environments per engineer. The cost appears prohibitive, and it should be.

The solution mirrors what the industry adopted for compute two decades ago. You do not allocate separate machines to every workload. You multiplex.

A single Kubernetes cluster can maintain one stable shared version of every service and host thousands of lightweight test environments above it. Each test environment deploys only the modified service. Requests routed to that environment pass through the changed service, while every other hop resolves to the shared stable versions. The modified service communicates with real dependencies, which remain unaware of any change.

A test environment costs roughly the price of one pod and starts in seconds. Fifty parallel agents share one stable environment instead of cloning it fifty times. This multiplexing makes system-level verification inside the agent loop economically feasible.

Agents need governed verification, not only environments

Inexpensive, rapid environments are insufficient alone. Agents also require a structured method for using them: send this request, capture that log, assert this contract held, report the result. Without guidance, each agent develops its own checks, and no two runs become comparable.

The effective model has platform teams define those steps once, as an approved sequence that exercises a change against the live system and documents what occurred. Agents invoke them through the skills and hooks that Claude Code, Cursor, and comparable tools already provide, so verification integrates into the loop rather than following it. Governance matters as much as the steps themselves, since platform teams must prevent agents from executing unsafe actions in shared clusters.

The output carries equal weight. A record showing which requests were sent, which services were touched, and which contracts held becomes an artifact downstream layers can consume. Review tools and merge gates can observe that a change was tested against live services before human review, transforming continuous integration from the first place cross-service failures emerge into a confirmation step.

Verification belongs inside the agent's loop

Continuous integration vendors will continue accelerating, and coding agents will improve at running code in sandboxes. Both trends benefit everyone, and neither bridges the gap between a repository and a system. Organizations that gain the advantage will stop asking how quickly a pipeline can confirm that a repository still passes its own tests. Instead, they will ask how early an agent can demonstrate that a change works with everything around it. That question finds its answer in one location: inside the agent's loop, against the real system, before the pull request.

Source: The New Stack · Reporting supplemented by The Silicon Ledger staff.