Why remediation doesn't scale
The economics of correcting software have quietly inverted. Most enterprise AI adoption plans are still built on the old ones, and the bill arrives about eighteen months in.
Here is a scene playing out in a lot of enterprises right now. A team runs an agentic pilot. It goes well, the volume is real, the demo lands, the executive summary writes itself. Then someone points the existing static analysis and architecture tooling at the output, and it returns three hundred and forty findings.
What happens next is the entire argument. Not the findings. What happens next.
IThe old loop was surgical
For fifty years, software correction worked one way. Someone wrote code. Someone else, or a tool, found a problem. A person who understood the code made a small, targeted change, and the rest of the system stayed exactly where it was.
That last part is the load-bearing one, and it was so reliable that nobody named it. Correction was local. Fixing one thing did not move the other things. You could therefore run the loop as many times as you needed and be confident that each pass left you closer to correct than the last. Defect counts went down and stayed down. The whole apparatus of code review, QA and remediation backlogs assumes this property.
It assumes, in other words, that fixing converges.
IIThe new loop isn’t a loop
When an agent produces a block of forty thousand lines, the person reviewing it did not write it. Often nobody did, in any meaningful sense, there is no author to consult about intent, because the intent lived in a prompt and a context window that no longer exist.
So the natural correction is not to edit. It’s to re-run. Adjust the instruction, regenerate, re-check. Every team doing this arrives at the same reflex within about two weeks, and it feels efficient, because it is much faster than reading forty thousand lines.
Correction stopped being an edit and became a re-manufacture. That single change breaks the assumption everything downstream depends on.
Regeneration is not idempotent. Same specification, same model, same temperature, different output. Not wildly different, usually. Different enough. Structures move. Names change. A helper that three other blocks depended on gets inlined. The fix lands, and the neighbourhood around it shifts.
IIIWhy it stops converging
Think of it as a simple property rather than a metaphor. Each regeneration cycle has some probability of resolving the violation you targeted, and some probability of introducing violations you did not have before. As long as the second number is meaningfully above zero, and as long as the volume of code being regenerated scales with the volume being produced, the loop does not walk downhill. It oscillates.
You can run it forever and stay roughly where you started, a stable population of violations whose membership keeps changing. Every cycle produces a report showing progress against the specific findings you targeted, and every cycle quietly reseeds the pool.
This is why teams describe the same experience: the dashboard improves, the estate does not. Detection tools are excellent at telling you where you are. They are structurally incapable of telling you whether you are getting closer, because they measure the artefact, not the trajectory.
IVThe symptom is already measurable
There’s an early empirical signature of this. In a study of three hundred open-source projects, of which fifty were partially or entirely AI-generated, the most consistent anti-pattern was not a security flaw or a performance problem. It was avoidance of refactors, present in the great majority of AI-generated codebases.
Machines add. They very rarely restructure. Asked to fix something, a model will reliably produce new code that handles the case rather than reorganise existing code so the case cannot occur. Each correction is additive. Which means that under a regenerate-to-fix regime, the estate does not get simpler as it gets more correct. It gets larger and more correct-looking, which is a considerably worse place to be.
The same study found high rates of over-specification, defensive, duplicated logic written to satisfy an instruction rather than an architecture, and of recurring defects, the same class of bug reappearing across generations because nothing in the system remembers that it was already solved once.
V“Shift left” is not the answer people think it is
The standard response to all of this is to move detection earlier. Run the scanner in the IDE. Run it on every commit. Run it in the agent’s own inner loop.
This is worth doing and it does not solve the problem, because it changes when you detect without changing what happens after you detect. If the response to a finding is still regeneration, an earlier finding just means an earlier non-convergent cycle. You’ve made the oscillation higher-frequency. The population is still stable.
Shifting left is a scheduling change. What’s required is a categorical one.
VIConstraints have to be inputs, not criteria
The alternative is not more sophisticated detection. It’s making the architecture a production input, something the machine is held inside while it works, rather than something a checker compares the output against afterwards.
Concretely, that means the unit of production is bounded before anything is generated: which module, which interfaces, which patterns, which prior decisions apply and which are explicitly out of scope. The generator does not have the option to reach outside that boundary, so the class of violation you would otherwise be remediating cannot be expressed. You are not catching it. You are making it unrepresentable.
You cannot inspect your way to an architecture. You can only produce inside one.
This is not a novel idea in manufacturing, which is where the vocabulary comes from. No serious production line achieves tolerance by measuring finished parts and re-making the ones that miss. It achieves tolerance by constraining the machine, and then measures to confirm the constraint held and to detect when it stops holding. Inspection is how you learn the process drifted. It was never the mechanism of quality.
VIIWhat falls out of it
There’s a second-order effect worth naming, because it’s the part most teams don’t anticipate.
If you constrain at the point of production, you necessarily know, at that moment, without extra work, which requirement caused this block, which decision shaped it, which constraints it was held to, what was checked and who accepted it. Provenance is a by-product of producing this way. It costs nothing additional.
If you inspect instead, none of that exists. It has to be reconstructed later, by hand, usually by the least happy person in the organisation, usually in the four weeks before an audit, and usually to a standard everyone privately knows is a best-effort narrative rather than a record.
Which is why the teams that get production control right stop treating compliance as a project. The answer already exists before anyone asks the question.
VIIIWhat we are not saying
This is not an argument against static analysis, scanners, policy tools or code review. All of them remain necessary. Constraints fail. Envelopes are mis-specified. Models find edges nobody anticipated. You need detection precisely because prevention is never total, that’s what detection is for.
The argument is narrower and, we think, harder to dismiss: detection cannot be the primary control at machine volume. It was an adequate primary control when correction was surgical, because the loop converged. It stops being adequate the moment correction becomes regeneration. Keep the scanners. Stop asking them to do a job whose economics no longer work.
IXThe question worth asking
If you are evaluating an agentic delivery capability right now, a vendor’s, or your own team’s, the standard question is about pass rates. What percentage of generated code passes review, or clears the benchmark, or ships without intervention.
It’s the wrong question, because it measures the good case. Ask about the bad one instead:
- When a violation is found, what actually happens, an edit, or a regeneration?
- If it’s a regeneration, what stops the second attempt from breaking something the first attempt got right?
- How many cycles does a typical finding take to close, and is that number falling?
- Six months from now, will the estate be smaller and better understood, or larger and better documented?
A capability that can answer those is doing production control. One that can only quote a pass rate is doing inspection, and will discover the difference somewhere around month eighteen, which, going by the organisations that have already been there, is roughly when people stop saying that AI is accelerating their development and start saying they can no longer ship features because they don’t understand their own systems.