This article grew out of my WeAreDevelopers talk “A True Story About Speeding Up the Wrong Things.” The talk eventually turned into three articles. The first looked at an old pattern: systems become faster long before they become better at steering themselves. This second part stays with software engineering and looks at what happens when AI suddenly makes one of its most visible outputs much cheaper.
Estimated reading time: 8 min
The original 40-minute talk is available on WeAreDevelopers.
Software development has spent decades treating code as one of its expensive resources. That assumption shaped planning, tooling, hiring, estimates, and quite a bit of professional identity. AI changes the price.
A developer can now ask for an implementation, a refactoring, a test, a migration, or another approach and receive a plausible candidate within seconds. If the first attempt is wrong, generating another one costs very little.

This is genuinely useful, but it also exposes something that was easier to ignore while producing code required more effort: writing the implementation was never the whole engineering problem.
The difficult part increasingly sits around the code. A change has to belong in the system. An abstraction has to earn its place. An existing implementation may be ugly but sufficient. A requested feature may work and still not justify the additional surface area it creates. None of these questions are new, but their relative cost becomes more obvious once another implementation is almost free.
This also changes what productivity can look like. Generated code is visible and countable. Choosing not to build something leaves much less evidence of activity, although it may be the better engineering decision. The same is true when somebody avoids an unnecessary abstraction before it reaches the repository. Once production becomes cheap, volume becomes an increasingly poor proxy for progress.
If code becomes cheap and understanding stays expensive, systems do not become simple. They become easier to complicate. The factory got faster. The map did not get better.
The interesting problem with AI-generated code is not simply that it can be wrong. Developers have been quite capable of writing wrong code without help. Plausibility is more troublesome because language models are exceptionally good at giving unfinished reasoning a finished surface.
Naming can look sensible, explanations arrive neatly structured, comments appear in plausible places, and a proposed design may sound as though the alternatives have already been considered. Even uncertainty arrives organized.

Traditional software failures often provide something obvious to attack: the code does not compile, a test fails, an exception appears, or the behavior is visibly wrong. AI can produce something coherent enough to invite trust before that trust has been earned.
Fluent uncertainty is a dangerous user interface.
That changes the nature of review. The reviewer has to determine which assumptions entered the implementation, which constraints were available to the model, whether an abstraction solves a problem the system actually has, and whether a locally plausible solution fits the architecture around it. None of this becomes cheaper merely because the first draft appeared quickly.
In real projects, review time is finite and deadlines do not disappear when generation gets faster. An implementation that looks reasonable and already passes obvious checks therefore has a considerable advantage over one that visibly looks unfinished. “Good enough for now” can enter the repository easily, and repositories are very bad at remembering which parts were supposed to be temporary.
Agentic coding pushes the same problem further. The agent receives a goal, writes code, runs tests, inspects the failures, modifies the implementation, and repeats the process. It may call other tools or another agent along the way. Eventually, one of those attempts passes.
There is nothing unusual about feedback loops in software. Tests, static analysis, build systems, deployment checks, and monitoring all help us change something, observe the result, and adjustment.

The important question is what a successful loop actually establishes. A green test suite tells us that the implementation satisfied the checks encoded in that suite. That is useful evidence, but it does not provide a design rationale. It does not tell us whether the implementation fits the surrounding system, whether its assumptions are sound, whether the added complexity is justified, or whether the checks cover the part of the problem we actually care about.
An agent can therefore discard several failed approaches and eventually produce a passing one without arriving at a particularly good explanation for the final shape of the solution. “It eventually passed the tests” is not the same as “we know why this is the right solution.”
The conversational interface makes that distinction surprisingly easy to overlook. Every retry can arrive with a plausible explanation, which makes a search through alternatives feel more deliberate than it may actually be. The useful part may simply be that the system can explore possibilities cheaply and react quickly to feedback. There is nothing wrong with that, as long as we do not confuse successful exploration with understanding.
Agentic AI can turn software development into brute force with a conversational interface. The conversation makes the process easier to follow; it does not automatically turn the result into an engineering argument.
Assume the process works well. The agent produces a patch, reacts to failures, adjusts the implementation, and eventually reaches something acceptable. A developer reviews the result and merges it. At that point, the inexpensive part of the process is finished.
The code remains. Any abstractions, dependencies, assumptions, and relationships introduced by the patch become part of the system. Future work has to take them into account, and somebody eventually needs to understand enough of the new structure to change it safely.

AI can therefore reduce the cost of creating complexity without reducing the cost of carrying it. Complexity itself is not evidence of bad engineering. Seven dependencies may be entirely justified, and a new abstraction may solve a difficult problem elegantly.
What matters is whether the reason for that complexity exists and can be explained. A passing test can confirm observed behavior, but it cannot explain why a particular structure was necessary, why an additional dependency was worth introducing, or which constraint ruled out a simpler solution.
That distinction becomes more important as generation gets cheaper because costs can move rather than disappear. The implementation may take minutes instead of hours, while the resulting structure remains in review, maintenance, debugging, and every future attempt to understand that part of the system. Cheap code can still create expensive context.
There is another cost in AI-assisted workflows that I find more troubling than bad generated code. Bad code is familiar. We know how to recognize it, complain about it, replace it, or develop sufficient folklore around it that nobody touches the file on Fridays.
The disappearing reasoning path is different. Software engineering has always depended on traces left behind by other people. Mailing lists, issue trackers, forum threads, Stack Overflow answers, and old comments preserve much more than final solutions.

Their value often lies precisely in the mess: somebody tried the obvious approach and discovered why it failed; somebody misunderstood an API and was corrected; a workaround looked clever until another developer explained what it broke.
That accumulated path is part of engineering knowledge. The accepted answer alone is often less useful than the discussion around it.
AI-assisted work increasingly happens in places where that trail is easy to lose. An IDE assistant proposes one approach, an agent tries another, a test exposes an unexpected constraint, the implementation changes again, and eventually the repository receives a working patch. The sequence that produced it may never become part of anything the team can search or reuse.
The failed approaches disappear with it, along with the assumptions they disproved. Another developer or another agent can rediscover the same dead end tomorrow because the system retained the answer but not the lesson. Cheap iteration then becomes less impressive than it first appears: the organization may repeatedly spend computation and developer attention learning things it has already learned once.
If only the final patch survives, the workflow has generated activity, not necessarily knowledge. A system that repeatedly starts from zero can become very efficient at forgetting expensively.
The consequences become visible later, often when nothing is actually broken. Someone needs to modify a perfectly functional piece of code and discovers that nobody can explain why it has its current shape.
The implementation may have been generated by an agent, accepted by a developer, prompted by a vaguely specified ticket, and validated by a test suite that covered the expected behavior. Every step can have been locally reasonable.

Six months later, however, the reason for choosing this exact solution may be gone. A system that used to need a function may now contain seven dependencies, and “the tests were green” describes how the change was accepted without explaining why those dependencies were justified.
That is the point at which maintenance becomes archaeology.
Responsibility becomes harder to locate for the same reason. The important issue is not who physically typed the code. It is whether the engineering decision survived alongside the implementation. Cheap generation, automated retries, disappearing reasoning paths, and ordinary delivery pressure can produce changes that made perfect sense at the moment they entered the system while becoming increasingly difficult to justify later.
This is the bottleneck that moved. AI can produce possible answers at a rate that was unrealistic only a few years ago. Engineering still has to decide which of them belong in the system, understand the complexity they introduce, and retain enough context for those decisions to remain intelligible.
The next problem is therefore no longer how to generate more. It is how to build the layer around AI that can observe, challenge, remember, constrain, and eventually turn generated output into system state somebody can actually be responsible for.
That is where the final part begins.
AI can produce code and other technical artifacts at a fraction of the previous cost, while the surrounding engineering work does not shrink at the same rate. Deciding whether a change belongs in the system, understanding what it introduces, and maintaining the result remain expensive. Agentic workflows add another problem. They can reach working results through repeated attempts while preserving very little of the path that led there. The code survives more reliably than the reasoning.
© Copyright 2026 Cybercraft GmbH