+49 (0)5583 2829973 info(at)coders(dot)care
Shadow

A true story about speeding up the wrong things

This is the final part of a series based on my WeAreDevelopers talk A True Story About Speeding Up the Wrong Things. The first looked at speed without steerability; the second at software engineering once code became cheap. But the architectural problem does not stop at code. Models now call tools and change other systems as well. So the question is broader: what has to exist around AI before its output or actions become state we rely on?

Estimated reading time: 8 min

The original 40-minute talk is available on WeAreDevelopers.

Generation, Then Commitment

Somewhere, a proposal has to become state.

Suppose an agent is asked to handle a customer complaint. It reads account data, drafts a reply and may update the ticket or trigger a refund. A coding agent changing a repository is only one version of the same pattern. In both cases AI can produce something ready to use while the surrounding system still decides what may happen.

AI tooling compresses these steps. Results look finished and agents can call tools immediately, so suggestion and effect may be separated by one click or none at all. The distinction still matters.

I find it useful to call that transition commitment. Before commitment, the model can explore. It may draft alternatives, query other systems, test a patch or discard its first assumption. None of that necessarily needs to alter accepted state.

At some point something changes outside the model: a reply is sent, data is written, a payment or deployment is triggered, or a decision becomes binding. From then on, the surrounding system has consequences to own.

Digital systems already use boundaries like this. A database transaction can do work before it commits, and a CMS can hold a draft without publishing it. AI-assisted systems need an equally visible boundary whether the result is code, content or an action elsewhere. The model can generate. The system has to govern what generation is allowed to become. Once models can also act, the same layer has to govern what they are allowed to change.

Observation Before Judgment

The result is only the end of the story.

If we later want to understand why an AI system sent a message, changed a record or recommended an action, the visible result is not enough. We may need to know which model produced it and which version was running.

The context matters, as do the tools the model could call and the data it could access. It can matter who started the task and which intermediate results affected the outcome. Some of this looks like logging, but logs are not the point.

What we need is enough provenance to reconstruct why an outcome exists and on what basis we trusted or executed it. A sent customer response, for example, tells us very little about the source data, rules or assumptions behind it.

The same problem exists with a patch, a generated report or an automated account action. Validation gets weak when the thing being checked is detached from the process that produced it.

A system that accelerates work therefore also needs to observe the work it accelerates. Otherwise we get more activity while progressively losing the ability to explain where the resulting state came from.

Routing and Limits

Useful decisions start before the answer exists.

There is another reason not to wait for the final output before applying control. Different tasks may reasonably go to different models. Some can use a small local model, others justify a much more capable one, and some should reach a human before an agent is allowed to do anything interesting.

Access works the same way. An agent preparing a summary may only need to read data. Customer service might update a ticket but require approval for a refund. A coding agent may inspect a repository without being allowed to merge.

A task involving customer data has a different risk profile from one using public documentation. Publishing externally is different from drafting internally. Those choices belong in the architecture. The question is not merely whether an agent can call a tool.

We also care what kind of action it is taking, which resources it may use and whether approval is required. Some effects cannot be undone by reviewing the final answer. If confidential data has already been sent to the wrong place, checking the reasoning afterwards is a little late.

That is why routing, permissions, policy and risk classification belong around the work while it happens.

The Commit Point

Acceptance should leave a trace.

Eventually something has to cross the boundary from candidate to accepted state. A draft may be published, a recommendation accepted, a tool call executed or a code change merged. Low-risk cases may pass automatically. Others may require validation, security checks or explicit approval.

What matters architecturally is that the line exists. A generated result does not become trustworthy because it looks plausible, and successful execution does not by itself make an action legitimate.

There should be a record of what was accepted or executed, under which conditions and through which checks. That gives responsibility somewhere to attach.

If something later causes trouble, "the AI did it" is not a useful explanation. The interesting questions are what was allowed to happen, which validation ran, which policy applied and who or what authorized the transition. The model produced a candidate or requested an action. The surrounding system let it become real.

That also makes agent autonomy less mysterious. An agent can have considerable freedom to explore, read and compare without receiving the same freedom to alter state outside itself. We already know how to work with boundaries like that elsewhere in computing.

Memory Beyond the Session

Keeping the answer is not enough.

Our agent may have learned something useful before it reached the outcome that was finally accepted.

Perhaps it discovered that a policy exception applies, that a data source is stale, or that an apparently valid technical approach does not work. Some of those findings may matter more later than the final answer.

Once the response is sent, the action completed or the patch committed, most current workflows throw those discoveries away.

The next person can repeat the same failed experiment. So can the next agent. At scale, we can spend a surprising amount of compute rediscovering things the organization technically learned yesterday. Preserving every prompt and token would not fix that. A raw transcript is history, not necessarily knowledge.

The useful part is whatever survived verification: a disproved assumption, a failed approach, a constraint that mattered, or a solution tested under known conditions. Those can become reusable knowledge instead of disappearing with the session. Software development gives us an obvious example. Stack Overflow, issue trackers and public discussions left searchable traces of the route from confusion to understanding. Private AI sessions can reverse that process by leaving only the final result behind.

A replacement for that shared memory does not have to be another central website. I can imagine a federated mesh of knowledge nodes, with organizations or projects deciding what they trust, publish, import or keep internal. A validated observation in one place does not automatically become universal truth elsewhere. The talk only points in that direction; it does not contain a finished federation protocol.

Useful experience needs a durable form if agents are going to accumulate knowledge rather than merely consume it.

Topology Matters

Interfaces do not remove relationships.

The same problem appears once we connect agents to tools, data sources and services. MCP-style interfaces are useful.

Standardizing how a model can discover and call capabilities solves real problems. It does not tell us how many relationships we should create, which paths context may take or who is responsible for governing them.

If every agent talks directly to every relevant tool, service and other agent, the number of relationships grows quickly.

Context moves through different paths. Permissions are enforced in different places. Policies drift. Traces are split across systems. A common protocol can make those connections easier to implement while leaving the underlying control problem exactly where it was.

I would rather have fewer explicit relationships and put the important control functions at boundaries we can reason about. A governed service layer or node can handle context, access, policy, routing, traces, validation and commitment without requiring every participant to reinvent those decisions. The principle is the same whether the agent is working on code or operating another system.

That still leaves an awkward question.

Controlling the Control Layer

The steering machinery also needs an owner.

Once this layer handles permissions, routing, knowledge, validation and commitment across several systems, it becomes a fairly powerful piece of infrastructure. Putting all of that into something opaque would be a peculiar way to regain control.

An organization should be able to inspect the rules that determine what its agents may see, call, change or publish. It should be possible to understand why an output was accepted or an action allowed, export accumulated knowledge and replace components without losing the ability to operate the system.

That is why I think the basic steering mechanisms need to remain inspectable and replaceable. It does not follow that every service around them must be open source. Hosting, operations, enterprise integrations and specialized extensions can all be commercial products. The narrower requirement is that the parts on which actual control depends cannot themselves become inaccessible machinery.

Otherwise we have built a better black box around the first black box. There is also a fairly mundane efficiency argument. Agents that forget what previous agents discovered will repeat work. Failed approaches cost tokens, compute, time and energy every time they are rediscovered. Memory can reduce the amount of work the system needs to do at all.

I am not looking at this solely as an interesting architecture problem. Systems in this general direction are very close to what I want to build. Details are still missing, particularly around federation, trust and verification. Different domains will need different commitment rules. But the system boundary is already visible.

AI has made generation cheap enough that producing another plausible result is increasingly the easy part. Agents can also act on those results now. The question is no longer only what a model can produce, but what it may be allowed to affect in the systems around it.

That requires observation, memory and control around the model, with a recognizable point at which exploration becomes something other people and systems are expected to rely on.

In the first article, speed was not direction. This is the layer that has to provide the steering.

AI systems no longer just produce answers. Through agents and tools they can retrieve data, call services and cause changes in other systems. A serious AI architecture therefore needs a steering layer between the model and the systems it can affect. That layer controls access and routing, preserves provenance, validates results and makes explicit when a proposal or action becomes accepted state. That boundary is commitment: before it, models can explore and fail; after it, the surrounding system becomes responsible for what was accepted or allowed to happen. The steering layer itself must remain inspectable and replaceable, otherwise the control problem has merely moved into another black box.

Jo Hasenau