AI coding systems can produce more changes in an hour than an engineer can inspect in a day. This is often presented as proof that the engineer has become the slow part of software development. It proves something more limited. A person cannot remain the control system for a production process that runs at machine speed. If the engineer must follow every edit, compare every attempt, and recover every decision from an agent transcript, human judgment has been placed too late in the work.

The engineer has more leverage before the code is produced. A few decisions can define what the system is allowed to do, which data it may use, what must never happen, and what evidence is required before a result is accepted. Those decisions can govern many automated attempts without requiring a person to watch each one. Human attention is still limited, but the limit now shapes the system instead of creating a queue of code waiting to be reviewed.

The bottleneck is created after generation

Consider a café that asks a group of coding agents to build an online ordering system. The request sounds simple. Customers should choose a shop, select a drink, pay, and receive a pickup time. In practice, the agents must also deal with menus, sizes, milk options, sold-out ingredients, prices, tax, card payments, refunds, duplicate requests, receipts, and the queue seen by the barista.

If the agents are free to solve all of this in one shared codebase, they can produce a great deal of software overnight. One may calculate prices in the checkout page. Another may repeat the calculation in the receipt service. A third may create a kitchen ticket before payment is confirmed because it makes one test pass more quickly. By morning, the system may appear to work, but the engineer must discover which business rules the agents invented, where the same rule was implemented twice, and which temporary fix has become a dependency.

The engineer now looks like the bottleneck because production continued without a boundary and control was postponed until the end. The problem is not that the engineer reads too slowly. The problem is that the workflow asks one person to reconstruct the meaning of a system after many automated processes have already changed it.

Decide what a coffee order means first

The same project can begin with a small set of rules that everyone can understand. The price shown at checkout must come from the current menu for the selected shop. The app may not sell a drink marked as unavailable. A failed card payment must not send a ticket to the barista. If the payment provider sends the same confirmation twice, the kitchen must still receive only one order. The system may store a payment reference, but it may not store card details. An agent may change how these rules are implemented, but it may not change the rules themselves.

These are architectural decisions because they define the meaning and limits of the system. They are not instructions about variable names or code style. They answer practical questions. Can a customer be charged for a drink the shop cannot make? Can one payment produce two cappuccinos? Can a checkout page quietly apply a price that differs from the menu? Can payment data leak into the order database? The engineer, product owner, and café operator decide these matters before an agent starts searching for code.

Once those decisions are expressed in the interfaces and checks around the ordering capability, agents can work quickly inside them. They may try different database layouts, queue mechanisms, retry strategies, or user interfaces. They remain free to find an implementation. They are not free to redefine what a valid order is simply because a different rule would make the implementation easier.

Let the rules supervise the attempts

A small execution environment can exercise the ordering system before any version is accepted. It can place a flat white with oat milk and verify the displayed price. It can request a sold-out cold brew and verify that no payment begins. It can simulate a declined card and check that no kitchen ticket appears. It can send the same successful payment confirmation twice and check that the barista sees one drink, not two. It can also verify that the order record contains a payment reference rather than raw card information. This small environment is the harness: a place where an agent can work freely without being able to alter the rest of the system.

An agent may run through these cases dozens of times while it changes the implementation. The engineer does not need to watch the attempts. A version that changes a menu price, creates an order after a failed payment, writes card data into the wrong place, or produces a duplicate ticket is rejected before it can become part of the wider system.

This is more than testing added after the code is written. The boundary and its checks exist first. They define the space in which automated work is allowed to happen. If an agent cannot complete the task without changing one of those rules, it must return the decision to a person. It cannot solve a local difficulty by quietly giving itself more authority.

Stop temporary fixes from becoming the design

Without such boundaries, automated work can accumulate in plausible but damaging ways. Suppose an agent finds that payment confirmations sometimes arrive slowly. To make the interface feel faster, it sends the order to the kitchen before payment is confirmed. A later agent discovers unpaid drinks in the queue and adds a cancellation job. Another adds refund logic for drinks that were already prepared. Each patch addresses the system it inherited, but none returns to the original question of when an order should exist.

The same drift can happen with something as ordinary as an oat milk surcharge. If the price is copied into the menu page, checkout, receipt generator, and kitchen display, changing it in one place may leave three different totals in the system. The codebase begins to record the history of local fixes rather than one deliberate design.

A longer agent transcript does not solve this problem. More history can show how the patches appeared, but it does not make them correct. Rules that matter must live in stable interfaces, permissions, and executable checks. A new agent should be able to learn how an order works from the accepted system, not from a conversation explaining why five earlier attempts failed.

Keep the accepted result, not the entire search

During development, an agent may try many approaches. It may replace a queue, change a data model, or discard an unsuccessful payment flow. That search can be messy and temporary. Once one version satisfies the fixed rules, the system keeps the accepted code, its interface, its permitted dependencies, and the checks it passed. The working conversation may remain available for audit, but it does not become the foundation of the next task.

The next agent begins from a clear menu service, order service, payment boundary, and kitchen queue. It does not need to inherit every abandoned idea that led to them. Generation may remain probabilistic, but the accepted artifact is fixed, versioned, and testable again. This is how automated search can remain fast without allowing its uncertainty to spread through the whole codebase.

What still belongs to the engineer

Some questions cannot be settled by a test because they are decisions about the business. What happens when oat milk runs out after a customer has paid? May the shop substitute another milk, or must it ask first? Can a store manager give away a drink after a long delay? How close to closing time should the app stop accepting orders? Should a refund be automatic when the kitchen cannot prepare an item?

These questions belong to people because they involve service, risk, cost, and responsibility. An agent can show the available options and implement the chosen rule. It should not make the rule through an incidental coding choice. When a decision returns to the engineer, it should look like a real decision, not a request to read thousands of changed lines in order to discover that the question exists.

The engineer may still inspect source code when security, performance, or unusual risk requires it. The difference is that source inspection is no longer the only thing standing between uncontrolled generation and release. The main control has already been applied through the boundary that the implementation must respect.

Finite attention can shape the system

A person should be able to explain the ordering system in a few plain statements. The menu publishes what each shop can sell and at what price. The order service accepts only available items. The payment service confirms whether money was received. The kitchen queue receives one ticket for one successful order. If those responsibilities cannot be stated clearly, the system has become too broad or its boundaries have become unclear.

This is where human limitation becomes useful. Since an engineer cannot govern an indefinitely expanding body of agent-generated code, automated work must be divided into parts that remain understandable and independently checkable. The system is not kept small because people dislike complexity. It is kept bounded because complexity that no one can explain cannot be responsibly accepted.

AI can multiply the number of implementations a team can try. Human judgment determines what those implementations are allowed to mean. The engineer is a bottleneck only when placed at the end of an unbounded production stream. Placed before generation, human attention becomes the force that keeps automated development coherent, limited, and under deliberate control.

Research basis

David L. Parnas, 1972. On the Criteria to Be Used in Decomposing Systems into Modules. A foundational account of modularity as a way to separate design decisions and keep systems comprehensible as they change.

Lisanne Bainbridge, 1983. Ironies of Automation. An analysis of how automation can leave people responsible for passive monitoring and difficult exceptions.

Raja Parasuraman, Thomas B. Sheridan, and Christopher D. Wickens, 2000. A Model for Types and Levels of Human Interaction with Automation. A framework for allocating functions between people and automated systems rather than treating maximum autonomy as the only goal.

Nelson F. Liu and colleagues, 2024. Lost in the Middle: How Language Models Use Long Contexts. Evidence that access to a long input does not guarantee reliable use of all information within it.

Carlos E. Jimenez and colleagues, 2024. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? A benchmark built around repository navigation, coordinated edits, and executable validation.

John Yang and colleagues, 2024. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. A study showing that the tools and interface given to a coding agent affect its behavior and results.

Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, and Lingming Zhang, 2024. Agentless: Demystifying LLM-Based Software Engineering Agents. A constrained workflow based on locating a problem, producing a repair, and validating it.