AI Moves the Bottleneck to Review

July 23, 2026

A vast river of artifacts narrowing through a basalt gate beside a tiny reviewer

I can now ask several agents to work in parallel and come back to a small mountain of output: pull requests, research notes, interface variants, tests, launch copy. The production feels almost free.

Then I have to look at it.

That second step is where the fantasy of unlimited productivity breaks. I can generate ten times as much code, but I cannot understand ten times as many architectural decisions. I can produce a hundred interface variants, but I cannot develop a hundred times more taste. I can draft a month of content in an afternoon, but my audience does not gain another month of attention.

AI expands our capacity to produce much faster than it expands our capacity to judge. The bottleneck does not disappear. It moves from generation to review.

I do not mean that humans will keep reading every line forever. Human-in-the-loop will shrink dramatically. But the shape of that shrinkage depends on who—or what—consumes the result. Agent-facing systems can chase almost pure efficiency. Human-facing products still have to answer for experience.

That split may define more of the AI economy than the models themselves.

Production is not throughput

It helps to separate three things we currently collapse into “output”:

  • Generated output: everything a model produces
  • Reviewed output: everything someone or something has checked
  • Trusted output: everything we are willing to ship and take responsibility for

Only the third category creates durable value.

The distinction is easy to miss because generated output is so visible. We count tokens, pull requests, prototypes, campaigns, and completed agent tasks. The numbers rise, so productivity appears to rise with them.

But a pile of plausible artifacts is inventory, not throughput. If nobody can tell which artifacts are correct, coherent, safe, or worth using, generation has merely moved work downstream.

This is one reason AI can feel faster before it makes a system faster. In a 2025 randomized trial, METR found that experienced open-source developers using then-current AI tools took longer on the studied tasks, even while they believed AI had sped them up. That result is a snapshot, not a law about future models. But the gap between perceived production speed and measured end-to-end throughput is the important part.

The unit that matters is not “things made.” It is things we trust enough to use.

Review is bigger than code review

When I say review, I do not mean a senior engineer reading every diff line by line.

Review happens at different scopes:

  • An engineer inspects the implementation.
  • A tech lead checks the architecture and failure modes.
  • A designer walks through the interaction.
  • A product manager accepts the behavior against the intent.
  • A team watches metrics and support tickets after release.

As agents become more capable, review moves upward through these layers. First I stop writing every line. Then I stop reading every line and rely on tests. Later I may stop looking at individual changes and inspect only system behavior, exceptions, and outcomes.

The human loop gets thinner, but it does not necessarily vanish. It migrates from implementation to intent, behavior, and consequences.

This is the progression I expect:

human-in-the-loop → human-on-the-loop → human-at-the-boundary

First, a person approves each action. Then a person supervises an autonomous process and handles exceptions. Finally, a person defines the objective, the risk budget, and the conditions under which the system must stop.

We review fewer artifacts, but take responsibility for larger systems.

Validation capacity is the real multiplier

The obvious response is that review can also be automated. I agree. That is how AI productivity will keep compounding.

Tests, simulations, type systems, formal constraints, independent agents, observability, staged rollouts, and automatic rollback all expand validation bandwidth. They let one person supervise much more production without examining every artifact.

The next major productivity gains may come less from generating faster and more from making results cheaper to trust. A coding agent backed by a strong test suite can operate far more independently than one working in a repository where “looks right” is the only acceptance criterion.

This is why evals, telemetry, and rollback are not secondary infrastructure. They are the machinery that converts abundant generation into reliable throughput. Even formal risk frameworks such as NIST’s AI RMF organize the problem around governing, mapping, measuring, and managing—not simply producing.

But automated validation has a boundary. It works best when success can be specified from outside the task. A parser can be checked against a schema. A migration can be checked against invariants. A trading strategy can be tested against a risk limit.

“Is this delightful?” is harder.

A shallow-focus row of analog gauges, indicator lights, and controls on an industrial panel

Every output has a proximate consumer

I think every product or service can be classified by its proximate consumer: the next entity that has to use the output is either a human or another agent.

The word proximate matters. Almost every economic system is ultimately anchored to human wants. An agent may consume an API response, pass its result to another agent, and trigger ten more automated steps, but somewhere the chain began with a delegated human objective.

Still, the immediate consumer changes how the product should be built.

If another agent consumes the output, it does not care whether the interface feels elegant. It does not need a dashboard, a charming empty state, or an explanation written for a general audience. It needs structured inputs, predictable outputs, low latency, low cost, and machine-verifiable success.

If a human consumes the output, functional correctness is only the floor. The product also has to be legible, trustworthy, responsive, and sometimes beautiful. A game can be bug-free and boring. A video can satisfy every requirement and still not be worth watching. A medical portal can return the correct result while making a patient feel lost.

The same model may power both worlds, but the optimization regimes are different.

Agent-native systems will remove the loop

For agent-native work, I expect humans to disappear from most intermediate steps.

This is already visible in the shift from chat toward longer agentic tasks. Anthropic’s Economic Index has tracked the rise of directive automation, where people delegate complete tasks with less back-and-forth. As reliability improves, more software will be built for agents to call rather than for humans to operate.

Many interfaces will collapse into protocols. Reports that nobody needed to read will become structured state. Internal dashboards may become queries an agent runs only when an anomaly occurs. Documentation written solely to help one machine operate another may be generated on demand—or disappear into contracts and tests.

In this world, human review is often waste.

If failures are contained, automatically detectable, and reversible, requiring a person to approve every step lowers quality by adding latency and reducing scale. The right goal is not to preserve human involvement for sentimental reasons. It is to build enough validation that involvement is unnecessary.

Agent-native systems will optimize for:

  • efficiency
  • scale
  • reliability
  • reversibility
  • machine-verifiable success

Their ideal output may be completely inscrutable to me. That is fine if I can trust the boundary around it.

Human-native products cannot escape experience

The other world is made for people.

This includes obvious entertainment products: games, films, music, short video, social media. Their value exists in the experience itself. Producing more of them does not create more human attention. It creates more competition for the same attention.

But “human-facing” is broader than entertainment. Education, healthcare, travel, food, companionship, public services, and physical spaces all contain qualities that cannot be reduced to task completion. People experience products; they do not merely execute specifications.

This is where review becomes taste, empathy, and responsibility.

An agent can generate a thousand game levels. A human still has to want to play one. It can make a personalized film for every viewer. The viewer still decides whether the film moved them. It can optimize a learning path, but the student has to feel capable of continuing.

Human judgment may be assisted, sampled, or inferred from behavior. It will not always appear as a person clicking an approval button. But when the output is an experience, the human response is part of the definition of quality.

That is why unlimited generation does not imply unlimited value. In human-native markets, attention is finite and preference is the final eval.

Museum visitors sitting and standing in a dark gallery while viewing a large framed painting

Humans will move from production to preference

There is a tempting endpoint to this argument: machines do the useful work, while humans create and consume entertainment.

I think that is directionally right but too narrow.

Humans may spend far less time producing what is instrumentally necessary. Agents will write the glue code, reconcile the accounts, schedule the logistics, negotiate with other systems, and maintain much of the machinery behind daily life.

But what remains is not only entertainment. Humans also choose goals, form relationships, contest values, build identity, and participate in activities whose value comes from doing them rather than optimizing their output.

I may still play piano even if an agent can compose better music. I may still cook when a machine can make a more consistent meal. I may still build something with friends when an agent could finish it alone. Participation is itself a kind of consumption.

The long-term shift is not from work to entertainment. It is from production to preference: from making what is required to deciding what is worth making, what is worth experiencing, and what kind of systems we want around us.

Machines can optimize inside a goal. Humans still live at the layer where goals acquire meaning.

Two economies, one boundary

I expect the AI economy to separate into two large domains.

The agent-native domain will become increasingly invisible, autonomous, and efficient. Human-in-the-loop will approach zero across long chains of machine-to-machine production. The winners will remove latency, standardize interfaces, contain failures, and turn judgment into executable constraints.

The human-native domain will become increasingly abundant but remain brutally selective. The winners will understand attention, trust, identity, emotion, and lived experience. Their constraint will not be the ability to produce another option. It will be the ability to produce an option that a person chooses.

Many products will sit across both domains. A travel agent may negotiate flights, hotels, and schedules entirely through machine-facing systems, then present three meaningful choices to a person. A game studio may automate most asset production while humans playtest the experience. A healthcare agent may handle administrative work autonomously while keeping a clinician at the boundary of consequential decisions.

The most important product decision will be knowing where the boundary belongs.

Put it too early, and humans become a bottleneck for work machines can safely verify. Put it too late, and we optimize away the context, experience, or accountability that made the product valuable.

The future is not human-in-the-loop everywhere. It is human judgment placed deliberately at the point where machine-verifiable correctness ends and human value begins.

Closing thought

AI can make production nearly unlimited, but it cannot make attention, judgment, or responsibility unlimited. Machines may produce most of the world’s outputs; humans will still decide which outputs are worth producing—and which experiences are worth having.