AI is making software output cheaper. That makes judgment more valuable, not less.

“Factory” has become the word of the moment in AI and software development. Cursor says developers are beginning to build “the factory that creates their software,” with fleets of agents operating as teammates. Warp is more explicit, arguing that engineers are becoming “factory engineers” responsible for improving the machine that builds products. Microsoft now has an Agent Factory positioned around scaling AI agents across the enterprise.

There is a lot to like in this direction. I believe every serious enterprise technology organization should be exploring it. Agents that can triage issues, create specifications, implement changes, review code, verify behavior, monitor production, and learn from failures will fundamentally change the economics of software development. Warp is already describing semi-automated SDLC flows, while Cursor says more than one-third of the pull requests it merges are created by agents running independently in the cloud.

But I keep coming back to an uncomfortable question.

Is the factory the purpose? Is output even the purpose?

Because a factory that efficiently produces things nobody needs is not a great factory. It is a highly optimized waste machine.

We are copying the conveyor belt, not the factory

The technology industry has developed a strangely shallow view of manufacturing. We talk about factories as if their defining characteristic is repetition: inputs enter, products emerge, unit cost falls, executives applaud.

That is not how the best factories work. Toyota’s own description of the Toyota Production System emphasizes eliminating waste, stopping when abnormalities are detected, maintaining quality, and making only what is needed, when it is needed, in the amount needed. The production line is only one part of the system. Demand, quality, flow, feedback, and economics determine whether the line should move at all.

That distinction matters enormously for AI software development. A mature factory does not celebrate producing twice as much inventory that customers do not want. Yet many AI engineering strategies are drifting dangerously close to doing exactly that with software.

More pull requests. More commits. More features. More autonomous tasks. More code per engineer. Lower cost per change.

Those can be useful operating metrics, but they are not the reason the company exists.

Software does not begin with requirements. It begins with uncertainty.

This is where the manufacturing metaphor starts to break.

In traditional manufacturing, you generally know what a finished unit is before production begins. In product development, the “raw material” is often a hypothesis. We believe a customer has a problem. We believe a new workflow will change behavior. We believe an automation will reduce cost. We believe a feature will improve retention.

Then reality gets a vote.

Microsoft Research’s experimentation guidance begins with a falsifiable hypothesis and explicit success and guardrail metrics. That is not process bureaucracy. It reflects a basic truth about product development: shipping something does not prove it was worth building.

This is why I worry about the emerging AI factory conversation. We are making the answer cheaper before improving the question.

In fact, a 2025 ACM Queue paper on the commercial value of software makes the point bluntly: software has value when it is used, and sustainable organizations must balance development cost with product success, innovation, operational efficiency, and risk. Code production is one input into that system, not the outcome of it.

That should change how executives think about AI engineering.

The bottleneck is moving upstream

For decades, engineering capacity was expensive enough to act as a crude filter. A questionable idea had to compete for scarce developers, roadmap space, architecture attention, and release capacity. That filtering process was often political and imperfect, but scarcity forced choices.

AI weakens that constraint.

When implementation becomes dramatically cheaper, organizations can suddenly afford to build far more questionable ideas. The backlog that once represented years of work can become months of agent execution. Every executive request becomes technically feasible. Every product manager can generate another experiment. Every engineer can launch another agent.

This sounds like abundance. Without a corresponding improvement in judgment, it becomes congestion.

DORA’s research is an early warning. Its analysis found that a 25 percent increase in AI adoption was associated with a 1.5 percent decrease in delivery throughput and a 7.2 percent decrease in delivery stability, with faster code generation contributing to larger batches that were harder to review and more destabilizing. DORA separately describes a “verification tax,” noting that 30 percent of developers reported little to no trust in AI-generated code.

The lesson is not that AI coding fails. The evidence is moving too quickly for that simplistic conclusion. METR’s early 2025 controlled study found experienced open-source developers were slower with then-current AI tools, while its 2026 follow-up found signs that newer tools likely improve speed but also concluded that selection effects made the magnitude difficult to measure reliably. More interestingly, METR’s 2026 work explicitly separates speed from value, because doing more tasks faster does not necessarily mean producing proportionally more valuable work.

That distinction may be the most important metric conversation technology executives have this year.

The scarce resource is becoming judgment

Warp’s recent factory-engineering memo is worth taking seriously because it articulates the new model with unusual clarity. It proposes measuring the percentage of changes shipped automatically and the cost of completing them, while acknowledging that shipped product is an imperfect measure of value.

I agree with the direction and disagree with where the measurement can lead.

The percentage of autonomous changes is an excellent measure of factory capability. It is a dangerous measure of factory purpose.

A team that autonomously ships 80 percent of its changes may be technologically extraordinary and commercially irrelevant. A team that ships fewer changes may be transforming customer retention, opening a new market, eliminating operational cost, or materially reducing risk.

The next generation of technology leadership must therefore connect two systems that organizations have historically managed separately.

The first is the production system: agents, tools, harnesses, context, evaluation, CI/CD, observability, security, and cost.

The second is the value system: customer problems, strategy, discovery, commercial hypotheses, adoption, experimentation, quality, risk, and measurable outcomes.

The competitive advantage will not come from building the best AI conveyor belt. Those capabilities will diffuse rapidly through Cursor, Warp, GitHub, Microsoft, open-source frameworks, and whatever arrives next. The advantage will come from building a closed loop between sensing demand, choosing where to invest, producing safely, measuring impact, and feeding that learning back into the next decision.

That is a real factory.

Build the factory. Just remember why.

I am bullish on software factories. I expect agentic development to become a core enterprise capability, and I believe engineering leaders who dismiss it will find themselves structurally disadvantaged.

But the executive question is not, “How much more software can we produce?”

It is, “How much more valuable change can we create, how quickly can we learn whether it worked, and how cheaply can we stop when it did not?”

That is a very different operating model. It places product management, engineering, design, data, finance, and AI inside the same feedback system. It treats a requirement as a hypothesis, a release as an experiment, adoption as evidence, quality as a constraint, and business value as the final measure.

The great irony is that the technology industry may need to study factories more seriously. Toyota did not become Toyota by maximizing the number of cars moving down a conveyor belt. Its production philosophy is built around demand, quality, waste reduction, flow, and the ability to stop when something is wrong.

As software production approaches abundance, that lesson becomes more important, not less.

Build the factory. Automate it aggressively. Measure its economics. But never confuse the machine with the mission, because producing the wrong thing faster is not transformation. It is just waste at machine speed.