Most technology organizations are trying to move into agentic development with the wrong operating model. They are buying tools, running pilots, asking engineers to use AI more often, and hoping productivity turns into transformation. That is not a strategy. It is procurement dressed up as ambition.

The real question is not whether AI agents can help build software. They already can. OpenAI describes Codex as a cloud-based software engineering agent that can work across codebases, run tasks in parallel, fix bugs, answer questions, and propose pull requests in isolated environments. Cisco has also described using Codex inside enterprise engineering workflows, citing improvements such as faster build times, higher defect remediation throughput, and more than 1,500 engineering hours saved per month in global environments. The harder question is how leaders decide which work should be explored, which work should be executed, and which work should be protected from premature automation.

This is where the scout vs. strike team framework becomes useful. It gives leaders a clean way to classify every initiative before they assign people, budgets, AI agents, or executive expectations. In traditional product development, organizations often treated most initiatives as delivery problems. A roadmap item appeared, product wrote requirements, design created flows, engineering estimated effort, and the machine moved forward. That model was already showing its age before AI. In an agentic world, it becomes actively dangerous.

AI does not remove uncertainty. It accelerates whatever system it enters. The 2025 DORA report makes this point clearly by framing AI as an amplifier that magnifies an organization’s existing strengths and weaknesses. Google’s summary of the same research notes that AI adoption can improve throughput and product performance, but can also create stability issues when testing, architecture, feedback loops, and workflow clarity are weak.

That means the first leadership discipline of agentic development is classification. Before you ask, “Can AI help us build this faster?” ask, “Is this a scout mission or a strike mission?”

A scout mission exists when the problem, value, workflow, architecture, or adoption path is still unclear. The goal is not to ship. The goal is to learn enough to decide whether shipping is worth it. A strike mission exists when the problem is validated, the outcome is clear, the technical path is bounded, and the organization needs concentrated execution. The goal is not exploration. The goal is production impact.

This distinction matters because the two modes require different people, governance, metrics, and AI usage. When leaders confuse them, they create theater. Scouts get punished for not delivering production code. Strike teams get slowed down by unresolved discovery. AI agents get thrown into messy systems where speed only increases rework, security exposure, and review burden.

The scout team should be small, senior, and uncomfortably close to the customer or user. A strong scout team is usually two to four people: a product lead, a senior engineer, a designer or researcher, and sometimes a domain expert, data lead, security partner, or forward-deployed engineer. The product trio model has been pushing organizations away from linear handoffs for years because product, design, and engineering need shared ownership of discovery. In AI-heavy work, that shared ownership becomes non-negotiable.

Scout teams should be time-boxed, not capacity-boxed. Give them two to six weeks, a clear learning thesis, and permission to stop work if the evidence is weak. Their assignment is to answer a narrow set of questions: Is the customer pain real? Can the workflow change? Is the codebase or data environment ready? Can an agent perform meaningful work safely? What evidence would justify a strike team?

AI agents belong in scout work, but they should be treated like interns with superpowers and no judgment. They can map codebases, generate prototypes, summarize customer research, create synthetic test cases, compare workflow variants, analyze support tickets, or draft technical spikes. They should not be granted production authority, broad credentials, or unsupervised access to sensitive systems. Anthropic’s engineering writing on how it contains Claude across products is instructive here: as agents become more capable, their potential blast radius grows, and deterministic boundaries such as sandboxes, virtual machines, filesystem limits, and egress controls become more important than trust prompts alone.

The output of a scout team is not a demo. It is an evidence package. That package should include the customer insight, business case, workflow hypothesis, technical feasibility assessment, data and security risks, prototype findings, agent suitability, recommended operating model, and a kill, hold, or strike recommendation. If the scout cannot produce evidence, the initiative should not graduate. Weak evidence is not a reason to extend discovery forever. It is a reason to stop pretending.

The strike team is different. It is a temporary, focused, cross-functional team built to deliver a validated outcome at speed. A good strike team is usually four to seven people with dedicated capacity: product owner, tech lead, two or three engineers, design support, QA or SRE, and security or data support when needed. It should have direct access to decision makers, a defined production goal, a fixed duration, and authority to remove local blockers without turning every decision into a steering committee.

Strike teams are where agentic development can look dramatic. Once the problem is known and the guardrails are real, AI agents can generate code, run tests, modernize repetitive patterns, draft migrations, produce documentation, and prepare pull requests in parallel. This is where the economics start to change. McKinsey’s research on AI-driven software organizations found that top performers saw meaningful gains across productivity, customer experience, time to market, and software quality. But the same conclusion matters more than the statistics: value requires an overhaul of roles, processes, and ways of working rather than tool adoption alone.

The strike team should be measured by business and engineering outcomes, not activity. Did adoption move? Did cycle time fall? Did escaped defects remain controlled? Did the team reduce operational load? Did the work create a reusable capability, or just another local exception? In an AI-native organization, speed without reuse is just a faster way to rebuild the same fragmentation.

This is why some work should graduate from scout to strike, while other work should become platform work. If a scout discovers a repeatable capability that many teams need, such as agent-safe test generation, customer-specific workflow configuration, identity integration, document extraction, or deployment automation, it should not become another one-off strike mission. It should move into a platform backlog. Team Topologies is useful here because it forces leaders to distinguish stream-aligned teams, platform teams, enabling teams, and complicated subsystem teams instead of inventing vague hybrid structures that create dependency fog.

The transition to agentic development should therefore look less like a grand reorganization and more like a portfolio triage discipline. Every initiative should be classified into one of four lanes. Scout work explores uncertainty. Strike work executes validated outcomes. Platform work creates reusable leverage. Core product work stays with stream-aligned teams once the pattern is known and repeatable.

That classification gives executives a practical migration path. Start by selecting a small portfolio of scout missions around agentic development. Pick areas where the organization feels real pain but can contain the blast radius: test automation, legacy code understanding, support workflow analysis, integration mapping, internal tooling, documentation modernization, or low-risk feature acceleration. Give each scout a clear thesis and a short clock. Do not let the first phase become an innovation lab with better branding.

Then graduate only the strongest findings into strike teams. A strike team should start with a written decision memo, similar in spirit to Amazon’s working backwards PR/FAQ discipline, where the customer value, scope, business outcome, and open questions are made explicit before the organization commits to build. This is especially important with AI because prototypes are cheap enough to fool executives into believing the hard work is done.

The scout vs. strike model also changes what executive recruiters should look for in technology and product leaders. The next generation of technology leaders will not be differentiated by whether they “use AI.” Everyone will use AI. The differentiator will be whether they can redesign the operating system around it. That means leaders who can classify uncertainty, deploy scarce senior talent surgically, protect teams from performative roadmaps, and build the control systems that allow AI agents to operate safely.

The best leaders will have a field instinct as much as a platform instinct. They will know when to send senior builders directly into the customer environment, the way forward-deployed engineering models have long argued, because some problems cannot be understood through dashboards, tickets, or quarterly business reviews. Palantir helped popularize this pattern with forward-deployed software engineers who work close to customer problems, but the lesson now applies far beyond Palantir. In an AI era, the fastest path to learning is often not another roadmap meeting. It is putting technical product builders into the real workflow and forcing the organization to confront what actually breaks.

Agentic development will punish organizations that cannot tell the difference between learning and execution. It will make strong teams faster, but it will also make confused teams louder. The answer is not to slow down. The answer is to become more precise about the shape of the work.

Scouts create conviction. Strike teams create impact. Platform teams create leverage. Stream-aligned teams create durable product value. The executive job is to stop treating every initiative like the same kind of work and start assigning the right structure before the work begins.

That may sound simple, but it is the management discipline most organizations skip. In the AI era, skipping it will be expensive. The companies that win will not merely have better models or more licenses. They will have leaders who know when to scout, when to strike, and when to turn one hard-won lesson into a system the whole organization can use.