Right about the bottleneck. The gate has to scale.
Anthropic's AI-native SDLC playbook already automates most of the work between decisions. The remaining step is an adjudication layer that decides, at every transition, what the evidence permits.
Anthropic has published the clearest account yet of the AI-native software development lifecycle. It already describes automated handoffs, agentic review and a closed operational loop. The remaining step is narrower but consequential: applying explicit, policy-driven delegation to every transition.
The diagnosis is right
Anthropic’s AI-Native SDLC playbook begins with the diagnosis we would have chosen: code is no longer the bottleneck.
When agents compress implementation from weeks to hours, planning, design, review, security and release do not disappear. If those activities continue at human speed for every change, the queue simply moves.
Anthropic replaces linear handoffs with a chain of committed artefacts: intent, specification, plan, code, tests, findings and incident records. Accepted artefacts trigger subsequent stages. Skills carry institutional knowledge, hooks enforce action boundaries and agentic reviews reduce what people must inspect directly.
This is already more than a collection of prompts. It is a lifecycle model with automation across six stages.
The unanswered question is not whether the lifecycle should become a loop. Anthropic has answered that.
The question is what decides whether work may advance through it.
Automation is not delegation
The playbook automates much of the work between decisions, but several transitions still require a person to approve each change.
A product owner accepts the intent and reviews the specification. Engineers accept the plan. A code owner approves the pull request. Named authorities approve consequential releases.
These checkpoints are appropriate when judgement is required. They become a bottleneck when routine, conforming work follows the same route.
Skills, hooks, agentic reviews and branch protection each solve part of the problem. They do not, by themselves, form a single decision layer that evaluates every transition against the organisation’s delegation policy.
Without that layer, more agentic production still creates more material for people to approve.
The maintenance loop shows the stronger pattern
The playbook’s most complete delegation model appears in its Maintain stage.
A deterministic script observes production metrics. Version-controlled thresholds determine when a condition has been breached. Only then is the model invoked. A response tier constrains what it may do: record the event, investigate without making changes, open a pull request or trigger a pre-approved runbook.
The pattern is strong:
- The organisation defines the condition.
- A deterministic mechanism produces evidence.
- Versioned policy determines the permitted response.
- The agent acts within bounded permissions.
- A person receives the exceptions and consequential decisions.
- Every input, verdict and action remains reconstructable.
Anthropic also introduces an independent confidence gate when this autonomous maintenance loop passes work through subsequent stages.
The components therefore already exist. The next step is to make this policy-driven delegation the operating model for every lifecycle transition.
Every transition needs an adjudicator
An AI-native lifecycle needs agents that produce work and controls that inspect it. It also needs an adjudication layer that decides what the evidence permits.
From intent to design, the system should establish whether mandatory constraints, ownership and open questions have been addressed.
From design to build, it should evaluate architectural boundaries, approved dependencies, security requirements and unresolved exceptions.
From build to test, it should determine whether the implementation matches the accepted design and whether the required evidence exists.
From test to release, it should combine test results, security findings, change classification and rollout policy.
At every gate, the outcome should be explicit: advance, repair, stop or escalate.
That outcome cannot rest solely on the model’s confidence. Models can assess context and produce findings, but authority must come from approved policy. Objective conditions should be tested deterministically. Contextual assessments should be bounded, recorded and open to independent challenge.
Agents produce the work. Controls produce evidence. Policy adjudicates. People own the rules and decide the exceptions.
That is delegation.
Human authority is not human throughput
Human accountability and human intervention are not synonyms.
Architecture authorities can approve the boundaries every system must respect. Security can define which findings block, which can be repaired automatically and which require review. Release authorities can establish the conditions under which deployments may advance.
Those people remain accountable because they own the policy and retain authority over exceptions. They do not need to inspect every conforming output individually.
This is how any large organisation scales. Leaders define mandates, decision rights and escalation rules. Teams operate within those boundaries. Only exceptions and consequential decisions move back up the hierarchy.
Agentic software production requires the same operating model, made executable and auditable.
Advice, assessment and adjudication are different
Anthropic distinguishes between skills and hooks. Skills advise the model. Hooks enforce deterministic boundaries around its actions.
A governed factory needs a third category.
Advice tells the agent how it should behave.
Assessment determines what it produced and whether observable requirements were met.
Adjudication applies organisational policy to that evidence and decides what may happen next.
A prompt cannot certify that its instruction was followed. A model reporting compliance is not independent evidence. A hook may block a specific action, but a collection of local hooks does not automatically become a coherent lifecycle policy.
This matters especially for architecture. Documentation, skills and review instructions can guide generation, but they do not make architecture measurable.
Approved patterns, dependency direction, interface rules, ownership boundaries and structural invariants need enforceable representations. The rules that guide generation must also produce the controls that judge the result.
Measure the gate, not only the flow
Review time, acceptance rate and deployment frequency describe flow. They do not establish whether a gate made the correct decision.
A rising acceptance rate may mean generation improved. It may also mean the control became permissive.
A governed production system must therefore measure both velocity and control quality: false positives, false negatives, overrides, escapes and the cost of a successful production outcome.
The judge must be measurable too. Otherwise, an automated verdict is no more falsifiable than an unexplained human rejection.
What we keep and extend
We keep Anthropic’s committed artefact chain, automated handoffs, explicit plans, separation of authorship and approval, continuous evaluation, agentic review and version-controlled response tiers.
We extend that control pattern across the complete lifecycle.
At Solario, this is already the operating model.
The Solario Governance Harness binds every Project before execution and remains active from signed Intent to deployment. Factory and Project policy define the architecture, approved stacks, ADRs, quality standards, test obligations and authorised exceptions before agents begin work.
Solario Core owns the state machine and every legal lifecycle transition. Agents execute bounded work inside it. They do not define the rules or authorise their own progression.
At each gate, deterministic controls verify the artefacts and declared outcomes. Policy determines whether the Workload advances, enters a repair loop, stops or reaches a named human authority. Human approval is reserved for moments where judgement cannot be delegated.
Every verdict is retained with its inputs, controls, findings, repairs, exceptions and approvals. The production record is created with the software rather than reconstructed after the fact.
Model access sits behind a controlled SDK supporting frontier and open-weight endpoints. Models remain replaceable production resources. The organisation’s policies, gates and decision history do not change when the provider does.
This is the difference between adding agents to a delivery process and operating a governed AI software factory.
Anthropic is right about the bottleneck. It is also right about the shape of the loop.
The gate now has to operate at the same scale.
Solario already does.