Skip to content
/03Loop engineering9 min

What makes an agent loop converge

The expensive tasks did not verify too much. They lost the ability to act. Across 7,565 local Qwen3.6 turns, retained evidence separated productive loops from a costly reconstruction spiral.

The expensive tasks did not spend too many turns verifying. They lost the ability to act.

In tasks taking 60 turns or more, reading rose to 82.6% of turns while writes fell to 1%. Three-quarters of all turns were spent reopening a file already read during the same task. By contrast, tasks completing within 20 turns spent almost four times as much of their activity running commands and checks.

This action-to-reacquisition transition is the distinctive result of our analysis of 7,565 agent turns captured during a recent Solario factory campaign. We call the observed behaviour the read spiral. Our working diagnosis for its dominant mechanism is reference-surface eviction: the evidence needed for the next action falls out of retained context and has to be reconstructed.

Bounded tasks show the healthy profile. With a clear file target, the local Qwen3.6 model running on an H100 completed them in 8 to 20 turns, at roughly one minute per task. It read a bounded surface, wrote or edited, verified continuously and closed the task.

This distinction matters because the normal explanations point in the wrong direction. The corpus does not suggest that verification is the tax, that a larger turn budget is the answer or that open-weight models have reached a general ceiling.

It points instead to a working-set retention failure that becomes visible as a transition from action to reacquisition. Full execution telemetry allowed us to see that transition, measure its cost and identify where the next control belongs.

Context pressure and compaction are familiar agent-engineering problems. The contribution here is more operational: a production failure signature, a measured contrast between bounded and grinding tasks, and controls that can move upstream into workload design.

The production record made the problem visible

The corpus contains 7,565 paired request and response records from an application build campaign on 24 and 25 August 2026. Solario captured the model inputs, outputs, tool calls, result sizes, reasoning traces, token use and completion reasons into a 619 MB append-only debug record.

The analysis focuses on 4,844 implementation turns and 1,917 remediation turns. Each turn was attributed to its owning task from the execution contract in the system context.

That level of detail is important. Aggregate token use can tell us that a task was expensive. It cannot tell us what the model was doing while the cost accumulated. The production record can.

The distribution was strongly right-tailed. Across 56 implementation tasks, the median was 34 turns and the mean was 78.2. A small number of tasks reached 144 turns or more, with one accumulating 791 turns across attempts and relaunches.

Block cost followed the number of these grinding tasks more closely than the total number of tasks. A block with many bounded tasks could remain efficient. A block with one or two tasks carrying a large working set could dominate the campaign.

Healthy loops spend their turns acting

We compared cheap tasks, those completing within 20 turns, with grinders taking 60 turns or more.

Primary action Cheap tasks Grinders
Read a file 48.2% 82.6%
Run a command 27.2% 7.3%
Edit a file 3.3% 2.3%
Write a file 6.9% 1.0%
Complete the task 8.3% 0.5%

The healthy loop has a recognisable rhythm:

  1. Read the declared surface.
  2. Make a bounded change.
  3. Run the relevant check.
  4. Correct the result if required.
  5. Complete against the task contract.

Verification is not the tax. Cheap tasks spent 27.2% of their turns running commands, almost four times the share observed in grinders. They checked their work constantly and still completed quickly.

The grinder profile is defined by the disappearance of action. Writes fall to 1% of turns, edits to 2.3% and completion signals to 0.5%. Reading expands to fill the loop.

Good agent loops do not minimise verification. They preserve enough context to keep moving between evidence and action.

Reference-surface eviction and the read spiral

Three-quarters of every grinder turn was spent re-reading a file the model had already read during the same task. For cheap tasks, re-reads represented 6.2% of turns.

The repetition was not arbitrary. The most frequently re-read files were the reference surface for the hardest work: service implementations, domain entities, dependency-injection wiring and integration code.

The traces are consistent with a feedback loop:

  1. The base context, reference files and transcript approach the effective context limit.
  2. Transcript compaction removes older file contents.
  3. The model correctly re-reads those files to recover its working state.
  4. The transcript grows again.
  5. Compaction fires again and removes the same reference surface.

Cheap tasks experienced around 11 transcript compactions. Grinder traces commonly experienced between 100 and 139. In the 791-turn task, the model performed 691 re-reads and crossed 139 compaction events.

Tool errors appear to feed the same loop. Grinders carried roughly seven times the tool-error density of cheap tasks. The traces suggest that a failed action often sent the model back to reading for grounding, increasing context pressure before the next compaction.

Reasoning followed the same pattern. Grinders generated about three times more reasoning text per turn while taking fewer effective actions. In this corpus, that pattern is more consistent with repeated state reconstruction than with productive depth.

Why remediation converges

The same model behaved differently inside Solario’s remediation stage.

Remediation starts from an explicit finding and receives relevant evidence in its context. The loop does not have to rediscover the problem before acting. Its action mix reflects that difference:

  • 46.6% of turns read files, compared with 73.7% across implementation;
  • 14.2% edited files, compared with 4.9%;
  • 5.8% completed, compared with 1.4%.

Remediation still re-read paths. Inlining evidence did not eliminate repetition. It did, however, preserve enough of the problem definition and relevant surface for the model to keep acting.

This supports a broader factory principle. An agent loop is more likely to converge when the work arrives with a bounded objective, the evidence needed to act and an independently checkable completion condition.

More turns cannot manufacture those properties. They have to be designed into the workload.

Move the control earlier

The production loop already contains controls for several common taxes. Sampling intents adapt generation posture to the task. Degeneration guards detect collapsing output. Retry classes distinguish recoverable failures from terminal ones. Remediation receives a funded path back to conformity.

Those mechanisms appeared in the corpus doing the work they were designed to do. The read spiral was the dominant tax without a corresponding control.

The measurements point to five concrete changes.

Retain the declared reference surface

Files declared by the task should survive transcript compaction, either as pinned content or as a stable summary. This directly addresses the 75.1% re-read share in grinders.

Read signatures before bodies

Large files should first expose their exported symbols, structure and requested ranges. Full content remains available when required, but it no longer has to occupy the working set by default.

Promote batch reads

The available batch-read operation accounted for no observed turns. A declared task surface can be pre-read or loaded in one operation, reducing both tool-call overhead and transcript fragmentation.

Estimate working set during Design

Task generation can estimate the likely reference surface from declared files, file sizes and dependency fan-in. A task that exceeds a defined context budget can be split before it reaches Build.

This converts a runtime pathology into a design-time control.

Dampen error feedback

After a failed tool call, the loop should act on the returned error before reopening the entire reference surface. Existing nudges can intervene earlier, before the task reaches spiral depth.

Open-weight models are productive inside the right loop

The positive result is not merely that we found a problem.

Most bounded tasks never entered the spiral. Qwen3.6 worked effectively on local compute, used tools, verified its changes and closed clear tasks quickly. In this campaign, remediation was substantially more action-oriented when evidence was assembled around a specific finding.

The factory also made the remaining cost legible. Because every action, result and transition was recorded, we could distinguish verification from re-reading, useful reasoning from reconstruction, and ordinary task cost from the few loops dominating the campaign.

This is what governance contributes to agent engineering. It does more than stop invalid output. It makes the production system measurable enough to improve.

Our article on structured artefacts and executable contracts described how intent survives the boundary between stages. This analysis adds the runtime counterpart: the evidence required by a task must survive long enough for the agent to act on it.

Convergent loops preserve the evidence they need and keep acting. Solario now has the production record required to engineer for that outcome.

Method note

This is a descriptive analysis of one application campaign, not a controlled comparison between models or loop configurations. The debug capture begins after the campaign started, so the first block and most of the second are not represented. Task totals accumulate attempts and relaunches, which reflects total factory cost but not the duration of one uninterrupted attempt.

A turn is one paired model request and response record. The action table assigns each turn once to the extractor’s primary tool-call category. A re-read means that the same file path had already been read within the same attributed task, including earlier attempts and relaunches. A compaction is a recorded transcript-shrink event; in the diagnosed configuration, compaction retained 11 messages inside a nominal 65k-token context window. The analysis observes a strong association between compaction, re-reading and action collapse. The eviction loop is our causal interpretation of the traces, not the result of a controlled intervention.

The thresholds used for the comparison, 20 turns or fewer for cheap tasks and 60 or more for grinders, describe this corpus and should be tested on subsequent campaigns. Before treating this as a general result, we intend to publish the exact primary-action classifier, task-level distributions and anonymised trace evidence in the Solario claims ledger.

Solario · Point of View /03

← All articles