Skip to main content

An Agent's Life Ends at Its Push

What: dispatch a skilled, bounded implementation life; move passive observation out of that life. Budget the work in measured usage classes and report enough current evidence for safe continuation. These are adjustable operating budgets, not new merge gates.

Evidence — one pilot, not a general benchmark. The pilot partner's de-identified measurement supplied on 13 September 2026 reports the following weekend fleet. Raw usage/billing records are not distributed with this kit, so these are reported measurements, not independently reproduced kit results.

Reported measureCoordinator (one session)Sixteen fixer agents
Model turns1,58817,418
Output tokens2.1 million2.2 million
Cache-read tokens0.9 billion9.5 billion
Median cache-read context per turn570,000430,000; largest reported agent 694,000
Turns reported as only watching CI/reviewNot supplied10%; two agents approximately 25%
Bash calls per edit, median landed PRNot supplied36; worst reported agent 81
Most expensive PR versus six cheapest landed PRs combinedNot suppliedReported 4.5× cost; billing reconciliation not supplied
Share of fleet cache reads in four agents reported as most expensiveNot supplied69%

Why — usage is turns × mean context, not output alone. More precisely, input/cache usage is the sum of per-turn usage records; multiplying turns by a median is not that sum. The fleet's rounded cache total implies approximately 545,000 cache-read tokens per turn. Combined output is approximately 0.04% of the supplied cache-read-plus-output total, not 0.4%; absent uncached-input/cache-write totals do not establish the source's “99% of all tokens” statement. Token classes have different prices: cache volume is not billed money. A growing retained context can make cumulative usage superlinear in lifetime, but compaction, caching and model changes mean it is not a universal law.

The coordinator produced almost as much output as the fleet in far fewer turns; that does not establish that it was output-bound in price or latency. Nor does a high Bash : edit ratio prove that tests were wasteful. Both are investigation signals requiring task and command evidence.

How — six rules and their reference budgets. Freeze an authorized profile revision before comparing windows. Defaults come from the pilot proposal; their benefits remain to be measured.

RuleReference threshold / actionRequired receipt
1. Workers do not watch CI or reviews: a life ends at its authorized push and report, or a checkpoint if push is not authorized.Watching-only turns / model turns <5%; verified observer handles later eventsWatching share plus review-response latency and observation gaps
2. Handoff oversized context through the milestone report.250,000 cache-read tokens per turn, explicitly a proxy rather than total contextMeasured class, profile revision and successor handoff
3. Declare a turn budget for every life.Alarm 800; stop 1,500 turns; checkpoint and stop new effects, propose decomposition where appropriateTurns per life and per PR, including all replacement lives
4. Avoid purposeless rerun loops; target the meaningful edit cycle.Bash calls / edit operations <10 is an investigation signalActual check selections, exits and justified repeated runs
5. Keep the coordinator's context focused; delegate independent rendering when authorized and batch independent calls.No universal numerical cap; avoid whole-transcript dumpsTask-scoped usage and inspected evidence, not shorter output as a quality proxy
6. Stop new dispatch before exhausting the approved budget.Compare remaining budget at a measured six-hour pace with time to a known resetSame-unit horizon, known reset, remaining budget and dispatch decision

At an alarm, finish a safe current edit cycle and checkpoint. At stop/handoff, reconcile child jobs and fence old writers before a fresh life starts. A replacement inherits remaining mandate budget; respawning cannot reset it. A threshold does not prove that splitting a coupled PR is safe. Splitting, additional spending or continuing integration needs the appropriate authority. Missing usage, unknown reset and zero denominators are cannot-assess, not unlimited budget or zero cost.

The report file is the successor's memory, not its authority. Write at root-cause, fix, checks and push milestones: task/run/head, owned changes, actual exits, unresolved gaps, active child jobs and next bounded action. Revalidate current sources, mandate/revocation, ownership and permissions on takeover (§5.4–5.5). Do not replay a stale report as permission. Reserve shutdown/reporting allowance inside the mandate before dispatch; aggregate usage across all lives and agents rather than counting only the final worker. Freeze turn/edit definitions in the profile. Read mandatory instructions completely; relevant sections are preferred for ordinary source/log inspection, not as a reason to omit necessary context or visual checks. RED/GREEN, baseline, final comprehensive, environment and justified flakiness runs remain required even when they repeat a command.

Resolve response latency structurally. Earlier advice against tight polling and pressure to respond quickly to reviews need not conflict: an actual non-LLM observer can poll REST, for example every ten minutes, and dispatch on a relevant event. That consumes no model tokens for the polling itself, but still has API/runtime cost. A red check supplies its exact head and failing log; a review supplies findings and reviewed head. Both need current authority and remaining budget, not just the log. The kit does not ship this watcher or promise unattended resumption. If the host only offers model-driven waiting, use its supported bounded mechanism within budget and disclose the limitation.

The authorized controller owns ready labels and review requests, not the implementer. Prefer one review at the believed-final head; later substantive changes require reassessment. A label is not CI evidence, superseded is not verified-fixed, and a human authority is not inferred from an agent's “coordinator” label. Duplicate/stale events and old writers must be checked before new effects.

Context optimisation — worked case. Investigate five levers in this order as a hypothesis, not a measured ranking: (1) remove passive model-driven watching; (2) bound task context with safe handoffs; (3) remove redundant commands without removing verification; (4) choose fitting roles and delegate independent work within authority; (5) evaluate optional compression. The first four avoid processing unnecessary material; compression optimises what remains. Permission-filter before retrieval/compression and retain authoritative rules, glossary definitions, acceptance criteria, scope and permissions without lossy transformation. Preserve retrievable originals and isolated caches. Compare actual cost per correctly completed, traceable task, retries, citations and correction effort—not token reduction alone. See business intent lifecycle and governed factory; product adapter/budget contracts remain separate from installed kit enforcement.

Detectors and receipts, with boundaries:

  • Pre-dispatch review checks role + reason, paired reviewer, four budget units, report path and no-watching instruction. The kit's presence probes check these artifacts, not an agent's conduct.
  • An adopter's collector deduplicates provider usage records by message ID and associates life/task/head. These establish measured usage; money remains estimated until billing reconciliation.
  • Watching-only turns require separately authorized, sanitized tool-event metadata and a frozen classifier. Usage records alone cannot identify watching or edits. Distinct watching-only turns / turns differs from matching calls / turns, which can exceed 100%. Do not surface raw transcript text.
  • An adopter may implement lifecycle state, missing-milestone and per-PR cache-volume signals. At ≥2× a comparable landed-PR median, investigate cost; missing/zero baseline is cannot-assess. No runtime collector, budget enforcement or board is shipped by these templates.
  • Next-five-landings receipt: median turns per PR, context proxy, watching share and Bash : edit; include sample size, unfinished work and review-response latency. After twenty landings, compare role cohorts and cost/correctness with frozen workload/profile and report confounders. Neither receipt is achieved by this revision.

Operational home: skill:s4u-loop-dispatch; paste-ready lifecycle block: templates/agents/dispatch-brief-lifecycle.md (installed under .claude/templates/, not as an agent). Reference-only configuration templates: templates/dispatch-budget-v1.yaml, templates/diagnostic-rules-v1.yaml, templates/agent-usage-profile-v1.yaml; adopt them explicitly with project review, not as preinstalled enforcement. See §5.3–5.5, §7.9 and the retired appendix-c / appendix-k pointers.