Skip to main content

Appendix P — The delivery-flow dashboard (reference)

A reference view of the delivery-flow evidence described in §7.9 (cadence is part of what a gate costs), §7.10 (an agent's life ends at its push) and the governed factory operating model: which figures a delivery surface may show, what each one is allowed to claim, and where a figure must say not assessed instead. Every value reproduced here is rendered from a synthetic, de-identified fixture built for illustration. This appendix is not a shipped runtime, not a productivity claim, and not a release authorization.

When to use it

Use this reference when deciding what a delivery-flow view should display and how each figure must be qualified — designing an internal board, a weekly delivery review, or a gate-cadence exercise under §7.9. It describes a reference layout and an evidence discipline, not a product the kit installs.

Do not use it as a source of measurements. The screens render a synthetic fixture: a fourteen-day window of fictional work items in an invented domain, with de-identified identifiers and no real repository, person, customer or vendor. Every number in the screenshots is demo data. Where this appendix cites a threshold it is a canon reference figure carrying canon's own hedging — §7.10's budgets (alarm 800, stop 1,500 turns; a 250,000 cache-read-tokens-per-turn handoff proxy) come from one pilot proposal whose benefits remain to be measured, not from a benchmark.

Do not use it as a control. A board is a decision surface, not a gate: it displays what checks and platform events recorded, adds no observation of its own, and carries no approval authority.

What each screen answers

Business intent. Answers what was approved, by whom, at which revision. Approved intent revisions with approver and package digest; the execution mandate bound to that revision; the mandate's result receipt with its domain labelled simulation or actual. Business validation is its own record, separate from merge. Readiness is reported as individually named blockers — missing owner, ownership conflict, absent approval, missing evidence, stale publication, missing review — and deliberately not as a score, because one number lets six different obligations cancel each other out.

Business intent screen, synthetic demo: approved intent revisions with approver roles and package digests, an execution mandate whose receipt is labelled simulation, a business-validation record, and six named readiness blockers with no aggregate score

Delivery today. Answers what landed and how it moved. Landings in the window, change lead time at median and p85, rework share, merge-forward cycles, CI minutes, and landings by hour. Each tile carries its own provenance label and sample size; in the fixture several are not assessed because the underlying check-run stamps were never collected.

Delivery today screen, synthetic demo: tiles for landings, median and p85 lead time, rework share, merge-forwards and CI minutes, several reading not assessed instead of a number, above an hourly landings chart

Where time goes. Answers which stage the calendar time sits in. For each stage it shows both the p50/p85 lead-time distribution and that stage's share of total lead time. Held on purpose is a first-class state, not a gap. Work with no event coverage is shown separately as unknown. The two rankings can disagree, and the screen shows both rather than picking one.

Where time goes screen, synthetic demo: per-stage p50 and p85 lead times beside each stage's share of total lead time, a held-on-purpose band, a separate unknown-coverage band, and a largest-p85 stage that differs from the largest-share stage

Decisions and receipts. Answers what was proposed, what a named person authorized, and what came back. Each entry is a bounded mandate: evidence → action → expected receipt → unit budget → named authority. Receipts either verify or refute the expectation, and refuted is a result, displayed as prominently as verified: the receipt contract in the governed-factory model treats a missing or empty receipt as unassessed, and a refuting receipt as a completed observation.

Decisions and receipts screen, synthetic demo: bounded mandates showing evidence, action, expected receipt, unit budget and named authority, lifecycle badges reading proposed, applied, receipt and archived, and one mandate marked refuted

Cost. Answers what the window consumed, in units first: token classes (uncached input, cache-read, cache-write, output), CI minutes, review credits and wall-clock. Money appears only behind a fold and is labelled estimated until reconciled to billing — §7.9's point that there is no single per-minute rate, and §7.10's that cache volume is not billed money.

Cost screen, synthetic demo: unit totals for token classes, CI minutes, review credits and wall-clock, with a collapsed fold holding a money figure labelled estimated until reconciled to billing

Fleet. A detail screen: agents working or waiting, held items, and the queue. Producer state that an agent reports about itself is labelled self-reported, never promoted to observed.

Fleet detail screen, synthetic demo: counts of agents working and waiting, held work items and queue depth, with producer states marked self-reported

Trace of one work item. The events of a single item in order, with stage attribution where the event stream supports it and not assessed where it does not. No transcript content.

Trace detail screen, synthetic demo: one fictional work item's events in order, each row either attributed to a stage or marked not assessed, with no transcript text

Model usage. Weekly allowance consumed, tokens by day and hour, and the measured provider usage records behind them, deduplicated by message ID.

Model usage detail screen, synthetic demo: weekly allowance consumption, token totals by day and hour by class, and a table of measured provider usage records

Fourteen days. Landings per day and daily median lead time, each point carrying its own n so a one-item day cannot read as a trend.

Fourteen-day history screen, synthetic demo: landings per day and daily median lead time, each point annotated with its sample size n

Same board at twelve teams. An enterprise projection: what stays per-team and what aggregates. Nothing is pooled across incompatible flow profiles; profile mismatch is shown as a reason for absence, not silently averaged away. No aggregate on this screen is evidence about any one team; it is a projection of how the same board would be laid out, not a measurement.

Enterprise projection screen, synthetic demo: the same board repeated for twelve fictional teams, per-team figures kept separate and aggregates withheld where flow profiles differ

The provenance vocabulary

Every figure carries exactly one label. The label is part of the figure; a number without one is not renderable.

LabelAdmissible basisLimit
observedread directly from a system of record in this passScoped to what that record covers at that revision; not a claim about paths the record does not see
self-reportedstated by an actor (including an agent about itself)Not independently checked; never promoted to observed by repetition
measuredcounted by an instrument, e.g. provider token usage recordsEstablishes usage volume only; cannot classify watching, edits or correctness
estimatedmoney, and any figure reconstructed from a rateRemains estimated until reconciled to billing; mixed runner classes make a single-rate total wrong
derivedcomputed from observed inputsCarries the provenance and the limits of its weakest input
not assessedthe input was not collectedRendered as the words or a dash — never as zero, never omitted

The last row is the load-bearing one. Absence of events is unknown, not idle: missing telemetry must not become zero defects, and a zero denominator is cannot-assess rather than a passing ratio.

Flow profile v1

The stage vocabulary the screens render is a declared profile, frozen by revision before any two windows are compared. It fixes:

  • Stage boundaries — the event that opens and the event that closes each stage, by event type, so a stage cannot be redefined between windows.
  • Held — an explicit state for work deliberately paused. Held time is reported, not redistributed into a neighbouring stage.
  • Unknown — items whose event coverage does not support attribution. Reported as its own band with its own count.
  • Overlap allocation — the rule for time covered by two stages at once. The profile names one rule and applies it to every item; an unallocatable overlap is unknown.
  • Landing definition — what counts as landed, since lead-time percentiles and every per-landing rate depend on it.
  • Minimum comparable sample — below it, a percentile is reported as cannot-assess rather than drawn.
  • The p85-versus-share rule — the stage with the largest p85 and the stage holding the largest share of total lead time are different questions and may name different stages. The profile requires both to be shown; neither is the headline.

Mandate lifecycle and receipts

A diagnostic rule may draft an evidence → action → expected receipt → budget proposal. A named human authority fires it. The board displays the lifecycle; it does not advance it.

A worked synthetic example, from the demo fixture's mandate A. Evidence: a queue-wait figure and thirteen concurrent items observed against a profile target of three; provenance self-reported, sample size shown. Action: pause new dispatch until work-in-progress drains. Expected receipt: over the next five landings, concurrent items at or below target, reported with sample size and unfinished work. Unit budget: the §7.10 reference budgets — alarm 800 turns, stop 1,500 — with the remainder inherited by any replacement life rather than reset by respawning. Named authority: the delivery-flow owner recorded on the mandate. In the fixture the receipt came back refuted: the count stayed above target. That is a completed cycle, not a failure of the board.

Cost and usage accounting

Units first, money last. Token classes stay apart (uncached input, cache-read, cache-write, output): they price differently, and cache volume is a usage signal rather than billed money. CI minutes are counted per runner class, since runner classes bill differently and self-hosted bills at zero. Provider usage records, deduplicated by message ID and associated to a life, task and head, are measured usage. Any money figure built from those units is estimated until reconciled against the billing record, the only authoritative rather than reconstructed figure. A per-item cost comparison against a comparable-landing median — the reference diagnostic uses a 2× multiplier — is an investigation signal, not a verdict about waste, and a missing or zero baseline is cannot-assess.

What this does not claim

It does not claim productivity. It does not rank employees and renders no transcript content. It does not establish that a control is enforced: a tile reporting a gate is a view of that gate's recorded outcome, not the gate. It does not authorize release, merge, spending or dispatch — each needs its own grant. It does not prove its own coverage: a figure is scoped to the emitter coverage and window stated beside it. And it is not independently reproduced evidence; the demo is synthetic by construction.

Adopting it

Three inputs are human-authored and cannot be inferred:

  1. A flow profile — the stage boundaries, held and unknown definitions, overlap rule, landing definition and minimum comparable sample, frozen by revision.
  2. Expected receipts — for each mandate class, what would verify and what would refute it, written before the action.
  3. Budget authority — who may spend which units, up to what, and what happens at alarm and stop.

Run first in demo mode against the synthetic fixture, so the layout and the provenance labelling can be reviewed with no live data and no possibility of a real figure being read as approved. The reference configuration templates are in templates/: dispatch-budget-v1.yaml (units, alarm/stop, replacement inheritance, receipt window), diagnostic-rules-v1.yaml (signals, comparison window, minimum sample, and automatic_authority_or_merge_block: false) and agent-usage-profile-v1.yaml (the field inventory and frozen definitions). They are reference-only: no collector, budget enforcement or board is installed by adopting them.

Relationship to the Business Intent Workspace

The workspace owns identity, approvals and the authoritative delivery-flow record; this appendix describes a reference projection of it, and the generated documentation site is a further audience-filtered projection of the same facts. Where the two disagree, the workspace's delivery-flow view is authoritative and the rendered page is stale. Detailed product design for the workspace is a private record, not cited here; the public obligations it must meet are in the governed factory and the business intent lifecycle.