S4U Methodology — Changelog
All normative changes to the canon are recorded here. The canon had no version identity before 3.0.0 (assessment 2026-06-11, finding PW-5); v2.x history below is reconstructed from downstream pointers.
4.0.5
Reference appendix for delivery-flow evidence and a hosted synthetic demo. No new rule, no operating-card change, no kit file change; adopters are not migrated by this version label.
- Appendix P — The delivery-flow dashboard (reference): which figures a delivery surface may show and what each may claim; the provenance vocabulary (observed · self-reported · measured · estimated · derived · not assessed); flow profile v1 (stage boundaries, held, unknown, overlap allocation, landing definition, minimum comparable sample, the p85-versus-share rule); the mandate lifecycle in which a diagnostic drafts and a named human authority fires; units-first cost accounting with money estimated until reconciled to billing. It extends §7.9, §7.10 and the governed factory model as reference material; it homes no control.
- Synthetic demo:
website/static/demo/delivery-flow.html, rendered from a de-identified, fully synthetic fixture (fictional member-services domain, stableWI-nnnidentifiers, no repository, person, customer or vendor). Ten screens in story order: business intent, delivery today, where time goes, decisions and receipts, cost, then fleet, trace, model usage, fourteen days and the twelve-team projection. Every figure carries its provenance label; not assessed is a state, never zero. Screenshots in Appendix P come only from this demo. - Same-day revision of the demo (published under the same version label; no canon text changed): every screen gained its detail layer behind folds — approval and evidence records, the 40 most recent landings by check verdict, landing sizes, four-window comparison and tuning log, mandate lifecycle steps and the applied-but-unreadable mandates, per-day and per-job cost rows, the landing order and landed blocks, a trace item picker, token classes, daily stage medians and ledgers, the merge ladder — plus a "How to read this screen" fold per screen.
- Not supplied: a collector, a runtime board, budget enforcement, measured savings or any
adopter figure. The reference templates
dispatch-budget-v1.yaml,diagnostic-rules-v1.yamlandagent-usage-profile-v1.yamlare unchanged.
4.0.4
Universal release-validation guidance, adopter-owned evidence records and a read-only structural checker. Review adoption before accepting the pinned kit; no project is migrated by this version label.
- Universal protocol: independently reviewed expectations, authoritative record → payload → actual UI/API/CLI/artifact/event output, adverse states, alternate-path protection, accountable exceptions and exact-candidate evidence.
- V1 reference pack: separately supplied
candidate/profile/corpus identities, explicit surface inventory and RV1–RV6 checklist.
scripts/release_validation.pyvalidates structure/accounting and confined local evidence only; it never performs semantic comparison or fetches external references. Exit 0 can accompany producer-reported failed/cannot-assess outcomes. - Canon, operating card, rule inventory and review/testing/dispatch procedures carry the universal controls and evidence limits. The decision-support addendum is explicit opt-in, nonnormative and independently removable; no core installation or checker fixture depends on it. Internal lessons, implementation plans and designs are not admitted to the public projection.
Not supplied: semantic certification, UI crawler, automatic adopter migration, runtime coordinator, guaranteed savings or automatic release authorization. Local synthetic fixture evidence is not captured production validation, live host/server assessment, remote exact-head CI, independent semantic approval or deployment. Security/export and independent-review gaps require separate disposition; existing website dependency-security limits remain applicable.
Demand-driven CI is a selectable cadence. One new rule, homed in §7.9, with a shipped workflow shape and one operating-card line. Register rows 173–174.
- The selectable rule: a project may choose no schedule, no
pushtrigger and nopull_requestopen/synchronize trigger; a label (ci:full/ci:gates/ci:security) orworkflow_dispatchis the demand. One full run on the final head before a release includes gates, the complete suite and every applicable configured security tier. Merges under an authorised fast mode retain the[skip ci]caveat and use dispatch for a marked final head. The project records its selection and costs in its own ADR. - New template
templates/workflows/ci-on-demand.yml: label- and dispatch-triggered, three tiers, exact-scope cancellation and fail-closed placeholders. For a fresh on-demand-only install usebash setup/bootstrap.sh <target> --with-ci --decline .github/workflows/pr-review.yml; for automatic PR CI use--decline .github/workflows/ci-on-demand.yml. A decline preserves an existing file and relinquishes kit checksum ownership; it does not disable an installed workflow, so existing adopters need a separately reviewed owner-authorized migration. - Surfaces: operating card (cadence selection),
skill:s4u-testing-standard(evidence scope), andtemplates/project-claude.md(the adopting project declares its model and ADR). - Card budget: admitted by the consolidation pass the rule inventory demanded, not by a third cap raise — roughly 400 bytes of prose compressed out of eleven existing bullets, no rule retired, cap unchanged at 12,000 bytes. The combined 4.0.4 card measures 11,966 bytes (34 bytes headroom); the next card rule needs a retirement.
- This repository's own workflows move to the selected demand model:
canon-ci.ymlloses itspushandpull_requesttriggers and gains label/dispatch demand plus exact-scope cancellation; independent repository-policy assessments no longer cancel one another. The measurement is recorded in the repository lesson.
Enforcement accounting. Check implementations, thresholds and coverage floors are unchanged. Merge enforcement is not: ADR-0261 records the repository owner's authorization to remove required_status_checks from ruleset 20224854; deletion and non-fast-forward rules remain, while required-check blocking does not. The human obligation to demand a complete final-head run is not enforcement-equivalent, and no detector verifies it; that is an unresolved recommended (not enforced) gap. Other adopters are not migrated automatically: cadence and any required-context change remain owner decisions recorded against effective server policy. The figures are two repositories measured on one date, not a benchmark or promised saving.
Scope: the six release-validation patterns and selectable demand cadence are incorporated as universal guidance and reference tooling, not a semantic detector, runtime service or adopter acceptance receipt. The original lesson remains internal and is not a public source. R170–R172 reassessment and non-waivable controls remain binding.
4.0.3
One new canon section, §2.11 — A Review Loop Is a Design Signal, and rules R170–R172. This release adds review guidance and corrects navigation; it does not install a round collector or new runtime enforcement. See the adoption guidance.
- Repeated findings: on the second confirmed same-shape finding, enumerate affected sites within the approved boundary and test discovery of unregistered consumers. Preserve structural coverage gaps and independent behavioral tests.
- Round-four reassessment: count per predicate family, not total PR rounds. Obtain an accountable disposition: structural correction, justified bounded correction or authorized deferral. Existing incident authority can permit urgent containment; neither a counter nor a tracking issue waives blockers.
- Pre-review evidence: stable family IDs and review/head references, reader discovery scope and a pre-request verification pass in
skill:s4u-code-review. No shipped round collector or new runtime enforcement. - Testing appendix §14.2 distinguishes independently discovered readers from hand-maintained lists, includes positive and adverse fixtures, and limits AST claims to the access forms actually inspected.
- Operating card summarizes reassessment without raising its 12,000-byte cap. R170–R172 follow the 4.0.2 worker rules R168–R169. The next card addition needs consolidation, not a cap increase.
- Navigation includes the new Philosophy page; generated source alone had left it outside the explicit sidebar and failed CI coverage.
Evidence limits: the pilot partner reports one project/day with three PRs taking 10, 12 and 11 review rounds, roughly eight hours each, versus other PRs reportedly taking 1–2. Convergence after policy consolidation supports investigation, not proof of equal difficulty, round-one availability or causal cost savings. Four is a chosen reference threshold, adjustable through a reviewed profile. No claim of universal reader discovery, lowest cost, achieved improvement or automatic design correctness.
4.0.2
One combined kit revision for bounded worker lives and skilled-worker selection.
- Agent lifecycle and token budget (§7.10): a life ends at its authorized push/report; turn and context budgets, milestone handoffs and event-driven observation. The worked case qualifies the pilot partner's 13 September 2026 measurements: usage is a sum of per-turn classes, not output alone; cache volume is not billed cost.
- Skilled workers by default (§5.3–5.4): six implementer cards, type + reason in each brief and diff-matched reviewer pairing. Appendix C and Appendix K remain pointers to the canon and updated loop-dispatch/code-review skills, not duplicated procedures.
- Operating card: six-line token-discipline summary within the existing byte budget. Budgets never waive required tests, complete instruction reads, current authority or stale-writer checks.
- Adoption: the installer receipts the lifecycle brief as a template, not an invocable agent, and preserves local adaptations. Reference budget, diagnostic and usage templates distinguish measured usage, inferred activity and unreconciled money. Packaging and installer probes cover their stated structural/adoption scope.
What this revision is not: no existing merge gate is changed; no watcher, collector, board or runtime budget enforcement is installed. It does not claim lower cost or better results from typed workers. The next-five-landings check and batch-2 comparison after twenty landings remain pending; a successful kit test is not either receipt. contracts/v1 is unchanged. Reviewers remain non-editing in policy; the existing migration card's shell capability still needs actual host restriction.
4.0.1
Adoption hardening for existing repositories. This patch changes installation selection and missing-policy assessment; review the migration guide before updating an adopted kit. It does not install a factory orchestration runtime or certify an application's readiness.
- Explicit CI consent: workflows require
--with-ci, including when the target already has.github. The 4.0.0 directory-based selection was not true opt-in. - Selective adoption: repeatable
--decline <installed-path>options preserve files, omit proposals and exclude declined paths from checksum ownership. No manually composed receipt is needed for that selected scope. Options are per-run; retain the reviewed command for subsequent upgrades. Fresh settings omit declined hook callers; existing settings remain untouched. See selective adoption. - Configured project gates: the PR integration guidance
explains the template's
S4U_PROJECT_GATEvalue. Unconfigured or missing executables fail closed, and the selected gate's status propagates. The workflow no longer assumes a nonexistent project script. Template checkout steps explicitly disable persisted credentials. - Honest diagnostics: the memory hook distinguishes missing, nonregular, unreadable and invalid UTF-8 inputs while remaining advisory and avoiding private-content disclosure.
- Recovery and activation: add an unstamped-installation recovery flow. Byte equality is not historical provenance. Bootstrap explicitly reports unverified host registration and the staleness hook's project-specific caller requirement. No existing host settings are silently activated.
- Missing documentation policy: actual doc-sync assessment without its committed mapping now reports UNASSESSED and exits 2, not an inert success. Deleted or dirty-only mappings cannot be waived. Explicit non-adoption belongs in the reviewed caller/profile policy. Historical dry-run analysis remains separate.
- Discoverable memory helper: the operating card links the optional catalogue workflow. Decline it when unused; it is navigation tooling, not agent retrieval.
Regression coverage exercises real installer plans, writes, receipts, retained settings, configured gate results, input diagnostics and committed-policy selection. Local fixtures do not prove host invocation, server-side required checks or adopter acceptance. The 4.0.0 dependency-risk boundary is unchanged by this patch; it does not transfer to another project.
4.0.0
This release covers the methodology, adoption kit and public documentation. It is not an application-readiness claim or proof of an adopter's migration. Pin the exact Git revision when adopting it; the version label alone does not identify uncommitted changes, local configuration or effective server controls.
Operating model and business-to-engineering handoff
- Add the governed factory model, retaining the Business Intent Lifecycle: AI proposes, the BA reviews first, accountable owners accept business meaning, and engineering returns evidence for an exact candidate. Approval, execution permission and release acceptance remain separate.
- Define scoped context, glossary/taxonomy, versioned method profiles, execution
mandates and result receipts. The
contracts/v1pack supplies structural schemas and synthetic examples, not an orchestration runtime, authorization service or proof of receiver-side semantic validation. Protocol and method versions are independent. - Cover legacy knowledge mining, slow-moving baseline changes, integration contracts and in-flight amendment handling. Configuration is appropriate only within an already implemented, tested and authorized capability; an unforeseen capability can still require engineering.
- Add primary-source factory research with explicit limitations. Type 3 is an operating-model description, not a certification or evidence that autonomous delivery is safe for every adopter.
Control and tooling corrections
- Reconcile canon, operating card, inventory, skills and templates around actual control scope. A local hook is not a server-side enforcement boundary; absence of findings or inability to assess does not become evidence of compliance. Correct privilege, fixture-isolation, review-independence and historical measurement interpretations without deleting their history.
- Harden bootstrap planning and recovery, preserving local instructions, settings, edited files and upgrade proposals. Installation, host registration and effective CI/repository policy need separate verification.
- Protect the ADR mirror's first promotion rename with interruption recovery, and compare the live output against its staging snapshot before replacement. Concurrent authored edits stop generation rather than being overwritten. Cooperative locking remains necessary: the final comparison and rename are not an atomic transaction with unrelated editors.
- Validate gate inputs and checked I/O; replace passing CI placeholders with an explicitly configured executable gate; use committed documentation mappings and scoped mutation receipts. Preserve visible unassessed outcomes and recovery conflicts rather than reporting them as clean results. Compare the complete supported inventory/canon version, including a prerelease suffix; an RC and a final release are not interchangeable audit identities.
- Use one admitted-source publication path with owned generated files, bounded locking/recovery and freshness checks. Keep authored pages and public routes. Correct production-page fragment drift during asynchronous diagram rendering without serializing local machine paths into public assets.
- Add the migration guide, including wrapper/helper
compatibility, the unused
DOC_SYNC_BLOCKINGflag, upgrade conflicts and first-green evidence. These are adoption instructions, not activation receipts.
Evidence and remaining acceptance
Website dependency security
Update colord, Joi, SVGO, qs, fast-uri, Nano ID, DOMPurify and Mermaid to
patched releases, and update js-yaml for the additional advisory found by a
fresh audit. Keep Mermaid on the tested 11.x line with an explicit 11.16.1
override; override qs to 6.16.0 because its consumers' ~6.15.1 range does
not admit the fix. These are reviewed compatibility constraints, not permanent
exemptions from future security updates. See the
Mermaid release
and YAML advisory.
Two upstream image-size 2.0.2 advisories remain unresolved:
ICNS parsing and
JXL/HEIF parsing.
Docusaurus uses this package to measure local Markdown images during a build;
the published site serves static output, not an image-upload/parsing service.
Malformed source images can nevertheless exhaust a build process. A successful
build of the current corpus does not remediate the dependency. The upstream
package's root disableTypes API does not configure the separately bundled
fromFile API used here. No such mitigation or clean dependency audit is claimed.
The website CI job now has an explicit ten-minute timeout to bound individual
runner exposure; a regression checks this configured budget. This is not a
parser fix, an aggregate cost limit, or a timeout for local builds. Do not build
untrusted content locally. Maintainers must reassess when the parser or
Docusaurus dependency changes, or before accepting external image inputs.
Candidate verification
The version-integrated local candidate passed 207 Python, 192 shell and 16 Node tests, site typecheck/build, clean locked site dependency verification, generation/freshness checks and bounded desktop/mobile diagram review. Scoped Aikido scans and independent read-only reviews informed the corrections; neither is a whole-repository security clearance. Historical entries below retain their original assessment context and do not define today's scanner status.
The final Git revision still requires remote CI and authenticated publication/readback before being called a published release. Real adopter host registration, effective server controls and application acceptance remain separate. Neither this entry nor the structural contract pack claims that the future business-intent product has been implemented.
Interpretation corrections — 2026-09-12
The dated entries below preserve what was reported at the time. Current controls and the limitations documented in the canon supersede historical interpretations:
- Doc-sync probe (3.2.0): 2 would-block cases among 40 commits is 5% trigger
incidence, not an established false-positive rate. The original
docs/probes/doc-sync-blocking-2026-06-14.txtremains unchanged; itsPASSdoes not authorize blocking activation. The adjacent-correction.mdrecords the withdrawal. Classify expected/actual cases and gaps before a policy decision. - Generated freshness (3.2.2): reproducibility is a bounded comparison, not an "ungameable" control or a measured near-zero false-positive rate. Input identity, ownership, errors, content/modes and invocation matter; semantic fidelity and deployed access remain separate checks. See Appendix H.
- Review evidence (3.6.0–3.7.0): complementary catches in a small sample do not demonstrate statistical independence or uncorrelated blind spots. Neither agreement nor silence proves correctness. Historical harm comparisons and throughput observations are not universal priorities or performance guarantees.
- De-projection (3.1.x): this history and the showcase retain bounded project evidence. The older "project-free repo-wide" wording is not a claim about today's publication; historical numbers must not be borrowed as expected outcomes for another project.
- Current reference behavior: the Stop adapter is an advisory/debug reminder, skills are instructions, and local push interception is bounded. Neither their presence nor the older adoption narrative proves server enforcement. See Appendix E and the governed operating model.
These corrections are not a 4.0 release or an application-readiness declaration.
Unreleased
-
Publish the Business Intent Lifecycle white paper as a proposed reference model, reachable from the lifecycle sidebar and Codex/Astra page. Correct generated operating-page links to use absolute site routes so Cloudflare trailing-slash redirects do not break them.
-
Restore traceability for the already-present §2.10 rules R165–R167, explicitly classified as canon prose rather than shipped enforcement. Regenerate the standalone site with its own generator and carry public classification into generated YAML frontmatter.
-
Document the working TrustRelay Codex/Astra primary implementation workflow, separately from the established Codex review role. Record targeted local verification followed by comprehensive CI before merge; a full local run is optional diagnostic work, not a universal push prerequisite.
-
Add two workflow diagrams and a dated #1421 evidence checkpoint. Mark the capped backend run incomplete and current-head CI/review pending. Aikido and CodeRabbit remain unpurchased, unvalidated extension candidates; no performance or cost advantage is claimed.
-
This is documentation of a project operating configuration, not a new portable installer or a completed comparative evaluation.
3.8.0 (2026-09-08)
One new section, §2.9 — Claims That Borrow Their Authority, and four rules (161-164). Doc-only; no new executable, no change to the install surface. Every rule below survived an adversarial pass whose default was to reject and whose kill criteria were "already covered by an existing rule" and "could not have changed a decision made today" — six sibling candidates were killed under exactly those criteria, two of them falsified by their own evidence.
- The unifying claim (§2.9). The dangerous claim is not the unchecked one; it is the one standing next to something that was checked, close enough to inherit the feeling of having been measured. Three failures on unrelated subject matter share this shape.
- Row 161 — the substitution premise. Removing a carried value because something else re-derives it needs a test that runs the re-deriver on the input the removed value covered, watched failing. The test of the behaviour you changed passes whether or not the premise holds, so it is not evidence for the removal. Measured: a blocker dropped because a sweep "re-derives it" — the sweep read string entries only, so a structured malformed value raised nothing, trading an unwithdrawable blocker for a missing one.
- Row 162 — absence authority. Only an explicitly stated absence authorizes overwriting stored data with "absent"; a record's shape, a missing key, or a producer's default
Nonenever does, and where no producer states it the clear stays closed. Measured: four consecutive review rounds on one rule, each fix creating the next, because a document that found an owner but could not read a stake emitted exactly what an assessed absence emitted. - Row 163 — sequence claims. A number from a shared sequence (ADR id, migration revision) is claimed by the open pull-request set, not by mainline. A hole scan cannot serve: bounded by its own input, it sees only the interior, never the tip where every contested number lives. Measured: two collisions in two days — a migration revision, then three consecutive ADR numbers held by open PRs and invisible on mainline.
- Row 164 — the trigger, because judgement is what fails. The canon already carried the general form (row 160). It was broken three times in one day by its own author, hours after writing it, so the instruction is deliberately syntactic: a sentence carrying a consequence word — live, broken, deleted, idling, every, exposed — names the command whose output produced it, in the same breath, or is written as a hypothesis. Cheaper than a rule about care, because it does not require noticing that care is needed. In all three instances a nearby real measurement existed and its authority was spent on a question it never answered: proximity to a real measurement is what makes an unmeasured claim feel measured.
Declared gaps. The evidence under 161 and 162 is one repository over two days and must not be quoted as more. The canon grows by four here; two offsetting retirements (merging the two gate-cadence rows, and retiring the always-exit-0 doc-staleness rule with its hook) were deliberately separated because each touches a live tier table and a shipped kit file, and bundling them would be the scope creep this canon exists to prevent.
3.7.0 (2026-09-08)
One new section, from measurement on the flagship project: §5.9 — What Rate-Limits Two Sessions. It partly corrects §5.8. Doc-only; no new executable, no change to the install surface.
- The throughput case for two sessions is REFUTED, not merely unproven. §5.8 warned against pricing the pattern as double throughput; the measurement settles it. Over 60 pull requests: open → last review round 9 min (6% of PR life), CI critical path 15 min overlapping review, and last review → merge 63 min — 62%. Review and CI together occupy about a quarter of a PR's life; the reviewer answers in eight minutes and 15% of PRs merge with no review at all. The dominant term is idle time on work that is already reviewed, green and mergeable. At peak the project held 15 PRs open simultaneously at ~0.9 merges/hour, and the peak day — 35 merges — predates the second session's existence. The binding resource is agent attention across simultaneously-open PRs, not compute. What justifies §5.8 is therefore a severity asymmetry — unrecoverable harms avoided against recoverable hours — and never a rate.
- Complement on reading, collide on writing (row 157). Every catch that mattered came from one session re-deriving what the other had convinced itself of; every collision came from both writing — two stale-base near-misses, a squash-merge that broke a stacked PR, one wasted work-block. So: one session owns mainline and merges, the other reads, challenges, and never holds a base another stacks on. Split roles, not work.
- Peer agreement carries no evidential weight (row 158). Measured: 2 of 3 corrected false claims had already been "confirmed" by the second session. This is the mirror of §5.8's rule 155 — disagreement is settled by counter-measurement, and agreement settles nothing at all. A second session that agrees has told you only that it read the same way.
- Merge starvation is a property of the SEQUENCE, not the PR (row 159). The merge gate is a per-PR rule, applied correctly one PR at a time by both sessions, and neither asked what its own merge did to the other's in-flight run. Measured: mainline moved twice in forty minutes while a PR with zero open findings and 19/20 checks green sat blocked on CI predates current mainline; three re-run cycles gained nothing, because a 15-minute run cannot survive a mainline moving every 20–35 minutes — and the re-run was the correct lever, since merging mainline in resets review-at-head. Convention: no merge within ~20 minutes of another session pushing an otherwise-green head; when both want mainline, whoever is closer to done goes first.
- Adversarially review the PREMISES, not only the reasoning (row 160). Four independent analyses, each adversarially challenged, all built on one unquantified premise — "the deployed build is behind mainline" — and none asked for the number. It was seven commits, not the seventy-nine pull requests the framing implied, which turned a frightening decision into a twenty-minute one. Same failure family as a diagnosis carrying an unmeasured blast radius (§7.7), one level up: the fact everybody accepts is the one nobody measures, and consensus is where it hides.
- Two gaps declared. The idle tail is measured but its cause is inferred — the correlation with open-PR count is strong and no experiment held one variable fixed, so a team that bounds work-in-progress and sees no change should report that rather than assume the figure travels. And as with §5.8, this is one project over three days and must not be quoted as if it spanned months.
- Register 155 → 160 rows;
skill:s4u-loop-dispatch14 → 18,skill:s4u-code-review+1; skill subtotal 68 → 73.
3.6.0 (2026-09-07)
One new section, from measurement on the flagship project: §5.8 — Two Sessions on One Backlog. Doc-only; no new executable, no change to the install surface.
- §5.8 — a second session buys a second READ-REACH, not double throughput. Two full agent sessions (not subagents — §5.4 dispatches subagents that inherit the controller's brief and report into its context, so the controller's blind spots propagate down) run concurrently against the same tracker and remote from separate checkouts, with no shared memory. The argument is a limit on what a negative control can prove: a negative control proves a mechanism within the guard you looked at, and says nothing about the guard you did not read. A session that mutates its own change and watches the test go red has proved what it examined (appendix-a §14); it has proved nothing about the consumer it never enumerated. That gap is a property of the reading, not of the code, so it is undetectable from inside — more diligence searches the same neighbourhood harder. It is the §7 gate-admission rule turned on the reviewer: a reading verified only by the reader who made it is indistinguishable from a complete one. Measured over one working period, 2026-09-07: 5 paths in the merge instrument where an unreadable answer was read as green — found by the other session, in the instrument being used to merge at the time; 2 merges stopped after an independent re-reading (a missing retention guard; a register change that would have turned the mainline red); and 1 completed, tested and mutation-checked fix that proved to be a regression on enquiry and was never committed — it would have raised a high-severity finding on every clean case, and the verdict layer requires zero open high findings to release, so a correction to a false clear would have blocked every release in the product. Two further measured properties: the blind spots are not correlated (both sessions reported a wrong measurement on the same day, each refuted by the other with a counter-measurement rather than an opinion), and it corrects the review layer too — one finding the §7.7 adversarial reviewer had explicitly rejected proved real on re-measurement.
- The cost is stated with the benefit. The heaviest pull requests of that period carried 28, 11 and 8 review rounds. Effort moves from producing to adjudicating, and a team that plans this as twice the throughput has mispriced it. Three rules ship with the section (register rows 153-155): one issue is claimed by exactly one session with the git worktree as the isolation boundary; every load-bearing exchange is repeated on the issue or PR with its measurement; disagreement is settled by counter-measurement, never by seniority.
- Two gaps declared rather than implied. Rule 154 is asserted, not mechanized — the sessions have a direct private channel and nothing forces an exchange onto the board. Measured: one closed pull request whose reason exists only in that channel, so read from the tracker it is a closure with no stated reason. Every merge, review and finding is on the board; the intermediate reasoning is there by discipline, and wiring that repetition into the pipeline is the obvious next mechanism and does not exist yet. Second, the pattern is recent — these figures cover one working period, not the project's life, and unlike §7.7's adversarial-review measurements they must not be quoted as if they spanned months.
- Register 152 → 155 rows;
skill:s4u-loop-dispatch11 → 14; skill subtotal 65 → 68.
3.5.1 (2026-08-04)
Not doc-only — it started that way and did not stay that way. One anti-pattern named with cross-project measurement, one honest divergence flagged, AND a new executable: gen-memory-index.sh, installed into .claude/scripts/ by setup/bootstrap.sh and covered by 8 harness checks. Its --check is a real detector, but it is NOT wired as a gate — nothing invokes it automatically, so it is a tool an operator or a hook must call. Stated because a release that grows the install surface and calls itself doc-only hides exactly the change a maintainer needs to see.
- §6 + appendix-d — the hub is a topic index, never a chronological journal. The prescribed MEMORY.md format was always topic-keyed, but nothing said the journal shape was wrong, and it is the shape a long-running project drifts into: each session has something to record and appending is the obvious move. It fails structurally rather than stylistically — a topic index has a fixed number of slots, a journal grows by construction, so the hub crosses the 24,000-byte cap on a schedule set by how often you work. The subtler cost is orphaned spokes: a journal compacts by merging dated entries and takes the
[[pointers]]inside them along, whereas a topic hub compacts inside a slot that stays. Measured across three mature projects (2026-08-04), spokes with no hub reference: trust-relay-workflow 27% (155 spokes, topic sections), zol-rag 22% (131, topic sections + history last), ratiba 53% (113, chronological status log). The journal-shaped hub was also the smallest of the three (17.1 KB vs 21.5 KB), so the shape shows up as lost pointers, not as size — which is why a byte-budget hook alone never surfaces it. Shipstemplates/scripts/gen-memory-index.sh, installed by the bootstrap to.claude/scripts/(harness 181 → 189, and the kit-drift audit extended to cover installed project-side scripts), because the canon called the index auto-generated while no generator existed — adopters would hand-maintain it and it would drift, recreating the problem. Two remedies, both cheap: drain the narrative — promote each entry's durable lesson into the standing sections and discard the rest, because dated milestone and status narrative isgit log, PR history and current task status, all already forbidden by §6 (relocating it to ahistory_<year>.mdspoke moves the staleness out of the budgeted hub without satisfying the rule; a dated spoke is defensible only as a short-lived staging area with a recorded deletion date, and genuinely human-facing narrative belongs in repo-versioned docs); and keep an auto-generatedindex_all_topics.mdlisting every spoke with its frontmatterdescription:, linked from the hub ONCE, as a navigation and audit aid. Ratiba's hub went 17.1 KB → 11.4 KB (73% → 47% of cap) with no durable content lost. Scope correction: an unreferenced spoke is unreferenced, not unreachable — §4'sdescription:lives in the spoke itself and is what drives relevance matching, so a spoke the hub never names can still be recalled. Losing a pointer degrades curation and deliberate navigation, not retrievability; the earlier "61 unreachable spokes → 0" phrasing overclaimed and the index is not a loading mechanism. - appendix-d — the 250-character bullet rule diverges between prose and hook. §6 scopes the limit to the perishable active-work entries;
templates/hooks/memory-budget-check.shapplies^- .{250,}to every top-level bullet, durable sections included. The hook is the stricter reading and the one that actually fires, so it is documented as operative — a long durable entry becomes a lead line plus indented sub-bullets, which reads better in a hub anyway. Flagged rather than silently reconciled in either direction: the prose and the hook disagree, and only the author of §6 knows which was intended.
3.5.0 (2026-08-03)
Operations enters the canon — who reviews, how a deploy refuses, and what a gate costs to run — alongside the repair of four claims the repository itself refuted. Every section is written from measured evidence in this canon's own projects and carries its declared gaps rather than implying coverage it does not have.
- §7.7 — Codex is named as the review function, because the measurement names it (#42). Earlier revisions wrote "Codex, CodeRabbit or equivalent", which reads as a menu of interchangeable options. Measured over every inline review comment the flagship repository holds: 1,153 by
chatgpt-codex-connector[bot]against 13 by a human — 89:1 — across 348 pull requests that received findings (median 2 per PR, max 24), median 1 review round, 14% needing more than one and 4% more than three. On an agent-built codebase the bot is not a supplement to human review; it is the review function, and the rule that follows is "merge when the reviewer has gone quiet on the current head", not "merge when CI is green". The distinction is load-bearing: bot review is a review, not a check — it posts findings on the diff and reports no status, so it blocks nothing, and a bot that stalls or rate-limits looks exactly like a bot that found nothing. An adopter running a different reviewer should substitute freely, but should not read the old plural as evidence that the choice is unimportant or that a human is covering the same ground. Setup is four steps, none of whichsetup/bootstrap.shcan do for you (it makes zero GitHub API calls, the same limitation §7.6 records for rulesets) — so the canon says to record that the app was installed, because nothing downstream can detect it. - §7.8 — Deploying to a Single Server (#39). Two co-tenant systems on one VPS, no orchestrator, no registry, no platform team: the deploy script is the control plane, and the refusals are where the design lives (47 committed deploy logs end in success 41 times, in an explicit refusal twice, in an aborted build once). The load-bearing rule is health is reachability; verification is a separate step that can fail the deploy — one project's
/healthreturns 200 against a schema several revisions behind, so it runs migrations after health passes and treats "did not reach head" as a failed deploy. Four hazards are recorded because each was paid for live: Compose does not persist-facross invocations, so re-running with one file fewer silently reconfigured a service and unbound port 443 whiledocker psand the health check both looked fine; a shared external network gives every service its own name as an alias and analiases:list is additive, so two stacks each definingkeycloakcontended for the bare name and login traffic resolved into the wrong identity provider — invisible to both stacks' health checks, because both stacks were healthy; shipping an image over the wire cost ~8 GB and 30+ minutes against a ~5-minute build on the box, and building on the target also fails before anything is recreated; and rollback targets what was actually running (docker inspect --format '{{.Config.Image}}', captured before the new image goes live), never a:previoustag that does not exist. Gates run in the order they were learned — lock, CI green for the SHA being deployed, an in-flight-work pre-check, and an override ledger the monthly census reads. Declared gaps, stated rather than implied: no restore has ever been recorded, so these projects have backups and not a proven recovery; the shipped "zero-downtime" release is a short hard cut behind a health wait; alerting routes nowhere until real contact points are configured; and neither project deploys from CI. - §7.6 + §5.7 — three
Core-mechanizedrules had no mechanism (#37). v3.4.0 split the Core tier into mechanized and asserted; this closes the three whose mechanism turned out to be a repository setting nobody had been told to create. Newscripts/check-repo-config.sh(222 lines) reads the live repository configuration, plusCODEOWNERSand the CI wiring. The general rule is §5.7 — an agent configures, it does not instruct: a console click-path cannot be verified (nothing reads a UI), cannot be reviewed (there is no diff), and rots silently when a vendor moves a menu. Where a vendor ships a CLI, the agent runs it. Worked from this canon's own maintenance, including the trap that deploying is not publishing — a Cloudflare deploy without--branchmatching the production branch creates a preview at a hash URL and leaves the live site untouched, which looks like success. Also lands the controller is not tiered: everything in §5.6 allocates models to subagents, while the controlling session runs at the highest available capability and effort as a standing choice — not because quality is worth paying for, but because the controller's output is not code (it is decomposition, dispatch briefs, adjudication, and the judgement about what has actually been verified), and because cheaper models take more turns on multi-step judgement, so a per-call saving is routinely a wall-clock and total-spend loss. - §7.9 — gate cadence is part of what a gate costs (rules 149-150). The canon required that a gate observe something (§7.3) and not claim more than it checks (§14), and said nothing about how often it should run. Measured on a flagship project, 2026-08-03: ~84 billable minutes per push, ≈ $0.67 a push, ≈ $69 in seven days — roughly a third of it buying nothing. The test is whether the gate's input can change between runs. Three diagnoses, each running on every push since it was written: a full-history secret scan whose two known findings live in git history (~260 minutes a week); two SAST jobs uploading to a code-scanning API that returns 403 on a private repository without Advanced Security, their results surviving only as an artifact nobody opens (~1,470 minutes a week — while the project's own ADR already recorded that one "is not an active control", the register and the invoice disagreeing); and the same container image built three times per push, because two jobs each spun up a runner to look inside an image a third already had in its local daemon. Every gate now declares a cadence next to its blocking status, and "every push" is a choice to be earned rather than the default. Corrected in review before merge (Codex on #49): the first draft of the cadence table put the full-history secret scan and the licence/notice re-measure in the weekly tier, and both were wrong for the same reason — the subject was confused with the finding. Every new commit extends the history a full scan covers, and a PR can edit the dependency manifest a notice re-measure reads; what was invariant was the two findings already in history, not the scan. The correct move is the opposite of narrowing — keep the per-commit scan and widen it to the push range as well as the PR range, since a direct push under an admin bypass is exactly what a PR-range scan misses. The general form is now canon text: "this check keeps returning the same answer" is an observation about recent outputs, not a proof about the input. The cost figures also gained a re-derivation method (the
ghcommands, the round-up-to-the-minute billing rule) and an explicit statement that no run log is retained, so the totals are not auditable a month later and the method is what must survive. - The installer destroyed local edits, and four documents claimed things the repo refutes (#43-#48).
setup/bootstrap.shwas documented as "additive — never clobbers existing files"; it is now checksum-manifest driven (.claude/.s4u-manifest): a kit-owned file you have not edited is upgraded in place, one you have edited is left alone and the new version lands beside it as<file>.s4u-new,CLAUDE.mdand.claude/settings.jsonare never overwritten, and--dry-runprints every action without writing. Gate scripts are not installed into the target — they run from the kit, which the docs had also got wrong. Two Evidence lines were withdrawn rather than defended: the Brainstorm Gate does not "refuse to clear" (no detector enforces it; §14 recordsnone (human-asserted), and a missing Pre-Mortem Block is a person noticing an absence), and the pre-push hook did not enforce tests or coverage (sections 6 and 7 shipped commented out, and it exited 0 with nobackend/present). - Then the withdrawn claim was mechanized instead of left withdrawn (#47).
templates/hooks/pre-push-gate.shnow enforces tests and fails closed in both directions of not-knowing — exit 2 when it cannot find its source root, and exit 2 whenS4U_TEST_CMDis unset, rather than guessing a runner, finding nothing, and reporting success. Verifiable directly:S4U_TEST_CMD='sh -c "exit 1"'against the hook returns 2. Coverage stays genuinely optional and is labelled so. The pair is the point — #43-#48 made the documentation honest, #47 made the honest version stronger than the false one had claimed. - A version bump now requires its own changelog entry. This entry was itself missing:
check-version-header.shverified that the named changelog file exists and that the site navbar agrees with the canon — both true of v3.5.0 while nothing recorded it, so six commits shipped under an undocumented version with every check green. That is the adjacent-property shape §14 exists to name. The check now requires a## <version>entry in the changelog it names, and the fixture it runs against was itself found carrying the defect (acanon-goodmethodology declaring3.0.0-devbeside an empty changelog). - Harness 138 → 181 checks; the rule register 137 → 150 rows (138-150 added). Every new mechanism ships with a fixture that fails when the rule is broken.
3.4.0 (2026-08-02)
Canon self-audit remediation — the methodology's own §7 gate-admission rule, applied to the methodology (epic #23, findings #24 / #25 / #27). Three claims the canon made about itself were false; each is now either true or honestly labelled, and each is backed by a check that fails when the claim stops holding.
- §14 Core tier split into
Core-mechanizedandCore-asserted(#24). The Core row listed eight components and justified them with "trustworthy without relying on any individual's discipline" — while five of the eight had no detector anywhere in the kit. That is a claim about control SHAPE presented as a claim about control STATE. Core is now two tables, each row carrying the three fields §7 demands: per-occurrence cost, enforcement mechanism (a resolvable path orrepo-config:control, else the literalnone (human-asserted)), and retirement condition — the mechanism that would move an asserted row up. The rationale sentence no longer covers both halves uniformly. R1-R3 keep their single home on the operating card (correct under §4.5) and §14 now gives the pointer it was missing. Mechanism:scripts/check-tier-mechanisms.sh— fails if aCore-mechanizedrow names nothing that resolves, if aCore-assertedrow claims a mechanism it does not have, if any field is blank, or if a named mechanism file does not exist. - The single-source rule now has a mechanism that implements it (#25). §4.5 named
check-canon-consistency.shas the detector of project-copy drift, and repeated it under "Reproducible" — but that script only ever checked the canon's internal consistency and never opens a project file. Newscripts/check-single-source.shreads a project'sCLAUDE.md/AGENTS.md, shingles every declared rule home (operating card, appendix-m,skills/*/SKILL.md), and fails on a verbatim run of ≥12 words carrying no deviation ADR — namingfile:lineand the copied text. Allowed copies are explicit: anADR-NNNNreference in the same block, or an<!-- s4u-allow-copy: reason -->marker. Its limit (verbatim only, not paraphrase) is stated in the script header rather than implied by the canon. Run live against the flagship project on its first outing, it found a real violation.appendix-m's "copy the mandatory + default rows into ADR-0001" instruction — which mandated creating the copy §4.5 forbids — is corrected to pointer-plus-deviations. rule-inventory.mdis a register again, and is checked (#27). It was incomplete in both directions with no detector: §3.5 (Incident-Response) and all of §15 (Securing the AI Collaborator) had zero rows despite predating its own audited version; four live operating-card rules (the probe-gated flag flip, the public-doc classification gate, the mandatory STATE.md/memory cadence, the Core tier split) had none; row 9 still asserted a blocking doc-staleness gate that v3.1.2 had retired to advisory; the header was pinned at v3.1.2; the stated extraction count (120) disagreed with the table (123); and the size projection cited a ≤20,000-byte target against an enforced cap of 11,500. Twelve rows added (124-135), row 9 rewritten to the advisory form, rows 38/43/123 corrected, counts and the byte section re-derived from measurement. Sections with genuinely no normative rule (§1, §4.4, §11.5) are now declared with a reason instead of silently absent. Mechanism:scripts/check-rule-inventory.sh— six deterministic checks (version pin, leaf-section forward coverage or explicit declaration, dangling §-refs, card mechanisms registered, bolded card rules registered, arithmetic).templates/scripts/flip-flag.shgets a spine home (§7, "Gate: probe-gated flag flip") — it shipped and was mandated on the card, but appeared nowhere inmethodology.md, so no register row could cite it.- The four reviewer templates now carry a "Test evidence" checklist block (
MOCK APPROVEDpresent, no internal-class mocking, migrated-schema oracle). §8 claimed a bare mock "is flagged at review" while none of the reviewer definitions mentioned mocks; the block is the artifact §8 now names, and §8 states plainly that this is review, not a gate. - Harness 117 → 138 checks; three new gates wired into
canon-ci.yml. Every new mechanism ships with a mutation-proved fixture: break the rule, the named check fails.
3.3.0 (2026-08-01)
Two rules the canon was missing, both found by applying the methodology's own §7 gate-admission rule to the methodology (audit epic #23), plus a false-caveat correction:
- §7 field (d) — a gate must prove it observed something (#26). The gate-admission rule required a mechanism but never required the gate to see anything, so a named, green job that examined nothing was indistinguishable from one that examined everything. Field (d) has two halves. Demonstrated failure (admission time): every shippable gate script declares
# failing-case: <suite> "<test name>"(or a stated-reasonnoneexemption), andscripts/check-test-coverage.shFAILS when the named test is not present in that suite — so the claim cannot drift away from the test it cites (22 shippable scripts examined; 17 name a real firing test, 5 are declared-exempt with a stated reason). Work witness (run time): a gate that enumerates a subject set prints the count it examined and FAILS at zero — implemented incheck-adr-register.sh(examined 0 ADR files),check-doc-classification.sh(examined 0 markdown files) andcheck-test-coverage.sh(examined 0 shippable scripts), each with a fixture pinning the empty-subject failure. All three previously exited 0 over an empty register/corpus/script set. The corollary is now canon text: a guard verified only where it passes is indistinguishable from a guard that cannot fail. Card + PR template +s4u-code-review(name-the-count and name-the-failing-case become review lines; the oracle table gains a work-witness column) + rule-inventory row 136. - Mutation discipline (#29) — appendix-a §14, new. The testing standard (995 lines) did not mention mutation testing, and had no answer to "can this test fail?" beyond an unmechanized third-party mental check that PoC mode's tests-after ordering removes the RED phase from. Adds the three-outcome contract (KILLED / SURVIVED / INCONCLUSIVE, where INCONCLUSIVE is never a pass), the rule that DATA files are first-class mutation targets (21 code-side mutations once all killed their tests while the guard they protected was inert, because its subject was a JSON file), and the triage rule that a survivor means suspect the TEST first (nine of nine survivors in one session were defective tests). Mechanism:
templates/scripts/mutation-probe.shasserts five preconditions before scoring — baseline collected >0, baseline green, anchor occurs exactly once, bytes changed, mutated run still collects >0 — and--collect-cmdis required and fail-closed, because without a collection oracle "the test passed" cannot be told from "no test ran". Whole-suite mutation scoring is explicitly recommended (not yet enforced). Card +s4u-testing-standard+ §8 principle 5 + rule-inventory row 137. - ADR skill §9 caveats were false (#28). Both "matching caveats" were wrong against the current
check-adr-register.sh:ADR-NNNN-*.mdfiles are scanned (-name '[0-9]*.md' -o -name 'ADR-[0-9]*.md', lines 14/28/43) and the back-reference match does tolerate bold (\*{0,2}Supersedes\*{0,2}: ?ADR-<n>). Commit7aaf8b5fixed the script and wrote the pre-fix prose into the skill in the same commit; two later edits left it standing. The caveats steered agents to rename a register away from the flagship's own working convention — a stale warning that reports a working control as broken suppresses the control. Replaced with execution-verified behaviour plus a standing instruction to re-verify by running the script, never by reading. - Harness 117 → 138 checks (mutation-probe contract ×11, work witnesses ×4, meta-gate failing-case + zero-subject ×6). Verified by mutation: six mutations of the new mechanisms, six kills, no survivors.
3.2.3 (2026-06-16)
A worked-example generator for the §7.5 freshness gate — ADR publishing (canonical → site):
templates/scripts/generate-adr-mirror.sh— renders the canonical ADR corpus (docs/adr/ADR-NNNN-<slug>.md, owned by appendix-g) into a Docusaurus docs tree asNNNN-<slug>.mdpages (frontmatterid/sidebar_position= ADR number + 1 /titlefrom the H1, body verbatim). Closes the "published ADR mirror lags canon" drift class — observed in the field where a site's ADR section stopped at 0049 while canon reached 0066 (each new ADR needed a manual copy + sidebar edit). Run it as a build/start prestep with the ADR sidebar set to autogenerated, so new ADRs reach both the docs tree and the nav with no manual step.- §7.5-conformant by construction: takes its output dir as
$1, rewrites only theNNNN-*.mdpages it owns (authored siblings cancel out of the diff), and is deterministic (C-locale glob order, title read from file, quote-escaped YAML, no timestamps/randomness) — so it is pinnable bycheck-generated-fresh.sh(verified: FRESH on a clean tree, FAIL with drift listed on a mutated page; quote-laden titles escape correctly across 65 real ADRs). Full-generate, single source of truth: canon is authoritative; a project with divergent hand-curated published ADR pages reconciles into canon first (appendix-g: accepted ADR bodies are immutable) rather than running a full-generate gate over divergent copies. - Documented in appendix-h (new "ADR Publishing (canonical → site)" section, sibling to the architecture index). Tier: Recommended (§14) — adopted where a project publishes ADRs to a living-doc site; the generator is generic (canon dir in, Docusaurus dir out). Fixture-tested in
tests/run-checks.sh(deterministic output + the §7.5 gate reporting fresh-vs-stale over the generator).
3.2.2 (2026-06-15)
New §7.5 blocking tier — generated-artifact freshness (durable fix for the recurring "generated copy lags source" class behind the 3.2.1 hotfix and the Codex findings on PRs #13/#16/#18/#19):
templates/scripts/check-generated-fresh.sh— a generic regenerate-and-diff gate: seed a throwaway dir from the committed generated tree, run the generator into it (generator takes its output dir as$1and rewrites only the files it owns, so authored siblings cancel out of the diff), then diff. Any difference = the source changed without regenerating → block. Ungameable (runs the real generator, not a hand-maintained file map) and self-correcting (the fix is "run the generator and commit"). Fail-closed on missing/!executable generator, missing out dir, generator error. Generic by construction — pins a Docusaurus site, generated protobufs, or an OpenAPI client alike.- Blocks immediately, no warm-up flag (unlike the heuristic
DOC_SYNC_BLOCKINGcode↔doc gate): regenerate-and-diff has a structurally near-zero false-positive rate provided the generator is deterministic.website/scripts/build-docs.shgained an optional output-dir argument (default unchanged) so the gate can target a temp dir without mutating the tree. - Compares content and the executable-bit set (Codex P2 on PR #21:
diff -qis content-only, so a generated script that lost its+xbit would pass on identical bytes; the gate now diffsfind -perm -u+xtoo, and seeds withcp -Rpso the comparison is faithful). Matters for generated scripts/wrappers, not the markdown site. - Wired into
.github/workflows/canon-ci.yml(real-tree, authoritative) and thepre-push-gate.shtemplate (documented optional block §8); fixture-tested intests/run-checks.sh(a trivial deterministic generator + matching/stale trees + a mode-drift case + fail-closed branches). Harness 102 → 113. Operating card + rule-inventory (row 123) updated.
3.2.1 (2026-06-15)
Bot-review follow-through (Codex review on PR #18, two findings merged-over — caught in a post-merge sweep):
- P1 — incident routing contradicted its own cycle. The §3.4 trigger table's Incident-Response entry point read
reproduce → root-cause → regression-pinning test, routing responders to diagnosis before service restoration — directly contradicting §3.5's "Mitigate before you diagnose." The entry point now leads with mitigate first (flag-off / revert), matching the cycle's reverse-lever-first posture. - P2 — stale generated site page.
website/docs/lifecycle/lifecycle-integration.mdstill showed the pre-v3.2 6-row trigger table (missing the Inception and Incident-Response rows) because PR #18 edited canon after PR #17's site regeneration without re-runningbuild-docs.sh. Regenerated from canon. Underlying class (canon-PR-merges-after-regen-PR) noted for a future same-PR doc-sync gate.
3.2.0 (2026-06-14)
Two structural additions (design: specs/2026-06-14-v3.2-inception-stage-and-doc-sync-design.md; second-party reviewed across two rounds; deep-research-backed):
- §3.6 The Inception Cycle — an optional fifth lifecycle for new projects / new bounded contexts: arc42 inception canvas + ≥3 quality-attribute scenarios + risk-storming (security lens) + per-system trust-boundary/STRIDE threat enumeration + ≥1 fitness function; gated on artifact presence (substance via the §3.1 second-party threshold). Templates in
templates/; full reference inappendix-o-inception.md. Fixed the stale §3 intro ("three lifecycles" → five) and fanned the cycle intos4u-lifecycle(which had also been missing the §3.5 Incident-Response row). - Tiered documentation-sync (§7.5 + §11.4) — doc-sync becomes a tiered fitness-function: scoped code↔doc-pointer drift + link integrity + ADR-register integrity BLOCK (fail-closed, via
templates/scripts/check-doc-sync.sh+ adoc-pointersmanifest + a[skip-docs:]escape hatch logged for the census); blanket 30-day staleness + prose/Diátaxis stay advisory — preserving the v3.1.2 honesty fix (a blanket age-gate surprises contributors).s4u-doc-excellencerewritten to the tiered model;canon-pinned-defects.tsvpins the retired "advisory-only" universal claim, disjoint from the kept blanket-tier text (the blocker the two-round review caught). The §2.8 census now surfaces the skip-docs log + manifest coverage (silent-decay defense). The scoped gate ships behindDOC_SYNC_BLOCKINGuntil a sub-10% false-positive dry-run; probedocs/probes/doc-sync-blocking-2026-06-14.txtreports 5.0% (PASS). - Operating card + rule-inventory (rows 121-122) updated to match; byte budget held. The Docusaurus site-format template ships separately (
specs/2026-06-14-docusaurus-site-template-design.md).
3.1.2 (2026-06-13)
Adversarial-review fixes (regressions from the heavy 3.1.x editing + structural hardening):
- Fan-out regression fixed: the 6→7 Decision-Cost axis change had reached the
canon spine but NOT the skills agents load (
s4u-lifecycle,s4u-adr) — their Pre-Mortem / Decision-context block formats omitted theCost:line, so an agent following the skill emitted a 6-axis block. Now seven everywhere; therule-inventoryand CHANGELOG editorial note are aligned too. - Fails closed now:
check-canon-consistency.shscansskills/*/SKILL.md+rule-inventory.md(was methodology.md-only — the hole the 6/7 drift fell through), with a pinned-defect on stale axis counts. The meta-gate's "tested" check now requires a non-comment invocation (a false# no-fixture:can no longer hide an untested script). - Honesty drift corrected: prose that asserted a blocking doc-staleness gate
(§2.5, appendix-h) downgraded to advisory, matching the shipped exit-0 hook;
the STATE.md staleness claim is now true —
generate-state-md.shemits a staleness date andcheck-doc-staleness.shsurfaces a STATE.md >30 days old. - Two real script bugs:
methodology-health.shgrep -c || echo 0crash fixed;consolidation-census.shborn-date pickaxe anchored to a word boundary (was matchingauth_enabledinsideoauth_enabled, inverting the age-rank). Both now invoked by the suite. Harness 92 → 102. - New §3.5 Incident-Response Cycle (detect → mitigate via flag-off/revert → root-cause → regression-pin → named-class postmortem → memory+STATE.md) — the missing reverse lever for a production posture. Worked Cost-axis example added.
pr-review.ymlAI-review re-gated via aneeds-output job (the job-levelif: secrets.*pattern is unreliable); bootstrap now installs the doc-staleness hook; product-scale-planning skill gains a §5.6 crosswalk.
3.1.1 (2026-06-13)
Audit-remainder polish (no rule changes; consistency, coverage, and enforcement):
- Canon links: repointed 18 dead/circular links into the redirect-stub appendices (b/c/f/i/k/n) to their real homes (§5, §13, the relevant skills, showcase.md). §7.1 reframed "Three-Layer Defense" → "Three Core Gates + Optional Security Layer" (count now consistent). §5.2 column rekeyed Model → Capability Tier (tier is the contract; model names are point-in-time bindings). Operating-card budget stated byte-cap-primary (the enforced unit).
- Skills now carry their mapped procedures: reviewer-dispatch mapping (s4u-code-review); R1–R3 silent-failure discipline + the 90%-in-PoC core floor (s4u-testing-standard); stale-test classification + convergent-design signal + a Superpowers prerequisite line (s4u-lifecycle); always-on subagent-dispatch hygiene + the four-status protocol (s4u-loop-dispatch); single-source pointers fixed to survive a bootstrapped project.
- Gates that now gate:
expect_errtest helper (asserts the diagnostic, not just the exit code) + fixtures for previously-untested high-stakes branches; a shipped advisorytemplates/hooks/check-doc-staleness.sh; CI (canon-ci.yml) now runs the operating-card version header + the adr-register / doc-classification / kit-vs-project integrity gates it previously skipped; canon-consistency gains a redirect-stub-link flag; consolidation-census age-ranks flags by born-date + documents the manual retirement gate (honest over a flaky stateful auto-fail). Harness 77 → 92 checks. - Adoption/site: bootstrap writes a
.s4u-kit-versionstamp (upgrade story)- points adopters at
.claude/agents/*.md; reviewer-agent frontmatter notes the capability tier above eachmodel:line; the site gains local search (@easyops-cn/docusaurus-search-local) and a navbar version indicator.
- points adopters at
- Mandated STATE.md + memory cadence (explicit): the operating card now carries a crisp rule — every finished branch regenerates STATE.md and updates the relevant memory files in the same commit; memory also updates on any decision/correction/new-pattern mid-session; STATE.md stale >30 days is a defect (flagged by the doc-staleness hook). §3.1 Finish & Merge and the §11.4 documentation-commit pattern (now a 5th required artifact) state it as non-optional, confirmed at review.
3.1.0 (2026-06-13)
- §15 "Securing the AI Collaborator" (NEW spine section) — a threat model of
the collaborator: 15.1 untrusted-input-is-data (the behavioral rule marked
recommended (not yet enforced), with §15.4's permission gate as the enforced
backstop), 15.2 secret hygiene (→
check-doc-classification.sh), 15.3 tool/MCP least-privilege + provenance (→ advisoryscripts/check-tool-provenance.sh), 15.4 permission-mode tied to blast-radius. Safety sign-off: Tsunami-max, 2026-06-13 — trigger 6 cleared (seespecs/2026-06-13-methodology-hardening-design.md). - Decision-Cost Rubric → seven axes: added Cost (compute/token spend) to
§2.7, the §3.1 Pre-Mortem Block format,
templates/spec-template.md, and the rubric diagram. Closes the audit's "no economic axis for subagent-heavy work". - Machinery: self-policing meta-gate
scripts/check-test-coverage.sh(every shippable script must be tested or carry# no-fixture:); fixtures added for the 5 previously-untested scripts (harness 53→71); new project-agnosticscripts/methodology-health.sh(effectiveness trend instrument, no borrowed metrics) +scripts/adoption-smoke.sh(stack-agnosticism probe). The probe caught and we fixed a real external-validity gap:templates/project-claude.mdhard-coded Python/Postgres commands outside the{{...}}placeholders. - SETUP-GUIDE.md retired to
docs/archive/(it taught three killed patterns); inbound pointers repointed tobootstrap.sh.setup/ADOPTION-TRIAL.mdadded (external-validity protocol for outside adopters).
Unreleased — editorial (2026-06-13)
- Non-normative: added six Mermaid diagrams to the spine, visualizing existing prose (no rule changes): the documentation knowledge-flywheel (§2.5), the the Decision-Cost Rubric diagram (§2.7), the consolidation/subtraction loop (§2.8), the four memory types feeding session context (§6.2), and — in §3.4 — the skill-chaining pipeline and the subagent orchestration model. Authored for the standalone methodology documentation site; they propagate to all sync targets.
- Non-normative: de-projecting + readability pass across the canon. Removed
all references to specific real projects and their metrics/incidents
(named projects, frameworks, and incident anecdotes). Every
**Evidence:**paragraph that cited a project metric was replaced with a generic, mechanism-based statement (name the enforcing hook/script/gate; reproducible). Each section gained an inverted-pyramid lead (**In one line:**/**Do this:**); body prose tightened. No rules changed; all four canon gates stay green. Rationale: the methodology should read as a reusable, engaging, actionable standard — not a project case-study novel. Evidence is now demonstrated by runnable machinery, not borrowed metrics. - Audit-driven finishing pass (multi-agent audit, 57 verified findings). Gates:
fixed the BP-5 defect in the blocking
pre-push-gate.sh(wasruff check app/ --quiet; now fullruff check .+ruff format --check .), corrected its block exit code (1→2, the only code Claude Code PreToolUse blocks on), added python3 interpreter guards; broadenedcheck-canon-consistency.shso it actually validatess4u-*skill refs (the old regex excluded the digit4); boundary-anchoredcheck-adr-register.sh(was a substring match); extendedcheck-doc-classification.shto catch IPv6 / URL-embedded creds / cloud keys (AKIA/ghp_/glpat-/sk-/AIza), tuned to not false-match ISO-8601 timestamps; re-budgetedcontext-budgets.tsvonto the actually-loaded files. Test harness grew 40→53 (negative fixtures for each new branch). De-projection completed beyonddocs/: README, root CLAUDE.md, the five canon-mirroring skills, and appendix-l/j/a are now project-free (the earlier "removed all references" claim is now true repo-wide); front doors repointed from the superseded SETUP-GUIDE tobootstrap.sh. Spine: removed phantom/designing,/planningskill names, de-duplicated §13, fixed the §4.3 Python-floor contradiction. Site: removed dead Docusaurus scaffold + fixed the broken social-card ref.
3.0.0 (2026-06-12)
-
Added machine-readable version header + this changelog (PW-5).
-
§2.8 Consolidation Review written as a mechanism (census script + retirement mandate); was a dangling "(planned)" reference; cadence set to monthly (AF-1).
-
Canon consistency check shipped (the "mechanical greps" §5.6/appendix-l promised): phantom skill names fixed (/designing, /planning, /worktree, /finish, /code-review → installed Superpowers names, incl. both mermaid diagrams); freezegun three-way contradiction resolved (§8 defers to §4.5 default); §5.2 model-naming self-contradiction removed; §2.7 case-study timeline de-inflated ("month three"/"for months" → days, matching its own facts); /writing-skills added to the §4.4 table (was listed as 14, contained 13).
-
§3.1: SIXTH Brainstorm-Gate trigger — changes to safety policy / guard / refusal behavior gate hardest and require a literal human sign-off line (assessment meta-pattern C: the jury dosing incident entered through an APPROVED policy relaxation no gate covered).
-
§4.5 rule 1: verbatim-copy mandate REPLACED by single-source-plus-pointers; per-project canon mirrors retired; drift checking is mechanical (CE-5, PW-5).
-
§3.1 lifecycle v3: spec+plan merged into ONE design artifact (templates/spec-template.md); Plan Walkthrough RETIRED (zero recorded completions in 3 months) in favor of a second-party scrutiny threshold (safety path / schema / public API / auth -> fresh-context or human review); §7 gains three wrong-oracle defenses: live contract smoke (I13), migrated- schema oracle (I35), and the name-the-oracle review line (meta-pattern A).
-
§14 rewritten: Core tier = the enforced gates with documented saves (required CI checks, CODEOWNERS safety review, safety-floor evals, R1-R3); named-project specifics demoted to project-specific-with-ADR; permission mode reclassified preference -> security control; NEW §14.1 multi-dev operating model (org repo, deploy lock, generated STATE.md, incident roles) (TA-02/05/06).
-
§7 gate-admission meta-rule: every gate declares cost + enforcement mechanism + retirement condition; enforced via the canon PR template (assessment §6 standing meta-rule).
-
§6 + appendix-d: hub budget restated in BYTES (24,000 — the loader's unit); durable-first section ordering is policy under truncation; advisory memory-budget Stop hook shipped; MEMORY-template reordered (CE-2).
-
Operating card extracted (docs/operating-card.md, ~7.6KB / 42 rules of 119 inventoried in docs/rule-inventory.md) — the only always-loaded surface; methodology.md demoted to reference (CE-1/CE-4/PW-1).
-
Showcase split + appendix dispositions: §12 + appendix-f -> showcase.md; appendix-c merged into §5 (specialists now OPTIONAL — resolves the flagship-forbids-them conflict); appendix-i tables merged into §13; appendices b/k/n retired to skills; appendix-e rewritten — prompt-type blocking Stop hooks REMOVED at all sites (completion-loop failure mode), command-type diff-aware advisory is the standard; appendix-m's false project-stack claim corrected (CE-5); a/d/g/l carry superseded-by notes.
-
STATE.md policy: generated from git/gh or absent, never hand-maintained; generator shipped at templates/scripts/generate-state-md.sh (CE-8, TA-05).
2.3 (2026-05-12) — reconstructed
- §2.7 Decision-Cost Rubric; §3.1 Brainstorm Gate (Pre-Mortem Block).
2.2 — reconstructed
- §2.6 Investigative Discipline; appendix-n Documentation Excellence Passes.
2.1 / 2.1.x — reconstructed
- §4.5 canonical tech stack; §5.5 /loop pattern; §5.6 product-scale planning; appendices k/l/m.