Governed AI software factories
Executive assessment
AI-assisted software factories are a credible engineering direction, but neither unrestricted autonomy nor universal adaptation without programming is established by the available evidence. Specification-driven generation, external scenario evaluation, agent orchestration and configurable business rules already have public implementations. An integrated product must therefore be justified by the quality of its collaboration and control model, not by claiming to have invented these components.
The recommended architecture combines a business-intent workspace, a versioned engineering method and replaceable delivery adapters. Business participants approve meaning; agents perform bounded analysis and engineering; independent checks establish what was actually verified. Customer differences become configuration where the platform already supports the required behaviour, and reviewed extensions where it does not. This is an architectural recommendation, not a report of an already implemented system.
The principal opportunity is the continuity between evidence, stakeholder contributions, approved rules, engineering work and observed outcomes. A useful product makes that continuity easy to inspect and cheap to maintain. It also admits unknowns, preserves private exploration and identifies who can authorize a consequential change.
1. Terminology and scope
“Type 3” should be defined locally rather than presented as an industry certification. Dan Shapiro's published ladder calls its highest category Level 5, while StrongDM describes non-interactive development without adopting that same numbered taxonomy. These are practitioner descriptions, not interchangeable conformance standards. Shapiro's framework, StrongDM's account.
In this paper, a governed AI software factory is a delivery system that accepts an authorized, versioned mandate; allows agents to perform supported work within declared permissions; evaluates results against protected criteria; and returns evidence for an explicitly authorized release decision. An autonomous execution interval does not grant autonomy over business policy, access rights or production release.
Two separate dimensions matter. How software is built may be highly automated. What the delivered application does may be deterministic, agentic or mixed. An AI-built application need not contain agents, and an agentic application need not be built by an autonomous factory. Each needs its own assurance evidence.
2. Evidence and its limitations
Factory implementations
StrongDM reports a factory driven by specifications and scenarios, with evaluation cases separated from ordinary implementation tests and behavioural twins of external services. This is direct evidence that a team has described and demonstrated this approach, not independent evidence of universal reliability or economics. Its prescriptive token-spending and no-human-code-review positions should not become enterprise requirements without local validation. StrongDM.
Anthropic's long-running application experiments report improved outcomes from planning and evaluation around generation, but also remaining defects, substantial execution cost and the need to tune evaluators. The authors simplify the harness as models improve. The practical implication is to evaluate each orchestration component rather than assume that more agents always improve delivery. Harness design.
Paperclip's current documentation describes agent adapters, budgets, task ownership, activity history and human governance. Its heartbeat protocol also describes confirmation of a plan bound to the latest document revision. Consequently, generic agent management and revision-bound confirmation alone are not a defensible novelty claim. Whether Paperclip adequately supports confidential stakeholder exploration, domain-scoped approval and business-to-delivery fidelity requires a concrete fit test; absence from these pages is not proof the capability is absent. Core concepts, heartbeat protocol.
Productivity and organizational evidence
METR's early-2025 randomized study involved 16 experienced contributors and 246 tasks on familiar open-source repositories. AI access increased completion time by 19% in that setting. That result is neither a contemporary verdict on all frontier tools nor a measurement of the proposed factory. Study and limitations.
Its February 2026 follow-up explicitly reports selection and measurement problems that weaken estimates of current speedup. The later May 2026 survey of 349 technical workers reports median perceived value gains of 1.4–2×, while warning that self-reported counterfactuals and selection bias limit interpretation. These studies must be presented together rather than selecting either the slowdown or the optimistic survey as a universal multiplier. Experiment update, 2026 survey.
DORA's 2025 research characterizes AI as amplifying organizational strengths and weaknesses. Its findings support investing in the delivery system and user focus, but survey associations are not causal proof that adopting a particular platform will improve a customer's results. DORA report.
Evidence classification
| Evidence class | Useful for | Insufficient for |
|---|---|---|
| Practitioner implementation | Establishing feasibility of a technique in a described setting | General accuracy, enterprise safety or comparative ROI |
| Controlled task study | Estimating effects within its population, tools and workload | Every later model, team or workflow |
| Survey or observational study | Adoption, perceptions and organizational associations | A guaranteed causal productivity multiplier |
| Technical specification | Defining an interface, mechanism or conformance requirement | Proving a particular installation complies |
| Local acceptance exercise | Assessing the tested build and reference scenarios | Whole-market demand or untested production conditions |
3. Business intent and delivery architecture
The factory should not begin with an unrestricted prompt and discover its authority while working. It should begin with a bounded capability, intended outcomes, known constraints, source provenance, unresolved questions and named responsibilities. A business analyst reviews extracted interpretations before they become stakeholder requests. Domain owners approve the exact rule and example revisions that the implementation will be judged against.
The execution mandate is the handover contract, not simply a generated PRD. It identifies the approved intent and terminology, permitted repository and environment, method/profile version, integration contracts, expected evidence, budget, stop conditions and release authority. Its consumer must acknowledge the exact version. A changed baseline requires an eligibility assessment; it must not silently replace the inputs of an active run.
The return path is equally important. Delivery supplies artifact identity, changes, verification receipts, known gaps and proposed corrections. An implementation discrepancy does not authorize an agent to rewrite the approved specification. Business intent, technical design and actual behaviour remain distinct records with linked views.
A shared interface need not replace every enterprise system. It gives each fact and decision a designated authority. Documentation, diagrams and search are projections of selected revisions. Existing enterprise terminology, identity and delivery platforms can remain authoritative through explicit adapters.
4. Adaptability without arbitrary self-modification
Dynamic behaviour is achievable when the variation is expressible using capabilities the runtime already implements. A rule threshold, required field, review routing condition or supported workflow transition can change without changing application source. A new protocol, a new class of side effect or an unsupported algorithm usually needs engineering.
This is not a new foundational technology. OMG's DMN separates business decision modelling from surrounding processes and is designed to complement process and case models. CEL provides a constrained expression environment with parsing and checking before runtime evaluation. These are useful precedents; neither removes the need to define correct business meaning, constrain resources or secure the host application. OMG DMN, CEL overview.
| Change | Preferred mechanism | Boundary |
|---|---|---|
| Change terminology, thresholds, required evidence or routing within existing capabilities | Versioned configuration and business review | No source-code change, but still a tested behavioural release |
| Introduce a new approved workflow step using an existing action | Validated workflow definition | Must preserve authorization, recovery and in-flight semantics |
| Add a new provider/protocol or unsupported data operation | Reviewed adapter or product extension | Requires code, testing, compatibility and deployment |
| Require an unobservable fact or an unavailable model capability | Explicit unsupported state and investigation | No fabricated integration or silent fallback |
| Change security boundaries or approval authority | Controlled policy change with designated approval | Ordinary project configuration cannot weaken core protections |
The preferred product is a stable core with narrow, versioned extension points. Avoid implementing a universal BPMN editor, policy language, agent marketplace and generic application builder before a useful vertical slice exists. Begin with a few well-defined review primitives and add expressiveness when a real variation demonstrates its need.
The system should distinguish configuration-only adaptation of the collaboration workspace from configuration-only adaptation of the application it helps deliver. A workspace can express a changed rule, yet the target application's design may still require new code. Marketing must not collapse these two claims.
Controlled configuration lifecycle
Configuration should have a schema version, digest, parent revision, owner, effective time, compatibility constraints and evidence. Reject unknown fields and unauthorized actions rather than silently ignoring them. Approvals bind the exact package; edits invalidate pending approval. Missing inputs produce explicit unknown/error outcomes, not permissive defaults.
Signed bundles can establish origin and integrity, not business correctness. OPA supports signature verification on configured bundle-loading paths, but its documentation distinguishes those paths from commands that do not verify signatures. An adoption test must exercise the actual activation path. OPA bundles.
Existing cases should retain their selected workflow and rule versions by default. Moving them to a new version requires an explicit compatibility/migration decision. Runtime rollback does not reverse external emails, payments, disclosures or other irreversible effects; recovery may require compensating actions and human handling. Temporal's worker-versioning guidance illustrates the importance of preserving workflow-code compatibility across deployments, without prescribing Temporal for every product. Worker deployments.
5. Existing systems and integration boundaries
Existing-system discovery is an intake option, not the centre of the generic product. An organization replacing a slowly changing application can begin with a baseline, an urgent-fix notification agreement, release-triggered reinspection and scheduled reconciliation. Continuous streaming or change-data capture is justified only where the rate, risk and observability of change warrant it.
Distinguish source inspection, runtime observation, expert testimony and intended policy. Source code can reveal apparent rules without proving that the inspected revision is deployed or that those rules are correct. Incomplete excerpts must preserve their coverage limits. Existing tests characterize behaviour; business-approved examples determine which behaviour should be retained or changed.
An integration register should cover provider/consumer ownership, schemas and semantics, identifiers, authorization, timeout/retry rules, idempotency, ordering, duplicate handling, data classification, expected failures and deprecation. External API shape is only one part of the contract. An enum can remain syntactically valid while its business meaning changes.
Consumer/provider contract verification can help decide whether specific versions may be deployed together. Pact documents a compatibility-matrix-based deployment check; it is not a substitute for business-semantic acceptance, security or complete end-to-end testing. Pact deployment checks.
Behavioural twins can make failure testing cheaper and safer, but their fidelity is a separate claim. StrongDM describes comparing twins against real dependencies. Use a bounded, authorized sample of actual-provider contract checks to detect divergence; never call a simulation a verified live integration. Digital Twin Universe.
6. Verification across the complete lifecycle
Agent-generated tests are useful, but the same mutable context must not supply the only specification, implementation and verdict. Anthropic describes combining deterministic, model-based and human grading, calibrating model judges and testing both desired and undesired behaviour. Evaluation environments and graders can themselves be defective. Agent evaluation guidance.
The following is a proposed verification portfolio. Apply it proportionally to the capability and risk; do not interpret it as a mandate to run every expensive check on every edit.
| Lifecycle area | Evidence to establish | Important counterexample |
|---|---|---|
| Inception and BA | Source provenance, scoped terminology, owner-approved outcomes | A plausible rule with no supporting source or owner |
| Architecture | Boundaries, failure behaviour, data authority, quality scenarios | Valid API schema with conflicting business semantics |
| Implementation | Unit/property tests, types, static/security analysis, targeted mutation tests | Passing tests that never exercise the changed branch |
| Storage and migration | Real migrated schema, runtime identity, isolation, concurrency | Elevated test identity bypasses the production control |
| Integration | Contracts, actual-provider samples and labelled substitutes | Provider outage converted into a successful/clean result |
| Product acceptance | Independent expected facts and real role journeys | A polished screen presents the wrong outcome or hides an exception |
| Agent behaviour | Grounding, unauthorized-action attempts, uncertainty, bounded repeated trials | A judge rewards a persuasive but unsupported answer |
| Documentation | Complete selected bundle, stable links, reviewed semantic coverage | Fresh files contradict the approved decision |
| Release and operations | Artifact identity, configuration, checks, recovery and observation | The deployed artifact differs from the tested one |
Maintain an independent acceptance bank containing approved examples and adverse cases the implementer cannot alter merely to pass. Tests visible during development and protected holdouts serve different purposes. Holdouts must still measure disclosed requirements, not secret requirements. Disclose their coverage and refresh them as capabilities evolve.
For important controls, demonstrate a known failing case, a known passing case and evidence that the intended subject was examined. Report pass, fail, not assessed and not applicable separately. An empty scan or unavailable dependency is not a clean result. Review false-positive harms as well as missed defects.
Build provenance helps connect an artifact to its inputs and builder; it does not prove functional or business correctness. SLSA v1.2 provides source/build tracks and attestation guidance that can inform this boundary. NIST SP 800-218A complements the SSDF with considerations for AI model and AI-system producers; applicability should be mapped to the actual role rather than copied as a blanket certification claim. SLSA v1.2, build provenance, NIST profile.
7. Agent and context design
Begin with responsibilities, not an agent headcount. Evidence extraction, clarification/impact analysis, artifact preparation and private contextual assistance are four useful scopes for a collaboration pilot. They may use the same approved model service through distinct permissions and evaluations. Workflow transitions, identity, approval and budgets belong in deterministic application controls.
The engineering side may add a planner, implementation worker and independent verification function. These are not necessarily permanently running services, and independent verification does not mean merely assigning another persona to the same information. Anthropic recommends starting with simple composable patterns and adding agentic complexity where it demonstrably helps. Building effective agents.
A run manifest should identify repository/worktree, code revision, selected intent baseline, source permissions, method and prompt versions, approved model route, tool grants, budgets and handoff state. Retrieval should supply relevant detail on demand. A large startup file containing every historical decision makes it harder to distinguish current obligations from history; a compact manifest with authoritative links is the proposed alternative.
Tool receipts should distinguish attempted, completed and failed actions. Auditability means retaining authorized evidence and decision rationale, not demanding hidden chain-of-thought or indiscriminately storing every conversation. Private drafts do not become shared knowledge simply because their summaries are convenient to retrieve.
8. European deployment and privacy
An enterprise frontier-model service with private networking is a viable candidate deployment pattern, but “private cloud,” “EU processing,” “no training,” “no retention” and “no operator access” are different commitments. Azure documents different processing geography for regional/geography, DataZone and Global deployments. AWS similarly distinguishes geographic from global cross-region inference, and PrivateLink addresses network connectivity rather than every processing obligation. Azure data processing, Bedrock geographic inference, PrivateLink.
The proposed deployment profile records the actual service, model/version, permitted processing geography, network route, storage and retention, abuse-monitoring arrangements, subprocessors, access controls and permitted fallback. Verify these for the chosen offering; do not assume every model supports every region or contractual arrangement. When the approved route is unavailable, preserve human work and pause model-dependent actions. Do not silently send content to a public alternative.
GDPR principles include purpose limitation, minimization, accuracy and storage limitation. Apply a separate retention/access policy to recordings, transcripts, private exploration, shared contributions and operational evidence. OPA's decision-log documentation provides a concrete example of masking sensitive policy inputs before logging. A detailed trace is not automatically a lawful or proportionate trace. Commission GDPR principles, OPA decision logs.
AI Act obligations are risk- and role-dependent. Classify the development tools, the collaboration product and each delivered AI application's intended use separately, with qualified review of the applicable rules and effective dates. An agentic-factory label does not establish either high-risk status or exemption. This paper is engineering guidance, not a legal opinion or conformity assessment. Commission AI Act overview.
9. Methodology as a maintained product
A method intended for agents should have a normative control register, readable guidance, tested adoption tooling and explicit release compatibility. Each control needs an ID, purpose, applicability, owner, mechanism, evidence scope, known limitation and exception route. Mechanized, manually reviewed, configured and independently verified are different states.
Generate repetitive projections where possible, but do not pretend that a schema captures every semantic obligation. Keep one normative home per control; connect operating cards, skills, examples and public pages to it. Rule changes should trigger a propagation review across all consuming surfaces. Historical examples should not resemble current copy-and-paste instructions.
A documentation build is an executable pipeline and part of the supply chain. Treat generation scripts, path handling, shell quoting, template safety, dependencies and publication access as code. Test clean regeneration, preservation of authored content, negative fixtures, link compatibility and rendering in light/dark and narrow layouts. Diagrams should explain ownership, approvals, state changes and failure paths, with textual equivalents.
Methodology release, adapter readiness and application readiness should have separate evidence. A verified starter kit cannot certify every adopter. A successful application pilot cannot validate every method control. A new major version should supply a migration guide and pinned adoption manifests rather than silently upgrading active projects.
10. Product strategy and validation
Three approaches remain viable: configure existing orchestration/collaboration products; build a focused business-intent layer over standard infrastructure; or build a general-purpose autonomous platform. The middle approach is recommended because it concentrates engineering on the proposed differentiator while preserving replaceable execution machinery. Paperclip remains a candidate adapter or future comparison baseline, not a mandatory dependency or a dismissed competitor.
| First-customer option | Core value hypothesis | Validation obstacle | Decisive experiment |
|---|---|---|---|
| Software consultancy | Reuse a consistent BA-to-delivery workflow across customers; reduce reconstruction and handoff effort | Different customer tool, data and IP constraints may force customization | Repeat the same workflow in two differently constrained projects and measure configuration versus code changes |
| Internal enterprise team | Reduce clarification delay while preserving policy, terminology and auditability | Integration, procurement, authority mapping and business availability | Run one real capability through existing governance and compare review effort and defects with the current process |
Compare both segments using the same outcomes. Measure accepted capability lead time, BA and stakeholder effort, unsupported-claim corrections, escaped semantic defects, recovery, configuration-only changes, extension effort and operating cost. Do not use generated code volume, agent count or tokens consumed as substitutes for value. DORA's platform guidance likewise recommends balancing delivery performance with experience, adoption and task success. Platform measurement.
The first proof should stay small: one fictional capability; a second differently configured project; BA-first review; private selected submission; exact-revision approval; a living PRD and diagrams; a configuration change that needs no core code; and an unsupported request that is honestly routed to engineering. A later adapter exercise can return simulated verification receipts clearly labelled as simulated. Real autonomous customer delivery is a separate acceptance stage.
The commercial conclusion is conditional: there is enough substance to build and evaluate a product seed, but not enough evidence to claim breakthrough novelty or established demand. The strongest position is a repeatable, inspectable operating model intended to improve business-to-IT collaboration and earn broader autonomy through measured results.
References
Technical documentation is version-sensitive. Named versions identify the specifications used here; living documentation should be checked again when selecting an implementation. Publication dates below belong to the cited sources, not to an implied historical release of this proposal.
- Dan Shapiro. The Five Levels: from Spicy Autocomplete to the Dark Factory. January 23, 2026. Practitioner taxonomy.
- Justin McCarthy / StrongDM. Software Factories and the Agentic Moment. February 6, 2026. First-party factory account.
- StrongDM. Digital Twin Universe. Living technical account.
- Prithvi Rajasekaran / Anthropic. Harness design for long-running application development. March 24, 2026.
- Anthropic. Demystifying evals for AI agents. January 9, 2026.
- Anthropic. Building effective agents. December 19, 2024.
- Paperclip maintainers. Core concepts and heartbeat protocol. Living project documentation; no pinned product-version conformance claimed.
- Joel Becker, Nate Rush, Beth Barnes and David Rein / METR. Early-2025 developer productivity study. July 10, 2025.
- METR. Changing the developer productivity experiment design. February 24, 2026.
- Joel Becker / METR. Self-reported impact of early-2026 AI on technical worker productivity. May 11, 2026.
- DORA / Google Cloud. State of AI-assisted Software Development 2025. 2025 research report landing page.
- DORA. Platform engineering capability. Updated January 12, 2026.
- OMG. Decision Model and Notation and DMN 1.5 specification register. Formal version adopted August 2024.
- CEL project. Common Expression Language overview. Updated September 4, 2026.
- Open Policy Agent. Bundles and decision logs. Living documentation.
- Temporal. Worker deployments. Living deployment guidance.
- Pact. Can I Deploy. Living contract-verification guidance.
- SLSA. Approved specification v1.2 and build provenance.
- NIST. SP 800-218A: Secure Software Development Practices for Generative AI and Dual-Use Foundation Models. Final, July 2024.
- Microsoft. Data, privacy and security for Foundry Models sold by Azure. Living service documentation.
- AWS. Geographic cross-region inference and VPC / PrivateLink. Living service documentation.
- European Commission. GDPR processing principles and AI Act overview. Living official guidance; deployment-specific legal assessment remains necessary.