The front door
One API, exposed to agents as a tool surface over MCP. Portals and command lines are views on it, never separate paths with their own rules.
MCP tool server · Backstage · a CLI
Platform engineering in the agentic era · Working concept · October 2026 · Max Körbächer
Software factory is what everyone wants to have. But this is what it runs on.
Coding agents now write most new code at Google, Uber and Anthropic. Code became cheap. The environment that code needs still works in tickets. The Infrastructure Factory is the production system that closes that gap.
Definition
An Infrastructure Factory is a platform that turns intents, from humans or agents, into compliant, owned, observable, cost-attributed environments, with no ticket in the path.
Live model · where the queue forms
What it is
An Infrastructure Factory is a platform1 that turns intents2, from humans or agents3, into compliant, owned, observable, cost-attributed4 environments5, with no ticket in the path6.
A production system that a team runs as a product, with customers, a backlog and a service level. Not a project, not a folder of modules, not a committee of approvers. Its operators build the line.
A request that says what the environment is for, who owns it, what data it touches and how long it lives. The factory translates the intent into resources. The requester never has to.
The same contract serves a developer in a portal, a delivery pipeline, an operations agent remediating an incident, and a coding agent validating a change in the middle of the night. If an agent cannot use it unaided, it is a portal, not a factory.
Four properties present on arrival, never added later. Compliant: every policy evaluated, every decision recorded. Owned: a cost center, a compliance owner and an on-call owner that resolve to real people. Observable: telemetry flowing before the first request arrives. Cost-attributed: every euro lands on the owner's line.
The factory's unit of product: a bounded set of resources with a name, an owner, a lifetime, a data classification, a network position, an identity, telemetry and a cost line. An account, a namespace, a cluster, a sandbox or a set of managed services can each be an environment. A virtual machine is a part.
No human action between the intent and the environment. Humans decide at the contract, once, and the decisions live on as defaults, delegations and policies. A human may be notified, asked to approve a renewal above a threshold, or paged when a policy fails. A human is never a station.
Each of the parts exists as open source or as a cloud primitive today. The factory is the arrangement, not any one of them.
One API, exposed to agents as a tool surface over MCP. Portals and command lines are views on it, never separate paths with their own rules.
MCP tool server · Backstage · a CLI
Versioned, machine-readable contracts that state what the requester may choose and what the platform decides. The requester reads the contract before asking, so the first request is fulfillable.
Kratix Promises · Crossplane Compositions · Backstage templates
The source of truth for ownership: teams, cost centers, compliance contacts, on-call rotations, data classifications, and for physical estates the assets themselves.
Backstage catalog · NetBox · the HR and finance systems behind them
The pre-check that answers allow or deny with a reason, and the delegation rules that fill the fields no agent can sign.
Open Policy Agent · Kyverno · Gatekeeper
One per terrain: cloud accounts and projects, virtual machines, clusters, namespaces, sandboxes, bare metal. The contract hides which line runs.
Account Factory for Terraform · Cluster API · KubeVirt · Metal3 · Agent Sandbox
Capacity held as stock: warm pools, quotas, reservations and, off the cloud, procurement forecasts with a published fill rate.
SandboxWarmPool · node pools · hardware pools
Short-lived, scoped identities for every requester and every workload, agents included. Credentials move through it and nowhere else.
SPIFFE/SPIRE · Keycloak · cloud IAM
Telemetry, cost attribution and the decision record for every environment, kept as the factory's output rather than compiled for an audit afterwards.
OpenTelemetry · OpenCost · Prometheus · policy audit logs
The table reads left to right as a history: most organizations have built the first three and call the result a platform.
| Dimension | Ticket desk with scripts | Developer portal | Platform orchestrator | Infrastructure Factory |
|---|---|---|---|---|
| Who can ask | Humans, in a queue. | Humans, through a form. | Humans and pipelines, through an API. | Humans and agents, through a contract the agent can read. |
| What they send | Free text. | Form fields. | A specification or a claim. | An intent, with purpose, owner, data and lifetime. |
| Who decides | A person, per request. | A person, per request, behind the form. | Partly encoded; the rest escalates. | Encoded once: policy, defaults, delegations. Nobody per request. |
| What comes back | Resources. | Resources and a catalog entry. | An environment. | An environment with owner, evidence and cost attached. |
| How long it lives | Forever. | Forever. | Sometimes until deleted. | Until its lifetime ends, then it is reclaimed. |
| Whose problem is capacity | Procurement's. | Procurement's. | The cloud's. | The factory's, held as inventory with a fill rate. |
| How it is measured | Ticket SLA. | Adoption. | Deployments. | Lead time, zero-touch rate, time to owner, orphan rate. |
It is also not an AI factory in NVIDIA's sense, not an IaC generator, and not a product from any one vendor. A portal is its front door. An orchestrator is one of its stations. An IaC generator is a faster way to write the input the factory exists to make unnecessary.
Most organizations are somewhere on this ladder. Each stage keeps everything the previous one built and removes one kind of human from the path. The test at every stage is the same: can an agent with a scoped identity obtain an owned, compliant environment at any time?
Requests arrive as text. People build what they understand. Lead time is days to weeks, and the organization measures the queue.
Fails the test: nobody is awake.
Scripts and modules do the building, but a person still approves, classifies and assigns cost. Lead time drops; the queue stays.
Fails the test: the approval is a person.
Humans fill forms. Decisions are still made per request, behind the form, and the form itself is unreadable to an agent.
Fails the test: the agent cannot fill the form, and the form cannot sign the owner.
Machine-readable contracts, a policy pre-check with readable denials, defaults for ownership. Agents can ask. Evidence, lifetime and capacity are still somebody else's problem.
Fails the test at volume: orphans accumulate and the audit is compiled by hand.
Ownership by default, identity for every requester, a lifetime on everything, evidence shipped with the product, capacity held as inventory, and the factory publishing its own metrics.
Passes the test, at production volume, on every terrain.
Evidence
In software factories, agents write most new code, around the clock, and engineers build the factories that build the software. But the delivery side has not kept up, and the gap shows in the telemetry.
Share of new code at Google written by AI
Percent of new code, as stated by Google executives. The definition shifted between statements.
| Date | Share stated |
|---|---|
| October 2024 | more than 25% |
| April 2025 | well over 30% |
| Fall 2025 | about 50% |
| April 2026 | 75% |
More runs, less shipped on main
CircleCI, year-over-year change for the median team, 2026 report
| Measure | Change |
|---|---|
| Daily workflow runs | +59% |
| Feature-branch throughput (median team) | +15% |
| Main-branch throughput (median team) | −7% |
| Main-branch success rate | 70.8% |
The acceleration whiplash
Faros AI, AI Engineering Report 2026, change across 4,000+ teams
| Measure | Change |
|---|---|
| Task throughput per developer | +33.7% |
| PRs merged with no review | +31.3% |
| Median time in review | +441.5% |
Every figure above is reported by the organization that benefits from it, and Faros covers only its own customers. Read them as direction. The direction is consistent across sources that do not share incentives: code volume rises, delivery capacity stays flat, and queues form after the commit.
The consultancies assume the delivery side will simply keep up. BCG Platinion's agentic SDLC ends in an "operation" phase where agents "automate deployment, monitor production, and remediate incidents". Factory.ai, which sells software factories, concedes that "almost no one has meaningfully instrumented this loop to be fully AI-driven".
Review is the station everyone measures because it is the station that leaves a timestamp in the repository. The stations after it leave timestamps in ticketing systems: the staging environment, the new account, the firewall rule, the cost center, the compliance owner. Bunnyshell's summary of the same Faros data reads "AI writes code in minutes, your team waits days for staging".
No software-factory definition owns those stations. That absence is what this concept names.
Symmetry
Factory.ai, BCG Platinion, StrongDM and Uber describe the software factory in the same terms: a production system with inputs, agent workers, quality gates and a throughput number. The Infrastructure Factory has the same shape. Its product is an environment instead of a pull request, and the software factory is one of its customers.
| Element | Software factory | Infrastructure factory |
|---|---|---|
| Product | Pull requests, releases | Environments: compliant, owned, observable, cost-attributed |
| Inputs | Business intent, issues, signals, specifications | Intents from humans or agents, policies, budgets, owners |
| Workers | Coding agents, orchestrated by engineers who "build the factories that build the software" | Provisioning agents and orchestrators, run by platform engineers who build the factory that builds the environments |
| Production line | Plan, write, test, review, merge, deploy | Request, policy pre-check, vend, attach identity and telemetry and cost, hand over, reclaim |
| Quality gates | Tests, code review, CI, security scanning | Policy as code, admission control, compliance evidence, cost guardrails |
| Unit of trust | A reviewed, tested change | A signed owner, a classified dataset, an attributed cost center |
| Throughput | PRs merged per week, tokens per engineer, DORA metrics | Environments per day, lead time from intent to compliant environment, zero-touch rate |
| Failure mode | Review backlog, unverified code, main-branch regressions | Ticket queues, orphaned resources, unowned spend, policy drift |
| Who already has one | NVIDIA, EY, Adobe, Palo Alto Networks, Adyen, Uber, StrongDM, per their own accounts | Platform One, AWS Control Tower users with Account Factory, landing-zone factories. None of them names it. |
Some definitions say yes. Encore's AI software factory "provisions and deploys to the team's own cloud account" and pays off "once standing up infrastructure for each experiment is the thing slowing the agents down". Port writes that platform engineering moved "from building a portal to building the software factory", and that the portal work "is largely finished". (voice from the off: I don't agree on how Port defines platform engineering, that was never about building a portal)
Those definitions treat infrastructure as one step inside one code pipeline. The Infrastructure Factory is a separate production system with its own inputs (intents, policies, budgets, owners), its own throughput and its own failure modes. It serves workloads that no single software factory owns: the data platform, the shared cluster fleet, the compliance boundary, and the sandbox an agent asks for at 3am on behalf of nobody's pipeline.
Why now
The market moves fast and we are swamped by changes in the way we do things. The first is the one no software-factory definition addresses, and the one the "AI writes your infrastructure code" vendors skip.
An environment request carries fields that no agent can sign: a cost center, a compliance owner, a data classification, an on-call owner, a network zone. Each is a human decision, made by hand, usually in a ticket, usually after the code already exists. StackGen's own example is a React application generated in three hours followed by two to three days of hand-written Terraform, and that is before anyone asks who pays for it.
A factory cannot wait for those signatures. It has to encode them as defaults, delegations and policies, so that a request from an agent arrives already owned. Generating the Terraform faster does nothing for the fields a human still fills in. The table shows the same request twice: as it is resolved today, and as a factory resolves it.
| Field in the request | Who resolves it today | How the factory resolves it |
|---|---|---|
| compute · network · database | Terraform written by hand, or by an AI that freestyles it. human reviews | A golden-path contract exposes the choices the requester may make. The platform decides everything else. contract |
| cost_center | A manager approves it in a ticket. human signs | Inherited from the owning team's catalog record, with a delegated budget ceiling the agent can spend within. default |
| compliance_owner | Named by a human after a review meeting. human signs | A delegation rule: the service owner's registered compliance contact, enforced as policy. policy |
| data_classification | Someone guesses "internal". human guesses | Derived from the declared data sources. Higher classes need a human exactly once, when the contract is published. policy |
| on_call_owner | Blank until something breaks. nobody | Inherited from the catalog. No reachable owner, no environment. default |
| network_zone · egress | A security-team ticket. human signs | Decided by policy from the classification and the contract. policy |
| lifetime | Forever. nobody | A TTL by default. Renewal is a request like any other and carries the same owner. default |
Megaport asks whether coding agents are "the new CI bottleneck": validation capacity, meaning compute, caching and data movement, has to absorb agent spikes that arrive in bursts rather than at the pace of a working day. Bunnyshell makes the same point about staging. CircleCI's telemetry shows what that looks like in aggregate: pipelines run 59% more often, and the median team ships 7% less on main.
Qodo's line, quoted by platformengineering.com, is that the bottleneck shifted "from writing code to trusting it". Gartner, cited by StackGen, expects better coding efficiency to create "a bigger backlog for code reviews and security reviews". StrongDM went furthest: "Code must not be written by humans. Code must not be reviewed by humans." To make that work it built behavioral clones of Okta, Jira and Slack to validate what the agents produced. Simulated environments on demand is an infrastructure-factory problem, whatever you call it.
Humanitec ships scoped, least-privilege service users for AI agents per project and environment. Kubernetes SIG Apps added Agent Sandbox, with Sandbox, SandboxTemplate, SandboxWarmPool and SandboxClaim resources for "isolated environments for executing untrusted, LLM-generated code". The CNCF blog followed with "why sandboxing your agent is not enough". Gartner's 2026 Hype Cycle for Platform Engineering, as summarized by TrueFoundry, names "Agent Experience" and rates it transformational: back-end systems prepared so that APIs, data, documentation and workflows are machine-readable and discoverable.
If your documentation is trapped in PDFs or your security policies are inconsistent, an AI agent won't fix it and it will simply fail more efficiently.Rickey Zachary, Thoughtworks
BCG's practitioners report productivity gains of 3 to 5x and factories where "as few as three engineers" run a line in which humans no longer write code. Their lifecycle ends with agents that "automate deployment, monitor production, and remediate incidents". TechTarget's survey of Robusta, Komodor and Akamas gives the operators' view: "platform engineering is about building platforms for AI agents, and the role of humans is to build a platform that agents can use effectively." Both sides describe the Infrastructure Factory from a distance. Neither builds it, because neither owns it.
Lineage
The software factory is a fifty-year-old idea. Its most instructive chapter is the last one before the agents: the US Air Force built software factories, then discovered that they only scaled on top of a shared infrastructure platform.
Bob Bemer proposes the "software factory".
System Development Corporation operates its Software Factory.
Jack Greenfield and Keith Short publish Software Factories with Wiley and Microsoft.
The US Air Force launches "DOD's first software factory".
Established to "assist software factories by helping them focus on building mission applications": DevSecOps as a service, Kubernetes, Iron Bank hardened images and the Big Bang tooling.
It was called a software factory. It was an infrastructure platform.
Air Force CIO Lauren Knausenberger: "starting with Kessel Run smuggling DevSecOps into the DOD, and continuing with Platform One leading the way to Kubernetes, a common repo, and a desire to bring the entire community together and leverage common enterprise services."
Factory.ai's "Factory 2.0" (June 2026), BCG Platinion's "Agentic Software Factory" (March 2026), StrongDM's no-human rules (February 2026) and Uber's "Running a Software Factory Efficiently at Uber Scale" (August 2026). In Factory.ai's words, "the incremental units of this system are AI agents."
The platform the agentic software factory runs on, named as a production system of its own, with its own line and its own metrics.
Platform engineering was never about Kubernetes. It was about letting factories run.
Production line
An agent working a backlog at night needs an environment to validate a change against a real database and a public endpoint. Nobody is awake to open a ticket. This is the line that serves it, built from open source building blocks that exist today. Each station names the projects that implement it.
The agent calls a tool exposed over MCP or a plain API: an environment for the service orders-api, with Postgres, a public endpoint, classification derived from its data sources, for 48 hours. A human in a portal sends the same request through a form. Nobody writes Terraform here.
A structured intent, as YAML or JSON, validated against a published schema.
Building blocksThe intent resolves to a published abstraction: a Kratix Promise or a Crossplane Composition. The contract is versioned and machine-readable. It lists what the requester may choose and what the platform decides, which is exactly the information an agent needs to ask for something fulfillable the first time.
A claim against a named contract version.
Building blocksOPA or Kyverno evaluates the claim before anything is created and returns allow or deny with a reason the agent can act on. Defaults fill the fields the agent cannot sign: the cost center from the owning team's catalog record, the compliance owner by delegation, the classification from the declared data. A denial is a usable answer, and never a ticket.
An admitted claim with ownership, classification and budget attached, plus the policy decision kept as evidence.
Building blocksThe vending lines that already exist do their work: an account or project from AWS Account Factory for Terraform, Azure subscription vending or Google's project factory; a cluster or namespace from Cluster API or a virtual cluster; the Postgres and the endpoint from OpenTofu modules. GitOps reconciles the result and keeps reconciling it.
A reconciled environment, with endpoints and a record of what was created.
Building blocksWhen the environment exists to run code an LLM wrote, a Kubernetes Agent Sandbox supplies the isolated runtime from a warm pool, claimed in seconds. The sandbox is one station and never the whole factory. Stations three, six and seven are what the CNCF means by "sandboxing your agent is not enough".
A SandboxClaim bound to a running, isolated workload.
Building blocksWorkload identity for the environment, and a scoped, least-privilege identity for the agent that asked for it. Telemetry wired to the platform's collectors. Cost attribution tags that match the cost center from station three. The ownership record registered in the catalog. None of this needs a person, and all of it used to wait for one.
An environment that is observable, attributable, and reachable only by the identities the policy allowed.
Building blocksThe requester receives endpoints, credentials through the identity system, and the evidence bundle: policy decisions, owner, classification, budget. The TTL and the owner record make reclamation automatic. The factory emits its own metrics at this station: lead time, zero-touch, time to owner.
An environment in use, a metrics record, and a scheduled reclaim.
Building blocksCNCF project Other open source, Kubernetes SIG or cloud-provider building block
Terrain
Everything above quietly assumed a hyperscaler underneath: an API for every resource, capacity that appears in seconds, a meter on everything. That is the easy case. Many of the infrastructure that matters sits on something less cloud-like: a virtualization estate, bare metal, a sovereign or regulated cloud, factory floors and branch sites. The factory model holds there too. It has more stations and a harder first mile.
NIST's 2011 definition of cloud computing lists five essential characteristics: on-demand self-service, broad network access, resource pooling, rapid elasticity and measured service. An Infrastructure Factory consumes all five. A hyperscaler supplies them as APIs. Wherever one is missing, the factory has to manufacture it before it can vend anything, and that is the whole difference between the cloud factory and every other kind.
| NIST characteristic | What a cloud factory gets for free | What a less cloud-like factory builds first | Building blocks |
|---|---|---|---|
| On-demand self-service | Every resource has an API: accounts, identity, networks, databases. Vending lines such as Account Factory already exist. | APIs exist for compute (vSphere, OpenStack, KubeVirt), but network, firewall, IP addresses, DNS and storage sit behind tickets. The factory needs a source of truth and an API façade over every one of them. | NetBox or Nautobot, Ansible, OpenTofu providers, Crossplane providers, Cluster API providers for vSphere, OpenStack and Metal3 |
| Broad network access | A global backbone, managed load balancers, DNS and certificates on request. | VLANs per site, firewalls with change windows, load balancers as appliances. The network has to become a policy the factory enforces instead of a ticket it waits for. | MetalLB, Cilium or Calico network policy, cert-manager, external-dns |
| Resource pooling | Multi-tenant by construction. Isolation is an account boundary. | Dedicated hardware per team, licenses per socket, snowflake clusters. The factory has to create the pool: a virtualization layer, and tenancy inside clusters. | KubeVirt, OpenStack, vCluster, Capsule, Kamaji, namespaces with quotas |
| Rapid elasticity | Capacity appears in seconds and leaves the bill when released. | Capacity arrives on a truck. Lead time has a floor set by procurement, so the factory keeps inventory: warm pools, buffer stock, reservations and forecasts. | Metal3, Tinkerbell, Ironic, MAAS, Cluster API, SandboxWarmPool |
| Measured service | A billing API, tags, and cost allocation down to the resource. | No meter. Cost is a depreciation line in a spreadsheet. The factory has to build metering and showback before cost attribution can be one of its outputs. | OpenCost with on-premises pricing, Kepler, Prometheus |
Two things NIST does not list but a hyperscaler also supplies: one identity system across every resource, and a certified compliance posture under a shared-responsibility model. Off the cloud, the factory owns both. Keycloak and SPIFFE/SPIRE take the first; the evidence bundle from station seven takes the second.
The first three stations do not change. Intent, contract and pre-check look the same whether the environment lands in an AWS account or on a rack in Nuremberg, and that is the point of the contract: the requester cannot tell where the environment will run, and should not have to. Everything between the pre-check and the hand-over is where the terrain shows.
Lead time has a floor the factory cannot remove: a VM from a warm hypervisor in minutes, a bare-metal node with firmware, RAID and an operating system image in tens of minutes to hours, a new rack in the weeks to months that 2026 component lead times dictate. The factory's job is to hide that floor behind inventory, never to pretend it is zero.
| Station | Cloud factory | Less cloud-like factory |
|---|---|---|
| 1 to 3 · Intent, contract, pre-check | Unchanged. | Unchanged. The contract hides the terrain. |
| 4 · Vend | One call to a vending line. | Splits into capacity reservation, node provisioning (firmware, RAID, image), network and storage attach, cluster bootstrap, and the GitOps hand-off. |
| 5 · Sandbox | Agent Sandbox on a managed cluster. | The same, on a cluster the factory built. Warm pools matter more, because nodes do not appear on demand. |
| 6 · Attach | Identity, telemetry and cost tags exist as cloud primitives. | Identity federation and metering have to be built before they can be attached. |
| 7 · Hand over, reclaim | Delete the resources and the bill stops. | Return the hardware to the pool, wipe it, and make it visible to station zero again. |
| 0 · Supply (new) | None. | Procurement, inventory and forecasting. Capacity as stock, with a published fill rate. |
Hyperscaler cloud · the easy case
The vending lines exist: Account Factory, subscription vending, project factory. The hard stations are the non-technical ones, the fields no agent can sign. Factory effort goes into contracts, defaults and delegations, and almost none into supply.
Account Factory for Terraform · Crossplane · Kratix · OPA · Kyverno · Agent Sandbox
Virtualization estates · an estate in motion
vSphere, Nutanix, OpenStack, Proxmox and KubeVirt all have APIs. Few are exposed as self-service, because change advisory boards and per-socket licensing sit in front of them. The Broadcom licensing changes turned this terrain into a construction site: 48% of VMware customers plan to shrink their footprint by 2028, while only 4% have finished leaving. KubeVirt, a CNCF incubating project heading for graduation with release 1.8, and Forklift for moving VMs are the open source route out. An IDC-documented enterprise cut provisioning from three to four days to about thirty minutes once self-service replaced its approval queue.
KubeVirt · Forklift · Cluster API for vSphere and OpenStack · Crossplane · NetBox · Ansible
Bare metal and edge · physics in the loop
Factory floors, stores, telco sites and GPU clusters. In 2026, component lead times stretched to 35 to 40 weeks for power-management chips and 21 to 26 weeks for baseboard controllers, and HGX-class GPU systems quoted 8 to 20 weeks. Thousands of sites with intermittent connectivity mean pull-based GitOps and immutable images. Here the factory ships images, not environments, and station zero is the one that decides whether a request at 3am can be served at all.
Metal3 (CNCF incubating since August 2025) · Tinkerbell · Ironic · Talos or Flatcar · Flux · Argo CD · Kepler
Sovereign and regulated · evidence is the product
NIS2 applies since October 2024, the Digital Operational Resilience Act since January 2025 (the EU regulation, not the metrics) and the Data Act since September 2025. Gartner expects European sovereign cloud spending to more than triple to $23bn by 2027. Sovereign providers expose thinner API surfaces and fewer managed services, so more of the five characteristics fall to the factory. Data residency and operator independence become contract fields, and the evidence bundle from station seven is the artifact the regulator reads.
OPA · Kyverno · SPIFFE/SPIRE · Harbor · Zarf · Sigstore · OpenCost
The contract, the metrics and the first three stations are the same on every terrain. What changes is how much of the cloud the factory has to build before it can sell an environment. In the cloud, the factory is policy and ownership. Everywhere else, it is policy and ownership plus a supply chain. The hardest station in the cloud is a signature. The hardest station on bare metal is a truck.
Metrics
The software-factory crowd brings numbers: pull requests per week, tokens per engineer, 3 to 5x. An Infrastructure Factory needs its own set, the equivalent of DORA for environments. Six are enough to start, and each has a counterpart on the software side.
Little's law for environment requests
At 40 requests a day and a lead time of 3 days, 120 environment requests are open at any moment. Through a factory with a lead time of 20 min, fewer than one is.
Work in progress = arrival rate × lead time. Little's law holds for any stable queue, whatever is in it. The inputs are yours; the arithmetic is the only claim.
Factors
The Twelve-Factor App of 2011 told a generation how to write software that a platform could run. These twelve say how to build the platform that agents can run on. Each one is written as a test. A factory either passes it, or there is a ticket in the path somewhere.
Intents over resource lists.
Defaults over approvals.
Contracts over tickets.
Evidence over reports.
There is value on the right of every line. The factory encodes it once, so that nobody has to supply it per request.
The input is an intent: what the environment is for, who owns it, what data it touches, how long it lives. Never a resource list. A resource list is the factory's output, translated back into its inputs by a human.
The output is an environment that is owned, compliant, observable and cost-attributed on arrival. Anything less is a resource, and resources are what the ticket queue used to deliver.
What the requester may choose and what the platform decides is a versioned schema. The API is the interface; portals and CLIs are views on it. Agents are first-class customers, and a contract an agent cannot read is a contract a human will be asked to interpret.
Humans decide once, at the contract, or encode the decision as a default, a delegation or a policy. A human may be notified. A human is never a blocking station. Platform engineers build the line; they do not fulfil requests.
The pre-check returns allow or deny, with a reason the requester can act on, before anything is created. Nothing is built to be torn down by a reviewer later, and a denial is an answer rather than an escalation.
Cost center, compliance owner, on-call owner and data classification are inherited from the catalog, never typed into a form. No reachable owner, no environment.
Every request runs under a short-lived, least-privilege identity, agents included. Credentials move through the identity system and never through a prompt, a ticket or a chat message.
A lifetime is part of every intent. Renewal is a request like any other and carries the same owner. Reclaim is automatic, and an orphan is a defect of the factory, not of the requester.
The decision record travels with the environment: policies evaluated, owner, classification, budget, jurisdiction. An audit reads the factory's output, not a spreadsheet assembled afterwards.
Warm pools, quotas and buffer stock absorb bursts. Where elasticity is missing, the factory keeps stock and publishes its fill rate instead of pretending that lead time is zero.
Lead time to a compliant environment, environments per platform engineer, zero-touch rate, policy-violation rate, time to owner, orphan rate. Published to the factory's customers the way an SLO is, and argued about the same way.
Contracts outlive tools. Every station is a building block that can be swapped, Crossplane for Kratix, Kyverno for OPA, a hyperscaler for a rack, without changing what the requester asks for or what the requester gets back.
Prior art
The term is open. Any may have said since mid-2025 that the bottleneck moved from code to infrastructure delivery, and a few projects already use the words.
| Usage | Who | What it means | Relation |
|---|---|---|---|
| InfraFactory | MozeBaltyk, open source on GitHub | "A factory for building cloud infrastructure, provider-agnostic and modular": OpenTofu, cloud-init, K3s or RKE2 and Flux to stamp out VMs and clusters. | A vending line |
| C.I. Factory, Capability Infrastructure Factory | bonometti718, open source on GitHub | "Machine-native infrastructure for autonomous AI agents": small verifiable capabilities such as an "Agent Economic Action Gate". | Adjacent, tiny |
| AI Factory | NVIDIA, Presidio and others | Physical accelerated datacenters that "convert energy into tokens". | Different thing |
| Account Factory, Account Factory for Terraform | AWS Control Tower, with HashiCorp | A Terraform pipeline that vends and customizes governed AWS accounts inside a landing zone. | A vending line |
| Subscription vending | Microsoft Azure Architecture Center | Bicep and Terraform modules that create workload landing zones at scale. | A vending line |
| Project factory | Google Cloud Foundation Fabric | A Terraform module that vends GCP projects from YAML. | A vending line |
| "A factory with multiple production lines" | Cellenza, Azure consultancy | Industrialized landing-zone delivery. | Closest precedent |
| Platform One | US Air Force | A "software factory" that in practice provides DevSecOps as a service, Kubernetes, hardened images and tooling to other software factories. | Lineage |
Both factories have inputs, production lines, quality gates and a throughput number. Nobody has drawn the second line next to the first.
Cost center, compliance owner, data classification, on-call owner. Agents cannot sign them, so the factory encodes them as defaults, delegations and policies.
It was about letting factories run. Platform One proved it before the agents arrived.
Sources
Everything quoted on this page comes from the publications below. Links are left out on purpose until each one has been re-verified against the original.