Platform engineering in the agentic era · Working concept · October 2026 · Max Körbächer

The Infrastructure Factory

Software factory is what everyone wants to have. But this is what it runs on.

Coding agents now write most new code at Google, Uber and Anthropic. Code became cheap. The environment that code needs still works in tickets. The Infrastructure Factory is the production system that closes that gap.

Definition

An Infrastructure Factory is a platform that turns intents, from humans or agents, into compliant, owned, observable, cost-attributed environments, with no ticket in the path.

Live model · where the queue forms

Pull requests shipped0
Waiting for an environment0
Environments delivered0
Both lines receive work at the same rate. The software factory ships it at agent speed. Every change also needs an environment, and today that request stops at a gate a human opens a few times a day. Built an Infrastructure Factory and the gate becomes a policy check that answers in seconds. The queue is the thesis.

What it is

What an Infrastructure Factory is

The definition, taken apart

An Infrastructure Factory is a platform1 that turns intents2, from humans or agents3, into compliant, owned, observable, cost-attributed4 environments5, with no ticket in the path6.

  1. 1

    Platform

    A production system that a team runs as a product, with customers, a backlog and a service level. Not a project, not a folder of modules, not a committee of approvers. Its operators build the line.

  2. 2

    Intents

    A request that says what the environment is for, who owns it, what data it touches and how long it lives. The factory translates the intent into resources. The requester never has to.

  3. 3

    Humans or agents

    The same contract serves a developer in a portal, a delivery pipeline, an operations agent remediating an incident, and a coding agent validating a change in the middle of the night. If an agent cannot use it unaided, it is a portal, not a factory.

  4. 4

    Compliant, owned, observable, cost-attributed

    Four properties present on arrival, never added later. Compliant: every policy evaluated, every decision recorded. Owned: a cost center, a compliance owner and an on-call owner that resolve to real people. Observable: telemetry flowing before the first request arrives. Cost-attributed: every euro lands on the owner's line.

  5. 5

    Environments

    The factory's unit of product: a bounded set of resources with a name, an owner, a lifetime, a data classification, a network position, an identity, telemetry and a cost line. An account, a namespace, a cluster, a sandbox or a set of managed services can each be an environment. A virtual machine is a part.

  6. 6

    No ticket in the path

    No human action between the intent and the environment. Humans decide at the contract, once, and the decisions live on as defaults, delegations and policies. A human may be notified, asked to approve a renewal above a threshold, or paged when a policy fails. A human is never a station.

What it is made of

Each of the parts exists as open source or as a cloud primitive today. The factory is the arrangement, not any one of them.

The front door

One API, exposed to agents as a tool surface over MCP. Portals and command lines are views on it, never separate paths with their own rules.

MCP tool server · Backstage · a CLI

The contract registry

Versioned, machine-readable contracts that state what the requester may choose and what the platform decides. The requester reads the contract before asking, so the first request is fulfillable.

Kratix Promises · Crossplane Compositions · Backstage templates

The catalog

The source of truth for ownership: teams, cost centers, compliance contacts, on-call rotations, data classifications, and for physical estates the assets themselves.

Backstage catalog · NetBox · the HR and finance systems behind them

The policy engine and the defaults

The pre-check that answers allow or deny with a reason, and the delegation rules that fill the fields no agent can sign.

Open Policy Agent · Kyverno · Gatekeeper

The vending lines

One per terrain: cloud accounts and projects, virtual machines, clusters, namespaces, sandboxes, bare metal. The contract hides which line runs.

Account Factory for Terraform · Cluster API · KubeVirt · Metal3 · Agent Sandbox

The inventory

Capacity held as stock: warm pools, quotas, reservations and, off the cloud, procurement forecasts with a published fill rate.

SandboxWarmPool · node pools · hardware pools

The identity system

Short-lived, scoped identities for every requester and every workload, agents included. Credentials move through it and nowhere else.

SPIFFE/SPIRE · Keycloak · cloud IAM

The meter and the evidence store

Telemetry, cost attribution and the decision record for every environment, kept as the factory's output rather than compiled for an audit afterwards.

OpenTelemetry · OpenCost · Prometheus · policy audit logs

What it is not

The table reads left to right as a history: most organizations have built the first three and call the result a platform.

DimensionTicket desk with scriptsDeveloper portalPlatform orchestratorInfrastructure Factory
Who can askHumans, in a queue.Humans, through a form.Humans and pipelines, through an API.Humans and agents, through a contract the agent can read.
What they sendFree text.Form fields.A specification or a claim.An intent, with purpose, owner, data and lifetime.
Who decidesA person, per request.A person, per request, behind the form.Partly encoded; the rest escalates.Encoded once: policy, defaults, delegations. Nobody per request.
What comes backResources.Resources and a catalog entry.An environment.An environment with owner, evidence and cost attached.
How long it livesForever.Forever.Sometimes until deleted.Until its lifetime ends, then it is reclaimed.
Whose problem is capacityProcurement's.Procurement's.The cloud's.The factory's, held as inventory with a fill rate.
How it is measuredTicket SLA.Adoption.Deployments.Lead time, zero-touch rate, time to owner, orphan rate.

It is also not an AI factory in NVIDIA's sense, not an IaC generator, and not a product from any one vendor. A portal is its front door. An orchestrator is one of its stations. An IaC generator is a faster way to write the input the factory exists to make unnecessary.

From ticket desk to factory

Most organizations are somewhere on this ladder. Each stage keeps everything the previous one built and removes one kind of human from the path. The test at every stage is the same: can an agent with a scoped identity obtain an owned, compliant environment at any time?

  1. 0

    Ticket desk

    Requests arrive as text. People build what they understand. Lead time is days to weeks, and the organization measures the queue.

    Fails the test: nobody is awake.

  2. 1

    Automated fulfilment

    Scripts and modules do the building, but a person still approves, classifies and assigns cost. Lead time drops; the queue stays.

    Fails the test: the approval is a person.

  3. 2

    Self-service portal

    Humans fill forms. Decisions are still made per request, behind the form, and the form itself is unreadable to an agent.

    Fails the test: the agent cannot fill the form, and the form cannot sign the owner.

  4. 3

    Contracted API

    Machine-readable contracts, a policy pre-check with readable denials, defaults for ownership. Agents can ask. Evidence, lifetime and capacity are still somebody else's problem.

    Fails the test at volume: orphans accumulate and the audit is compiled by hand.

  5. 4

    Infrastructure Factory

    Ownership by default, identity for every requester, a lifetime on everything, evidence shipped with the product, capacity held as inventory, and the factory publishing its own metrics.

    Passes the test, at production volume, on every terrain.

You have an Infrastructure Factory when

  • An agent with a scoped identity can obtain an owned, compliant environment with nobody awake.
  • Every environment has a reachable owner, a cost center and a classification before it exists.
  • Every denial carries a reason the requester can act on without opening a ticket.
  • Every environment has a lifetime, and reclaim runs without a human.
  • Every environment ships with its decision record.
  • The factory publishes its own lead time, zero-touch rate, time to owner and orphan rate.
  • A contract can move to another terrain without changing what the requester asks for.
  • Platform engineers spend their time on contracts, policies and lines, and none of it fulfilling requests.

Evidence

The bottleneck has moved

In software factories, agents write most new code, around the clock, and engineers build the factories that build the software. But the delivery side has not kept up, and the gap shows in the telemetry.

Code supply75%of new code at Google is AI-generated and approved by engineers, up from 25% in October 2024.Sundar Pichai, Cloud Next, April 2026
Code supply70%+of pull requests at Uber are attributed to local or cloud agents. Engineers have built more than 3,600 agent skills.Uber Engineering, August 2026
Code supply90%+of Anthropic's code is written by Claude Code, according to its CFO. Self-reported, and definitions vary between companies.Krishna Rao, May 2026
Delivery side+441%median time a change spends in review, while task throughput per developer rose 33.7%.Faros AI, AI Engineering Report 2026, 22,000 developers
Delivery side−7%main-branch throughput for the median team, while daily workflow runs rose 59% year over year. Main-branch success fell to 70.8%, a five-year low.CircleCI, February 2026, 28.7 million workflows
Delivery side16×growth in sandboxes on GKE in under five months after the Agent Sandbox preview, with 300 sandboxes per second claimed at GA.Google Cloud, May 2026

Share of new code at Google written by AI

Percent of new code, as stated by Google executives. The definition shifted between statements.

100% 75 50 25 0 Oct 2024 Apr 2025 Fall 2025 Apr 2026 >25% >30% ~50% 75%
The April 2025 point is plotted at 30 for a statement of "well over 30%". Always quote the latest figure with "new" and "approved by engineers".
View as table
DateShare stated
October 2024more than 25%
April 2025well over 30%
Fall 2025about 50%
April 202675%

More runs, less shipped on main

CircleCI, year-over-year change for the median team, 2026 report

−20%0+20+40+60%
Pipelines run more often and feature branches move faster. What reaches main fell. More code is not more shipped.
View as table
MeasureChange
Daily workflow runs+59%
Feature-branch throughput (median team)+15%
Main-branch throughput (median team)−7%
Main-branch success rate70.8%

The acceleration whiplash

Faros AI, AI Engineering Report 2026, change across 4,000+ teams

0+100+200+300+400%
Output per developer grew by a third. The wait after the code exists grew by a factor of five. Review is one station of that wait. Environments are the rest.
View as table
MeasureChange
Task throughput per developer+33.7%
PRs merged with no review+31.3%
Median time in review+441.5%

How to read these numbers

Every figure above is reported by the organization that benefits from it, and Faros covers only its own customers. Read them as direction. The direction is consistent across sources that do not share incentives: code volume rises, delivery capacity stays flat, and queues form after the commit.

The consultancies assume the delivery side will simply keep up. BCG Platinion's agentic SDLC ends in an "operation" phase where agents "automate deployment, monitor production, and remediate incidents". Factory.ai, which sells software factories, concedes that "almost no one has meaningfully instrumented this loop to be fully AI-driven".

Where the queue forms is the point

Review is the station everyone measures because it is the station that leaves a timestamp in the repository. The stations after it leave timestamps in ticketing systems: the staging environment, the new account, the firewall rule, the cost center, the compliance owner. Bunnyshell's summary of the same Faros data reads "AI writes code in minutes, your team waits days for staging".

No software-factory definition owns those stations. That absence is what this concept names.

Symmetry

The Infrastructure Factory mirrors the software factory

Factory.ai, BCG Platinion, StrongDM and Uber describe the software factory in the same terms: a production system with inputs, agent workers, quality gates and a throughput number. The Infrastructure Factory has the same shape. Its product is an environment instead of a pull request, and the software factory is one of its customers.

Software factory intent · issues · signals pull requests releases Plan Write Test Review Merge Deploy other consumers: ops agents · data platform · sandboxes needs an environment · today a ticket, here an API call environment, with owner and evidence attached Infrastructure factory intents · policies budgets · owners environments owned · compliant cost-attributed Request Pre-check Vend Attach Hand over Reclaim Dashed lines are requests and deliveries between the two factories. Solid lines are each factory's own conveyor.
The software factory's deploy station is a customer of the infrastructure factory, and so are consumers that belong to no code pipeline. That is why the second line is a peer, and why its product needs an owner and evidence attached before it is handed back.
ElementSoftware factoryInfrastructure factory
ProductPull requests, releasesEnvironments: compliant, owned, observable, cost-attributed
InputsBusiness intent, issues, signals, specificationsIntents from humans or agents, policies, budgets, owners
WorkersCoding agents, orchestrated by engineers who "build the factories that build the software"Provisioning agents and orchestrators, run by platform engineers who build the factory that builds the environments
Production linePlan, write, test, review, merge, deployRequest, policy pre-check, vend, attach identity and telemetry and cost, hand over, reclaim
Quality gatesTests, code review, CI, security scanningPolicy as code, admission control, compliance evidence, cost guardrails
Unit of trustA reviewed, tested changeA signed owner, a classified dataset, an attributed cost center
ThroughputPRs merged per week, tokens per engineer, DORA metricsEnvironments per day, lead time from intent to compliant environment, zero-touch rate
Failure modeReview backlog, unverified code, main-branch regressionsTicket queues, orphaned resources, unowned spend, policy drift
Who already has oneNVIDIA, EY, Adobe, Palo Alto Networks, Adyen, Uber, StrongDM, per their own accountsPlatform One, AWS Control Tower users with Account Factory, landing-zone factories. None of them names it.

"Isn't infrastructure just a station in the software factory?"

Some definitions say yes. Encore's AI software factory "provisions and deploys to the team's own cloud account" and pays off "once standing up infrastructure for each experiment is the thing slowing the agents down". Port writes that platform engineering moved "from building a portal to building the software factory", and that the portal work "is largely finished". (voice from the off: I don't agree on how Port defines platform engineering, that was never about building a portal)

Those definitions treat infrastructure as one step inside one code pipeline. The Infrastructure Factory is a separate production system with its own inputs (intents, policies, budgets, owners), its own throughput and its own failure modes. It serves workloads that no single software factory owns: the data platform, the shared cluster fleet, the compliance boundary, and the sandbox an agent asks for at 3am on behalf of nobody's pipeline.

Why now

Why this is hard right now

The market moves fast and we are swamped by changes in the way we do things. The first is the one no software-factory definition addresses, and the one the "AI writes your infrastructure code" vendors skip.

It's all about ownership, and not Terraform.

An environment request carries fields that no agent can sign: a cost center, a compliance owner, a data classification, an on-call owner, a network zone. Each is a human decision, made by hand, usually in a ticket, usually after the code already exists. StackGen's own example is a React application generated in three hours followed by two to three days of hand-written Terraform, and that is before anyone asks who pays for it.

A factory cannot wait for those signatures. It has to encode them as defaults, delegations and policies, so that a request from an agent arrives already owned. Generating the Terraform faster does nothing for the fields a human still fills in. The table shows the same request twice: as it is resolved today, and as a factory resolves it.

Anatomy of an environment request: who resolves each field today and how an infrastructure factory resolves it
Field in the requestWho resolves it todayHow the factory resolves it
compute · network · databaseTerraform written by hand, or by an AI that freestyles it. human reviewsA golden-path contract exposes the choices the requester may make. The platform decides everything else. contract
cost_centerA manager approves it in a ticket. human signsInherited from the owning team's catalog record, with a delegated budget ceiling the agent can spend within. default
compliance_ownerNamed by a human after a review meeting. human signsA delegation rule: the service owner's registered compliance contact, enforced as policy. policy
data_classificationSomeone guesses "internal". human guessesDerived from the declared data sources. Higher classes need a human exactly once, when the contract is published. policy
on_call_ownerBlank until something breaks. nobodyInherited from the catalog. No reachable owner, no environment. default
network_zone · egressA security-team ticket. human signsDecided by policy from the classification and the contract. policy
lifetimeForever. nobodyA TTL by default. Renewal is a request like any other and carries the same owner. default

Capacity was sized for humans

Megaport asks whether coding agents are "the new CI bottleneck": validation capacity, meaning compute, caching and data movement, has to absorb agent spikes that arrive in bursts rather than at the pace of a working day. Bunnyshell makes the same point about staging. CircleCI's telemetry shows what that looks like in aggregate: pipelines run 59% more often, and the median team ships 7% less on main.

Trust moved downstream

Qodo's line, quoted by platformengineering.com, is that the bottleneck shifted "from writing code to trusting it". Gartner, cited by StackGen, expects better coding efficiency to create "a bigger backlog for code reviews and security reviews". StrongDM went furthest: "Code must not be written by humans. Code must not be reviewed by humans." To make that work it built behavioral clones of Okta, Jira and Slack to validate what the agents produced. Simulated environments on demand is an infrastructure-factory problem, whatever you call it.

Agents are a new kind of customer

Humanitec ships scoped, least-privilege service users for AI agents per project and environment. Kubernetes SIG Apps added Agent Sandbox, with Sandbox, SandboxTemplate, SandboxWarmPool and SandboxClaim resources for "isolated environments for executing untrusted, LLM-generated code". The CNCF blog followed with "why sandboxing your agent is not enough". Gartner's 2026 Hype Cycle for Platform Engineering, as summarized by TrueFoundry, names "Agent Experience" and rates it transformational: back-end systems prepared so that APIs, data, documentation and workflows are machine-readable and discoverable.

If your documentation is trapped in PDFs or your security policies are inconsistent, an AI agent won't fix it and it will simply fail more efficiently.Rickey Zachary, Thoughtworks

Everyone assumes the other side keeps up

BCG's practitioners report productivity gains of 3 to 5x and factories where "as few as three engineers" run a line in which humans no longer write code. Their lifecycle ends with agents that "automate deployment, monitor production, and remediate incidents". TechTarget's survey of Robusta, Komodor and Akamas gives the operators' view: "platform engineering is about building platforms for AI agents, and the role of humans is to build a platform that agents can use effectively." Both sides describe the Infrastructure Factory from a distance. Neither builds it, because neither owns it.

Lineage

The Department of Defense learned this before the agents arrived

The software factory is a fifty-year-old idea. Its most instructive chapter is the last one before the agents: the US Air Force built software factories, then discovered that they only scaled on top of a shared infrastructure platform.

  1. 1968

    The term is coined

    Bob Bemer proposes the "software factory".

  2. 1975

    The first one runs

    System Development Corporation operates its Software Factory.

  3. 2004

    The book

    Jack Greenfield and Keith Short publish Software Factories with Wiley and Microsoft.

  4. 2017

    Kessel Run

    The US Air Force launches "DOD's first software factory".

  5. Dec 2019

    Platform One

    Established to "assist software factories by helping them focus on building mission applications": DevSecOps as a service, Kubernetes, Iron Bank hardened images and the Big Bang tooling.

    It was called a software factory. It was an infrastructure platform.

  6. 2021

    The lesson, stated

    Air Force CIO Lauren Knausenberger: "starting with Kessel Run smuggling DevSecOps into the DOD, and continuing with Platform One leading the way to Kubernetes, a common repo, and a desire to bring the entire community together and leverage common enterprise services."

  7. 2025–26

    The agentic software factory

    Factory.ai's "Factory 2.0" (June 2026), BCG Platinion's "Agentic Software Factory" (March 2026), StrongDM's no-human rules (February 2026) and Uber's "Running a Software Factory Efficiently at Uber Scale" (August 2026). In Factory.ai's words, "the incremental units of this system are AI agents."

  8. 2026 →

    The Infrastructure Factory

    The platform the agentic software factory runs on, named as a production system of its own, with its own line and its own metrics.

Platform engineering was never about Kubernetes. It was about letting factories run.

Production line

An end to end line

An agent working a backlog at night needs an environment to validate a change against a real database and a public endpoint. Nobody is awake to open a ticket. This is the line that serves it, built from open source building blocks that exist today. Each station names the projects that implement it.

  1. Intent arrives

    The agent calls a tool exposed over MCP or a plain API: an environment for the service orders-api, with Postgres, a public endpoint, classification derived from its data sources, for 48 hours. A human in a portal sends the same request through a form. Nobody writes Terraform here.

    What moves to the next station

    A structured intent, as YAML or JSON, validated against a published schema.

    Building blocks
    • MCP tool server
    • Backstage
    • Portal or CLI
  2. Golden-path contract

    The intent resolves to a published abstraction: a Kratix Promise or a Crossplane Composition. The contract is versioned and machine-readable. It lists what the requester may choose and what the platform decides, which is exactly the information an agent needs to ask for something fulfillable the first time.

    What moves to the next station

    A claim against a named contract version.

    Building blocks
    • Crossplane
    • Kratix
    • Backstage software templates
  3. A policy pre-check the agent can read

    OPA or Kyverno evaluates the claim before anything is created and returns allow or deny with a reason the agent can act on. Defaults fill the fields the agent cannot sign: the cost center from the owning team's catalog record, the compliance owner by delegation, the classification from the declared data. A denial is a usable answer, and never a ticket.

    What moves to the next station

    An admitted claim with ownership, classification and budget attached, plus the policy decision kept as evidence.

    Building blocks
    • Open Policy Agent
    • Kyverno
    • Gatekeeper
  4. Vend

    The vending lines that already exist do their work: an account or project from AWS Account Factory for Terraform, Azure subscription vending or Google's project factory; a cluster or namespace from Cluster API or a virtual cluster; the Postgres and the endpoint from OpenTofu modules. GitOps reconciles the result and keeps reconciling it.

    What moves to the next station

    A reconciled environment, with endpoints and a record of what was created.

    Building blocks
    • OpenTofu
    • Argo CD
    • Flux
    • Cluster API (Kubernetes SIG)
    • Account Factory for Terraform
  5. Execution sandbox for untrusted code

    When the environment exists to run code an LLM wrote, a Kubernetes Agent Sandbox supplies the isolated runtime from a warm pool, claimed in seconds. The sandbox is one station and never the whole factory. Stations three, six and seven are what the CNCF means by "sandboxing your agent is not enough".

    What moves to the next station

    A SandboxClaim bound to a running, isolated workload.

    Building blocks
    • Agent Sandbox (SIG Apps)
    • gVisor or Kata Containers
  6. Attach what humans used to add by hand

    Workload identity for the environment, and a scoped, least-privilege identity for the agent that asked for it. Telemetry wired to the platform's collectors. Cost attribution tags that match the cost center from station three. The ownership record registered in the catalog. None of this needs a person, and all of it used to wait for one.

    What moves to the next station

    An environment that is observable, attributable, and reachable only by the identities the policy allowed.

    Building blocks
    • SPIFFE / SPIRE
    • OpenTelemetry
    • OpenCost
    • Backstage catalog
  7. Hand over, then reclaim

    The requester receives endpoints, credentials through the identity system, and the evidence bundle: policy decisions, owner, classification, budget. The TTL and the owner record make reclamation automatic. The factory emits its own metrics at this station: lead time, zero-touch, time to owner.

    What leaves the line

    An environment in use, a metrics record, and a scheduled reclaim.

    Building blocks
    • Prometheus
    • Kyverno cleanup policies
    • Argo CD or Flux

CNCF project    Other open source, Kubernetes SIG or cloud-provider building block

Terrain

A cloud Infrastructure Factory is the easy case

Everything above quietly assumed a hyperscaler underneath: an API for every resource, capacity that appears in seconds, a meter on everything. That is the easy case. Many of the infrastructure that matters sits on something less cloud-like: a virtualization estate, bare metal, a sovereign or regulated cloud, factory floors and branch sites. The factory model holds there too. It has more stations and a harder first mile.

Virtualization48%of VMware customers plan to reduce their footprint by 2028. Only 4% have completed a full migration. The estate most factories vend from is itself moving.Virtified, April 2026; CloudBolt, February 2026
Sovereign cloud$23bnforecast European spend on sovereign cloud infrastructure in 2027, up from $6.7bn in 2025. Europe is expected to pass North America that year.Gartner, February 2026
Supply chain35–40weeks of lead time for server power-management chips in 2026, up from 21 to 26 weeks. Capacity off the cloud arrives on a truck, on a schedule the factory does not control.TrendForce, April 2026
Self-service83%of organizations say they need to increase or improve self-service provisioning of on-premises storage. The demand for the factory is already there; the APIs are not.Enterprise Strategy Group, August 2025, 380 organizations

Cloud-likeness, defined

NIST's 2011 definition of cloud computing lists five essential characteristics: on-demand self-service, broad network access, resource pooling, rapid elasticity and measured service. An Infrastructure Factory consumes all five. A hyperscaler supplies them as APIs. Wherever one is missing, the factory has to manufacture it before it can vend anything, and that is the whole difference between the cloud factory and every other kind.

NIST characteristicWhat a cloud factory gets for freeWhat a less cloud-like factory builds firstBuilding blocks
On-demand self-serviceEvery resource has an API: accounts, identity, networks, databases. Vending lines such as Account Factory already exist.APIs exist for compute (vSphere, OpenStack, KubeVirt), but network, firewall, IP addresses, DNS and storage sit behind tickets. The factory needs a source of truth and an API façade over every one of them.NetBox or Nautobot, Ansible, OpenTofu providers, Crossplane providers, Cluster API providers for vSphere, OpenStack and Metal3
Broad network accessA global backbone, managed load balancers, DNS and certificates on request.VLANs per site, firewalls with change windows, load balancers as appliances. The network has to become a policy the factory enforces instead of a ticket it waits for.MetalLB, Cilium or Calico network policy, cert-manager, external-dns
Resource poolingMulti-tenant by construction. Isolation is an account boundary.Dedicated hardware per team, licenses per socket, snowflake clusters. The factory has to create the pool: a virtualization layer, and tenancy inside clusters.KubeVirt, OpenStack, vCluster, Capsule, Kamaji, namespaces with quotas
Rapid elasticityCapacity appears in seconds and leaves the bill when released.Capacity arrives on a truck. Lead time has a floor set by procurement, so the factory keeps inventory: warm pools, buffer stock, reservations and forecasts.Metal3, Tinkerbell, Ironic, MAAS, Cluster API, SandboxWarmPool
Measured serviceA billing API, tags, and cost allocation down to the resource.No meter. Cost is a depreciation line in a spreadsheet. The factory has to build metering and showback before cost attribution can be one of its outputs.OpenCost with on-premises pricing, Kepler, Prometheus

Two things NIST does not list but a hyperscaler also supplies: one identity system across every resource, and a certified compliance posture under a shared-responsibility model. Off the cloud, the factory owns both. Keycloak and SPIFFE/SPIRE take the first; the evidence bundle from station seven takes the second.

What changes on the line

The first three stations do not change. Intent, contract and pre-check look the same whether the environment lands in an AWS account or on a rack in Nuremberg, and that is the point of the contract: the requester cannot tell where the environment will run, and should not have to. Everything between the pre-check and the hand-over is where the terrain shows.

Lead time has a floor the factory cannot remove: a VM from a warm hypervisor in minutes, a bare-metal node with firmware, RAID and an operating system image in tens of minutes to hours, a new rack in the weeks to months that 2026 component lead times dictate. The factory's job is to hide that floor behind inventory, never to pretend it is zero.

StationCloud factoryLess cloud-like factory
1 to 3 · Intent, contract, pre-checkUnchanged.Unchanged. The contract hides the terrain.
4 · VendOne call to a vending line.Splits into capacity reservation, node provisioning (firmware, RAID, image), network and storage attach, cluster bootstrap, and the GitOps hand-off.
5 · SandboxAgent Sandbox on a managed cluster.The same, on a cluster the factory built. Warm pools matter more, because nodes do not appear on demand.
6 · AttachIdentity, telemetry and cost tags exist as cloud primitives.Identity federation and metering have to be built before they can be attached.
7 · Hand over, reclaimDelete the resources and the bill stops.Return the hardware to the pool, wipe it, and make it visible to station zero again.
0 · Supply (new)None.Procurement, inventory and forecasting. Capacity as stock, with a published fill rate.

Different terrains, require one contract

Hyperscaler cloud · the easy case

Cloud-like on all five counts

The vending lines exist: Account Factory, subscription vending, project factory. The hard stations are the non-technical ones, the fields no agent can sign. Factory effort goes into contracts, defaults and delegations, and almost none into supply.

Account Factory for Terraform · Crossplane · Kratix · OPA · Kyverno · Agent Sandbox

Virtualization estates · an estate in motion

APIs exist, self-service does not

vSphere, Nutanix, OpenStack, Proxmox and KubeVirt all have APIs. Few are exposed as self-service, because change advisory boards and per-socket licensing sit in front of them. The Broadcom licensing changes turned this terrain into a construction site: 48% of VMware customers plan to shrink their footprint by 2028, while only 4% have finished leaving. KubeVirt, a CNCF incubating project heading for graduation with release 1.8, and Forklift for moving VMs are the open source route out. An IDC-documented enterprise cut provisioning from three to four days to about thirty minutes once self-service replaced its approval queue.

KubeVirt · Forklift · Cluster API for vSphere and OpenStack · Crossplane · NetBox · Ansible

Bare metal and edge · physics in the loop

Capacity arrives on a truck

Factory floors, stores, telco sites and GPU clusters. In 2026, component lead times stretched to 35 to 40 weeks for power-management chips and 21 to 26 weeks for baseboard controllers, and HGX-class GPU systems quoted 8 to 20 weeks. Thousands of sites with intermittent connectivity mean pull-based GitOps and immutable images. Here the factory ships images, not environments, and station zero is the one that decides whether a request at 3am can be served at all.

Metal3 (CNCF incubating since August 2025) · Tinkerbell · Ironic · Talos or Flatcar · Flux · Argo CD · Kepler

Sovereign and regulated · evidence is the product

The contract carries the jurisdiction

NIS2 applies since October 2024, the Digital Operational Resilience Act since January 2025 (the EU regulation, not the metrics) and the Data Act since September 2025. Gartner expects European sovereign cloud spending to more than triple to $23bn by 2027. Sovereign providers expose thinner API surfaces and fewer managed services, so more of the five characteristics fall to the factory. Data residency and operator independence become contract fields, and the evidence bundle from station seven is the artifact the regulator reads.

OPA · Kyverno · SPIFFE/SPIRE · Harbor · Zarf · Sigstore · OpenCost

The contract, the metrics and the first three stations are the same on every terrain. What changes is how much of the cloud the factory has to build before it can sell an environment. In the cloud, the factory is policy and ownership. Everywhere else, it is policy and ownership plus a supply chain. The hardest station in the cloud is a signature. The hardest station on bare metal is a truck.

Metrics

Measure it like a factory

The software-factory crowd brings numbers: pull requests per week, tokens per engineer, 3 to 5x. An Infrastructure Factory needs its own set, the equivalent of DORA for environments. Six are enough to start, and each has a counterpart on the software side.

Counterpart: lead time for changesLead time to a compliant environmentFrom intent submitted to environment handed over, with its evidence attached.Factory target: hours, measured per contract. Days means a ticket is still in the path.
Counterpart: deployment frequencyEnvironments per day per platform engineerDelivered environments, normalized by the people who run the factory.Why it matters: it is the number that justifies platform investment as agent demand scales.
Counterpart: PRs merged with no human reviewZero-touch rateShare of requests fulfilled with no human action anywhere in the path.Why it matters: its inverse is the ticket count. An agent at 3am only ever gets the zero-touch share.
Counterpart: change failure ratePolicy-violation rate on agent requestsShare of agent-originated requests the pre-check refuses, broken down by reason.Read it carefully: a rising rate usually means the contract is unclear to the agent, and rarely that the agent is reckless.
Counterpart: review pickup timeTime to ownerTime until a cost center and a compliance owner are attached to the environment.Factory target: zero. Ownership is a default of the golden path, never a step after it.
Counterpart: age of unmerged workOrphan rateResources alive past their TTL with no reachable owner.Why it matters: it is the day-2 health of the factory, and the first number finance will ask for.

Little's law for environment requests

120requests open at any moment with a ticket in the path
120In flight · ticket
0.6In flight · factory

At 40 requests a day and a lead time of 3 days, 120 environment requests are open at any moment. Through a factory with a lead time of 20 min, fewer than one is.

Work in progress = arrival rate × lead time. Little's law holds for any stable queue, whatever is in it. The inputs are yours; the arithmetic is the only claim.

Factors

Twelve factors of an Infrastructure Factory

The Twelve-Factor App of 2011 told a generation how to write software that a platform could run. These twelve say how to build the platform that agents can run on. Each one is written as a test. A factory either passes it, or there is a ticket in the path somewhere.

Intents over resource lists.

Defaults over approvals.

Contracts over tickets.

Evidence over reports.

There is value on the right of every line. The factory encodes it once, so that nobody has to supply it per request.

  1. I

    Intent in

    The input is an intent: what the environment is for, who owns it, what data it touches, how long it lives. Never a resource list. A resource list is the factory's output, translated back into its inputs by a human.

  2. II

    Environment out

    The output is an environment that is owned, compliant, observable and cost-attributed on arrival. Anything less is a resource, and resources are what the ticket queue used to deliver.

  3. III

    Machine-readable contracts

    What the requester may choose and what the platform decides is a versioned schema. The API is the interface; portals and CLIs are views on it. Agents are first-class customers, and a contract an agent cannot read is a contract a human will be asked to interpret.

  4. IV

    No ticket in the path

    Humans decide once, at the contract, or encode the decision as a default, a delegation or a policy. A human may be notified. A human is never a blocking station. Platform engineers build the line; they do not fulfil requests.

  5. V

    Policy before provisioning

    The pre-check returns allow or deny, with a reason the requester can act on, before anything is created. Nothing is built to be torn down by a reviewer later, and a denial is an answer rather than an escalation.

  6. VI

    Ownership by default

    Cost center, compliance owner, on-call owner and data classification are inherited from the catalog, never typed into a form. No reachable owner, no environment.

  7. VII

    Scoped identity for every requester

    Every request runs under a short-lived, least-privilege identity, agents included. Credentials move through the identity system and never through a prompt, a ticket or a chat message.

  8. VIII

    Everything expires

    A lifetime is part of every intent. Renewal is a request like any other and carries the same owner. Reclaim is automatic, and an orphan is a defect of the factory, not of the requester.

  9. IX

    Evidence ships with the product

    The decision record travels with the environment: policies evaluated, owner, classification, budget, jurisdiction. An audit reads the factory's output, not a spreadsheet assembled afterwards.

  10. X

    Capacity is inventory

    Warm pools, quotas and buffer stock absorb bursts. Where elasticity is missing, the factory keeps stock and publishes its fill rate instead of pretending that lead time is zero.

  11. XI

    The factory measures itself

    Lead time to a compliant environment, environments per platform engineer, zero-touch rate, policy-violation rate, time to owner, orphan rate. Published to the factory's customers the way an SLO is, and argued about the same way.

  12. XII

    Built from replaceable parts

    Contracts outlive tools. Every station is a building block that can be swapped, Crossplane for Kratix, Kyverno for OPA, a hyperscaler for a rack, without changing what the requester asks for or what the requester gets back.

Prior art

What already exists, and what this is not

The term is open. Any may have said since mid-2025 that the bottleneck moved from code to infrastructure delivery, and a few projects already use the words.

Who uses the words

UsageWhoWhat it meansRelation
InfraFactoryMozeBaltyk, open source on GitHub"A factory for building cloud infrastructure, provider-agnostic and modular": OpenTofu, cloud-init, K3s or RKE2 and Flux to stamp out VMs and clusters.A vending line
C.I. Factory, Capability Infrastructure Factorybonometti718, open source on GitHub"Machine-native infrastructure for autonomous AI agents": small verifiable capabilities such as an "Agent Economic Action Gate".Adjacent, tiny
AI FactoryNVIDIA, Presidio and othersPhysical accelerated datacenters that "convert energy into tokens".Different thing
Account Factory, Account Factory for TerraformAWS Control Tower, with HashiCorpA Terraform pipeline that vends and customizes governed AWS accounts inside a landing zone.A vending line
Subscription vendingMicrosoft Azure Architecture CenterBicep and Terraform modules that create workload landing zones at scale.A vending line
Project factoryGoogle Cloud Foundation FabricA Terraform module that vends GCP projects from YAML.A vending line
"A factory with multiple production lines"Cellenza, Azure consultancyIndustrialized landing-zone delivery.Closest precedent
Platform OneUS Air ForceA "software factory" that in practice provides DevSecOps as a service, Kubernetes, hardened images and tooling to other software factories.Lineage

The symmetry is explicit

Both factories have inputs, production lines, quality gates and a throughput number. Nobody has drawn the second line next to the first.

The queue is non-technical

Cost center, compliance owner, data classification, on-call owner. Agents cannot sign them, so the factory encodes them as defaults, delegations and policies.

Platform engineering was never about Kubernetes

It was about letting factories run. Platform One proved it before the agents arrived.

Sources

Sources and caveats

Everything quoted on this page comes from the publications below. Links are left out on purpose until each one has been re-verified against the original.

  1. Factory.ai, Matan Grinberg and Eno Reyes, "Factory 2.0: From coding agents to software factories", 15 June 2026.
  2. BCG Platinion, "The Agentic Software Factory", 26 March 2026.
  3. StrongDM engineering rules, February 2026, as reported by Simon Willison and Stanford CodeX; Justin McCarthy on token spend per engineer.
  4. Sundar Pichai, Google Cloud Next keynote, 22 April 2026; earlier figures from Alphabet earnings calls, October 2024 and April 2025.
  5. Anthropic spokesperson, January 2026; Krishna Rao, May 2026.
  6. Uber Engineering, "Running a Software Factory Efficiently at Uber Scale", 27 August 2026.
  7. Faros AI, "AI Engineering Report 2026: The Acceleration Whiplash", April 2026. 22,000 developers across more than 4,000 teams, telemetry from Faros customers only.
  8. CircleCI, 2026 engineering telemetry, blog post of 18 February 2026. 28,738,317 workflows.
  9. Google Cloud, "Bringing you Agent Sandbox on GKE and Agent Substrate", May 2026; Kubernetes blog on Agent Sandbox, March 2026.
  10. CNCF blog, "Why sandboxing your agent is not enough", July 2026.
  11. StackGen, Asif Awan, "Intent-to-Infrastructure: Platform engineers break bottlenecks with AI", platformengineering.org, sponsored, June 2025.
  12. platformengineering.com, "AI Writes Code Faster Than Most Engineering Organizations Can Ship It".
  13. Port, "AI Software Factory: What It Is, Why You Need One, Who Owns It".
  14. Encore, on AI software factories and provisioning to the team's own cloud account.
  15. Syntasso, Kratix and Kratix Agentic documentation and customer quotes.
  16. Humanitec, Platform Orchestrator documentation on service users for AI agents.
  17. Bunnyshell, "AI Writes Code in Minutes, Your Team Waits Days for Staging"; Megaport, "Are AI Coding Agents the New CI Bottleneck?"; Augment Code, "AI-Native Engineering".
  18. TechTarget, "AI agents' role in IT infrastructure is expanding".
  19. AWS re:Invent 2025, session OPN303, "Building agentic AI platform engineering solutions with open source".
  20. Gartner, Hype Cycle for Platform Engineering 2026, as summarized by TrueFoundry.
  21. Thoughtworks, Rickey Zachary, on documentation and policy consistency for agents.
  22. US Air Force: Kessel Run (2017), Platform One (December 2019); Lauren Knausenberger (2021).
  23. Bob Bemer (1968); System Development Corporation Software Factory (1975); Greenfield and Short, Software Factories (2004).
  24. MozeBaltyk, InfraFactory; bonometti718, C.I. Factory; AWS Control Tower Account Factory and AFT; Azure subscription vending; Google Cloud Foundation Fabric project factory; Cellenza on landing-zone factories.
  25. NIST, Special Publication 800-145, "The NIST Definition of Cloud Computing", September 2011.
  26. Virtified, VMware customer research, April 2026 (48% plan to reduce their footprint by 2028); CloudBolt, VMware migration survey, February 2026 (86% reducing, 4% fully migrated).
  27. Gartner, forecast of European sovereign cloud infrastructure spending, February 2026, as reported by The Register and Data Center Dynamics.
  28. CNCF, "Metal3.io becomes a CNCF incubating project", 27 August 2025; CNCF blog on Metal3 at KubeCon Europe, 23 March 2026.
  29. SiliconANGLE, "Kubernetes virtualization approaches CNCF graduation", 30 March 2026; InfoQ on KubeVirt 1.8 and its hypervisor abstraction layer, March 2026.
  30. TrendForce, extended component lead times and server growth, April 2026, as reported by Evertiq and TelecomTV; HGX B200 lead times as quoted by hardware resellers in 2026.
  31. Enterprise Strategy Group, brief on on-premises hyperscaler-style infrastructure, August 2025, 380 North American organizations.
  32. IDC white paper on self-service private cloud, 2026, sponsored by VMware (provisioning from three to four days to about thirty minutes).
  33. European Union: NIS2 Directive, applicable from 18 October 2024; Digital Operational Resilience Act, applicable from 17 January 2025; Data Act, applicable from 12 September 2025.
  34. Adam Wiggins, "The Twelve-Factor App", 2011.

Caveats

  • Most software-factory statistics are self-reported by vendors or executives, and definitions shift. Google's figure changed definition several times. BCG's 3 to 5x and StackGen's 75% are practitioner and vendor claims, not independent benchmarks.
  • The Faros AI, CircleCI, Uber and GKE figures come from each organization's own publication. Faros covers only its own customers. Factory.ai's dashboard numbers circulate only through a secondary source and are left off this page.
  • The Gartner content came through a vendor summary. Confirm the "Agent Experience" name and its rating in the report itself before repeating them.
  • The VMware figures come from vendor surveys (Virtified, CloudBolt), the provisioning-time example from an IDC white paper sponsored by VMware, and the GPU lead times from reseller quotes. The TrendForce, Gartner and Enterprise Strategy Group figures were read through press coverage, not the original reports.
  • The prior-art check covers web-indexed content. Talks at internal or regional events may use "infrastructure factory" without being indexed.