Infraze is under development and not yet generally available. What follows is the architecture, a set of claims we think are defensible, and a 90-day plan with exit criteria, written down while the work is early so that the plan can be checked against the result. No measured results are published here, because there are none yet.
The short version: an infrastructure agent that runs in a terminal, provisions and monitors cloud infrastructure step by step, and treats the model underneath as a dial rather than a dependency. The same task can run on an expensive engine that is fast and reliable, or a cheap one on open weights, chosen per step rather than per project.
Facts at a glance:
- Maker: Velofy Lab, Delhi, India
- Status: in development as of 4 October 2026. Early access by conversation
- Category: agentic DevOps, a term Microsoft introduced at Build 2025 and describes as a paradigm rather than a product
- Form: a command line interface (CLI), run locally or in continuous integration
- Engines: Claude Code CLI, or Kestrel on open weights
- Memory: Trove, file based, inspectable, per project
- Targets: cloud provisioning, continuous integration and deployment (CI/CD), Grafana dashboards and alert rules
- What it does not own: your cloud account, your model weights, or your data
The problem is who sells you the agent
There is no shortage of agents that will provision infrastructure. There is a shortage of agents that are not also selling you the infrastructure.
Amazon Web Services (AWS) shipped the AWS DevOps Agent to general availability on 31 March 2026. It investigates incidents, handles site reliability tasks, and spans AWS, Azure and on-premises systems. It is good, and it is Amazon’s. Datadog’s Bits AI for site reliability engineering reached general availability in December 2025, and it is Datadog’s. StackGen provisions from intent across the lifecycle, as a software service. Each is competent inside the business model that produced it.
That business model has a consequence. An agent built by a cloud vendor will optimise for that cloud, an agent built by an observability vendor will route you through that observability product, and an agent sold as a hosted service will hold your infrastructure state on someone else’s servers. None of this is dishonest. It is just what those products are for.
The gap we see is a tool that provisions and watches infrastructure while having no stake in which cloud you pick, no stake in which model you run, and no copy of your state.
Prior art
This is a crowded category and the paper is weaker if it pretends otherwise.
| Product | What it does | Shape |
|---|---|---|
| AWS DevOps Agent | Incident response and reliability work, AWS and Azure and on-premises | Vendor service, billed per second |
| Datadog Bits AI SRE | Incident investigation inside Datadog | Vendor service |
| StackGen | Intent-based provisioning and remediation | Software as a service |
| Clanker | Open-source Go CLI across AWS, GCP, Azure, Kubernetes, Cloudflare and more | Self-hosted binary, MIT licence |
Clanker is the closest thing to what we are describing, and it is the honest comparison. It is an open-source command line agent for any cloud, it generates and applies infrastructure plans, and it exposes a Model Context Protocol server so that Claude Code or Cursor can drive it. If you want a vendor-neutral infrastructure CLI today, it exists and you should look at it.
The awesome-devops-ai list counted 474 tools, agents and servers in this space as of July 2026. Infraze is not first and this paper does not claim it is.
What we think is not yet covered is the combination of three things: the model as a per-step choice with published economics, a memory layer that outlives the session, and no commercial interest in the answer.
What Infraze proposes
One: the engine is a dial. Each step in a plan is routed to an engine. Claude Code for steps that need judgement, Kestrel on open weights for steps that are mechanical. The routing decision is explicit, logged, and configurable per step class rather than fixed for the whole run.
Two: it remembers. Trove holds what was built, why, which engine built it, and what broke afterwards, in files you can read, inside the repository. Most agents in this space are stateless per session or keep state in the vendor’s store. The second time Infraze touches an environment it should know what it did the first time.
Three: it owns nothing. Local binary, your cloud credentials, your repository, your Trove directory. No hosted control plane holding infrastructure state.
A run goes like this. The engineer states an intent: a staging environment on one cloud with a Postgres database, a Grafana dashboard and a deploy pipeline. The planner turns that into a graph of steps. The router assigns each step an engine. Steps execute through a tool layer that wraps Terraform, kubectl, the GitHub CLI and the Grafana API rather than letting a model drive a cloud console directly. Anything destructive stops at an approval gate. Every run writes a record to Trove.
The engine dial is the actual argument
This is the part worth testing, so it is worth stating precisely.
Infrastructure work is not uniformly hard. Designing a network topology for a multi-region failover is a judgement problem. Adding a label to a manifest, bumping a chart version, or wiring a Grafana panel to a datasource that already exists is a mechanical problem. Today both get sent to the same expensive model, because the model is chosen once at the start.
The economics of not doing that are large. Our decision economics survey collected measured figures on a typed-decision task showing a tenfold cost spread and a ten to sixteen times latency spread between the cheapest small model and frontier models in thinking mode. Infrastructure steps are bigger than a single typed decision, but the shape holds: most steps in a provisioning run do not need the best model available, and paying for it anyway is the default behaviour of every agent in the list above.
So the dial is a routing problem, and routing is a decision problem. Which puts Infraze downstream of the open question we wrote about in the open loop: a fast router has to know when it is unsure, and current small decision models are poorly calibrated. Our abstention survey reports one benchmark finding a typed-decision model admitting ignorance on 49.7% of forced-uncertainty items where frontier models admit on 97.3 to 100%, with expected calibration error of 0.246 against a frontier range of 0.039 to 0.122.
That matters more here than in most places, because infrastructure changes are not uniformly reversible. A mis-routed step that drops a database is not recoverable by apologising. So the proposal depends on a mitigation rather than on the router being right: route by declared step class first, use model confidence only to escalate upward and never to downgrade, and gate anything destructive behind a human. Escalation is cheap, confidence is not trustworthy enough to save money with, and the approval gate is the backstop.
On the engines themselves: Claude Code is the expensive, fast and reliable option. Kestrel is the cheap option and runs on open weights. We are not naming the specific open-weight model in this paper because we have not finished benchmarking it, and naming it before we can publish the numbers would be the kind of claim this lab tries not to make. The model and its measured cost per completed step will be published together.
A project holds environments, one per provider. A plan changes an environment and is made of steps checked against policies. Each step routes to an engine, may require an approval, and executes as one or more runs that produce artifacts. Runs write memory entries scoped to the project, which is what makes the second visit cheaper than the first.
Scope
Three surfaces, in order of how confident we are that they are tractable.
Monitoring setup. Grafana dashboards, datasources and alert rules. This is the safest place to start: the artefacts are declarative, the blast radius of a mistake is a wrong graph, and the work is repetitive enough that the cheap engine should handle most of it.
Continuous integration and deployment. Pipeline configuration, build and deploy steps, environment promotion. Moderate risk, mostly file generation, verifiable by running the pipeline.
Cloud provisioning. Networks, compute, managed databases, identity and access. Highest value and highest risk, which is why it goes last and why every destructive operation stops at a gate.
What is running today, and what we are building
| Component | State |
|---|---|
| Kestrel, the cheap engine | Exists, documented, kestrel.velofy.co |
| Trove, the memory layer | Exists, open source, in use |
| Claude Code, the reliable engine | Exists, third party |
| Infraze planner, router, tool layer | In build |
| Published benchmarks | Not yet, by design |
The three components Infraze composes are real and in use. The layer that composes them is what is being built now. Nothing in this paper should be read as a measured result, because the numbers are not in yet.
A 90-day plan
Days 1 to 21, the monitoring surface only. Build the CLI skeleton, the Trove schema and a tool layer for the Grafana API. Exit condition: Infraze can stand up a dashboard and an alert rule from an intent, twice, in a clean project and then an existing one, with the second run demonstrably reading the first run’s memory.
Days 22 to 49, the dial. Add the second engine and the step router. Exit condition: a published table of cost and wall-clock time per completed step for the same plan run entirely on Claude Code, entirely on Kestrel, and routed. If routing is not cheaper than all-Claude Code at equal completion, the dial is not worth the complexity and we say so.
Days 50 to 70, CI/CD and approval gates. Exit condition: a pipeline generated and green, with every destructive operation having stopped at a gate, measured as a count of gates hit rather than an impression.
Days 71 to 90, one cloud provisioning path end to end, in a disposable account, with a written rollback. Exit condition: the environment comes up, gets torn down, and the Trove record is sufficient for a second engineer to understand what happened without reading the Terraform.
Whichever way each stage falls, it gets published with the numbers attached. Day 49 is the one that decides the shape of the product.
Open questions
- Whether per-step routing beats a single good model once you count the cost of a wrong route. We do not know and day 49 is designed to find out.
- Whether a file-based memory stays useful past a few dozen runs, or becomes noise that nobody reads.
- Whether approval gates land somewhere between useless and unbearable. Too many and people automate past them, too few and the gate is theatre.
- Whether vendor neutrality is something buyers want or something engineers say they want. The vendor agents have distribution we do not.
- Whether the open-weight engine is good enough at the mechanical steps for the dial to have two usable positions rather than one.
What we could not verify
- We have not yet run any benchmark described here. Every number in this paper comes from previously published Velofy papers or from the cited third parties.
- We have not used Clanker, the closest comparable. Its capabilities are taken from its own documentation and should be checked before anyone relies on the comparison.
- The AWS and Datadog capability summaries come from vendor announcements, not from our own testing.
- Infraze has not yet been tested against a customer cloud account.
Sources
Read on 4 October 2026.
- Announcing general availability of AWS DevOps Agent, Amazon Web Services, 31 March 2026
- Agentic DevOps: evolving software development with GitHub Copilot and Microsoft Azure, Microsoft
- clanker, an autonomous systems engineering CLI agent, MIT licence
- awesome-devops-ai, a curated list of 474 tools, updated July 2026
- What a million decisions costs, Velofy Lab
- Abstention and honest uncertainty for small decision models, Velofy Lab
- The open loop: fast decision models in cloud operations, Velofy Lab