The Coe Lab
← Back to Blog

Google's AX: When the Search Giant Reinvented Infrastructure for AI Agents

September 21, 20266 min read
AIGoogleinfrastructureagentsopen-source

Google's open-source Agent Executor (AX) treats AI agents as a fundamentally new workload class — not microservices, not batch jobs, but stateful actors that need sub-second suspend/resume at billion-scale. Here's what it changes.

Google just dropped something that doesn't look like a model release, doesn't sound like a chatbot update, and isn't trying to be either. It's called AX — Agent Executor — and it might be the most consequential piece of AI infrastructure released this year.

AX is an open-source, declarative control plane for running AI agents at scale. Not one agent. Not ten. Billions of them, potentially, across a single cluster. And it's built on a premise that sounds obvious in hindsight but that nobody else has articulated this clearly: agents are a completely new kind of workload.

The Problem: Agents Don't Fit Existing Infrastructure

Here's the thing about AI agents that traditional infrastructure wasn't built for: they're stateful, bursty, and long-running. An agent might compute intensely for thirty seconds — running code, calling tools, reasoning through a problem — and then sit idle for minutes waiting for a model API response, a human approval, or an external tool to finish.

If you run that agent in a standard Kubernetes pod, you're paying for the entire pod while it sits there doing nothing. Scale that to thousands or millions of agents, and the economics break down fast. Traditional orchestrators built for stateless microservices or predictable batch jobs become cost-prohibitive when keeping idle sandboxes running.

AX's answer is to treat agents as what they actually are: lightweight actors that can be checkpointed, suspended, and resumed in under a second. When an agent is waiting on a model response, it gets suspended. When the response comes back, it resumes — with full filesystem state and working memory intact — in less than 500 milliseconds.

How AX Works: Four Primitives

AX is built on top of Agent Substrate, a secure-by-default execution runtime that Google also open-sourced. Together, they provide four core primitives that handle the entire agentic lifecycle declaratively:

  • Tasks — the unit of agentic work. Each task runs as a lightweight actor with its own sandbox, network policy, and lifecycle. You declare what the agent should do; AX handles the rest.
  • Workspaces — reproducible environments that can be described in plain English. AX can hand a workspace description to an agent on first boot to install toolchains and verify dependencies automatically.
  • Network Policies — zero-trust isolation built in. Every task gets fenced network access by default, preventing agents from making unauthorized external calls.
  • Models — a unified interface for connecting agents to any LLM provider, abstracted away from the execution layer.

The whole thing is configured via YAML, Kubernetes-style. You write a task.yaml file describing your workspace and goal, run ax apply, and AX handles scheduling, isolation, and lifecycle management. You can SSH into running tasks, suspend and resume them, and watch their progress in real time.

The Density Numbers Are Staggering

Agent Substrate's demo shows ~250 stateful actors running across just 8 physical pods. That's 30x oversubscription — meaning the system is juggling 30 times more agents than it has physical containers, by suspending idle ones and resuming them on demand.

The runtime achieves sub-500ms resume operations at over 500 suspend/resume activations per second. That's not theoretical — it's demonstrated. And because suspended agents consume minimal resources, the cost of running a million idle agents drops to a fraction of what it would cost on traditional infrastructure.

For context: if you tried to run a million concurrent AI agents on standard Kubernetes pods, you'd need a massive cluster running at full capacity even when 95% of those agents were waiting on API responses. AX's actor multiplexing turns that idle time into spare compute capacity.

Framework-Agnostic by Design

One of AX's most important design decisions is that it doesn't care which agent framework you use. Because it manages standard OCI containers at the kernel level via gVisor, it can host agents built on any stack:

  • Google's own Agent Development Kit (ADK) with session state preservation
  • LangChain agents and tool calls
  • Claude Code, OpenAI Codex, and other coding agents with stateful filesystem persistence
  • Model Context Protocol (MCP) servers deployed as sandboxed actors
  • kagent, the CNCF sandbox project for Kubernetes-native AI agents

This matters because the agent ecosystem is fragmented. Locking into a single framework's runtime limits your options. AX's framework-agnostic approach means you can mix and match — or switch frameworks entirely — without rebuilding your infrastructure.

Built for Research, Not Just Production

AX explicitly targets two audiences: developers building production agent systems and researchers running large-scale experiments. For researchers, the ability to spin up massive numbers of reproducible sandboxes is a game-changer for collecting training trajectories, running reinforcement learning loops, and evaluating agents at scale.

The generative workspace feature is particularly interesting for research: describe the environment you need in plain English, and AX prepares it automatically before your task starts. No more manually scripting Dockerfiles for every experiment variant.

The Catch: It's Early

AX and Agent Substrate are explicitly in early development. The README states plainly that APIs are "almost guaranteed to change," there are no backward compatibility guarantees, and it's not ready for production use. The project is also marked as not an officially supported Google product and is ineligible for Google's Open Source Software Vulnerability Rewards Program.

But the ideas are what matter here. The conceptual shift — treating agents as a distinct workload class that needs suspend/resume, dense multiplexing, and declarative lifecycle management — is the kind of architectural insight that tends to become industry standard once articulated clearly.

Why This Matters

Every major tech company is grappling with the same problem: how do you run AI agents at scale without going bankrupt on compute costs? Google's answer — open-source the infrastructure, make it framework-agnostic, and treat agent execution as a first-class systems problem — is a signal that the agent infrastructure layer is about to become as important as the model layer.

If you're building agent systems, AX is worth watching. Not because it's production-ready today, but because it's articulating the architectural patterns that everyone will eventually need: actor-based scheduling, sub-second suspend/resume, zero-trust isolation, and declarative configuration. The agent infrastructure wars are just getting started, and Google just fired the opening shot.

Related Posts

Strata: When a 125B AI Model Ran on a Gaming PC at 100 Tokens Per Second

A new open-source tool called Strata lets you run Qwen 3.8 Flash Next — a 125-billion-parameter model — on an ordinary gaming PC with an RTX 4090. Nothing leaves your machine, and it's faster than you can read.

Oct 5, 2026• 6 min

When AI Agents Spend Your Money While You Sleep: Why Hard Budget Caps Are Becoming Non-Negotiable

AWS and Google Cloud finally launched hard spending limits in the same month. It's not a coincidence — it's a response to AI agents that can rack up thousands of dollars before you wake up.

Oct 4, 2026• 6 min

When Utah Banned VPNs: How a Court Stopped a Law That Demanded the Technically Impossible

A federal judge just blocked Utah's unprecedented anti-VPN law, ruling that lawmakers cannot mandate perfect geolocation — a technical impossibility. The case reveals a deeper problem: when legislation outruns engineering.

Oct 3, 2026• 7 min