Back to blog posts

11 min

Execution substrate for AI agents: what & why it matters

Learn what an execution substrate is for AI agents, why traditional compute fails, and how to evaluate isolation, state persistence, and cost for production.

Nicolas LecomteNico is a founder of Blaxel, who usually writes about AI, agentics, and the future of AI runtimes.

In staging, your team's agent handles document parsing and code generation throughout multi-step workflow execution without complaint. Then you push to production. Within a week you're fielding Slack messages about slow cold starts. Sandbox state vanishes between invocations. Your security team's routine audit turns up a container escape.

The agent's infrastructure causes the problem. Most teams deploy agents onto compute designed for stateless web requests or long-running monoliths. Neither model fits how many agents operate. Many agents run short bursts of isolated compute. Stateful agents also hold memory across sessions. Agents executing untrusted code need strict security boundaries. That gap has a name: the execution substrate.

TL;DR

  • Definition: An execution substrate absorbs provisioning and isolation for agent workloads while preserving state. Teams ship agent logic instead of infrastructure.
  • Why traditional compute fails: Containers commonly assume a long-running PID 1 process, while default container runtimes are generally intended for trusted internal workloads; VMs provide stronger isolation for untrusted workloads. Serverless functions don't guarantee state persistence between invocations; execution environments may be reused, so in-memory or temporary filesystem state can sometimes persist, but durable state should be stored externally. Neither matches bursty, stateful agent sessions.
  • Core capabilities: Hardware-level isolation for untrusted workloads, resume from standby without cold starts, persistent filesystem state, and shared storage. Enterprise networking controls also matter.
  • Production mechanics: Standby-based substrates snapshot environment state instead of destroying it. The next invocation resumes without cold starts or idle compute charges.
  • How to evaluate: Judge the isolation model, state persistence guarantees, compliance certifications, and total cost. Include the infrastructure headcount you avoid.

What is an execution substrate for AI agents

An execution substrate for AI agents is the managed infrastructure layer beneath agent workloads. It provides compute, storage, networking, and lifecycle management. It sits between the agent's reasoning logic and raw cloud primitives. Those primitives include VMs and containers, along with object storage. The substrate absorbs provisioning and isolation, and it persists state while enforcing network policy.

Databases did this for storage. They abstracted raw disk I/O behind a query interface, and nobody builds B-trees by hand anymore. An execution substrate does the same for agent compute. It turns VM provisioning and snapshot management into primitives an agent consumes directly, with security isolation built in. Create an environment, execute code, persist state, resume later.

Agent orchestration frameworks define the agent's orchestration flow and its reasoning logic, which contains the tool definitions. Sequencing multi-step processes belongs to workflow engines. Model endpoints serve inference. The substrate sits beneath all three and determines where and how the agent's work executes. Optimize the wrong layer, and you spend months on problems the substrate should absorb.

Why traditional infrastructure breaks down for agent workloads

Enterprise teams default to VMs and Kubernetes because they already run them. Serverless functions are another common choice. Each one creates specific technical friction for agent workloads.

Conventional per-agent VM and container deployments assume long-running processes

Conventional per-agent VM and container deployment patterns are built around processes that start once and run continuously. Agent workloads often follow the opposite pattern. AgentSysBench covers production agent sessions. The median session executes for only 20% of its lifetime. A dedicated VM per agent burns paid compute during the remaining time. For Kubernetes, keeping a dedicated Pod alive for every potential agent quickly becomes wasteful.

Sharing hosts across agents trades cost for isolation risk. Containers share the host kernel. For agents executing untrusted or AI-generated code, kernel-level escape is a documented reality. NIST SP 800-190 is direct about it. "The use of a shared kernel invariably results in a larger inter-object attack surface than seen with hypervisors." CVE-2019-5736 gave attackers host root through a runc escape. CVE-2024-21626 allowed host binary overwrites from inside the container. Containers remain the standard for trusted first-party services. Arbitrary AI-generated code presents a different threat model.

Serverless functions lose state between invocations

Serverless platforms were designed for stateless request handling. Treat AWS Lambda environments as single-use. Follow the AWS Lambda documentation: "you should assume that the environment exists only for a single invocation." An agent may need filesystem state and loaded datasets. It may also need a cloned repository. Without persistence, it repeats expensive initialization on every invocation.

Each call incurs the re-initialization tax. A coding agent invokes tools 10.8 times per session on average. A repository clone or dependency install then repeats on every invocation inside one task.

Cold starts compound the problem. Typical Lambda cold starts range from under 100 milliseconds to over one second. Heavier runtimes fare far worse. An unoptimized Java Spring Boot function cold-starts at a p50 of 5,047 milliseconds. For interactive agent products, delays at that scale break the experience.

Some teams build bespoke sandbox infrastructure in-house on microVM stacks. The work demands specialized expertise in microVM configuration and kernel tuning, plus memory snapshot management. Building this infrastructure can tie up multiple senior engineers for months before the first production deployment.

The spend continues after launch. Patching and snapshot reliability work create an on-call surface the product team did not previously own.

Salary math matters at Series A through C scale. Bureau of Labor Statistics data puts the software developer mean at $148,100, with the 90th percentile at $214,670. Benefits add another 31.6% of total compensation in professional and technical services. A small sandbox team is therefore a substantial recurring infrastructure cost for an early-stage agent company.

Core capabilities of an execution substrate

A purpose-built substrate delivers compute and storage. Networking completes the infrastructure layer. Lifecycle management and security cut across all three.

Isolated compute with instant availability

For agents executing untrusted or AI-generated code, isolation belongs at the hardware level. Each workload gets its own kernel and memory space through microVM technology such as Firecracker. Firecracker has run production Lambda workloads since 2018. It uses under 5 MB of memory overhead per microVM. Separate kernels shrink the shared-kernel lateral movement risk documented above. Hypervisor bugs remain a residual surface.

Availability has a hard threshold. Jakob Nielsen's response time research identifies 100 milliseconds as an important limit. It is "about the limit for having the user feel that the system is reacting instantaneously." Coding agents invoke tools repeatedly. Every tool call pays any resume latency above that ceiling. Teams also need control over what is installed in each environment. They should not have to rebuild it per invocation.

Persistent state across sessions

Agents that lose context between invocations can't build on previous work. A code review agent would otherwise re-clone the repository before every review. That wastes both minutes and money. The substrate maintains filesystem and memory state. Agents resume exactly where they stopped.

Standby persistence pauses the full environment and its running processes for later resumption. Block storage volumes persist across environment destruction and recreation. Datasets and models belong there, along with artifacts that must survive environment deletion. Enterprise architectures usually need both. Standby preserves state only for the environment's lifetime.

Shared storage adds a third layer for multi-agent systems. Distributed filesystems let multiple agents read and write the same data concurrently. A planning agent can produce artifacts that execution agents consume without copying files between environments. To test a vendor's claim, ask whether running processes and files survive the pause.

Production-grade networking and security controls

Security reviews stall more substrate evaluations than benchmarks do. Enterprise security teams look for these networking controls in an agent execution layer:

  • Custom domains: Agent endpoints resolve under your own hostname. Customers never see a vendor URL in a preview link.
  • Static outbound IPs: Your traffic leaves from a fixed address. Partners can allowlist it without opening their network to a cloud provider's whole range.
  • Proxy routing: The proxy injects credentials server-side on outbound requests. Raw secrets never reach agent code or logs.

Compliance certifications are table stakes for procurement. SOC 2 Type II attests that controls operated effectively over a period. It confirms that they operated in practice over time. ISO 27001 certifies the management system behind them.

Some buyers use contract terms to add Zero Data Retention. For health data, a Business Associate Agreement is a legal requirement. HHS guidance treats a cloud provider processing electronic protected health information (ePHI) as a business associate. That holds even when the provider handles only encrypted data without the key.

How to run execution substrates in production

Modern substrates often build on hardware-assisted hypervisors like AWS Nitro to enforce isolation at the silicon level.

Follow the agent execution lifecycle

On a substrate, the path is short. The agent receives a task, and the substrate provisions an isolated microVM. The agent executes code while reading and writing files. It also makes network calls. After a short period of inactivity, the substrate snapshots the environment's full state. That snapshot covers memory and the filesystem. It preserves running processes. It then transitions the environment to standby. The next invocation resumes from that snapshot instead of rebuilding.

A Firecracker snapshot pairs a guest memory file with a microVM state file. Warm-cache restore measures 4.0–4.9 milliseconds. That speed shows snapshot restoration can fit within an interactive latency budget.

Contrast the traditional path. A script spins up a container and the agent waits out the cold start. Code then runs before the container is destroyed. The next invocation starts from scratch. On a standby-based substrate, the environment instead retains its state while idle; snapshot storage replaces active compute charges.

Scale from prototype to a production fleet

Differences surface across a production fleet. Each agent needs an isolated environment and consistent resume times. Shared storage and centralized network policy also apply across the fleet.

Blaxel is one managed implementation of this category. It is an infrastructure foundation for autonomous agents and a perpetual sandbox platform. Compute comes through Sandboxes, plus Batch Jobs.

Storage comes through Agent Drive, a shared filesystem for cross-session context. It is in private preview in the us-was-1 region. Volumes handle long-term persistence. Networking covers custom domains for sandbox previews. Dedicated egress gateways are in private preview. Proxy secrets injection is available.

Idle environments are where fleet economics break down. Across infrastructure-as-a-service and platform-as-a-service, 29% of cloud spend is wasted. Agent fleets amplify this problem because utilization varies by session. Standby-based billing aligns compute charges with active execution while snapshot storage covers paused state.

How to evaluate an execution substrate for your architecture

Over 40% of projects in agentic AI will be canceled by the end of 2027. Escalating costs and unclear business value drive those cancellations. Inadequate risk controls add another failure mode. Infrastructure choices affect all three.

Assess the isolation model and security posture

Start with the isolation boundary. Hardware-level isolation gives each workload a separate kernel through microVMs, while process-level isolation shares the host kernel across containers. gVisor sits between the two by intercepting system calls in user space. That improves on containers but stops short of a hardware-enforced boundary.

Ask which hypervisor or sandboxing technology enforces the boundary. Can one tenant's workload ever share a kernel with another tenant's? Push on the failure case. What happens when AI-generated code forks processes or writes to /proc? Review tenant isolation and encrypted transmission. Geographic policies and data residency also belong in the review alongside the sandbox boundary. Also verify whether custom images or private registries alter the enforced boundary.

Then match certifications to your industry. SOC 2 Type II and ISO 27001 are the baseline for general enterprise procurement. Health-adjacent workloads need the BAA obligations covered earlier. Request the actual audit reports on the first call.

Check state persistence and the data lifecycle

Find out how long state persists and what guarantees back it. Some platforms delete idle environments after a fixed window. Indefinite standby with fast resume is the benchmark for agents that maintain long-running context. Ask for resume latency measured after days or weeks of standby. Find out what happens to volume data when your workspace quota changes.

Distinguish standby persistence from block storage persistence before you map workloads. Assign session context to standby. Datasets and artifacts belong on volumes, while cross-agent shared data belongs on a distributed filesystem. A vendor that offers only one of these tiers forces workarounds later.

Begin the proof of concept with standby. Pause an environment with a process running, then confirm that both the process and its filesystem state return after the expected standby window. Next, destroy and recreate another environment to verify that volume data survives independently. Mount shared storage from multiple isolated environments as a separate test and check concurrent reads and writes. These tests separate a resumable session from durable storage and shared data access before production architecture depends on any of them.

Weigh operational complexity and total cost

Total cost for agent execution infrastructure extends well past per-second compute pricing. Snapshot storage accrues for every environment sitting in standby. Engineering time to operate whatever the vendor does not manage lands on the same budget.

Establish whether your team manages the infrastructure. Alternatively, determine whether the vendor handles lifecycle and scaling, along with monitoring. The build-vs-buy math shifts sharply when a managed substrate removes senior infrastructure hires from the plan. Those salary and benefits figures recur every year the system runs.

Check networking controls before signing. Custom domains and egress policies are hard to retrofit. Proxy routing creates the same problem. A substrate that lacks them natively leaves you building them yourself or switching vendors mid-roadmap. Finish with a short proof of concept. Run your real workload's concurrency and measure resume latency under load. Include your security team's questionnaire from day one.

Build your agent architecture on the right execution substrate

Choosing the right execution substrate determines whether your agent architecture reaches production or stalls at prototype.

Teams that deploy agents onto infrastructure built for traditional web workloads pay for it in engineering cycles. They must separately engineer responsive execution and retained state. Isolation for untrusted code adds another infrastructure requirement. A purpose-built substrate absorbs that work at the infrastructure layer.

Blaxel provides the execution substrate for autonomous agents. That means microVM isolation and persistent sandboxes that resume in under 25ms from perpetual standby. That latency supports repeated tool calls in interactive coding-agent sessions without a noticeable resume delay.

Agent Drive, in private preview, carries shared context across sessions. Blaxel manages custom domains and proxy routing. It also handles network policy at the infrastructure layer. Dedicated egress remains in private preview. Sandboxes remain in standby indefinitely with zero compute cost while idle. They carry only snapshot storage. Explore the platform or talk to the team.

FAQ

What should teams know before choosing an execution substrate?

Use these questions to separate the execution layer from adjacent tools and evaluate the operational tradeoffs.

What is the difference between an execution substrate and an agent framework?

Use portability as the test. If changing frameworks forces you to rebuild provisioning and persistent state, or requires rebuilding network policy, the execution layer is too tightly coupled. A substrate should let reasoning logic, tools, and orchestration change independently from managed compute, storage, networking, isolation, and lifecycle controls.

Why can't I use Kubernetes as my execution substrate?

Compare the work Kubernetes leaves to your team. Ask for an estimate covering hardware isolation and memory snapshots. Include session persistence, standby recovery, security hardening, and on-call ownership. Kubernetes may schedule the workload, but the decision turns on whether you want to build and operate those surrounding layers.

How does an execution substrate reduce infrastructure engineering costs?

Build a recurring-cost model. Include platform engineers and benefits, along with snapshot storage. Monitoring, patching, and on-call ownership also belong in the model. Then compare that total with a managed service that handles provisioning and scaling. The service also handles isolation, snapshots, and lifecycle operations.

What compliance certifications should an execution substrate provide?

Start with SOC 2 Type II and ISO 27001 audit reports, then map requirements to the workload. Health data may require a Business Associate Agreement, while retention may require contractual Zero Data Retention terms. Confirm geographic policy and data residency during the first call; certifications do not provide complete coverage.

Related articles