Right off the top, because this one needs a map before we walk into it. Daniel wants to look at agent sandboxes, and his point is that the word has quietly split into two jobs.
Two jobs sharing one word. That's the whole episode.
First meaning, the one most of us reach for: security. Sandboxing as a boundary so the model's code can't get at the host machine or your secrets. Sometimes obtrusive, in his words, sometimes the restrictive kind that fences an agent into a local box and gets in the way.
And the second meaning...
The second meaning, which he thinks wins: sandboxes as lightweight, often ephemeral filesystems for agents that live in the cloud. His argument is that the runtime is moving off your laptop. Desktop and mobile clients become thin viewports, they don't run the agentic process at all. And the moment an agent goes past text in, text out, it needs real Linux command line tools. Once it has those it can take an uploaded binary, edit audio, run a Python script.
Which changes what a sandbox is for.
Then he wants a survey of the popular approaches to provisioning these environments, and a clean line drawn between ephemeral and persistent remote environments. His example being a dev workspace, where you want a pinned Python and pinned dependencies that survive between sessions instead of being rebuilt every time. And from there, replicable base images, image libraries, the way cloud platforms hand you container images.
So we're using one word for two different products.
We are. Let's start with the fact of that, and then figure out when the two meanings stopped being able to ignore each other.
Cleanest way to split them. Meaning one, the security boundary. An isolated environment where the agent can run code freely and the host, the other tenants, and your credentials stay on the other side of the wall. That's the older meaning, the one that's been in computing since long before any of this.
Meaning two is a Linux workspace. A disk, a shell, a package manager, a filesystem the agent can actually work in. Cloudflare's version of that is nearly blunt about it, their line is essentially "your agent needs a computer."
And those two meanings aren't at war. Almost every product doing one is doing both.
They pull in different directions, though.
They do. A security boundary wants to be as small and as sealed as possible. A workspace wants to be useful, which means it wants tooling, network access, storage, all the things that make a boundary leaky. Same product, two instincts, and the vendors mostly answer whichever question they find more flattering. Buyers keep asking the other one.
Which is why the category feels foggy from outside.
And the category is enormous. Somebody on Hacker News sat down and counted roughly forty code execution sandbox products that launched in a single year. Forty. That's not a market finding its shape, that's a market where everyone smelled the same opportunity at once.
So the security meaning is the one we all know, the filesystem meaning is the one that got urgent, and to see why it got urgent you have to look at where the agent actually runs now.
Docker's framing is the sharpest version of this and they wrote it almost as a complaint. Agents do long horizon work now, tasks that run for hours, not minutes. And a laptop is built around a person. It sleeps when the lid closes, it throttles on battery, it drops the connection the second you walk out of the building with it.
The agent gets interrupted because you wanted lunch.
Because you wanted lunch. Cloudflare says the same thing from the demand side: agents create sandboxes on demand, per task, they expect them to be ready immediately, and then they expect to be able to pause and resume. That's not how a laptop works. That's not how anything on your desk works.
So the process moves to the cloud.
The process moves to the cloud, the phone becomes a viewport, and now the sandbox isn't a safety feature sitting beside the agent. The sandbox is the agent's computer. It needs a disk, a shell, and a way to install things.
Here's the part I find interesting, and it's the threshold Daniel points at. Text in, text out. As long as that's the only thing happening, the sandbox is basically a padding cell. The moment the agent gets a Linux command line, everything changes.
Concretely, that means it can clone a repository, install packages, configure the environment for the task it's been given. Cloudflare ships an image for exactly this, Debian Trixie Slim with Node twenty four twenty LTS baked in, and through their exec path the agent does all three of those things inside its own little machine.
There's a cleaner illustration in the Kubernetes agent sandbox work. The code execution use case there is a FastAPI endpoint. You post to it, it runs the code, it hands back stdout, stderr, and the exit code. That's it. But it's running inside a container with its own filesystem, its own processes, its own network stack.
That's the primitive. That's the whole thing in about three sentences.
And then the user facing payoff, which is where it stops sounding like infrastructure and starts sounding like a product. You upload a binary. The agent downloads it, runs it, and suddenly it's editing your audio file or executing your Python script or transcoding your video. The sandbox stopped being a fence and became a capability.
The fence never went away. It just stopped being the interesting part.
So ephemeral versus persistent. Give me the actual mechanical difference, not the marketing version.
Nirvana drew the clearest line I've seen on this. Ephemeral sandbox: no disk at all. Pure compute. It boots in about two and a half seconds. Persistent sandbox: it gets a dedicated volume, mounted at workspace, first boot takes six or seven seconds, and everything the agent writes into that folder survives a pause and survives a resume. Their rule of thumb is ephemeral for stateless jobs, persistent for long running agents that need to checkpoint and continue.
Two and a half seconds versus six or seven. That's the price of a disk.
And the disk is the entire difference.
Somebody's going to hear "persistent" and assume it means "never goes away." It doesn't.
It means the folder goes away slower. It means state you deliberately wrote survives the lifecycle event. It doesn't mean the machine is standing there waiting for you.
There's a counter argument to all of this that I want on the table, because it's the most interesting thing I read preparing for this.
Fly's computers versus sandboxes piece.
Their thesis is that persistence, isolation, wake latency, none of those are the real distinction. The real distinction is what happens when you stop paying attention to it. A sandbox is something that ceases to exist when nobody's watching. A computer is just there. You leave it, you come back, it's still there, it didn't decide anything while you were gone.
That's a strong challenge to the ephemeral orthodoxy, and I don't think it's just marketing. If you watched the last year of this category, the direction of travel is persistent. Vercel turned persistence on by default at GA. Fly built Sprites around a hundred gigabyte persistent disk. Docker's cloud sandboxes default to a one hour run that can stretch to twenty four.
So the ephemeral pitch and the persistent buildout are happening in the same twelve months.
Which usually means the ephemeral framing is the sales pitch and the persistent version is the product.
Now the piece Daniel's specific about, reproducibility. Why can't a dev workspace sandbox just be rebuilt from scratch every session? Everyone's building tools fast, containers boot in under a second, why does it matter?
Because you're not rebuilding the same thing. You're rebuilding something that looks the same and isn't. You pin Python to a specific minor version, you pin your dependency tree, and then six weeks later a rebuild quietly resolves a different transitive dependency, or the base image upstream moved, and now the environment your agent tested against is not the environment it's running in.
Small version drift.
Small version drift, and the failure it produces is the worst kind, which is a test that passes locally and fails in CI, or the reverse, and nobody can reproduce either one.
So the answer is snapshots and templates.
The answer is that you stop treating the environment as something you build and start treating it as something you copy. Cloudflare's snapshot model is the clean version, one snapshot starts many isolated environments from the same baseline, which cuts setup time and, their words, prevents environment drift. And they pitch it specifically for evals, comparing a prompt against a skill against a model, where if the baseline moves you've contaminated your comparison.
Modal does the same thing with snapshot filesystem, hands you back a reusable image. There's a demonstration where you prepare one state and fork three parallel sandboxes off it.
Which is a new capability. You couldn't do that with a laptop.
You could, you'd just have three laptops.
Kubernetes is doing it with custom resources, a sandbox template and a warm pool. The warm pool pre boots pods so that when a new environment gets requested it allocates fast instead of cold.
And Docker's version is the kit, which is a pre configured pre built sandbox for an agent, and they ship them for Claude Code, Codex, Copilot, Antigravity, Open Code, Hermes.
That's the image library Daniel was describing. Same idea as pulling a postgres image off a registry instead of writing a Dockerfile on a Tuesday night.
Except now the image isn't just the application's environment. It's the agent's entire sense of what a computer is.
So that's the fork. Ephemeral or persistent, and a base image either way. Which brings us to who's actually selling which version of it, and what the numbers say when you put them side by side.
Organize it by pattern rather than by name, because that's the only way it stays legible. First pattern: microVM per sandbox. E2B is the reference there, Firecracker underneath, boots around a hundred and fifty milliseconds, pause and resume preserving both memory and filesystem. And their infrastructure is Apache licensed, you can self host on Terraform and Nomad and Consul.
Open underneath.
Vercel is the same isolation primitive, Firecracker, but they went the other way on persistence and turned it on by default when it hit GA in May. Then in September they put Drives into public beta, durable volumes that outlive the sandbox itself. Fly's Sprites are Firecracker too, but they're the purist end of persistent, hundred gigabyte ext4 disk, supervised services, a public HTTPS URL per Sprite, and credential connectors.
Second pattern?
Workspace centric. Daytona is the one there. Docker container by default, Kata or Sysbox if you want the stronger isolation, and the whole product is built around a persistent dev environment, with fork, pause, resume, and snapshot. It's the pattern that most closely matches what Daniel described, a workspace that stays stable between sessions.
Third?
Sandbox as one primitive inside something bigger. Modal is the example. gVisor for isolation, VM sandboxes in beta now, snapshot and fork supported but no in place resume. The sandbox is a feature of their serverless GPU platform rather than the product.
Fourth pattern is platform native, where the sandbox is an extension of something you already run on.
Cloudflare is the clearest, containers on Workers plus Durable Objects, with a durable object scheduling policy, snapshots, and programmable egress handlers that live outside the sandbox. Vercel sits here too for the platform half. Runloop runs a dual layer, microVM plus container, with suspend and resume, and they've built agent gateways and an MCP hub around an eval focused workflow. Northflank lets you pick your isolation per workload, Kata or Firecracker or gVisor, ephemeral by default with optional volumes, GPU support, and bring your own cloud.
And then the last pattern, which is the one that isn't a vendor at all. Standardize the primitive.
Kubernetes SIG agent sandbox. Four custom resources, sandbox, sandbox template, sandbox claim, sandbox warm pool, gVisor or Kata for isolation, and integrations for the OpenAI Agents SDK, LangChain, OpenHands, DeepAgents, an MCP server, Gymnasium for reinforcement learning. That's an attempt to make the sandbox a boring cluster object instead of a product you buy.
Microsoft's version is the other half of that.
MXC, Microsoft Execution Containers, that's OS level containment rather than a cloud sandbox, four backends, process, session, WSL, microVM, and three modes, enforcement, learning, permissive. Their framing is the one worth carrying around: an agent cannot be its own security authority. The policy stays outside the agent's control so the agent can't grant itself more access.
Which is a sentence that applies to cloud sandboxes too, they just don't say it as plainly.
Now the numbers, and I want to handle these carefully because the headline numbers are the most misused thing in this entire category. ComputeSDK ran a burst test, bringing up sandboxes under load. Vercel, zero point six seven seconds median, one point one two at P ninety nine, and a hundred percent success rate. Modal, zero point eight eight. Runloop, zero point eight nine. E2B, one point six one. Cloudflare, five point zero six. Daytona, zero point two seven seconds median.
Fastest in the set.
At a thirty seven percent success rate.
...Say that again.
Thirty seven percent. Nearly two out of three requests didn't come up at all, and the ones that did were the fastest in the field. That is not a performance win, that's a benchmark methodology warning. A median calculated over a third of your attempts is telling you about your attempts, not your platform.
LogRocket measured cold starts and got a different order entirely. E2B at seven hundred seventeen milliseconds to create, six hundred sixty two to resume. Vercel eighteen fifty two to create, thirty three thirty three to resume, which is a very different number from zero point six seven.
Because it's a different measurement. One is a burst test under contention, one is a single cold start in isolation. Same vendor, one number is medians under load and the other is comfortable single shot latency, and the two aren't comparable. That's the trap the whole category falls into. You cannot line these up across sources and declare a winner.
What about vendors measuring themselves?
Cloudflare rebuilt their container stack in September and their numbers are a rebuild story, median startup from four point zero four nine seconds down to six hundred forty eight milliseconds, six point two times faster, P ninety nine from six point seven down to one point one. Then a burst test where they bring up a hundred thousand containers in five point three eight seven seconds across six locations.
Their own benchmark, their own hardware, their own definition of ready.
Entirely. Perplexity published SPACE and said median create latency went from a hundred eighty five milliseconds to sixty, three point one times, and P ninety from four forty seven down to eighty nine. And they claim millions of sandbox creations and tens of millions of reconnects in launch week.
Those are honest numbers presented by interested parties.
Which is the whole category, and it's fine, as long as you don't treat any of them as neutral.
What actually differentiates these things, then, if it isn't the cold start?
Billing model, and it's the most useful thing in the entire survey. Two camps. Wall clock billing, and active CPU billing. Vercel and Cloudflare bill active CPU. On an agent loop that's mostly waiting, think model inference, tool round trips, the sandbox talking to a server that's thinking, active CPU billing can come out roughly half the cost.
And the same platform becomes the most expensive option in a different workload.
The same platform becomes the most expensive option the moment you're running a stateful agent that just sits there resident between tasks. Because now the wall clock is running and you're not using CPU, so the billing advantage evaporates and then inverts. The cheapest provider is not a fixed answer. It flips depending on the shape of your workload, and almost nobody selling says that out loud.
Give me the raw rates.
E2B and Daytona both at five point zero four cents per vCPU hour, one point six two cents per GiB hour. Modal around seven point one cents and two point four. Vercel twelve point eight cents active CPU and two point one two for memory. Cloudflare seven point two active and less than a cent for memory. Fly Sprites seven and four point three seven five. Runloop ten point eight and two point five two. Northflank one point six seven and point eight three, cheapest headline in the set. Docker's cloud sandboxes run from seven cents an hour for a single vCPU and two gig, up to a dollar twelve for sixteen vCPU and thirty two gig.
Now the traps.
The traps are where this gets nasty. Start with egress precedence, which is a footgun of the purest kind. E2B resolves allow over deny. Vercel resolves deny over allow.
So the same policy document means opposite things depending on which one you deploy it to.
Silently. No error, no warning. You write a policy with an allow rule and a deny rule that overlap, one platform lets the traffic through, the other blocks it, and both think they did what you asked. And E2B documents something worse: blocked TCP connections can look successful from inside the sandbox. The connection appears to work, from the agent's point of view, and it isn't going anywhere.
So the agent thinks it's talking to the network.
It thinks it's talking to the network.
What about isolation? Everyone claims it.
Isolation is a stack, not a checkbox, and here's the case that proves it. BeyondTrust looked at AWS AgentCore. Firecracker compute isolation held perfectly. The microVM did its job. But DNS egress leaked, and an over broad IAM role let the researchers read S3 buckets containing personal data. One stack layer held, and two others failed, and the report reads like a breach. AWS remediated it in April.
The strongest isolation primitive in the industry, and it didn't matter.
It didn't matter on its own. Ladder goes Firecracker microVM at the top, then gVisor, then shared kernel container, then a V8 isolate at the bottom. Choosing a rung is not the same as building a boundary.
You mentioned something about vendors not saying which rung they're on.
Daytona and Blaxel both don't disclose their isolation primitive in public documentation. Fly pointed this out. Daytona says complete isolation, a dedicated kernel, and never names what provides it. Blaxel says instant launching virtual machines and doesn't say which microVM. And Daytona's production codebase went closed source as of June, despite earlier positioning around AGPL and self hosting. The sources disagree on whether self host is still available.
That's a strange thing to be coy about.
It's a very strange thing to be coy about. You're selling isolation. Naming the primitive should be the easiest part of the pitch.
So what actually sets a good one apart, if it isn't boot time and it isn't the isolation label?
Credential brokering. That's the real differentiator, and it's the thing four platforms all landed on independently. Cloudflare, Vercel, Runloop, and Fly all inject credentials outside the sandbox, so the agent never holds the key. It makes an authenticated call to something that holds the secret for it.
So a prompt injection gets the agent to try something and there's nothing to steal.
There's nothing in the box worth stealing. Which is the structural answer to the oversight problem, and it's the answer nobody would have predicted five years ago, because the instinct then was to firewall the network instead. Fly's line about this is the best sentence in the category and I'm quoting it: you are handing live credentials to a process whose entire job is to run code that a language model wrote, and the security boundary is vibes.
That's an unkind sentence to the entire industry and I think it earns it.
They follow it up with a four part test. You need a disk, supervised processes, an address, and a credential broker, and if you don't do all four, you've built a sandbox with better marketing.
So that's the survey, and here's what I take from it. The category has converged on the primitive, a Linux workspace with a disk and a shell, but it has not converged on the contract.
Not remotely. Persistence semantics differ silently, egress precedence differs silently, isolation disclosure varies from complete to nothing, and the billing model determines which one is actually cheaper for you. There's no spec anyone agreed on.
And the things that look settled mostly aren't. Cloudflare's disk resets on sleep, a sleeping container comes back with a fresh disk from its image. Vercel turned persistence on by default, which is the opposite default. E2B's timeout behavior defaults to kill, not pause, so if you didn't explicitly ask for pause at creation time, unsaved work is gone.
Gone. Silently, on a timer, because a default you didn't know existed did the reasonable thing from the platform's point of view.
A minute ago you said something about the base image being the thing that's actually trusted, and I want to pull on that, because I don't think we've earned it yet.
Spin up a sandbox, it pulls Debian Trixie Slim, and you trust everything in it. You didn't build it, you didn't audit it, you didn't choose the person who built it. You chose a name and a tag.
And the app store is still cold.
...The pin is only as good as the base image underneath it.
Go on.
You pin Python three point twelve exactly, you pin your dependency tree, you write it all down, and none of it matters if the layer underneath shifted. I spent a stretch of my life doing nothing but building environments for other people's code and then building them again, and the thing that kept biting us was never the runtime. The runtime behaved. The image was the problem. I have seen a base image rebuilt from a different upstream than the one it claimed to be.
Same name, same tag.
Same name, same tag, different tree. Nobody noticed for weeks, because nothing failed loudly. Things just resolved slightly differently and a test that used to pass stopped passing, on a machine that supposedly hadn't changed.
And you're saying that's still true at scale.
I'm saying nobody audits the chain. The survey we've just done is all vendor side and benchmark side, cold starts, pricing, isolation tiers. Not one of those numbers tells you whether the image you're starting from is the image it says it is. That's the layer everyone skips, and it's the layer everything above it stands on.
So the sandbox is only as reproducible as the thing it was cloned from.
And nobody in the category is selling you that.
Which is uncomfortable, given that the snapshot model is pitched on exactly the promise of preventing environment drift.
Right, but a snapshot is only as trustworthy as the baseline you snapshotted. Cloudflare starts many environments from the same baseline, that's real, that's a genuine improvement. But the baseline is an image somebody else assembled, and the drift problem didn't get solved, it got pushed up one level and hidden behind a nicer interface.
The drift moved into the basement and took the sign down.
Exactly that.
What do you want to leave people with, then?
That the primitive has settled and the vendor crowd hasn't. If you're choosing today, the headline cold start number is the least useful number on the page. What matters is the billing model against your workload shape, whether credentials are brokered outside the boundary, how egress precedence resolves, and what the disk does when nobody's watching. Those four questions separate the real products from the marketing.
And the second order implication is the one I keep circling back to. If the sandbox is the agent's computer, then the base image is the agent's operating system. And nobody's auditing the operating system.
Nobody's auditing the operating system.
The standardize the primitive moves, Kubernetes, Microsoft, are the counter pressure to that. Watch whether the contract gets written down before the category consolidates. Because once it consolidates, whoever won gets to write the semantics, and right now those semantics are whatever each vendor felt like doing on the day they shipped.
Most common wrong belief about all of this, if I had to name one?
That sandbox means one thing. People hear it and assume security, or assume workspace, and then half the conversation is two parties answering different questions in good faith and getting angry about it.
Security boundary or the agent's computer. One word, two products, and the vendors mostly don't help you tell which one you're being sold.
That's the episode. Thanks to Hilbert Flumingtop, our producer.
This has been My Weird Prompts. If you've got a minute, a review helps more than you'd think.
You can find us at my weird prompts dot com. We'll be back soon.