Christian Lehnert — Linux, Hacking & Faith

Where Should Agent Code Run

Christian Lehnert • 2026-09-29 • ~8 min read

Where Should Agent Code Run

An AI agent writes code, and now you have a problem you did not have
before: where does that code run. If it runs in your application
process, a single hallucinated import os; os.remove('/'), or a single
prompt injection through a tool result, reaches your connection strings,
your tokens, and every system your app can talk to. The question of
where untrusted, machine-generated code executes is not a detail to
solve after the demo. It is the whole security design.

There are three broad answers I want to compare, because they sit at
three different points on the one axis that actually matters:
isolation strength. Plain Docker, my own tank-os-with-llms container,
and Azure Container Apps Sandboxes. They are not really competitors so
much as three answers to "how strong a wall, at what cost, run by whom."

The Axis That Matters: What Is the Wall Made Of

Before comparing, the one thing to understand is that not all container
isolation is the same wall, and the difference decides everything else.

A standard container is isolated by the Linux kernel: namespaces and
cgroups. Your code and the host share one kernel, and the container is a
set of restrictions the kernel enforces on a normal process. This is
fast and light, and it has a hard ceiling: a kernel vulnerability, and
there is a steady supply of them, is a container escape. One compromised
container executing a kernel exploit is on the host, and on a shared
node that means every other workload on it too. For your own trusted
code this is fine. For untrusted or agent-generated code, a shared
kernel is a thin wall.

The stronger answer is a microVM: a real, lightweight virtual machine
with its own kernel, isolated from the host by the hardware
virtualization boundary, the same boundary that separates two tenants
on a cloud. To escape it, an attacker has to break out of hardware
virtualization, which is a vastly higher bar than a kernel bug. This is
what Hyper-V isolation, Firecracker, and Kata Containers provide, and it
is the boundary the serious "run untrusted code" platforms are built on.

The middle option, gVisor, intercepts syscalls in userspace to shrink
the kernel attack surface without a full VM; Google's agent sandboxes
use it. It sits between the two on both strength and overhead.

Hold that axis in mind: namespace isolation (thin, fast) to gVisor
(medium) to microVM (thick, hardware-backed). Everything below is a
position on it plus a decision about who operates it.

Plain Docker: The Thin Wall You Already Have

Running an agent in a plain docker run container is where most people
start, and it is better than running it on the host directly, but it is
the weakest of the three walls and it is important to be honest about
why.

A default Docker container shares the host kernel. It is a namespace
boundary, not a security boundary against a determined escape. For
untrusted code specifically, this is the case the isolation experts
keep flagging: standard containers do not provide sufficient isolation
for agents executing code you do not trust, because a kernel exploit
walks straight through. Docker also gives you nothing about network
egress, filesystem exposure, or identity by default; a bare container
can reach your whole network and inherits whatever credentials you
hand it.

Plain Docker is the right tool for isolating your own code from your
host's mess, for reproducibility, for shipping. It is the wrong tool,
on its own, for running code an adversary might have influenced,
because the wall is one kernel bug thick and everything else you would
want, egress control, an unprivileged identity, a bounded blast radius,
you have to build yourself.

tank-os: Hardening the Thin Wall Yourself

Which is exactly what my tank-os-with-llms setup does. It does not
change what kind of wall a container is, it is still a shared-kernel
container, but it wraps that container in every layer Docker leaves
off, and it does so on hardware I own.

The setup runs Claude Code and Gemini CLI inside a hardened image, as an
unprivileged user, behind a default-deny iptables egress firewall that
drops the RFC 1918 ranges, the cloud metadata endpoint, and the host's
loopback, with only a bind-mounted workspace exposed. The base image is
read-only, setuid is stripped, and the corporate network is simply
unreachable from inside. The point is not that it makes the kernel wall
thicker, it does not, it is that it bounds what a compromise inside that
wall can reach: not the network, not the host files, not the metadata
endpoint, not extra privileges. If the agent gets hijacked through a
poisoned tool output, the iptables perimeter and the unprivileged user
are what stand between that and anything valuable.

The trade is explicit. tank-os is self-hosted: I run it, I patch it, I
own the box, and I own the isolation strength, which is namespace-level,
not microVM. For my threat model, an individual running coding agents on
my own hardware where I want no cloud dependency and full control, that
is the right trade. It is sovereignty over convenience, and the wall is
as good as my hardening plus the Linux kernel. What it is not is
hardware-isolated multi-tenancy. I would not run other people's
untrusted code next to mine in it and call that safe.

ACA Sandboxes: The Thick Wall, Someone Else's Machine

Azure Container Apps Sandboxes, which went generally available on 23
September 2026, sit at the other end of the axis. Each piece of
untrusted code, or each agent, gets its own Hyper-V-isolated microVM,
allocated from a warm pool in well under a second, with policy-controlled
network egress, and destroyed or reset when done. The related dynamic
sessions flavour is the pure ephemeral code-interpreter version; the new
Sandboxes resource adds a lifecycle you control and state that survives a
suspend-and-resume.

This is a genuinely stronger wall than anything a shared-kernel
container gives you. The isolation is hardware-backed, so a kernel
exploit inside a session does not reach the host or the other tenants;
each session is walled off at the virtualization boundary. Microsoft
runs over a million of these a day behind GitHub Copilot and its own
agent services, which is the scale argument: they operate the microVM
fleet, the warm pool, the patching, the escape-hardening, so you do not.
For running genuinely untrusted, per-user, multi-tenant code, an LLM
snippet from a random web user, this is the model that fits, because the
isolation is strong and you are not the one maintaining it.

The costs are the mirror image of tank-os. It is Azure, so it is a cloud
dependency, a bill, and your untrusted code running in Microsoft's fabric
under Microsoft's operational control. Much of the actual security still
depends on how you configure it, the network policy, what data you mount,
which identity a session maps to; the microVM wall is strong but a
sandbox with your production secrets mounted into it is still exposed to
whatever runs there. And it is a managed primitive: you get the isolation
model Azure offers, not one you shape end to end.

Putting Them Side by Side

The three line up cleanly once you see the axis.

Plain Docker: shared-kernel namespace isolation, self-run, no egress or
identity controls by default. Fine for your own trusted code. Not enough,
alone, for untrusted or agent code.

tank-os-with-llms: shared-kernel container, but hardened, unprivileged,
default-deny egress, self-hosted on hardware you own. The right answer
when you are running your own agents, you want full control and no cloud
dependency, and your threat model is "bound what my own hijacked agent
can reach," not "safely host strangers' code."

ACA Sandboxes: hardware-backed microVM isolation, cloud-managed, warm
pool, per-session walls. The right answer when the code is genuinely
untrusted and multi-tenant, when you need the strongest boundary, and
when you would rather Microsoft operate the isolation than build it. The
trade is the cloud dependency and trusting their fabric.

The pattern to notice: isolation strength and operational ownership pull
against each other. The microVM wall is stronger, but you rent it and run
your untrusted code on someone else's machine. The self-hosted container
is a thinner wall, but it is your wall, on your metal, with no third
party in the trust path. There is no universally correct pick, only the
one that matches your threat model and how much you value sovereignty
against how strong the boundary needs to be.

The Point

Where agent code runs is a security decision, and the axis is how strong
the wall is: a shared kernel that a bug walks through, or a hardware
virtualization boundary that it does not. Plain Docker gives you the thin
wall and nothing else, and it is not enough on its own for code you do
not trust. tank-os takes that thin wall and hardens everything around it,
egress, identity, blast radius, on hardware you own, which is the right
call for running your own agents with full control and no cloud in the
loop. ACA Sandboxes give you the thick microVM wall, managed at a scale
you could not match, which is the right call for genuinely untrusted,
multi-tenant code when you accept the cloud dependency.

Decide where the code runs before you decide which model writes it.
Match the wall to the threat: your own agents on your own metal, harden a
container and own it; strangers' code at scale, rent the microVMs. And in
every case, remember the wall only bounds what it surrounds, so do not
mount the secrets you cannot afford to lose into the box where the
untrusted code runs.

Tagged:
#security #azure #containers #docker
← Back to posts