Run untrusted AI-agent code in Firecracker microVMs.
Minimal, self-hostable, on-prem first — cloud-ready by construction.
Airlock gives an AI agent a place to run generated code that assumes the code is hostile: each session is a Firecracker microVM with its own kernel, so a compromised guest cannot reach the host. Sessions are durable — pause to a snapshot, resume in milliseconds, scale to zero when idle — and it is driven from a CLI. It runs on one KVM box today, and on cloud instances tomorrow without changing a line above the substrate.
A container shares the host kernel. A kernel exploit from inside one is a host compromise. Untrusted code generated by a model is exactly the workload where that matters, so Airlock's boundary is a microVM with its own kernel — a fork bomb, a memory bomb, a disk-fill, or a kernel exploit inside a guest can at worst take down that guest, which we destroy after the session anyway.
airlock run python:3.11-slim -- python3 -c "print('hi!')"
airlock session create python-ds # durable, pausable session
airlock exec <id> -- pytest -q
airlock pause <id> # snapshot, free the VM
airlock resume <id> # restore in ~ms
airlock fork <id> # clone from a snapshot
airlock ls ; airlock ps # inventory ; live VMMs
airlock metrics --watch # live guest CPU/mem
airlock logs -f <id> # output, even when stopped
The CLI is the first-class interface. SDKs for js, python, go, and rust wrap the same gRPC API.
A read-only web dashboard watches a fleet from another machine. It is a client of the same gRPC API the CLI uses, authenticates with its own certificate, and stores nothing — the only history is a four-minute ring in memory that dies with the process.
The workload Airlock exists for is code an agent just wrote, so an agent has to be able to drive it without a human in the loop. Two pieces ship for that.
An MCP server — airlock mcp — serving fourteen tools over
stdio: one-shot runs, durable sessions, pause, resume, fork, plus inspection, logs and
live metrics. It reuses the exact helpers the CLI commands call, so there is no second
implementation to drift. Point a client at it and the sandbox becomes something the
model can reach for on its own.
{
"mcpServers": {
"airlock": { "command": "airlock", "args": ["mcp"] }
}
}
One detail matters more than it looks. MCP speaks over stdin and stdout, and the guest is untrusted — so guest output never touches that channel. It is captured into size-capped buffers, returned inside the tool result with an explicit truncation flag, and progress goes to stderr. Untrusted output cannot interleave with the protocol carrying it.
An agent skill, shipped alongside it, teaching an agent when to reach for a sandbox rather than only how: generated programs, third-party scripts, adversarial input, anything that wants hard limits and no network. Knowing the tools exist is not the same as knowing when the host is worth protecting from.
What runs in production is our own static Go and Rust binaries, Firecracker, and a rootfs — the control plane adds an embedded SQLite. No Postgres, no Redis, no Kubernetes, no bundled observability stack. Everything shipped is built from source or vendored, so builds are reproducible and offline.
Everything environment-specific sits behind three Go interfaces — host provider, blob store, network provider. Today they have local implementations; a cloud deployment is new structs behind the same interfaces, and nothing above them changes.
An idle session is snapshotted and its VM freed; when a host has no VMs it is released. Zero VMs and zero hosts when there is no work, and a request rehydrates from the snapshot on demand.
The host emits OpenTelemetry — metrics, traces, and structured audit records — to an OTLP endpoint you supply. We ship the instrumentation, never the backend.
A guest gets no egress by default. Reaching the network at all means an operator allowlisted the destination, which is the opposite of the usual arrangement where code is loose until somebody notices. Alongside it are hard ceilings — CPU, memory, process count, disk, and wall clock — because a workload that will not stop is as much of a problem as one that escapes, and the VM is destroyed when the session ends rather than reused.
Exactly one component reads untrusted input: the guest daemon that receives code over
vsock and runs it. That one is Rust, deliberately tiny, with a minimal
dependency list — the rest of the system is Go, and the language boundary is drawn
where the hostile input is. The host never shells out to nft,
ip, or mke2fs either; it talks to the kernel directly, so
there is no command line for guest-influenced data to be interpolated into.
SOC 2 is a design input, not a retrofit. Every RPC resolves a principal from peer credentials, a client certificate, or a forwarded OIDC token, and security-relevant actions leave an audit record naming who did them. Storage is scoped by tenant, credentials come from the environment and are never committed, and the control plane refuses to start in an unauthenticated network configuration rather than warning about it.
What is not yet true is written down too: tenant isolation today is control-plane scoping, and the remaining shared surfaces are tracked in the open rather than implied away. Ask us for the current state — we would rather tell you than have you find out.
Airlock is early, and the beta is intentionally small — we would rather find the sharp edges with a handful of people than with everyone at once. Tell us what you would run in it.
Request an invite → Or mail getairlock@mist-os.com.