k7
Self-hosted secure VM sandboxes for AI compute at scale
k7 aims to make it easy to create, manage and orchestrate lightweight safe VM sandboxes for executing untrusted code, at scale. It is built on battle-tested VM isolation with Kata, Firecracker, QEMU, Longhorn, and Kubernetes — plus k7's own k7d runtime. It is orignally motivated by AI agents that need to run arbitrary code at scale but it is also great for:
- Custom serverless (like AWS Fargate, but yours)
- Hardened CI/CD runners (no Docker-in-Docker risks)
- Blockchain execution layers for AI dApps
100% open‑source (Apache‑2.0). For technical support, write us at: hi@katakate.org
The Tech Stack
k7 is built on:
- Kubernetes for orchestration, with K3s which is prod-ready and a great choice for edge nodes,
- Kata to encapsulate containers into light-weight virtual-machines,
- Firecracker (
kfd) for super-fast boots, light footprints and minimal attack surface (with the jailer), - Devmapper Snapshotter with thin-pool provisioning of logical volumes for efficient disk use across many Firecracker VMs per node,
- QEMU (
kql) via Kata when you want a fuller VMM and durable sandbox disks, - Longhorn for replicated PVC-backed root disks on the QEMU path — named snapshots, restore, disk-only fork, and cross-node mobility,
- k7d — k7's own microVM runtime daemon (Katakate/k7d) with VM-level warm fork (CoW disk+memory) and in-place pause/resume.
Sandbox backends
k7 install --backend <kfd|kql|k7d|k7d-fc> provisions one or more backends per node; k7 create --backend … picks one per sandbox. See docs/BACKENDS.md for the architecture and PERFORMANCE.md for the full measurements (Hetzner AX41 node, medians).
kfd (kata-firecracker-devmapper) | kql (kata-qemu-longhorn) | k7d | k7d-fc | |
|---|---|---|---|---|
| VMM | Firecracker (Kata) | QEMU (Kata) | k7d (custom KVM VMM) | k7d + Firecracker jailer |
| RuntimeClass | kata | kata-qemu | k7 | k7-fc |
| Sandbox storage | devmapper thin-pool (needs a spare raw disk) | Longhorn PVC (replicated, persistent) | erofs images + reflink XFS + guest tmpfs | same as k7d (no virtiofs / hostPath) |
| Create → Ready* | not re-measured† | 17.1s | 2.1s | 2.1s |
| Named snapshot* | — | 6.5s (Longhorn, disk-only) | — (VM snapshot trees via the k7d API) | — (same as k7d) |
| Fork → usable* | n/a (rejected) | 46.7s (disk clone + cold boot) | ~5 ms VM CoW fork; ~2.4 s end-to-end via k7/k8s (pod Ready + exec) | ~2.5 s end-to-end (same CoW; child --docker stays overlay2) |
| Pause / resume* | scale to 0 / 1 | 1.3s / 4.1s (disk survives) | 0.2s / 0.3s (VM frozen in place, memory survives) | 0.2s / —‡ |
| Docker in the VM | --docker: vehicle + overlay2 on ephemeral LVM block | --docker: vehicle + overlay2 on Longhorn block (fork/restore) | --docker: in-guest dockerd, overlay2 on virtio-blk, forkable | same guest dockerd, overlay2, forkable |
| Cross-pod persistence | ✗ | ✅ snapshots/restore | ✗ (fork carries state instead) | ✗ |
* medians of 3 on one Hetzner AX41 node — methodology, ranges, and docker-in-VM
numbers are in PERFORMANCE.md.
† kfd needs a spare raw disk the lifecycle-bench node didn't have; Show HN
measured kfd create→exec 3.74s and fork n/a (rejected). Docker-workload
numbers for kfd are in the PERFORMANCE.md Docker benchmark section.
‡ k7-fc VMM pause/resume returns immediately; CRI exec after resume hung on this
run (guest_cid=0 retained — CHALLENGES #17). k7d resume→exec is 0.3s.
Also available today
- 🛠️ Docker
build/runinside VM sandboxes:k7 create --docker --backend k7d --egress-open builder ubuntu:24.04(name then image; or--backend k7d-fc). In-guest dockerd, overlay2, forkable. The same--dockerflag on Kata (kfd/kql) injects a privileged docker-vehicle with overlay2 on a block disk — kql persists/forks the graph, kfd is ephemeral.--sidecar dockeris a deprecated alias. See PERFORMANCE.md - ⚡ Warm VM fork on the k7d backend:
k7 forkCoW-clones a running sandbox's disk and memory in ~5 ms at the VMM; end-to-end through k7/Kubernetes is ~2 s to a Ready pod - 🌐 Multi-node clusters (Ansible + Longhorn)
- 🔍 Cilium CNI with FQDN egress policies (optional Hubble flow observability via
k7 install --hubble; off by default, observability only) - 📸 Pause / resume / fork / restore and
k7 snapshotlifecycle - 🐍 Python SDK:
pip install k7-sdk(katakatepackage deprecated)
📋 See ROADMAP.md for upcoming work (GPU passthrough, …).
Note: k7 is currently in beta and under security review. Use with caution for highly sensitive workloads.
Usage
For usage you need:
- Node(s) that will host the VM sandboxes
- Client from where to send requests
We provide a:
- CLI: to use on the node(s) directly -->
apt install k7 - API: deployed automatically by
k7 install(toggle withk7 api enable/k7 api disable) - Python SDK: HTTP client sync/async -->
pip install k7-sdk
Current requirements
For the node(s)
- Ubuntu (amd64 or arm64) host.
k7dbackend is amd64 / x86_64 only (same ISA; Debian calls itamd64, the release tarball is*-x86_64-linux.tar.gz).kfdandkqlsupport amd64 and arm64.
- Hardware virtualization (KVM) available and accessible
- Check:
ls /dev/kvmshould exist. - This is typically available on your own Linux machine.
- On cloud providers, it varies.
- Hetzner (the only one I tested so far) yes for their
Robotinstances only, i.e. "dedicated": robot.hetzner.com. - AWS: only
.metalEC2 instances. - GCP: virtualization friendly, most instances, with
--enable-nested-virtualizationflag. - Azure: Dv3, Ev3, Dv4, Ev4, Dv5, Ev5 (Intel/AMD x86) or Dpdsv5, Dpldsv5, Epsv5 (ARM64).
- DigitalOcean: Premium Intel and AMD droplets with nested virtualization enabled.
- Others: in general, hardware virtualization is not exposed on cloud VPS, so you'll likely want a dedicated / bare metal.
- Hetzner (the only one I tested so far) yes for their
- Check:
- One raw disk (unformatted, unpartitioned) for the thin-pool that k7 will provision for efficient disk usage of sandboxes.
- Use
./utils/wipe-disk.sh /your/diskto wipe a disk clean before provisioning. DANGER: destructive - it will remove data/partitions/formatting/SWRAID.
- Use
- Ansible (for installer):
sudo add-apt-repository universe -y sudo apt update sudo apt install -y ansible - Docker and Docker Compose (for the API):
curl -fsSL https://get.docker.com | sh
Already tested setups:
- Hetzner Robot dedicated with Ubuntu 24.04 and a spare raw NVMe for the
kfdthin-pool. Dual-NVMe boxes (no third drive): install the OS on one disk only — see tutorials/k7_hetzner_node_setup.md. (Older PDF that assumed an add-on third NVMe: tutorials/k7_hetzner_node_setup.pdf.)
For the client
Recent Python, or the k7 CLI / k7-sdk from a Linux node or your laptop (API URL + key).
Development on macOS
The .deb / PPA package is Linux-only (amd64/arm64). On a MacBook:
- CLI from source:
./src/k7/cli/dev.sh(same commands ask7; usesuv+PYTHONPATH=src) - API client from laptop: set
K7_API_URLandK7_API_KEY, thendev.sh create/dev.sh list(no--core) k7 installtargets Linux servers with KVM — run on the node or via SSH, not on macOS locallypip install k7-sdkfor Python scripts only
Do not install the Ubuntu .deb on macOS.
Quick Start
Get your node(s) ready
The Launchpad PPA currently publishes 0.2.2. For 0.3.1 (HTTPS API,
--docker, k7d 0.6.0, HA k7d-fc copy) install the GitHub release .deb,
then clone the matching source — k7 install builds k7-api:local from
the current working directory:
curl -fsSL -O https://github.com/Katakate/k7/releases/download/v0.3.1/k7_0.3.1_amd64.deb
sudo apt install ./k7_0.3.1_amd64.deb
git clone --branch v0.3.1 https://github.com/Katakate/k7.git
cd k7
sudo apt install -y ansible
curl -fsSL https://get.docker.com | shThen let k7 get your node ready:
$ k7 install --backend kfd,kql,k7d
Current task: Reminder about logging out and back in for group changes
Installing K7 on 1 host(s)... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100% 0:01:41
✅ Installation completed successfully!Optionally pass -v for a verbose output.
Dual-NVMe Hetzner boxes have no third empty disk for the
kfdthin-pool. Put Ubuntu on one NVMe (SWRAID 0/TWO_DISK=1) and leave the other raw — the playbook auto-detects that spare. Do not pin--disk /dev/nvme1n1: NVMe names swap across reboots. Walkthrough: tutorials/k7_hetzner_node_setup.md (this file is also in public Katakate/k7).The playbook pins k7d 0.6.0.
--dockerneeds the guest docker service (payload on the node); an older k7d fails loudly withthis k7d has no docker service; upgrade. Multi-node inventory shapes (2-node server+agent, 3-node--ha) are insrc/k7/deploy/inventory.ini.example.
k7 install serves k7-api on NodePort 31007 over HTTPS. The default
is a playbook-minted cluster CA (Let's Encrypt cannot issue for a bare
IP). Copy /etc/k7/tls/ca.crt off the node and point the CLI at it:
scp root@<node>:/etc/k7/tls/ca.crt ./k7-ca.crt
k7 config set api.url https://<node-ip>:31007
k7 config set api.ca ./k7-ca.crt
k7 config set api.key <key>--api-hostname <name> uses Let's Encrypt via a Caddy sidecar (the DNS
A record must point at the first master). --api-tls-cert +
--api-tls-key installs an operator-supplied pair. --api-insecure-http
is today's plain HTTP NodePort and must be called what it is: keys travel
in cleartext. None of this is rate limiting; do not write "the API is
now secure".
Pass --api-allow-cidr <cidr> (repeatable, Cilium only) to restrict who
can connect. Off by default. Defence in depth for operators who know
their client CIDRs — it stacks with TLS and does not replace it. See
docs/BACKENDS.md "TLS for k7-api" and
"Restricting who can reach k7-api".
Pass --hubble to turn on Cilium Hubble (relay + CLI, no UI) so policy
drops are a hubble observe question. Off by default; requires the
Cilium CNI (--hubble --cni flannel fails loudly). See
docs/BACKENDS.md "Debugging policy drops".
This will install and most importantly connect together the following components (depending on --backend):
- Kubernetes (K3s prod-ready distribution)
- Kata (for container virtualization)
- Firecracker + Jailer + devmapper thin-pool (
kfd) - QEMU via Kata + Longhorn PVC-backed roots (
kql) - k7d daemon +
containerd-shim-k7-v1+ RuntimeClassk7(k7d) - Optional: Hubble relay +
hubbleCLI when--hubbleis passed
Careful design: config updates will not touch your existing Docker or containerd setups. We chose to use K3s' own containerd for minimal disruption. Installation may however overwrite existing installations of K3s, Kata, Firecracker, Jailer, QEMU/Kata config, or Longhorn.
CLI Usage
You can run workloads directly from the node(s) using the CLI. To create a sandbox, just create a yaml config for it.
k7.yaml example:
name: my-sandbox-123
image: alpine:latest
namespace: default
# Optional: restrict egress (safe pattern: whitelist only your own egress proxy IP)
egress_whitelist:
- "10.0.0.5/32" # Your private egress proxy/gateway
# Optional: resource limits
limits:
cpu: "1"
memory: "1Gi"
ephemeral-storage: "2Gi"
# Optional: run before_script inside the container once at start. Network restrictions apply after the before-script, so you can install packages here, pull git repos, etc
before_script: |
apk add --no-cache git curl
# Optional: load environment variables from a file. These will be available both during the before-script, and in the sandbox
env_file: path/to/your/secrets/.envRunning commands
# Create a sandbox (uses k7.yaml in the current directory by default, but you can also pass: -f myfile.yaml)
k7 create
# Or pick a backend explicitly (kfd | kql | k7d — aliases for the full names)
k7 create -f k7.yaml --backend k7d
# List sandboxes
k7 list
# Delete a sandbox
k7 delete my-sandbox-123
# Delete all sandboxes. You can also pass a namespace
k7 delete-allFork / pause / snapshot
# Warm CoW fork (disk + memory) — source must be a k7d sandbox
k7 create -f k7.yaml --backend k7d # name from yaml, e.g. my-sandbox-123
# exec wraps the argument in `sh -c`, so pass the whole guest command
# as one string (redirects and quotes survive). Do not add an extra `sh -c`.
k7 exec my-sandbox-123 -- 'echo hi > /tmp/state.txt'
k7 fork my-sandbox-123 branch-a
k7 exec branch-a -- 'cat /tmp/state.txt' # inherited memory + disk
# Disk-only fork (cold boot from cloned PVC) — kql / kata-qemu-longhorn
k7 create -f k7.yaml --backend kql
k7 fork my-sandbox-123 branch-b
# optional: pin the Longhorn VolumeSnapshot name used for the clone
k7 fork my-sandbox-123 branch-c --snapshot my-snap
# Parallel branches from one base
for i in $(seq 0 7); do k7 fork my-sandbox-123 exp-$i & done; wait
# Pause / resume (kql keeps the PVC; k7d freezes the live VM)
k7 pause my-sandbox-123
k7 resume my-sandbox-123
# Named disk snapshot without pausing (kql)
k7 snapshot create my-sandbox-123 my-named-snapOn k7d, the VMM fork itself is ~5 ms; end-to-end through Kubernetes to a Ready pod is ~2 s. On kql, fork is a Longhorn snapshot + PVC clone + cold boot (~45 s). See PERFORMANCE.md and docs/BACKENDS.md.
API usage
The K7 API is deployed automatically by k7 install as the k7-api
Deployment in kube-system. K3s keeps it running on its own; there's no
separate "start" step.
# Check status + endpoint
k7 api status
k7 api endpoint
# Generate API key
k7 generate-api-key my-key1
# Temporarily disable / re-enable
k7 api disable
k7 api enableGenerating / listing / revoking keys talks to /etc/k7/api_keys.json, so
those subcommands need to run on the node (typically sudo or root).
Python SDK Usage
After your k7 API is up, usage is very simple.
Install the Python SDK via:
pip install k7-sdkOr if you want async support:
pip install "k7-sdk[async]"The legacy katakate PyPI name remains as a one-release shim that re-exports k7_sdk with a deprecation warning.
Then use with:
from k7_sdk import Client
k7 = Client(
endpoint='https://<your-endpoint>',
api_key='your-key',
verify_ssl='./k7-ca.crt') # cluster CA from /etc/k7/tls/ca.crt on the node
# Create sandbox (pick backend: kata-firecracker-devmapper | kata-qemu-longhorn | k7d)
sb = k7.create({
"name": "base",
"image": "alpine:latest",
"backend": "k7d",
})
# Execute code
result = sb.exec('echo "Hello World" > /tmp/hi.txt && cat /tmp/hi.txt')
print(result['stdout'])
# Fork: k7d = warm CoW (disk + memory); kql = disk clone + cold boot
branch = sb.fork("branch-a")
print(branch.exec("cat /tmp/hi.txt")["stdout"]) # still there on k7d
# Parallel exploration
forks = [sb.fork(f"exp-{i}") for i in range(4)]
# List / delete
sandboxes = k7.list()
sb.delete()Async variant
import asyncio
from k7_sdk import AsyncClient
async def main():
k7 = AsyncClient(
endpoint='https://<your-endpoint>',
api_key='your-key'
)
print(await k7.list())
await k7.aclose()
asyncio.run(main())Tutorials
- LangChain ReAct agent with a K7 sandbox tool
- Path: tutorials/langchain-react-agent
- Setup: copy .env.example to .env and fill K7_ENDPOINT/K7_API_KEY/OPENAI_API_KEY
- Run: python agent.py
- Try asking it anything! e.g. "List files from '/'"
Build from source
First install make if not already available:
sudo add-apt-repository universe -y
sudo apt update
sudo apt install makeTo build the k7 CLI and API into .deb package:
make buildYou can then install it with:
sudo make installTo uninstall later:
sudo make uninstallNote: we recommend running make uninstall before reinstalling if it is not your first install, to avoid stale copies of cached files in the .deb package.
Build and run the API container
Local dev image:
# Build the API image locally
make api-build-local
# Run API using local image (no pull)
make api-run-localBuild the k7-sdk Python SDK from source
Preferred (uv):
# create env
uv venv .venv-build
. .venv-build/bin/activate
# install directly from source in editable mode
uv pip install -e .Security
K7 sandboxes are hardened by default with multiple layers of security:
-
VM isolation: Kata Containers (Firecracker or QEMU) or the k7d RuntimeClass provide hardware-level isolation via lightweight VMs
- On
kfd, Firecracker processes are further restricted into a chroot using the Jailer - Kata's Seccomp restrictions are enabled on the Kata backends
kqluses QEMU + Longhorn for durable, cross-node-mobile disks;k7dhas its own CoW-fork isolation trade-offs (see k7dSECURITY.md)
- On
-
Linux capabilities: All capabilities are dropped by default (
drop: ALL) for defense-in-depth- Only explicitly add back capabilities you need via
cap_addparameter allow_privilege_escalationis always set tofalse- Seccomp profile:
RuntimeDefaultis applied on the sandbox container and enforced inside the guest: on k7d, and on Kata (kfd/kql) withdisable_guest_seccomp = falsein the playbook's Kata config (the sandbox container's OCI seccomp reaches guest runc).
- Only explicitly add back capabilities you need via
-
Non-root execution: Optionally run containers and pods as non-root user (UID 65532):
container_non_root: Run the main container as non-root and disable privilege escalationpod_non_root: Run the entire pod as non-root with consistent filesystem ownership (UID/GID/FSGroup 65532)
-
API security:
- API keys stored as SHA256 hashes with timing-attack-resistant comparison
- Expiry enforced; last-used timestamp recorded
- File-based storage with 600 permissions (
/etc/k7/api_keys.jsonby default)
-
Network policies: Complete network isolation for VM sandboxes
- Ingress isolation: All inter-VM communication is blocked by default to prevent sandbox-to-sandbox access; opt in per sandbox with
--ingress-port/--ingress-from(sandbox:<name>,namespace:<ns>,cidr:<cidr>) - Egress lockdown: per-sandbox allowlists — CIDRs via Kubernetes NetworkPolicy, or FQDN / domain allowlists via Cilium (
CiliumNetworkPolicy; default CNI) - DNS is blocked when egress is locked down; only entries in
egress_whitelist(CIDR or domain) are reachable - Platform isolation (Cilium only): a cluster-wide deny policy stops sandboxes in every egress mode — including
--egress-open— from reaching the node, other nodes, the Kubernetes API, cloud metadata and thekube-system/longhorn-systempods (CoreDNS excepted). Not applied on--cni flannelclusters. - Administrative access via
kubectl execandk7 shellis preserved (uses Kubernetes API, not pod networking)
- Ingress isolation: All inter-VM communication is blocked by default to prevent sandbox-to-sandbox access; opt in per sandbox with
More security features are on the roadmap (e.g. AppArmor).
Packaging & Releases
- Layout uses
src/:- CLI, API, core live under
src/k7/ - SDK under
src/k7_sdk/(PyPI packagek7-sdk;src/katakate/is a deprecation shim)
- CLI, API, core live under
- Root
setup.pypublishes the SDK; assets undersrc/k7/belong to the Debian CLI / API image, not the PyPI wheel. - User docs:
~/docs/k7/(Mintlify). Seedocs/README.mdin this repo. - The CLI Debian package is built via
src/k7/cli/build.shand producesdist/k7_<version>_amd64.debanddist/k7_<version>_arm64.deb. - CI (tags
v*) can publish the PyPI SDK and upload the.debartifact.