#Sandbox API v1
The one API every mode serves — local (macOS), self-hosted (Linux) and cloud —
and every client speaks: the CLI, the agent wrappers, Studio, the SDKs. A
behaviour is "the same in all three modes" exactly where the conformance suite
(internal/api/conformance) says so.
Status: M2. The core below is implemented by sandboxd against the fake
backend. Later additions are listed at the end; they are additions, not changes.
#Transport
- HTTP/1.1, JSON bodies (
Content-Type: application/json), UTF-8. - File contents and process stdin are raw bodies
(
Content-Type: application/octet-stream). - Byte fields inside JSON (
stdout,stderr,data) are base64, as Go'sencoding/jsonwrites[]byte. - Streams are newline-delimited JSON (
application/x-ndjson), one event per line, flushed as produced. - All paths are under
/v1.
Reaching it:
| Mode | Listen | Auth |
|---|---|---|
| Local | unix socket, file mode 0600 | the socket's permissions; a token is optional |
| Self-hosted | TCP behind TLS | Authorization: Bearer <token>, required |
| Cloud | the control plane's endpoint | API key as a bearer token |
A server listening on TCP refuses to start without a token, loopback included: a loopback port is reachable by every user on the machine. One reachable from other machines needs TLS as well.
Request guard, applied before any handler:
Hostmust be a loopback name or one the server was configured to answer to (DNS-rebinding defence).- A request carrying an
Originthat is not configured is refused, so a web page cannot drive the API. - Bodies are capped at 1 MiB for JSON and 64 MiB for file writes.
#Errors
Every non-2xx response has this body:
{"error": {"code": "refused", "message": "network mode \"open\" is above this server's ceiling (allowlist)"}}
| HTTP | code |
Meaning |
|---|---|---|
| 400 | invalid_request |
malformed, or a field out of range |
| 401 | unauthorized |
missing or wrong token |
| 403 | refused |
well-formed, and refused by policy. Retrying will not help. |
| 403 | forbidden_origin |
the request guard refused the Host or Origin |
| 404 | not_found |
no such sandbox, process or file — among this server's own |
| 409 | conflict |
a name in use, or a sandbox in the wrong state |
| 501 | unsupported |
this endpoint lacks the capability the request needs |
| 503 | unavailable |
this endpoint takes no new sandboxes for now (a cordoned node), or, through a gateway, no node is answering or the one holding the sandbox is down; try again, or elsewhere |
| 500 | internal |
a bug; the message is safe to show |
refused and unsupported are different on purpose. The first is a decision:
this server will not loosen its policy for anyone. The second is a limitation:
another endpoint might be able to do it.
#Capabilities
GET /v1/capabilities
{
"api_version": "v1",
"backend": "firecracker",
"capabilities": {"network_policy_update": true, "suspend": false, "memory_snapshot": false},
"limits": {"max_cpus": 8, "max_memory_mb": 16384, "max_disk_mb": 102400},
"network": {"default": {"mode": "allowlist", "allow": ["github.com"]}, "ceiling": "allowlist", "may_allow": ["*"]}
}
A client asks rather than assumes. A request needing a capability the endpoint
lacks gets 501 unsupported, never a quietly weaker result.
audit is a property of the server rather than the backend: it is true when
the server keeps an event log (below).
#Sandboxes
States: pending → running → terminated. suspended is added with
the suspend capability; stopped with persistent sandboxes.
POST /v1/sandboxes
{
"name": "fix-login",
"image": "sandbox-base",
"cpus": 2,
"memory_mb": 2048,
"disk_mb": 10240,
"env": {"GOFLAGS": "-mod=mod"},
"network": {"mode": "allowlist", "allow": ["registry.npmjs.org"], "deny": ["gist.github.com"]}
}
Every field is optional. The responses:
201with the sandbox.409 conflictwhen the name belongs to a sandbox that is not terminated.400when a resource is above the server's limits.400for a field the server does not know. That includesbind, which mounted a host directory before the repository model was removed: a client that still sends it is told so, rather than given a sandbox without the mount it asked for.403 refusedfor any of:- a reserved environment variable name (the ones that control the sandbox's own startup, the loader and the shell);
- a network policy looser than the server permits (below).
The sandbox object:
{
"id": "sbx_7f3a9c2e1b4d",
"name": "fix-login",
"state": "running",
"image": "sandbox-base",
"cpus": 2, "memory_mb": 2048, "disk_mb": 10240,
"env_names": ["GOFLAGS"],
"network": {"mode": "allowlist", "allow": ["registry.npmjs.org"], "deny": ["gist.github.com"]},
"created_at": "2026-10-02T09:14:03Z"
}
Environment values are never returned: env_names only. The API is
read by dashboards and logs, and a value echoed back is a secret on a screen.
GET /v1/sandboxeslists them:{"sandboxes": [...]}, newest first, terminated ones included until the server forgets them.?label=key=value, repeatable, keeps only sandboxes carrying every label named.GET /v1/sandboxes/{ref}returns one.refis an id or a name, matched against this server's own sandboxes only and never passed to the backend to resolve.PATCH /v1/sandboxes/{ref}changes a live sandbox; each field left out is left as it is. A request with one bad field changes nothing.network: the egress policy of a running sandbox, under the same rule as create. Needs thenetwork_policy_updatecapability.name: renames it;""removes the name. Unique among live sandboxes, as at create (409); through a gateway, among the caller's sandboxes on every node.labels: replaces them whole;{}removes them all. The same rules as at create. A gateway keeps its owngateway.*labels: one sent back unchanged is accepted, any other is refused (400).idle_timeout_secs: between 1 andlimits.max_idle_timeout_secs;0(never) only where that limit is 0. Unlike at create, 0 is not "the server's default". Counted from the sandbox's last activity.
name,labelsandidle_timeout_secsare this server's records of the sandbox and change on a running or suspended one with nothing asked of its VM. A terminated sandbox is409. A VM's vCPUs and memory are fixed while it runs: a new size is a new sandbox, from a disk snapshot where the backend takes them (Studio's Resize does that).DELETE /v1/sandboxes/{ref}terminates the sandbox, stopping its processes and discarding its disk. It is idempotent: deleting a terminated sandbox is204.
#Network policy
{"mode": "none" | "allowlist" | "open", "allow": ["host", "*.example.com"], "deny": ["host"]}
none: no egress at all.allowlist: only names inalloware reachable, and a name indenyis refused even ifallowmatches it. Names are matched by name, not by resolved address, and case-insensitively.*.example.commatches subdomains but notexample.comitself; nothing else is a pattern.open: everything exceptdeny.
Tighten, never loosen. The server holds a default policy (applied when a
request omits network) and a ceiling. A request:
- may pick any mode at or below the ceiling (
none<allowlist<open); - may add
denyentries freely; - may send its own
allowlist, which replaces the default one: narrowing is sending fewer names. Every name must be either in the server's default list or within itsmay_allowpatterns.["*"]permits any name; an emptymay_allowpermits only the default's names. A wildcard is permitted only by*or by a wildcard at least as broad, so*.a.example.comis allowed under*.example.comand*.example.comis not allowed undera.example.com; - may not send an
allowlistthat resolves to nothing.noneis how you ask to reach nothing; an empty allowlist would otherwise be the strictest request producing the loosest result.
On the wire, "allow": null (or no allow key) means "the default list", and
"allow": [] means "nothing", which is refused. Clients must keep the two
distinct. allow makes no sense with none or open, and sending it there is
400. deny is dropped under none, since nothing is reachable.
Names are names: lowercase after normalising, no scheme, port, path or address,
and a trailing dot is ignored. Anything else is 400.
A request that breaks any of these is 403 refused, naming the rule. It is
never clamped silently.
#Processes
POST /v1/sandboxes/{ref}/run runs a command to completion.
{"argv": ["sh", "-c", "go test ./..."], "env": {"CI": "1"}, "cwd": "/sandbox/home/app", "stdin": "<base64>", "timeout_secs": 600}
→ 200 {"exit_code": 0, "stdout": "<base64>", "stderr": "<base64>", "truncated": false, "timed_out": false}
- Each output stream is kept up to 8 MiB. Past that it is cut, and
truncatedsays so. - When the timeout passes, the process is killed and the response comes back with
timed_out: trueandexit_code: -1. argvis executed directly, never through a shell, unlessargv[0]is a shell.- A command that does not exist is
400 invalid_request, not an exit code: it failed to start, so there is no process to report on. cwd, if given, is an absolute guest path, checked like a file path. Without it the process starts in the sandbox user's home,/sandbox/home. There is no/workspace: a sandbox has no repository, and code gets in bygit cloneinside it or through the files API.envhere is merged over the sandbox's own, under the same reserved-name rule.
POST /v1/sandboxes/{ref}/processes takes the same body without timeout_secs.
It starts the process in the background and returns 201 with the process
object:
{"pid": 4, "argv": ["sleep", "30"], "state": "running", "exit_code": null, "started_at": "…"}
pid is the server's handle for the process, not the guest kernel's.
| Request | Effect |
|---|---|
GET /v1/sandboxes/{ref}/processes |
lists processes, running and exited |
GET …/processes/{pid} |
returns one process |
GET …/processes/{pid}/output |
streams output from the beginning, then follows until exit (format below) |
POST …/processes/{pid}/stdin |
writes the raw body to stdin; ?close=1 then closes it |
POST …/processes/{pid}/signal |
sends {"signal": "TERM"}; one of INT, TERM, KILL, HUP |
The output stream is one JSON object per line:
{"stream":"stdout","data":"aGkK"}
{"stream":"stderr","data":"…"}
{"exit_code":0}
#Files
Paths are absolute guest paths. They are never host paths, and the server never
resolves them on the host. A relative path, or one containing .. or a NUL, is
400.
| Request | Effect |
|---|---|
GET /v1/sandboxes/{ref}/files?path=/sandbox/home/app/go.mod |
returns the raw bytes; 404 if absent |
PUT /v1/sandboxes/{ref}/files?path=… |
writes the raw body, creating parent directories; 204 |
DELETE /v1/sandboxes/{ref}/files?path=… |
removes the file or empty directory; 204 |
GET /v1/sandboxes/{ref}/dirs?path=/sandbox/home |
lists the directory (format below) |
A directory listing:
{"entries": [{"name": "go.mod", "type": "file", "size": 412}, {"name": "cmd", "type": "dir", "size": 0}]}
#Attach
GET /v1/sandboxes/{ref}/processes/{pid}/attach with Connection: Upgrade and
Upgrade: sbx-stream/1 switches the connection (101) to a two-way stream of
frames: a type byte, a big-endian uint32 length, then the payload.
| Direction | Frame | Payload |
|---|---|---|
| client → server | i stdin |
bytes |
| client → server | e end of input |
none (a terminal gets ^D) |
| client → server | s signal |
INT, TERM, KILL or HUP |
| client → server | r resize |
rows, cols (uint16 each) |
| server → client | o stdout |
bytes |
| server → client | E stderr |
bytes |
| server → client | x exit, last |
int32 exit code |
- Output is replayed from the start, so a late or second attach sees everything.
- Disconnecting detaches; it does not stop the process.
- A process started with
"tty": true(background only) runs on a terminal ofrows×cols, and its output is one stream.
#Tunnels (capability tunnel)
GET /v1/sandboxes/{ref}/tunnel?port=N with Upgrade: sbx-tunnel/1 switches
the connection to the raw bytes of a TCP connection to 127.0.0.1:N inside the
guest. Only the guest's loopback is reachable this way, so a tunnel reaches a
server in the sandbox and is never a way around its egress policy. Half-closes
are passed through.
#Suspend, resume and snapshots
| Request | Capability | Effect |
|---|---|---|
POST …/suspend |
suspend |
stops the sandbox, keeping memory, processes and disk. Refused (409) while a process is running, because its stream would be cut. |
POST …/resume |
suspend |
brings it back as it was |
POST …/snapshots |
memory_snapshot or disk_snapshot |
captures the sandbox without stopping it: memory, processes and disk with memory_snapshot, its files only with disk_snapshot. Returns {id, sandbox, image, kind, bytes, created_at}; kind is memory or disk. |
GET /v1/snapshots, DELETE /v1/snapshots/{id} |
either | list and delete |
Creating with "snapshot_id" starts a fork, which diverges from there:
- from a
memorysnapshot it begins where the snapshot was, processes running; the snapshot fixes image and resources, so a request asking for different ones is400; - from a
disksnapshot it boots afresh with the snapshot's files: nothing that was running comes back. Only the image is fixed; resources are the request's, and default to the original's; - a backend may refuse a fork with a network (
501), since a memory snapshot carries the guest's network identity.
A disk snapshot is a copy of a live filesystem: what the sandbox writes while it is taken may or may not be in it.
#A snapshot being taken
While a snapshot is taken, by request or by schedule, the sandbox carries its
progress, and GET …/{ref} and the list show it:
"snapshotting": {"started_at": "2026-10-06T23:26:10Z", "phase": "capture",
"bytes": 958398464, "estimated_bytes": 2260828160}
phaseiscapture(reading the sandbox), thenstore(keeping it where a fork starts from).bytesis what has been read;estimated_bytes, where the backend has one, is about how much there is. It is an estimate — on macOS the guest's owndf— for drawing a bar, andbytesmay pass it."scheduled": truemarks one the schedule started.- Not every step reports as it goes. On macOS the runtime reads nothing for
the first half minute or so (
bytesstays 0), andstore, about half the time, says nothing until it is done; Firecracker reports neither. - The field is gone once the snapshot is made or has failed.
- One at a time: a snapshot asked for while one is being taken is
409. - Deleting the sandbox does not wait for its snapshot, nor stop it. The
sandbox is
terminatedat once and keepssnapshottinguntil the snapshot ends; the snapshot is then listed like any other, or, if it could not finish, is not there at all, with nothing partial left on the host. On macOS the runtime's export, once begun, reads the whole disk, so a snapshot asked for before the delete comes out complete. - A snapshot asked for is not cancelled when its request is: a client that closes the connection part way leaves it to finish, and to be listed.
#Scheduled snapshots
"snapshot_every_secs": 1800, "snapshot_keep": 3 on create, or
PUT …/snapshot-schedule with {"every_secs": 1800, "keep": 3} on a running
sandbox, has the server snapshot it every so often while it runs and keep the
newest few (keep defaults to 1; every_secs 0 stops it). The schedule comes
back on the sandbox as snapshot_every_secs and snapshot_keep.
- The server takes them, so they happen whether or not a client is watching. The interval counts from the end of one snapshot to the start of the next: a disk snapshot that takes a minute never overlaps the following one.
- They are listed with the rest, marked
"scheduled": true. Retention removes only those: a snapshot taken by request is never counted or removed. Each removal is an audit event,snapshot.deletedwith reasonretention. - The server's
limits.min_snapshot_every_secsandlimits.max_snapshot_keepbound it, and a schedule outside them is400, not narrowed to one inside. - A schedule where the endpoint takes no snapshots is
501, and on a sandbox with volumes409, before anything is made. - Behind a gateway, a scheduled snapshot belongs to whoever owns its sandbox. On macOS one takes about a minute and the size of the sandbox's files on the host's disk.
#Also on create
"idle_timeout_secs"terminates the sandbox after that long with no request naming it and no process running. Without it the server's default applies, bounded bylimits.max_idle_timeout_secs."snapshot_every_secs"and"snapshot_keep"give it a snapshot schedule (above)."labels": {"agent": "claude", "team": "infra"}is the client's own metadata. Labels decide nothing about the sandbox: they come back with it, filter the listing and are recorded in its audit events. That is how a client says why a sandbox exists without the server knowing what the reason means. Keys are lowercase letters, digits and. _ / -, at most 63 characters. Values are at most 256 bytes of printable UTF-8. There are at most 32 labels. A violation is400.
#Metrics (capability metrics)
GET …/metrics is the last hour of what a sandbox used, oldest first, sampled
by the server every interval_secs (5) while it runs:
{"interval_secs": 5, "samples": [{"time": "2026-10-06T09:00:10Z", "cpu_percent": 49.8,
"memory_bytes": 441450496, "memory_limit_bytes": 2147483648,
"net_rx_bytes": 1200, "net_tx_bytes": 800, "disk_read_bytes": 0, "disk_write_bytes": 4096,
"processes": 3}]}
- It is measured on the host — the runtime's counters on macOS, the VMM's process and tap device on Linux — never asked of the guest, which could say anything.
- The hour is kept in sandboxd's memory, with the sandbox, and nowhere else:
not in the state directory and not in the records
--keep-sandboxeswrites. A sample is 88 bytes, so an hour is about 63 KB a sandbox. It stays readable after the sandbox ends, until sandboxd forgets that sandbox (it keeps the 100 newest terminated ones), and is gone when sandboxd restarts: a sandbox taken back after a restart starts a new hour. For a longer history, poll this endpoint and keep the samples yourself; sandboxd's Prometheus/metricshas node totals, not each sandbox's usage. cpu_percentis the share of the vCPUs the sandbox was given since the previous sample, 0 to 100; the first sample has nothing to measure from and is 0. The byte counts are totals since the sandbox started; a rate is the difference between two samples.- On Linux, memory is the VMM process's resident set: what the guest has
touched, which it does not give back by freeing it.
processesis 0 where the backend cannot see into the guest. - Samples are kept in the server's memory, with the sandbox: a restarted server starts again from none. Reading them is not activity, so a dashboard watching a sandbox does not keep it from idling out.
- Behind a gateway, a sandbox's metrics are its owner's, with the read scope.
A volume is a named filesystem that outlives the sandboxes it is mounted in: a package cache, a dataset, a model's weights. What one sandbox writes, the next one to mount it reads.
| Request | Effect |
|---|---|
POST /v1/volumes {"name", "size_mb"} |
creates an empty volume. 409 if the name exists. size_mb defaults to the server's default disk size and is bounded by max_disk_mb. |
GET /v1/volumes |
{"volumes": [{"name", "size_mb", "created_at", "attached_to"}]} |
DELETE /v1/volumes/{name} |
deletes it and everything on it. 409 while it is mounted. |
Mount volumes on create with "volumes": [{"name": "cache", "path": "/sandbox/home/.cache"}, {"name": "models", "path": "/models", "read_only": true}]:
- A volume has one writer or any number of readers among live sandboxes
(
409otherwise): several may mount itread_onlytogether, but none may mount it writable while another holds it, nor read-only while a writer does. A suspended sandbox still holds it.attached_tonames one holder (the writer, if there is one). - Paths are absolute and plain (letters, digits,
. _ - /). They may not be/tmpor a system directory:/proc /sys /dev /run /etc /usr /bin /sbin /lib* /boot. Two mounts may not nest. At most 8. read_onlyis enforced below the guest: the drive itself is read-only, so a write fails with409"read-only file system" even for root.- A sandbox with volumes cannot be snapshotted (
409): its forks would share them. A sandbox started from a snapshot cannot mount volumes. - On Firecracker the guest flushes and closes its volumes cleanly before its VM stops. A VM that crashes keeps only what it had already flushed, as a disk does after a power cut.
#Audit events (capability audit)
GET /v1/sandboxes/{ref}/events returns {"events": [...], "truncated": bool},
oldest first, at most the newest 5000. A terminated sandbox the server has
already forgotten is still answered by id, because the log outlives the
record; a name is matched against live sandboxes only.
{"time": "2026-10-02T10:17:02Z", "type": "process.exited", "sandbox": "sbx_…", "pid": 1, "exit_code": 4, "duration_ms": 12}
type |
Carries |
|---|---|
sandbox.created |
name, image, labels, network, env_names, snapshot, volumes; reason: "pool" when it came from a pool |
sandbox.network_updated |
network |
sandbox.updated |
each of name, labels, idle_timeout_secs that changed, as it is after |
sandbox.suspended, sandbox.resumed |
— |
sandbox.terminated |
reason: request or idle |
snapshot.created |
snapshot, bytes |
process.started |
pid, program, arg_count, args_sha256, cwd, env_names |
process.exited |
pid, exit_code, duration_ms |
file.read, file.written, file.removed |
path, bytes |
tunnel.opened |
port |
An environment value is never recorded, and neither is a file's contents
or a process's arguments. A process is recorded by its program (argv[0]),
how many arguments followed it, and args_sha256: SHA-256 over those
arguments, each followed by a NUL byte. An agent's arguments are its prompt,
and a command line is where a token gets typed, while the log outlives the
sandbox. The hash still lets a known command be matched:
printf '%s\0' -p "fix the tests" | sha256sum # claude -p "fix the tests"
The program itself is kept, so sh -c "<script>" records sh, 2 arguments and
the hash. Labels and paths are recorded as given.
#Images (capability images)
An image is pulled the first time a sandbox asks for it. These let the
endpoint's operator install one ahead of that, see what is installed and what
uses it, and remove what nothing uses. On a plain sandboxd the operator is
whoever holds its token. A gateway answers all three 501 unsupported and does
not report the capability: an install fills a node's disk for every tenant on
it, and the gateway's own catalog is not done yet.
GET /v1/images lists them, by name:
{"images": [
{"image": "ghcr.io/amitgb14/sandbox-base:edge", "state": "installed", "digest": "sha256:…",
"bytes": 2147483648, "installed_at": "2026-10-09T06:00:00Z", "in_use": 2, "default": true},
{"image": "python:3.13-slim", "state": "installing", "in_use": 0,
"progress": {"phase": "pulling", "done": 31457280, "total": 52428800}},
{"image": "ghcr.io/acme/tool:2", "state": "failed", "in_use": 0, "error": "…: manifest unknown"}
]}
stateisinstalled,installing(withprogress:pullingits blobs,doneoftotalbytes, thenbuildingits root disk) orfailed(witherror). A failed one stays listed until it is installed or removed, or the server restarts.in_usecounts the sandboxes running, suspended or starting here from it. An image one runs from is listed as installed even where the backend's own list does not have it.bytesis the image's root disk on Linux, where layers shared with other images are not counted; macOS does not report it.defaultandpooledmark the server's default image and the images it keeps pools of.
POST /v1/images with {"image": REF} starts installing it and answers 202
with its state: the pull and, on Linux, the build of its root disk, which is
most of a first sandbox's wait. It runs in the background, outlives the
request, and is followed in the listing. A second request while one runs is the
same install. The image is held to the rule a create's is: an image reference
(400), and one the policy's images list permits where it has one (403).
Installing one already installed checks its tag against the registry.
DELETE /v1/images?image=REF removes it and answers {"image", "freed_bytes"}.
On Linux its root disk goes once no other tag of the same image needs it, and
then the cached layers no installed image still records. It is refused (409)
while a sandbox here starts from it, while it is installing, and for the default
or a pooled image. Removing a failed one clears it from the list. 404 when it
is not installed.
#Node
GET /v1/node is what a sandboxd reports about itself to a gateway in front
of many (NodeStatus): its node name (empty when standalone), version,
the capabilities body, capacity and free (cpus, memory_mb,
disk_mb; free is capacity less what every sandbox not terminated has been
given, pooled ones included, and never negative), running, pooled by
image, the images whose root disks are already built, labels and
cordoned.
POST /v1/node/cordon with {"cordoned": true} makes every create answer
503 unavailable until {"cordoned": false}; sandboxes already there are
untouched. It returns the NodeStatus. A restart clears it.
A node started with --node-id n17 makes ids of the form
sbx_n17_0123456789abcdef.
#Gateway
sandbox-gateway serves this API in front of many nodes
(fleet.md). Every endpoint above that is not under /v1/node
is served the same way through it; what differs is who the caller is.
Credentials. Every request but GET /v1/health carries an API key the
gateway issued, Authorization: Bearer sgk_…. A missing, unknown or revoked
key is 401 unauthorized. A key acts as one user, in an optional tenant, with
the scopes it was issued:
| Scope | Covers |
|---|---|
sandbox:read |
GET of a sandbox, the list, processes, output, files, dirs and events; GET /v1/volumes, GET /v1/snapshots |
sandbox:create |
POST /v1/sandboxes, PATCH, run, processes, stdin, signal, attach, suspend, resume, snapshots, tunnel, file writes and deletes; POST /v1/volumes |
sandbox:delete |
DELETE of a sandbox, a volume or a snapshot |
sandbox:ssh |
the SSH endpoints below, and SSH logins: a login by SSH key or token needs its user to hold an active key with this scope, and an open connection is closed once they no longer do |
secrets:write |
PUT and DELETE of /v1/secrets/{name} |
org:create |
POST /v1/orgs |
admin |
every scope, on every user's sandboxes, and /v1/admin/*; may select any organisation |
Attach and a tunnel are GETs that write, so they need sandbox:create. A
key without the scope a call needs gets 403 refused ("this API key does not
have the … scope") before anything is looked up.
The request guard keeps the node's Origin check (an Origin other than the
gateway's own host or a --cors-origin is 403 forbidden_origin, checked
before the key), the body caps and the content types. It drops the Host
check: the gateway is reached by name on a network, behind a key no page can
hold.
Organisation. Any request may carry X-Sandbox-Org: NAME to act in
organisation NAME instead of the key's own tenant (Organisations).
The gateway checks it once, after the key and before anything else, and the
whole request then acts in that tenant. Absent, or naming the key's own
tenant (default for the default one), it changes nothing. NAME must be an
organisation the key's user is a member of (any one, for an admin key);
otherwise 404 not_found "no such organization", the same whether or not it
exists. The header given twice is 400 invalid_request. Scopes never change
with it. A plain sandboxd ignores it.
Ownership. A sandbox, volume or snapshot belongs to the user and tenant
that created it. A call on one the caller does not own — or that does not
exist — is 404 not_found, the same answer either way. A name is looked up
among the caller's own sandboxes only. GET /v1/sandboxes lists the caller's
own; an admin's lists every sandbox. Volume names are one namespace across the
gateway: a name another user holds is 409 conflict.
What the gateway adds or refuses on create:
- It stamps
gateway.owner(the user) andgateway.tenant(when set) on every sandbox. A request that sets any label beginninggateway.is400 invalid_request, and at most 30 labels may be given. - A name is unique per user: one in use among the caller's live sandboxes is
409 conflict. - Over the tenant's quota is
403 refused. - A sandbox from a snapshot, or mounting volumes, goes to the node that holds
them; volumes on different nodes are
409 conflict.
Capabilities (GET /v1/capabilities, any key) are what every answering
node offers: a capability only if all have it, the smallest limits, the
strictest ceiling.
Errors of its own:
| HTTP | code |
When |
|---|---|---|
| 503 | unavailable |
no node is answering, none is taking new sandboxes, or the node holding this sandbox is not answering |
| 502 | internal |
the node did not answer this request |
| 404 | not_found |
an endpoint the gateway does not serve, including /v1/node |
Anything a node answers is relayed as the node sent it.
#Gateway-only endpoints
A plain sandboxd answers each of these 404; that is how a client tells the
two apart (GET /v1/whoami).
GET /v1/health — {"status": "ok"}, with no credential, for load balancers.
GET /v1/whoami (any key) — the caller:
{"user": "alice", "tenant": "team-a", "key_id": "key_…", "scopes": ["sandbox:read", …], "org": "team-a"}.
tenant is the key's own; org is the organisation this request acts in,
after X-Sandbox-Org (default for the default tenant).
GET /v1/ssh (any key) — where the SSH server listens and the host key to pin:
{"host": "gateway.example.internal", "port": 2222, "host_keys": ["ssh-ed25519 AAAA…"], "fingerprint": "SHA256:…"}.
404 unsupported when the gateway serves no SSH.
POST /v1/ssh-keys (sandbox:ssh) {"key": "ssh-ed25519 AAAA… comment", "sandbox": "demo"}
registers a public key for SSH logins; sandbox, optional, limits it to that
sandbox, resolved to its id now. 201 with
{"id": "ssh_…", "fingerprint": "SHA256:…", "key": "…", "sandbox": "sbx_…", "created": "…"}.
One authorized_keys line, no options (command=, from= …): anything else is
400 invalid_request. Accepted types: ed25519, ECDSA P-256/384/521, the two
security-key types, and RSA of 2048 bits or more. A key already registered for
the same sandbox, or registered by another user, is 409 conflict.
GET /v1/ssh-keys (sandbox:ssh) — {"keys": [...]}, the caller's own.
DELETE /v1/ssh-keys/{id} (sandbox:ssh) — 204; 404 not_found for a key
that is not the caller's. Open SSH connections that logged in with the key are
closed.
POST /v1/sandboxes/{ref}/ssh-access (sandbox:ssh, and the sandbox must be
the caller's) {"ttl_secs": 600} issues a short-lived SSH login for one
sandbox. 201 with
{"user": "sgt_…", "host": "…", "port": 2222, "expires_at": "…", "command": "ssh -p 2222 sgt_…@…"}.
The token is the SSH user name and the whole credential; it is returned once.
ttl_secs 0 or absent is 15 minutes; more than 24 hours, or negative, is
400 invalid_request. 404 unsupported when the gateway serves no SSH.
Admin (admin scope; anything else is 403 refused):
| Body | Answer | |
|---|---|---|
POST /v1/admin/keys |
{"user", "tenant"?, "scopes": [...]} |
201 {"id", "user", "tenant", "scopes", "created", "secret"}; the secret is shown only here. A bad user, tenant or scope is 400. |
GET /v1/admin/keys |
{"keys": [{"id", "user", "tenant", "scopes", "created", "revoked"?}]}, never a secret |
|
DELETE /v1/admin/keys/{id} |
204; 404 for an unknown id. Before it answers, open requests made with the key (an attach, a followed output stream, a tunnel, a waiting run) are ended (api.revoked in the audit log), open SSH connections whose user no longer holds an active sandbox:ssh key are closed, and running jobs whose owner holds no active key are cancelled ("error": "cancelled: the owner's access was revoked") |
|
GET /v1/admin/nodes |
{"nodes": [{"name", "endpoint", "healthy", "last_seen", "error"?, "status"?}]}; status is the node's last NodeStatus |
|
POST /v1/admin/nodes |
{"name", "endpoint", "token_file"?, "ca_file"?, "cert_file"?, "key_file"?} |
201 with the node, polled once. A bad name or endpoint is 400; a file outside the gateway's node files directory is 403 refused; a node named in its node file is 409. |
DELETE /v1/admin/nodes/{name} |
204; 404 unknown; 409 for a node from the node file. Its sandboxes' owners are kept. |
|
POST /v1/admin/nodes/{name}/cordon |
{"cordoned": true} |
forwarded to the node's POST /v1/node/cordon; 200 with the node. 404 unknown; 502 when the node does not answer. |
POST /v1/admin/nodes/{name}/drain |
{"terminate": false} |
cordons the node and kept so across its restarts; 200 {"node", "cordoned", "remaining", "terminated": [...], "failed"?: [{"id", "error"}]}. With terminate, its sandboxes are terminated. 503 for a node not answering. |
GET /v1/admin/lost |
{"sandboxes": [{"id", "user", "tenant", "node", "cpus"?, "memory_mb"?, "node_down_since"}]}: sandboxes on nodes not answering for longer than --node-lost-after |
|
GET /v1/admin/audit |
?since=RFC3339&limit=N (1–1000, default 100) |
{"entries": [AuditEntry], "truncated"?}, oldest first; 501 when the gateway keeps no audit log |
GET /v1/admin/ssh-keys |
?user=U (required) |
{"keys": [...]}: that user's SSH keys |
DELETE /v1/admin/ssh-keys/{id} |
204; 404 for an unknown id. Open SSH connections that logged in with the key are closed, as by DELETE /v1/ssh-keys/{id} |
|
GET /v1/admin/orgs |
{"orgs": [{"name", "created", "created_by", "created_by_tenant"?, "members", "owners"}]}: every organisation |
A node's endpoint is https://host:port, unix:///path, or http:// on
loopback only.
#Organisations
Tenants users create and share (organizations.md). Types
are in internal/api/types.go (Org, OrgMember, AdminOrg). A member is a
user named with the tenant of their own keys; tenant is left out for the
default tenant.
| Who | Body | Answer | |
|---|---|---|---|
GET /v1/orgs |
any key | {"orgs": [{"name", "role": "owner" | "member", "created"?, "current"}]}: the key's own tenant first (default when it has none), then every organisation its user is a member of; current marks the one this request acts in |
|
POST /v1/orgs |
org:create |
{"name"} |
201 Org, the caller its owner. A name is 1–30 of a-z 0-9 -, starting with a letter, not ending with -, with no --, not default or admin: else 400. A name already an organisation or a tenant in use (by a key, sandbox, volume, snapshot, secret, job or service; compared without case) is 409 conflict. Past --max-orgs-per-user organisations made or owned is 403 refused. |
GET /v1/orgs/{name}/members |
a member, or admin | {"members": [{"user", "tenant"?, "role", "added"}]} |
|
POST /v1/orgs/{name}/members |
an owner, or admin | {"user", "tenant"?, "role"?} |
201 for a new member, 200 for a role change, with the member. tenant defaults to the caller's own (default names the default tenant); role is owner or member (default). Demoting the last owner, or adding a user whose name another member of another tenant already has in it, is 409 conflict. |
DELETE /v1/orgs/{name}/members/{user} |
an owner, or admin | ?tenant=T |
204; 404 for no such member; 409 for the last owner. Before it answers, the member's open requests, SSH connections and running jobs in the organisation are ended as for a revoked key. |
Someone who is not a member gets 404 not_found "no such organization" on
every one of these, as for one that does not exist; a member who is not an
owner gets 403 refused for a change. A key's own tenant that nobody created
as an organisation has no member list (404). The audit log records
org.created, org.member_added, org.member_role and org.member_removed
(kind org; target the organisation or the member as tenant/user;
result the role given).
#Services
A sandbox spec and a count the gateway keeps true (services.md).
Types are in internal/api/services_types.go.
| Scope | Body | Answer | |
|---|---|---|---|
POST /v1/services |
sandbox:create |
ServiceSpec |
201 Service. 409 conflict when the tenant has a service by that name; 400 for a secret the tenant does not have, 501 unsupported for secrets on a gateway without a secrets key. |
GET /v1/services |
sandbox:read |
{"services": [Service]}, the caller's own (an admin's: every one) |
|
GET /v1/services/{name} |
sandbox:read |
Service; 404 for another user's |
|
PUT /v1/services/{name} |
sandbox:create |
ServiceSpec |
200 Service. A changed image, command, resources, port, health, env or network starts a rollout to a new revision. |
POST /v1/services/{name}/scale |
sandbox:create |
{"replicas": 5} |
200 Service |
DELETE /v1/services/{name} |
sandbox:create and sandbox:delete |
204; its replicas are terminated |
ServiceSpec is {"name", "image"?, "command"?: [...], "replicas", "resources"?: {"cpus", "memory_mb", "disk_mb"}, "port"?, "health"?: {"http": "/path" | "command": [...], "every_secs": 10, "timeout_secs": 5, "failures": 3}, "env"?: {}, "network"?, "placement"?: {"spread": "node"}, "public"?: false, "secrets"?: [...]}; unknown
fields are 400. name is a DNS label with no --, at most 100 replicas.
Service is {"spec", "env_names", "owner", "tenant", "revision", "serving", "desired", "ready", "restarts", "rollout"?: {"state": "in_progress" | "done" | "failed", "from", "to", "reason"?}, "error"?, "url"?, "replicas": [{"sandbox", "node", "revision", "state": "starting" | "healthy" | "unhealthy" | "lost", "healthy", "last_check"?, "last_error"?, "restarts", "created_at"}], "created_at", "updated_at"}. spec.env is left out; env_names lists its names. An admin
names another tenant's service with ?tenant=T.
Each replica is a sandbox the owner owns, labelled gateway.service and
gateway.service.rev; a request that sets either is refused, as for every
gateway.* label.
#Jobs, agent runs and secrets
Work the gateway runs after the request has gone, each run in a sandbox of its
own made through the same create path (jobs.md). Types are in
internal/api/jobs_types.go.
| Scope | Body | Answer | |
|---|---|---|---|
POST /v1/jobs |
sandbox:create |
JobSpec |
201 Job. A bad spec is 400, a reserved env name 403 refused, a secret the tenant does not have 400, secrets on a gateway without a secrets key 501. |
POST /v1/agent-runs |
sandbox:create |
{"agent", "prompt", "name"?, "image"?, "retries"?, "timeout_secs"?, "env"?, "secrets"?, "network"?, "resources"?, "from_snapshot"?, "keep"?, "notify"?} |
201 Job of one run |
GET /v1/jobs |
sandbox:read |
{"jobs": [Job]}, the caller's own, without runs |
|
GET /v1/jobs/{id} |
sandbox:read |
Job; 404 for another user's |
|
DELETE /v1/jobs/{id} |
sandbox:create |
cancels it; 200 Job once its runs have stopped (a minute at most) |
|
GET /v1/jobs/{id}/runs/{n}/output |
sandbox:read |
{"stdout", "stderr", "truncated"} (base64), kept when an attempt ends |
|
GET /v1/jobs/{id}/runs/{n}/files?path=P |
sandbox:read |
the file kept from keep.files, application/octet-stream |
|
PUT /v1/secrets/{name} |
secrets:write |
{"value": "…"} |
204. The name is an environment variable name, not reserved; the value at most 64 KiB, no NUL. 501 without a secrets key. |
GET /v1/secrets |
sandbox:read |
{"secrets": [{"name", "updated_at"}]}, the tenant's; never a value |
|
DELETE /v1/secrets/{name} |
secrets:write |
204; 404 unknown |
A job's notify URL is POSTed {"event": "run.finished" | "job.finished", "job", "name"?, "job_state", "run"?, "state"?, "exit_code"?} — never output
or a value. It must be https to a public address, checked again as the
gateway connects; a gateway started with --notify-allow-private also posts
to private and loopback addresses, and over http to loopback.
Secrets are the tenant's: any key of the tenant may name one in a job or a service, whose sandboxes then hold the value in their environment.
#Later additions
- Persistent (named) sandboxes that suspend on idle rather than terminate.
- Exposed ports on a public hostname (cloud).
Each arrives with its capability flag and its conformance tests.