API v1

Reference

#Sandbox API v1

The one API every mode serves — local (macOS), self-hosted (Linux) and cloud — and every client speaks: the CLI, the agent wrappers, Studio, the SDKs. A behaviour is "the same in all three modes" exactly where the conformance suite (internal/api/conformance) says so.

Status: M2. The core below is implemented by sandboxd against the fake backend. Later additions are listed at the end; they are additions, not changes.

#Transport

  • HTTP/1.1, JSON bodies (Content-Type: application/json), UTF-8.
  • File contents and process stdin are raw bodies (Content-Type: application/octet-stream).
  • Byte fields inside JSON (stdout, stderr, data) are base64, as Go's encoding/json writes []byte.
  • Streams are newline-delimited JSON (application/x-ndjson), one event per line, flushed as produced.
  • All paths are under /v1.

Reaching it:

Mode Listen Auth
Local unix socket, file mode 0600 the socket's permissions; a token is optional
Self-hosted TCP behind TLS Authorization: Bearer <token>, required
Cloud the control plane's endpoint API key as a bearer token

A server listening on TCP refuses to start without a token, loopback included: a loopback port is reachable by every user on the machine. One reachable from other machines needs TLS as well.

Request guard, applied before any handler:

  • Host must be a loopback name or one the server was configured to answer to (DNS-rebinding defence).
  • A request carrying an Origin that is not configured is refused, so a web page cannot drive the API.
  • Bodies are capped at 1 MiB for JSON and 64 MiB for file writes.

#Errors

Every non-2xx response has this body:

{"error": {"code": "refused", "message": "network mode \"open\" is above this server's ceiling (allowlist)"}}
HTTP code Meaning
400 invalid_request malformed, or a field out of range
401 unauthorized missing or wrong token
403 refused well-formed, and refused by policy. Retrying will not help.
403 forbidden_origin the request guard refused the Host or Origin
404 not_found no such sandbox, process or file — among this server's own
409 conflict a name in use, or a sandbox in the wrong state
501 unsupported this endpoint lacks the capability the request needs
503 unavailable this endpoint takes no new sandboxes for now (a cordoned node), or, through a gateway, no node is answering or the one holding the sandbox is down; try again, or elsewhere
500 internal a bug; the message is safe to show

refused and unsupported are different on purpose. The first is a decision: this server will not loosen its policy for anyone. The second is a limitation: another endpoint might be able to do it.

#Capabilities

GET /v1/capabilities

{
  "api_version": "v1",
  "backend": "firecracker",
  "capabilities": {"network_policy_update": true, "suspend": false, "memory_snapshot": false},
  "limits": {"max_cpus": 8, "max_memory_mb": 16384, "max_disk_mb": 102400},
  "network": {"default": {"mode": "allowlist", "allow": ["github.com"]}, "ceiling": "allowlist", "may_allow": ["*"]}
}

A client asks rather than assumes. A request needing a capability the endpoint lacks gets 501 unsupported, never a quietly weaker result.

audit is a property of the server rather than the backend: it is true when the server keeps an event log (below).

#Sandboxes

States: pending → running → terminated. suspended is added with the suspend capability; stopped with persistent sandboxes.

POST /v1/sandboxes

{
  "name": "fix-login",
  "image": "sandbox-base",
  "cpus": 2,
  "memory_mb": 2048,
  "disk_mb": 10240,
  "env": {"GOFLAGS": "-mod=mod"},
  "network": {"mode": "allowlist", "allow": ["registry.npmjs.org"], "deny": ["gist.github.com"]}
}

Every field is optional. The responses:

  • 201 with the sandbox.
  • 409 conflict when the name belongs to a sandbox that is not terminated.
  • 400 when a resource is above the server's limits.
  • 400 for a field the server does not know. That includes bind, which mounted a host directory before the repository model was removed: a client that still sends it is told so, rather than given a sandbox without the mount it asked for.
  • 403 refused for any of:
    • a reserved environment variable name (the ones that control the sandbox's own startup, the loader and the shell);
    • a network policy looser than the server permits (below).

The sandbox object:

{
  "id": "sbx_7f3a9c2e1b4d",
  "name": "fix-login",
  "state": "running",
  "image": "sandbox-base",
  "cpus": 2, "memory_mb": 2048, "disk_mb": 10240,
  "env_names": ["GOFLAGS"],
  "network": {"mode": "allowlist", "allow": ["registry.npmjs.org"], "deny": ["gist.github.com"]},
  "created_at": "2026-10-02T09:14:03Z"
}

Environment values are never returned: env_names only. The API is read by dashboards and logs, and a value echoed back is a secret on a screen.

  • GET /v1/sandboxes lists them: {"sandboxes": [...]}, newest first, terminated ones included until the server forgets them. ?label=key=value, repeatable, keeps only sandboxes carrying every label named.

  • GET /v1/sandboxes/{ref} returns one. ref is an id or a name, matched against this server's own sandboxes only and never passed to the backend to resolve.

  • PATCH /v1/sandboxes/{ref} changes a live sandbox; each field left out is left as it is. A request with one bad field changes nothing.

    • network: the egress policy of a running sandbox, under the same rule as create. Needs the network_policy_update capability.
    • name: renames it; "" removes the name. Unique among live sandboxes, as at create (409); through a gateway, among the caller's sandboxes on every node.
    • labels: replaces them whole; {} removes them all. The same rules as at create. A gateway keeps its own gateway.* labels: one sent back unchanged is accepted, any other is refused (400).
    • idle_timeout_secs: between 1 and limits.max_idle_timeout_secs; 0 (never) only where that limit is 0. Unlike at create, 0 is not "the server's default". Counted from the sandbox's last activity.

    name, labels and idle_timeout_secs are this server's records of the sandbox and change on a running or suspended one with nothing asked of its VM. A terminated sandbox is 409. A VM's vCPUs and memory are fixed while it runs: a new size is a new sandbox, from a disk snapshot where the backend takes them (Studio's Resize does that).

  • DELETE /v1/sandboxes/{ref} terminates the sandbox, stopping its processes and discarding its disk. It is idempotent: deleting a terminated sandbox is 204.

#Network policy

{"mode": "none" | "allowlist" | "open", "allow": ["host", "*.example.com"], "deny": ["host"]}
  • none: no egress at all.
  • allowlist: only names in allow are reachable, and a name in deny is refused even if allow matches it. Names are matched by name, not by resolved address, and case-insensitively. *.example.com matches subdomains but not example.com itself; nothing else is a pattern.
  • open: everything except deny.

Tighten, never loosen. The server holds a default policy (applied when a request omits network) and a ceiling. A request:

  • may pick any mode at or below the ceiling (none < allowlist < open);
  • may add deny entries freely;
  • may send its own allow list, which replaces the default one: narrowing is sending fewer names. Every name must be either in the server's default list or within its may_allow patterns. ["*"] permits any name; an empty may_allow permits only the default's names. A wildcard is permitted only by * or by a wildcard at least as broad, so *.a.example.com is allowed under *.example.com and *.example.com is not allowed under a.example.com;
  • may not send an allowlist that resolves to nothing. none is how you ask to reach nothing; an empty allowlist would otherwise be the strictest request producing the loosest result.

On the wire, "allow": null (or no allow key) means "the default list", and "allow": [] means "nothing", which is refused. Clients must keep the two distinct. allow makes no sense with none or open, and sending it there is 400. deny is dropped under none, since nothing is reachable.

Names are names: lowercase after normalising, no scheme, port, path or address, and a trailing dot is ignored. Anything else is 400.

A request that breaks any of these is 403 refused, naming the rule. It is never clamped silently.

#Processes

POST /v1/sandboxes/{ref}/run runs a command to completion.

{"argv": ["sh", "-c", "go test ./..."], "env": {"CI": "1"}, "cwd": "/sandbox/home/app", "stdin": "<base64>", "timeout_secs": 600}

→ 200 {"exit_code": 0, "stdout": "<base64>", "stderr": "<base64>", "truncated": false, "timed_out": false}

  • Each output stream is kept up to 8 MiB. Past that it is cut, and truncated says so.
  • When the timeout passes, the process is killed and the response comes back with timed_out: true and exit_code: -1.
  • argv is executed directly, never through a shell, unless argv[0] is a shell.
  • A command that does not exist is 400 invalid_request, not an exit code: it failed to start, so there is no process to report on.
  • cwd, if given, is an absolute guest path, checked like a file path. Without it the process starts in the sandbox user's home, /sandbox/home. There is no /workspace: a sandbox has no repository, and code gets in by git clone inside it or through the files API.
  • env here is merged over the sandbox's own, under the same reserved-name rule.

POST /v1/sandboxes/{ref}/processes takes the same body without timeout_secs. It starts the process in the background and returns 201 with the process object:

{"pid": 4, "argv": ["sleep", "30"], "state": "running", "exit_code": null, "started_at": "…"}

pid is the server's handle for the process, not the guest kernel's.

Request Effect
GET /v1/sandboxes/{ref}/processes lists processes, running and exited
GET …/processes/{pid} returns one process
GET …/processes/{pid}/output streams output from the beginning, then follows until exit (format below)
POST …/processes/{pid}/stdin writes the raw body to stdin; ?close=1 then closes it
POST …/processes/{pid}/signal sends {"signal": "TERM"}; one of INT, TERM, KILL, HUP

The output stream is one JSON object per line:

{"stream":"stdout","data":"aGkK"}
{"stream":"stderr","data":"…"}
{"exit_code":0}

#Files

Paths are absolute guest paths. They are never host paths, and the server never resolves them on the host. A relative path, or one containing .. or a NUL, is 400.

Request Effect
GET /v1/sandboxes/{ref}/files?path=/sandbox/home/app/go.mod returns the raw bytes; 404 if absent
PUT /v1/sandboxes/{ref}/files?path=… writes the raw body, creating parent directories; 204
DELETE /v1/sandboxes/{ref}/files?path=… removes the file or empty directory; 204
GET /v1/sandboxes/{ref}/dirs?path=/sandbox/home lists the directory (format below)

A directory listing:

{"entries": [{"name": "go.mod", "type": "file", "size": 412}, {"name": "cmd", "type": "dir", "size": 0}]}

#Attach

GET /v1/sandboxes/{ref}/processes/{pid}/attach with Connection: Upgrade and Upgrade: sbx-stream/1 switches the connection (101) to a two-way stream of frames: a type byte, a big-endian uint32 length, then the payload.

Direction Frame Payload
client → server i stdin bytes
client → server e end of input none (a terminal gets ^D)
client → server s signal INT, TERM, KILL or HUP
client → server r resize rows, cols (uint16 each)
server → client o stdout bytes
server → client E stderr bytes
server → client x exit, last int32 exit code
  • Output is replayed from the start, so a late or second attach sees everything.
  • Disconnecting detaches; it does not stop the process.
  • A process started with "tty": true (background only) runs on a terminal of rows × cols, and its output is one stream.

#Tunnels (capability tunnel)

GET /v1/sandboxes/{ref}/tunnel?port=N with Upgrade: sbx-tunnel/1 switches the connection to the raw bytes of a TCP connection to 127.0.0.1:N inside the guest. Only the guest's loopback is reachable this way, so a tunnel reaches a server in the sandbox and is never a way around its egress policy. Half-closes are passed through.

#Suspend, resume and snapshots

Request Capability Effect
POST …/suspend suspend stops the sandbox, keeping memory, processes and disk. Refused (409) while a process is running, because its stream would be cut.
POST …/resume suspend brings it back as it was
POST …/snapshots memory_snapshot or disk_snapshot captures the sandbox without stopping it: memory, processes and disk with memory_snapshot, its files only with disk_snapshot. Returns {id, sandbox, image, kind, bytes, created_at}; kind is memory or disk.
GET /v1/snapshots, DELETE /v1/snapshots/{id} either list and delete

Creating with "snapshot_id" starts a fork, which diverges from there:

  • from a memory snapshot it begins where the snapshot was, processes running; the snapshot fixes image and resources, so a request asking for different ones is 400;
  • from a disk snapshot it boots afresh with the snapshot's files: nothing that was running comes back. Only the image is fixed; resources are the request's, and default to the original's;
  • a backend may refuse a fork with a network (501), since a memory snapshot carries the guest's network identity.

A disk snapshot is a copy of a live filesystem: what the sandbox writes while it is taken may or may not be in it.

#A snapshot being taken

While a snapshot is taken, by request or by schedule, the sandbox carries its progress, and GET …/{ref} and the list show it:

"snapshotting": {"started_at": "2026-10-06T23:26:10Z", "phase": "capture",
                 "bytes": 958398464, "estimated_bytes": 2260828160}
  • phase is capture (reading the sandbox), then store (keeping it where a fork starts from). bytes is what has been read; estimated_bytes, where the backend has one, is about how much there is. It is an estimate — on macOS the guest's own df — for drawing a bar, and bytes may pass it. "scheduled": true marks one the schedule started.
  • Not every step reports as it goes. On macOS the runtime reads nothing for the first half minute or so (bytes stays 0), and store, about half the time, says nothing until it is done; Firecracker reports neither.
  • The field is gone once the snapshot is made or has failed.
  • One at a time: a snapshot asked for while one is being taken is 409.
  • Deleting the sandbox does not wait for its snapshot, nor stop it. The sandbox is terminated at once and keeps snapshotting until the snapshot ends; the snapshot is then listed like any other, or, if it could not finish, is not there at all, with nothing partial left on the host. On macOS the runtime's export, once begun, reads the whole disk, so a snapshot asked for before the delete comes out complete.
  • A snapshot asked for is not cancelled when its request is: a client that closes the connection part way leaves it to finish, and to be listed.

#Scheduled snapshots

"snapshot_every_secs": 1800, "snapshot_keep": 3 on create, or PUT …/snapshot-schedule with {"every_secs": 1800, "keep": 3} on a running sandbox, has the server snapshot it every so often while it runs and keep the newest few (keep defaults to 1; every_secs 0 stops it). The schedule comes back on the sandbox as snapshot_every_secs and snapshot_keep.

  • The server takes them, so they happen whether or not a client is watching. The interval counts from the end of one snapshot to the start of the next: a disk snapshot that takes a minute never overlaps the following one.
  • They are listed with the rest, marked "scheduled": true. Retention removes only those: a snapshot taken by request is never counted or removed. Each removal is an audit event, snapshot.deleted with reason retention.
  • The server's limits.min_snapshot_every_secs and limits.max_snapshot_keep bound it, and a schedule outside them is 400, not narrowed to one inside.
  • A schedule where the endpoint takes no snapshots is 501, and on a sandbox with volumes 409, before anything is made.
  • Behind a gateway, a scheduled snapshot belongs to whoever owns its sandbox. On macOS one takes about a minute and the size of the sandbox's files on the host's disk.

#Also on create

  • "idle_timeout_secs" terminates the sandbox after that long with no request naming it and no process running. Without it the server's default applies, bounded by limits.max_idle_timeout_secs.
  • "snapshot_every_secs" and "snapshot_keep" give it a snapshot schedule (above).
  • "labels": {"agent": "claude", "team": "infra"} is the client's own metadata. Labels decide nothing about the sandbox: they come back with it, filter the listing and are recorded in its audit events. That is how a client says why a sandbox exists without the server knowing what the reason means. Keys are lowercase letters, digits and . _ / -, at most 63 characters. Values are at most 256 bytes of printable UTF-8. There are at most 32 labels. A violation is 400.

#Metrics (capability metrics)

GET …/metrics is the last hour of what a sandbox used, oldest first, sampled by the server every interval_secs (5) while it runs:

{"interval_secs": 5, "samples": [{"time": "2026-10-06T09:00:10Z", "cpu_percent": 49.8,
  "memory_bytes": 441450496, "memory_limit_bytes": 2147483648,
  "net_rx_bytes": 1200, "net_tx_bytes": 800, "disk_read_bytes": 0, "disk_write_bytes": 4096,
  "processes": 3}]}
  • It is measured on the host — the runtime's counters on macOS, the VMM's process and tap device on Linux — never asked of the guest, which could say anything.
  • The hour is kept in sandboxd's memory, with the sandbox, and nowhere else: not in the state directory and not in the records --keep-sandboxes writes. A sample is 88 bytes, so an hour is about 63 KB a sandbox. It stays readable after the sandbox ends, until sandboxd forgets that sandbox (it keeps the 100 newest terminated ones), and is gone when sandboxd restarts: a sandbox taken back after a restart starts a new hour. For a longer history, poll this endpoint and keep the samples yourself; sandboxd's Prometheus /metrics has node totals, not each sandbox's usage.
  • cpu_percent is the share of the vCPUs the sandbox was given since the previous sample, 0 to 100; the first sample has nothing to measure from and is 0. The byte counts are totals since the sandbox started; a rate is the difference between two samples.
  • On Linux, memory is the VMM process's resident set: what the guest has touched, which it does not give back by freeing it. processes is 0 where the backend cannot see into the guest.
  • Samples are kept in the server's memory, with the sandbox: a restarted server starts again from none. Reading them is not activity, so a dashboard watching a sandbox does not keep it from idling out.
  • Behind a gateway, a sandbox's metrics are its owner's, with the read scope.

A volume is a named filesystem that outlives the sandboxes it is mounted in: a package cache, a dataset, a model's weights. What one sandbox writes, the next one to mount it reads.

Request Effect
POST /v1/volumes {"name", "size_mb"} creates an empty volume. 409 if the name exists. size_mb defaults to the server's default disk size and is bounded by max_disk_mb.
GET /v1/volumes {"volumes": [{"name", "size_mb", "created_at", "attached_to"}]}
DELETE /v1/volumes/{name} deletes it and everything on it. 409 while it is mounted.

Mount volumes on create with "volumes": [{"name": "cache", "path": "/sandbox/home/.cache"}, {"name": "models", "path": "/models", "read_only": true}]:

  • A volume has one writer or any number of readers among live sandboxes (409 otherwise): several may mount it read_only together, but none may mount it writable while another holds it, nor read-only while a writer does. A suspended sandbox still holds it. attached_to names one holder (the writer, if there is one).
  • Paths are absolute and plain (letters, digits, . _ - /). They may not be /tmp or a system directory: /proc /sys /dev /run /etc /usr /bin /sbin /lib* /boot. Two mounts may not nest. At most 8.
  • read_only is enforced below the guest: the drive itself is read-only, so a write fails with 409 "read-only file system" even for root.
  • A sandbox with volumes cannot be snapshotted (409): its forks would share them. A sandbox started from a snapshot cannot mount volumes.
  • On Firecracker the guest flushes and closes its volumes cleanly before its VM stops. A VM that crashes keeps only what it had already flushed, as a disk does after a power cut.

#Audit events (capability audit)

GET /v1/sandboxes/{ref}/events returns {"events": [...], "truncated": bool}, oldest first, at most the newest 5000. A terminated sandbox the server has already forgotten is still answered by id, because the log outlives the record; a name is matched against live sandboxes only.

{"time": "2026-10-02T10:17:02Z", "type": "process.exited", "sandbox": "sbx_…", "pid": 1, "exit_code": 4, "duration_ms": 12}
type Carries
sandbox.created name, image, labels, network, env_names, snapshot, volumes; reason: "pool" when it came from a pool
sandbox.network_updated network
sandbox.updated each of name, labels, idle_timeout_secs that changed, as it is after
sandbox.suspended, sandbox.resumed —
sandbox.terminated reason: request or idle
snapshot.created snapshot, bytes
process.started pid, program, arg_count, args_sha256, cwd, env_names
process.exited pid, exit_code, duration_ms
file.read, file.written, file.removed path, bytes
tunnel.opened port

An environment value is never recorded, and neither is a file's contents or a process's arguments. A process is recorded by its program (argv[0]), how many arguments followed it, and args_sha256: SHA-256 over those arguments, each followed by a NUL byte. An agent's arguments are its prompt, and a command line is where a token gets typed, while the log outlives the sandbox. The hash still lets a known command be matched:

printf '%s\0' -p "fix the tests" | sha256sum    # claude -p "fix the tests"

The program itself is kept, so sh -c "<script>" records sh, 2 arguments and the hash. Labels and paths are recorded as given.

#Images (capability images)

An image is pulled the first time a sandbox asks for it. These let the endpoint's operator install one ahead of that, see what is installed and what uses it, and remove what nothing uses. On a plain sandboxd the operator is whoever holds its token. A gateway answers all three 501 unsupported and does not report the capability: an install fills a node's disk for every tenant on it, and the gateway's own catalog is not done yet.

GET /v1/images lists them, by name:

{"images": [
  {"image": "ghcr.io/amitgb14/sandbox-base:edge", "state": "installed", "digest": "sha256:…",
   "bytes": 2147483648, "installed_at": "2026-10-09T06:00:00Z", "in_use": 2, "default": true},
  {"image": "python:3.13-slim", "state": "installing", "in_use": 0,
   "progress": {"phase": "pulling", "done": 31457280, "total": 52428800}},
  {"image": "ghcr.io/acme/tool:2", "state": "failed", "in_use": 0, "error": "…: manifest unknown"}
]}
  • state is installed, installing (with progress: pulling its blobs, done of total bytes, then building its root disk) or failed (with error). A failed one stays listed until it is installed or removed, or the server restarts.
  • in_use counts the sandboxes running, suspended or starting here from it. An image one runs from is listed as installed even where the backend's own list does not have it.
  • bytes is the image's root disk on Linux, where layers shared with other images are not counted; macOS does not report it. default and pooled mark the server's default image and the images it keeps pools of.

POST /v1/images with {"image": REF} starts installing it and answers 202 with its state: the pull and, on Linux, the build of its root disk, which is most of a first sandbox's wait. It runs in the background, outlives the request, and is followed in the listing. A second request while one runs is the same install. The image is held to the rule a create's is: an image reference (400), and one the policy's images list permits where it has one (403). Installing one already installed checks its tag against the registry.

DELETE /v1/images?image=REF removes it and answers {"image", "freed_bytes"}. On Linux its root disk goes once no other tag of the same image needs it, and then the cached layers no installed image still records. It is refused (409) while a sandbox here starts from it, while it is installing, and for the default or a pooled image. Removing a failed one clears it from the list. 404 when it is not installed.

#Node

GET /v1/node is what a sandboxd reports about itself to a gateway in front of many (NodeStatus): its node name (empty when standalone), version, the capabilities body, capacity and free (cpus, memory_mb, disk_mb; free is capacity less what every sandbox not terminated has been given, pooled ones included, and never negative), running, pooled by image, the images whose root disks are already built, labels and cordoned.

POST /v1/node/cordon with {"cordoned": true} makes every create answer 503 unavailable until {"cordoned": false}; sandboxes already there are untouched. It returns the NodeStatus. A restart clears it.

A node started with --node-id n17 makes ids of the form sbx_n17_0123456789abcdef.

#Gateway

sandbox-gateway serves this API in front of many nodes (fleet.md). Every endpoint above that is not under /v1/node is served the same way through it; what differs is who the caller is.

Credentials. Every request but GET /v1/health carries an API key the gateway issued, Authorization: Bearer sgk_…. A missing, unknown or revoked key is 401 unauthorized. A key acts as one user, in an optional tenant, with the scopes it was issued:

Scope Covers
sandbox:read GET of a sandbox, the list, processes, output, files, dirs and events; GET /v1/volumes, GET /v1/snapshots
sandbox:create POST /v1/sandboxes, PATCH, run, processes, stdin, signal, attach, suspend, resume, snapshots, tunnel, file writes and deletes; POST /v1/volumes
sandbox:delete DELETE of a sandbox, a volume or a snapshot
sandbox:ssh the SSH endpoints below, and SSH logins: a login by SSH key or token needs its user to hold an active key with this scope, and an open connection is closed once they no longer do
secrets:write PUT and DELETE of /v1/secrets/{name}
org:create POST /v1/orgs
admin every scope, on every user's sandboxes, and /v1/admin/*; may select any organisation

Attach and a tunnel are GETs that write, so they need sandbox:create. A key without the scope a call needs gets 403 refused ("this API key does not have the … scope") before anything is looked up.

The request guard keeps the node's Origin check (an Origin other than the gateway's own host or a --cors-origin is 403 forbidden_origin, checked before the key), the body caps and the content types. It drops the Host check: the gateway is reached by name on a network, behind a key no page can hold.

Organisation. Any request may carry X-Sandbox-Org: NAME to act in organisation NAME instead of the key's own tenant (Organisations). The gateway checks it once, after the key and before anything else, and the whole request then acts in that tenant. Absent, or naming the key's own tenant (default for the default one), it changes nothing. NAME must be an organisation the key's user is a member of (any one, for an admin key); otherwise 404 not_found "no such organization", the same whether or not it exists. The header given twice is 400 invalid_request. Scopes never change with it. A plain sandboxd ignores it.

Ownership. A sandbox, volume or snapshot belongs to the user and tenant that created it. A call on one the caller does not own — or that does not exist — is 404 not_found, the same answer either way. A name is looked up among the caller's own sandboxes only. GET /v1/sandboxes lists the caller's own; an admin's lists every sandbox. Volume names are one namespace across the gateway: a name another user holds is 409 conflict.

What the gateway adds or refuses on create:

  • It stamps gateway.owner (the user) and gateway.tenant (when set) on every sandbox. A request that sets any label beginning gateway. is 400 invalid_request, and at most 30 labels may be given.
  • A name is unique per user: one in use among the caller's live sandboxes is 409 conflict.
  • Over the tenant's quota is 403 refused.
  • A sandbox from a snapshot, or mounting volumes, goes to the node that holds them; volumes on different nodes are 409 conflict.

Capabilities (GET /v1/capabilities, any key) are what every answering node offers: a capability only if all have it, the smallest limits, the strictest ceiling.

Errors of its own:

HTTP code When
503 unavailable no node is answering, none is taking new sandboxes, or the node holding this sandbox is not answering
502 internal the node did not answer this request
404 not_found an endpoint the gateway does not serve, including /v1/node

Anything a node answers is relayed as the node sent it.

#Gateway-only endpoints

A plain sandboxd answers each of these 404; that is how a client tells the two apart (GET /v1/whoami).

GET /v1/health — {"status": "ok"}, with no credential, for load balancers.

GET /v1/whoami (any key) — the caller: {"user": "alice", "tenant": "team-a", "key_id": "key_…", "scopes": ["sandbox:read", …], "org": "team-a"}. tenant is the key's own; org is the organisation this request acts in, after X-Sandbox-Org (default for the default tenant).

GET /v1/ssh (any key) — where the SSH server listens and the host key to pin: {"host": "gateway.example.internal", "port": 2222, "host_keys": ["ssh-ed25519 AAAA…"], "fingerprint": "SHA256:…"}. 404 unsupported when the gateway serves no SSH.

POST /v1/ssh-keys (sandbox:ssh) {"key": "ssh-ed25519 AAAA… comment", "sandbox": "demo"} registers a public key for SSH logins; sandbox, optional, limits it to that sandbox, resolved to its id now. 201 with {"id": "ssh_…", "fingerprint": "SHA256:…", "key": "…", "sandbox": "sbx_…", "created": "…"}. One authorized_keys line, no options (command=, from= …): anything else is 400 invalid_request. Accepted types: ed25519, ECDSA P-256/384/521, the two security-key types, and RSA of 2048 bits or more. A key already registered for the same sandbox, or registered by another user, is 409 conflict.

GET /v1/ssh-keys (sandbox:ssh) — {"keys": [...]}, the caller's own.

DELETE /v1/ssh-keys/{id} (sandbox:ssh) — 204; 404 not_found for a key that is not the caller's. Open SSH connections that logged in with the key are closed.

POST /v1/sandboxes/{ref}/ssh-access (sandbox:ssh, and the sandbox must be the caller's) {"ttl_secs": 600} issues a short-lived SSH login for one sandbox. 201 with {"user": "sgt_…", "host": "…", "port": 2222, "expires_at": "…", "command": "ssh -p 2222 sgt_…@…"}. The token is the SSH user name and the whole credential; it is returned once. ttl_secs 0 or absent is 15 minutes; more than 24 hours, or negative, is 400 invalid_request. 404 unsupported when the gateway serves no SSH.

Admin (admin scope; anything else is 403 refused):

Body Answer
POST /v1/admin/keys {"user", "tenant"?, "scopes": [...]} 201 {"id", "user", "tenant", "scopes", "created", "secret"}; the secret is shown only here. A bad user, tenant or scope is 400.
GET /v1/admin/keys {"keys": [{"id", "user", "tenant", "scopes", "created", "revoked"?}]}, never a secret
DELETE /v1/admin/keys/{id} 204; 404 for an unknown id. Before it answers, open requests made with the key (an attach, a followed output stream, a tunnel, a waiting run) are ended (api.revoked in the audit log), open SSH connections whose user no longer holds an active sandbox:ssh key are closed, and running jobs whose owner holds no active key are cancelled ("error": "cancelled: the owner's access was revoked")
GET /v1/admin/nodes {"nodes": [{"name", "endpoint", "healthy", "last_seen", "error"?, "status"?}]}; status is the node's last NodeStatus
POST /v1/admin/nodes {"name", "endpoint", "token_file"?, "ca_file"?, "cert_file"?, "key_file"?} 201 with the node, polled once. A bad name or endpoint is 400; a file outside the gateway's node files directory is 403 refused; a node named in its node file is 409.
DELETE /v1/admin/nodes/{name} 204; 404 unknown; 409 for a node from the node file. Its sandboxes' owners are kept.
POST /v1/admin/nodes/{name}/cordon {"cordoned": true} forwarded to the node's POST /v1/node/cordon; 200 with the node. 404 unknown; 502 when the node does not answer.
POST /v1/admin/nodes/{name}/drain {"terminate": false} cordons the node and kept so across its restarts; 200 {"node", "cordoned", "remaining", "terminated": [...], "failed"?: [{"id", "error"}]}. With terminate, its sandboxes are terminated. 503 for a node not answering.
GET /v1/admin/lost {"sandboxes": [{"id", "user", "tenant", "node", "cpus"?, "memory_mb"?, "node_down_since"}]}: sandboxes on nodes not answering for longer than --node-lost-after
GET /v1/admin/audit ?since=RFC3339&limit=N (1–1000, default 100) {"entries": [AuditEntry], "truncated"?}, oldest first; 501 when the gateway keeps no audit log
GET /v1/admin/ssh-keys ?user=U (required) {"keys": [...]}: that user's SSH keys
DELETE /v1/admin/ssh-keys/{id} 204; 404 for an unknown id. Open SSH connections that logged in with the key are closed, as by DELETE /v1/ssh-keys/{id}
GET /v1/admin/orgs {"orgs": [{"name", "created", "created_by", "created_by_tenant"?, "members", "owners"}]}: every organisation

A node's endpoint is https://host:port, unix:///path, or http:// on loopback only.

#Organisations

Tenants users create and share (organizations.md). Types are in internal/api/types.go (Org, OrgMember, AdminOrg). A member is a user named with the tenant of their own keys; tenant is left out for the default tenant.

Who Body Answer
GET /v1/orgs any key {"orgs": [{"name", "role": "owner" | "member", "created"?, "current"}]}: the key's own tenant first (default when it has none), then every organisation its user is a member of; current marks the one this request acts in
POST /v1/orgs org:create {"name"} 201 Org, the caller its owner. A name is 1–30 of a-z 0-9 -, starting with a letter, not ending with -, with no --, not default or admin: else 400. A name already an organisation or a tenant in use (by a key, sandbox, volume, snapshot, secret, job or service; compared without case) is 409 conflict. Past --max-orgs-per-user organisations made or owned is 403 refused.
GET /v1/orgs/{name}/members a member, or admin {"members": [{"user", "tenant"?, "role", "added"}]}
POST /v1/orgs/{name}/members an owner, or admin {"user", "tenant"?, "role"?} 201 for a new member, 200 for a role change, with the member. tenant defaults to the caller's own (default names the default tenant); role is owner or member (default). Demoting the last owner, or adding a user whose name another member of another tenant already has in it, is 409 conflict.
DELETE /v1/orgs/{name}/members/{user} an owner, or admin ?tenant=T 204; 404 for no such member; 409 for the last owner. Before it answers, the member's open requests, SSH connections and running jobs in the organisation are ended as for a revoked key.

Someone who is not a member gets 404 not_found "no such organization" on every one of these, as for one that does not exist; a member who is not an owner gets 403 refused for a change. A key's own tenant that nobody created as an organisation has no member list (404). The audit log records org.created, org.member_added, org.member_role and org.member_removed (kind org; target the organisation or the member as tenant/user; result the role given).

#Services

A sandbox spec and a count the gateway keeps true (services.md). Types are in internal/api/services_types.go.

Scope Body Answer
POST /v1/services sandbox:create ServiceSpec 201 Service. 409 conflict when the tenant has a service by that name; 400 for a secret the tenant does not have, 501 unsupported for secrets on a gateway without a secrets key.
GET /v1/services sandbox:read {"services": [Service]}, the caller's own (an admin's: every one)
GET /v1/services/{name} sandbox:read Service; 404 for another user's
PUT /v1/services/{name} sandbox:create ServiceSpec 200 Service. A changed image, command, resources, port, health, env or network starts a rollout to a new revision.
POST /v1/services/{name}/scale sandbox:create {"replicas": 5} 200 Service
DELETE /v1/services/{name} sandbox:create and sandbox:delete 204; its replicas are terminated

ServiceSpec is {"name", "image"?, "command"?: [...], "replicas", "resources"?: {"cpus", "memory_mb", "disk_mb"}, "port"?, "health"?: {"http": "/path" | "command": [...], "every_secs": 10, "timeout_secs": 5, "failures": 3}, "env"?: {}, "network"?, "placement"?: {"spread": "node"}, "public"?: false, "secrets"?: [...]}; unknown fields are 400. name is a DNS label with no --, at most 100 replicas.

Service is {"spec", "env_names", "owner", "tenant", "revision", "serving", "desired", "ready", "restarts", "rollout"?: {"state": "in_progress" | "done" | "failed", "from", "to", "reason"?}, "error"?, "url"?, "replicas": [{"sandbox", "node", "revision", "state": "starting" | "healthy" | "unhealthy" | "lost", "healthy", "last_check"?, "last_error"?, "restarts", "created_at"}], "created_at", "updated_at"}. spec.env is left out; env_names lists its names. An admin names another tenant's service with ?tenant=T.

Each replica is a sandbox the owner owns, labelled gateway.service and gateway.service.rev; a request that sets either is refused, as for every gateway.* label.

#Jobs, agent runs and secrets

Work the gateway runs after the request has gone, each run in a sandbox of its own made through the same create path (jobs.md). Types are in internal/api/jobs_types.go.

Scope Body Answer
POST /v1/jobs sandbox:create JobSpec 201 Job. A bad spec is 400, a reserved env name 403 refused, a secret the tenant does not have 400, secrets on a gateway without a secrets key 501.
POST /v1/agent-runs sandbox:create {"agent", "prompt", "name"?, "image"?, "retries"?, "timeout_secs"?, "env"?, "secrets"?, "network"?, "resources"?, "from_snapshot"?, "keep"?, "notify"?} 201 Job of one run
GET /v1/jobs sandbox:read {"jobs": [Job]}, the caller's own, without runs
GET /v1/jobs/{id} sandbox:read Job; 404 for another user's
DELETE /v1/jobs/{id} sandbox:create cancels it; 200 Job once its runs have stopped (a minute at most)
GET /v1/jobs/{id}/runs/{n}/output sandbox:read {"stdout", "stderr", "truncated"} (base64), kept when an attempt ends
GET /v1/jobs/{id}/runs/{n}/files?path=P sandbox:read the file kept from keep.files, application/octet-stream
PUT /v1/secrets/{name} secrets:write {"value": "…"} 204. The name is an environment variable name, not reserved; the value at most 64 KiB, no NUL. 501 without a secrets key.
GET /v1/secrets sandbox:read {"secrets": [{"name", "updated_at"}]}, the tenant's; never a value
DELETE /v1/secrets/{name} secrets:write 204; 404 unknown

A job's notify URL is POSTed {"event": "run.finished" | "job.finished", "job", "name"?, "job_state", "run"?, "state"?, "exit_code"?} — never output or a value. It must be https to a public address, checked again as the gateway connects; a gateway started with --notify-allow-private also posts to private and loopback addresses, and over http to loopback.

Secrets are the tenant's: any key of the tenant may name one in a job or a service, whose sandboxes then hold the value in their environment.

#Later additions

  • Persistent (named) sandboxes that suspend on idle rather than terminate.
  • Exposed ports on a public hostname (cloud).

Each arrives with its capability flag and its conformance tests.

Edit on GitHubdocs/api/v1.md