#Secrets: what sandbox-cli protects, and what it does not
A statement of the posture, so you can decide what to hand an agent rather than
infer it from behaviour. The backlog for the parts still open is
open-items.md, item 2; this file is what is true today.
The short version: sandbox-cli keeps a secret's value out of every place it would otherwise be written down, and does not keep it away from the agent. Those are different promises, and only the first is one this tool can make.
#The three ways a secret reaches a sandbox
| how | where the value comes from | reaches the agent as |
|---|---|---|
secrets: in your own config |
a host file, a host command, or a host env var, resolved per run on the machine running the CLI | an environment variable |
--env NAME / an agent's EnvAllow |
your shell's environment, forwarded only if set | an environment variable |
| the agent's own login | the files saved at ~/.config/sandbox/agents/<name>, copied in when a run starts |
a file the agent reads |
The third is the one most people forget, and it is usually the most valuable: on the default auth path it is an OAuth refresh token, not an API key. It is scoped to your whole account, does not expire on its own, and the same saved login is restored into every run of that agent, in every project.
#What is guaranteed
These are properties of the code, not intentions.
- A secret value never appears in an argv, a VM's config or its kernel
command line.
internal/credsresolves references on the host. The values travel in the API request body tosandboxd, and from there to the guest's environment over the agent channel. Neither the Firecracker config nor the macOS runtime's argv carries them (TestBuildConfigCarriesNoEnvironmentValue,TestBuildRunArgsCarryNoEnvironmentValue), and the API never returns them (EnvironmentValuesAreNeverReturned, conformance). - A secret value is never written to a config file by us. You give a reference — a path, a command, an env var name — and the resolution happens at run time.
- The audit log records environment variables by name only, and a process by
its program, argument count and a hash.
api.Eventhas nowhere to put a value or an argument, deliberately: the broker exists to keep secrets out of files, a log is a file, and a token pasted into a prompt is an argument. - A project's own
.sandbox.yamlcannot introduce secrets.secrets,envandenv_alloware on the refused list inpolicy/trust.go— a repository you cloned cannot make your machine resolve a credential. Naming a config with--config <path>is the deliberate act that overrides this. - Under
--profile prodno saved login is restored or saved, andValidateProfilerefuses a prod config that turns it back on. The refresh token never enters a sandbox, so it cannot leak from one. - A saved login carries the login and nothing else. Files that are also an
agent's settings keep only their login keys (
agents.FilterAuth). An MCP server, or another command an agent wrote into them, does not reach later runs.
#What is not guaranteed
Stated as plainly as the guarantees, because a reader who assumes the opposite will hand over the wrong credential.
- The agent can read every forwarded secret.
printenvis enough. It needs the value to authenticate, and nothing between it and the value can hide it without terminating TLS — which sandbox-cli has decided not to do, because the proxy that hid every secret would also hold every secret, every prompt in plaintext and a CA private key. So this one is permanent, not pending: the answer is to make a leak cheap, which is what the next section is about. - A leaked secret can leave through a host you allowed. If
github.comis on the egress allowlist — as it is under the default baseline — a token can be pushed there in a commit message or a gist. The allowlist decides where, not what. - Nothing stops a compromised agent from using a credential it holds. This survives every mitigation, including the TLS-terminating proxy: there the agent does not know the secret and can still act with its full authority.
#What to do about it
The threat model is prompt injection, so treat a leak as something to survive rather than prevent. What decides the damage is scope, lifetime, breadth and recoverability — not how well the value was hidden.
1. Run --profile prod when the credentials matter. It removes the
account-wide, never-expiring, cross-project credential from every sandbox, and empties the egress baseline so the permitted set is exactly what
you named rather than one that already includes github.com.
2. Broker short-lived, narrowly-scoped secrets. A secrets: entry runs a
command, so this needs nothing new:
secrets:
GITHUB_TOKEN:
command: gh auth token # long-lived: leaks for months
secrets:
GITHUB_TOKEN:
command: gh auth token --scope repo # or an STS/fine-grained mint
The agent still reads the value either way. The difference is whether what leaked is worth anything ten minutes later.
sandbox-cli says something when it can tell. A brokered value that looks long-lived produces one line naming the secret and what was observed about it:
sandbox-cli: secret GITHUB_TOKEN begins with "gho_" — GitHub OAuth tokens, which
is what `gh auth token` returns. A leaked value stays usable until you revoke it;
brokering a short-lived credential bounds what that is worth. …
Two things are checked, and they are not equally good. The expiry a JWT carries
in its own payload is a measurement: it works for any issuer, including ones
that do not exist yet, and it cannot go out of date. A prefix — ghp_,
glpat-, xoxb- and a few more — is a lookup against a short list, and lists
about the outside world are wrong at the edges forever. That is why the message
reports the prefix it saw rather than announcing whose credential you hold: if a
prefix is later reused by someone else, the sentence stays true and you can
dismiss it.
Read what that does not say. It never refuses: for some credentials the
long-lived form is the only form there is, and ANTHROPIC_API_KEY has no
ten-minute variant. And no warning is not a pass — most credentials are
opaque strings with no lifetime encoded in them, and the prefix list will never
be complete, so silence means nothing was recognized, not that what you brokered
is short-lived. Only you know that.
The check reads the value on the host, reports the format marker and nothing
else of it, and covers secrets: only. A credential forwarded with --env or a
wrapper's EnvAllow is not examined — those are mostly API keys with no
short-lived form, so warning on them would fire on nearly every run and become a
line nobody reads.
3. Use a different credential per project. A token shared across projects
reintroduces the breadth that prod just removed: one repository's compromise
becomes every repository's.
4. Keep --allow short. Free under prod, since the baseline is already
empty. Fewer reachable hosts is fewer places a leaked secret can be posted.
Together those leave the realistic worst case as a short-lived token, scoped to one repository, leaked from one run — which you rotate, or wait out.
#One deployment caveat
A secrets: command runs where the CLI runs, not where sandboxd runs. A
remote or self-hosted sandboxd therefore needs nothing of your tooling or
login. The value is resolved on your machine and sent in the create request,
which is one more reason a TCP sandboxd needs a token and, off loopback, TLS.