plimsoll trainer · 01 / foundations

Concepts

A load line is the mark painted on the outside of a ship's hull that shows how deep it may sit in the water, where anyone can read it. plimsoll paints the isolation tier, how strong the wall around a run is, the same way. It is a value on every response. A caller can set a floor, the weakest tier it accepts, before anything runs. And the daemon has to pass a startup check before it claims a tier at all.

The line and the water

The painted mark is the tier the provider, the backend that runs the code, proved at startup. The water is the floor the caller sets with minimum_isolation. Water above the mark means the request is refused before any code runs. Change the controls and read the ticket: it lists the daemon's checks in the order it makes them. The settings, in plain words:

Vvm Kkernel Ccontainer Pprocess proved: kernel no line painted floor: none

Verified runsc paints the K mark. No floor is stated, so the water sits below every mark and the run goes ahead.

Provider SANDBOX_PROVIDER daemon setting: one provider per daemon, fixed at startup

Runtime SANDBOX_DOCKER_RUNTIME daemon setting

Guard E2B_GUARD_URL or SANDBOX_DOCKERCLOUD_GUARD_URL daemon setting

Payload exactly one per Run

Grant grant_profile on each request

Caller's floor minimum_isolation on each request

the daemon's checksplimsolld · Run

    Nothing runs until somebody paints a line

    With SANDBOX_PROVIDER unset the daemon starts as the Disabled provider: every payload gets ErrDisabled, and nothing falls back to running the code some other way, not local Node and not the in-process engine. The operator configures the provider, so a request can state a floor and be refused, but it can never move the daemon to a stronger provider than the one configured.

    Two decisions, two people, two moments

    Whoever configures the daemon chooses the tier; whoever sends the request demands one. The in-process tier exists for the edit-run-debug loop: no account, no key, no container daemon, no cost per run. The floor keeps that loop out of production. It is one field on the request, checked before any code runs, so a deployment still set to wasm fails closed: it refuses, instead of quietly running hostile code in the daemon's own process.

    Read the mark, not the name

    The same Docker provider paints two different marks depending on what it can prove. Every result reports the tier that actually ran, and the daemon can refuse to start below a tier with SANDBOX_MIN_ISOLATION. Weakest first.

    1. ·none
      Disabled safe refusal

      Nothing runs. A deployment may deliberately stay here.

    2. Pprocess
      WASM: QuickJS inside the daemon too weak for hostile code

      QuickJS runs on wazero, a WebAssembly runtime written in Go, inside plimsolld. Memory is capped for real (wazero limits the engine's memory pages; it is not a wrapper that only reports a limit). But an escape, a bug that lets code out of the engine, lands in the daemon's own process. Fast, free, and no operating-system wall.

    3. Ccontainer
      Docker on runc, or an OpenShell sandbox too weak alone

      A read-only root filesystem, small fixed-size writable areas from which no program can be run, the network off, and root's special powers (Linux capabilities) dropped. Optionally, the shipped seccomp profile (SANDBOX_DOCKER_SECCOMP) lists the only system calls, requests to the kernel, the container may make. An escape still reaches the host's kernel, which the container shares. An OpenShell sandbox on a gateway that uses docker gets the same mark: OpenShell adds its own network and filesystem rules, but no OpenShell API reports which runtime its docker uses, so plimsoll claims nothing higher.

    4. Kkernel
      Docker on verified runsc eligible

      The same container, but the runtime is gVisor: a stand-in kernel, running as an ordinary program, takes the system calls of the guest (the code in the sandbox) before the host's kernel ever sees them. The mark is painted only after preflight, the check a provider runs at startup and on every readiness poll, confirms that the docker daemon plimsoll uses has runsc registered. That is evidence from configuration and behaviour, not attestation (cryptographic proof from the hardware) of the kernel actually running.

    5. Vvm
      E2B or Docker Cloud Sandboxes, a microVM strongest

      A microVM is a small virtual machine with its own kernel, created for one run and destroyed after it. The wall is hardware virtualisation, the processor's own separation between virtual machines, on somebody else's servers. E2B runs its microVMs on Firecracker, AWS's open-source virtual machine monitor, and Docker Cloud Sandboxes runs them on Docker-managed machines. The strongest wall, and the only tier that costs money per run: both services charge for the VMs they start.

    What a reported tier is, exactly: configuration, plus what the provider reports, plus the startup checks in section 03. It is never attestation. Any page that advertises a tier has to say so, and this one does.

    What a provider must prove before it serves

    Every real provider runs a startup test that exercises behaviour, not just configuration, and the daemon serves nothing if it fails. The tests differ because the sandboxes differ.

    Docker earns its mark in the mount table

    1. Use one local docker daemon. One Unix socket address, looked up once and reused. A remote docker daemon is refused, because the broker's per-run socket and the cleanup after each run only work on the same machine.
    2. Check the runtime. Read docker's registration of each runtime, including the path of the runsc program. Only a confirmed registration paints K; if a later check fails, the daemon reports itself not ready.
    3. Inspect both images. Both must be present. An image that declares a VOLUME is rejected, because docker would create an unlimited writable folder on the host for it. Every run then starts the exact image that was inspected, by its image ID, a digest (SHA-256 hash) of its content, so pointing a tag at a different image cannot slip it past the check.
    4. One throwaway container per image. From its own list of mounted filesystems, it must show a read-only root, and every writable mount as a tmpfs (a filesystem held in memory) of the promised size, marked noexec (no program on it can run) and nosuid (no program on it can gain privileges). Then it tries a write at every mount point, to show those are the only places a write lands.
    5. Reach a host socket. The first test container mounts a throwaway Unix socket the way a run mounts its broker socket, and must connect to it. Under runsc that needs --host-uds=open. A runtime that cannot carry a grant's calls refuses to start, instead of failing every run that has a grant.

    E2B earns its mark by creating a real VM

    1. Preflight checks settings only. It proves nothing about whether E2B is reachable, whether the key works, or whether the VM image (E2B calls it a template) exists. That is why /readyz never creates a VM: a test there would cost money on every poll, and anyone can poll it without logging in.
    2. One throwaway microVM at startup. It is created with both access tokens; its resources are checked live against the configured limits; several files are copied into its project directory; and a test program runs through exactly the path a project step takes.
    3. Egress, traffic leaving the VM, is checked live. The test proves a VM without a grant cannot reach out. That is the deny-all starting point, where nothing is allowed unless a rule allows it. The one exception, the path to the guard, is proved separately by make e2b-guard-live, which fails rather than skips when it is not configured.
    4. Every VM is labelled. A label naming this daemon instance lets it find and delete VMs it no longer tracks, such as one left behind by a garbled create response or a failed cleanup.

    The create response comes back without the token for envd, the agent inside the VM that runs commands. The VM can be reached at its public URL, so a run could go ahead. It does not: knowing the URL is not permission, the incomplete response is rejected, and the VM is killed. The provider never lowers its access requirements to make a run succeed.

    Docker Cloud Sandboxes earns its mark the same way, and checks the seal on every run

    1. Preflight checks settings only, as for E2B, so /readyz never creates a sandbox.
    2. The token can do the job. Docker's permissions check must list the read, create, delete and network-policy-read permissions a run needs.
    3. One throwaway sandbox at startup. It must come up with its network policy in force at deny-all, accept an upload of a directory and a file, and run the test through exactly the path a project step takes. That proves the tools are present, the working directory is honoured, and outbound traffic is blocked from inside the guest.
    4. The seal is read back on every run. The account's cloud network policy must default to deny-all (sbx --cloud policy init deny-all), because a create request that carries its own policy fails. Each run reads back the policy in force on its sandbox and refuses to run unless it is deny-all with nothing allowed.
    5. Every sandbox is named. A per-instance name prefix is recorded before the create is sent, so the daemon can delete untracked sandboxes it finds later.

    OpenShell earns its mark from the gateway, and reads every sandbox back

    1. The driver is the evidence. The gateway must report that it creates sandboxes with docker (its docker compute driver), the only way plimsoll has tested; any other is refused. That report is what paints C.
    2. Readiness asks the gateway at most once every 5 seconds. /readyz needs no login, so its answer is reused for 5 seconds; otherwise anyone who can reach the port could make the daemon call the gateway on every poll.
    3. One throwaway sandbox at startup. From inside, it must show that writes land only under /tmp, that outbound connections are refused and loopback is the only interface, that its own cgroup (the kernel's record of a group of processes' limits) holds the requested memory and CPU limits, that a cancelled command's processes are gone, and that a child process cannot read its parent's memory.
    4. Every run reads its sandbox back. Labels, image, limits, the policy and its hash, and no credentials attached by OpenShell. Any difference, and nothing runs.
    5. Every sandbox declares its lifetime. OpenShell sandboxes never expire on their own, so each one carries the seconds until its run's deadline. A daemon deletes its own untracked sandboxes at once, and another instance's once that lifetime plus 5 minutes has passed.

    With sessions on, one real session before serving

    Each provider's startup test also proves which interpreters its image runs, Python or JavaScript, and Describe reports only those. When the operator turns on sandboxes kept for many calls (SANDBOX_MAX_SESSIONS, section 07), the daemon then opens one for real and refuses to serve unless it keeps what section 07 promises: a process the first call leaves running is gone by the next call, a file survives calls and a suspend, an interpreter keeps what an earlier call defined or says it started fresh, and closing it ends it.

    A grant is a route, not the internet

    A run without a grant has no network. A run with one can reach exactly the routes an operator listed, through the broker, which holds the credential outside the sandbox. Pick a design and read what the hostile code can see.

    A grant profile lives in PLIMSOLL_GRANTS_FILE. It holds the API's address, the allowed routes, an allowed_callers list of principals (the authenticated caller ids allowed to use it), the scopes (permission labels the credential carries), and how the credential is made: one shared token from an environment variable, or a fresh short-lived JWT minted (created) for each run, naming the calling principal as its subject. The caller sends only the profile's name. A grant sent in full with a request is refused, an empty route list reaches nothing, and the scopes may not include code:run or *. plimsoll asks for the credential once per run.

    Three channels, kept apart

    "Your code ran and failed", "we never ran your code" and "our plumbing broke" are three different facts, and each decides differently whether a retry is worth it.

    resultThe code ran. A non-zero exit is a normal answer with stdout, stderr and the tier it ran behind.
    typed refusalNothing ran. A named error raised before dispatch, the moment the daemon hands the request to the provider, and marked as not dispatched, with a reason a client can branch on.
    infrastructureSomething around the run broke. Never disguised as the guest's stderr.
    The guest calls process.exit(2) deliberately

    result Result.ExitCode is 2, no Go error. Output, duration, sandbox name and isolation all stay meaningful. Retrying changes nothing unless the code changes.

    The snippet is over 256 KiB shared validation

    typed refusal ErrInvalidRequest, InvalidArgument over RPC. Checked before any provider is touched. The same channel carries a malformed floor and an unknown profile name.

    The floor is above the mark section 01

    typed refusal ErrInsufficientIsolation, FailedPrecondition. No hostile code ran. The official client also re-checks the tier on the response; a mismatch there is ErrIsolationEvidenceMismatch, DataLoss, and it is never a safe automatic retry because execution may already have happened.

    Every run slot is busy section 06

    typed refusal ErrAtCapacity, ResourceExhausted. Refused at once rather than queued, and worth retrying after a short, capped wait.

    A project's lint step fails stop on first failure

    result The step's exit code lives in Steps; the run's conclusion is ProjectResult.Outcome, one of completed, setup_failed, timed_out or protocol_error, with human context in Detail. A stable classification, not a string to parse.

    The container daemon disappears mid-run a dependency fails

    infrastructure An unmatched error, Internal over RPC. Output that was cut short, by contrast, is never an error: results carry StdoutTruncated and StderrTruncated flags, and the output that was kept never has a note written into it.

    Admission: the smaller of two numbers

    Admission is the daemon's capacity check just before a run starts. Counting runs and reserving memory are different limits, and it applies both. The number of runs that can be in progress at once is the run-count cap or the number of per-run memory allowances that fit in the total memory budget, whichever is smaller. Anything past that is shed: refused on the spot, not queued. Twelve runs arrive at once; move the controls.

    The total budget and the per-run allowance are set together or not at all, and a single run that needs more than the whole budget is a startup error, not a run that waits forever. E2B and Docker Cloud Sandboxes ignore the memory budget because their VMs run on other machines; the count still applies. On top of these limits, every caller has its own cap on runs in progress (SANDBOX_PER_KEY_CONCURRENT, half the global cap by default) and a token-bucket rate limit (SANDBOX_RATE_PER_MIN), which allows short bursts but refills at a steady rate, so one busy caller cannot take every slot.

    One sandbox, many calls

    A run gets a fresh sandbox and loses it at the end. A session keeps one sandbox for many calls: the files a call writes are there for the next call, and a cleanup between calls ends what a call left running. An agent that edits, builds and tests in a loop pays for one sandbox, and its later steps see what its earlier steps wrote. Sessions are off until the operator sets SANDBOX_MAX_SESSIONS. Two providers keep them: docker with a project image configured, under either runtime, and openshell. Describe says whether this daemon does (supports_sessions).

    What persists, and what does not the contract

    Files in the session's work directory persist, up to the session's disk budget (SANDBOX_SESSION_DISK_MB). Processes are another matter: a cleanup between calls ends what a call left running, while the interpreters a session keeps for its code (next item) are meant to outlive it. The cleanup runs after the answer has gone back, so the next call waits for it and the caller does not. When it cannot prove the sandbox clean, the session ends rather than run the next call beside a leftover. An interpreter can run between calls (a timer, a thread), so a process it starts then lives until the next cleanup, inside the session's limits. Which processes the cleanup spares, and how it tells them apart, is in the sessions guide (EXTERNAL · source repo ↗).

    State between calls notebook-style

    A cell is code the session runs in a Python or Node.js interpreter it keeps alive between calls, so variables, imports and loaded data can survive from one call to the next, as in a notebook. The request names the language, the code and optional files, which are written into the work directory before the code runs. The value of the last expression is printed. The exit code is 0 when the code ran, 1 when it raised an error, and 124 when the deadline ended it, which also ends the interpreter. interpreter_started says the call began with a fresh interpreter, so nothing defined earlier exists. A cell carries no grant, and it exists only in a session: Run refuses one. Describe lists the languages the startup test proved (languages); Python needs python3 in the image.

    Docker, OpenShell and E2B same contract, three sandboxes

    On docker a session is one locked-down container from the project image, with a run's walls: a read-only root, no network, size-capped writable directories held in memory, every capability dropped. Every call is a docker exec into it, in /work, and before each call the session compares the container's security settings with what they were at open and ends on any difference. On openshell it is one gateway sandbox whose main process does nothing, read back from the gateway before every call. On e2b it is one E2B virtual machine where the session's code runs as a user that cannot become root, while plimsoll's cleanup runs as root; a suspend pauses the machine, which keeps the files but ends the interpreters, so the next cell starts a fresh one. All three run the same cleanup between calls. dockercloud and wasm keep no sessions.

    Opening one without waiting a warm pool

    On docker, SANDBOX_SESSION_POOL keeps that many containers ready before anyone asks. Each has never run anyone's code and already has an interpreter running for every language the image runs, so opening a session takes one of them, and the first cell sends its code at once. On a laptop, opening a session plus its first cell took about 20 ms instead of about 750 ms under runc, and 30 to 50 ms instead of about 850 ms under gVisor. Each container goes to one session and is deleted when that session ends, like any session's. The price is memory: a waiting container holds its interpreters (39 MiB under runc, 66 MiB under gVisor, measured), and they count inside the session's memory limit. A caller can say which languages its cells will use when it opens a session, a hint: it gets the waiting container running those, and the pool keeps more containers of the languages callers ask for, so a container need not run an interpreter nobody uses. A hint changes how fast the first cell answers, never what runs. The rules are in the sessions guide (EXTERNAL · source repo ↗).

    Who can use a session one owner, one trust domain

    The caller that opened it, and nobody else. The session ID is treated as a secret: an unknown ID and another caller's ID get the same refusal, and the ID is never logged. Calls run one at a time, in order.

    The caller is the credential that authenticated, not the customers behind it. Everything a call leaves in a session is there for every later call, so a session must hold one trust domain: one person or one job. A service that runs many customers' code through one plimsoll credential has to keep each customer in a session of their own, keyed by identities it verified rather than by anything a user or a model typed; the strongest setup is one credential per customer. A sandbox that has run code is never used for anyone else: it is deleted when its session ends. The rules are in the sessions guide (EXTERNAL · source repo ↗).

    What a session gives up calls share a sandbox

    The wall around the code is a run's. What changes is how separate one call is from the next. Code a call runs can change the interpreter, leave a timer running or rewrite files, so a call's result is only as trustworthy as every call before it in the session; after a call that ran code you do not trust, open a new session. While a call with API access runs, anything an earlier call left running can use that access too, so a session refuses such a call before anything runs unless its profile sets allow_in_sessions. Code also gets the whole session lifetime, not one run's timeout, to probe the sandbox.

    When it ends, and when it sleeps typed, like a refusal

    At its lifetime (SANDBOX_SESSION_LIFETIME), past its disk budget, when code kills the sandbox's main process, when the sandbox changed under it, when the cleanup between calls cannot prove it clean, when the daemon shuts down, or when its owner closes it. A call on an ended session is refused before dispatch with ErrSessionEnded, and SessionEndedError says why. An idle session (SANDBOX_SESSION_IDLE) is suspended, not deleted. On openshell its sandbox is stopped: the files stay, the interpreters end, and the session gives back its admission slot. On docker the container is paused (docker pause): files, interpreters and memory stay, so the session keeps its slot.

    How you know no call was dropped a record per run

    Every answered run carries a run record: digests of what was sent and what came back, the evidence, and the times. The end time is the start plus the time that passed on a clock that never moves backwards, so it cannot fall before the start. In a session each record also names the previous one, so the calls form a chain. The daemon only computes the hashes; the caller's harness, a program outside the daemon that drives the runs, checks each record and signs it. A missing, reordered or foreign call breaks the chain, and the Go, Python and TypeScript clients each check it (the Go client reports a break as record.ErrChain).

    Details and measurements: sessions guide (EXTERNAL · source repo ↗).

    Before you switch it on

    A provider is ready when every line below is true for the job in front of you, not when its name sounds strong enough. Check what you have.

    Keep reading