plimsoll trainer

Dependencies

Two lists, not one. What is compiled into the daemon, and what a provider, the backend that actually runs the code, needs while it is running, are different questions with different failure modes. This page is the whole list, short enough to read on a phone.

The two lists

Compiled in

Four Go modules, linked into plimsolld: Connect, the RPC framework that serves plimsoll's procedures over ordinary HTTP; protobuf (Protocol Buffers), Google's schema-defined binary format for the messages; wazero, a runtime written in pure Go for WebAssembly, a portable bytecode format that runs inside a host program; and golang.org/x/net, Go's extra HTTP/2 code. They are in the process that handles hostile input, so each one is part of the trusted computing base, the code that has to be correct for plimsoll's security promises to hold, whether or not a particular run touches it.

Needed at run time

Whatever the selected provider reaches outside the binary: a container daemon and its images, or a cloud API and a template or image. Nothing here is linked into plimsoll. A missing one stops runs from starting; it does not weaken the boundary of a run that did start.

Those two failure modes are why the daemon checks them separately. A dependency that is absent makes the service refuse to serve. A dependency that is present but weaker than promised is the thing the startup smoke test, a real throwaway sandbox the daemon starts and checks before it serves, exists to catch.

Compiled in: four modules

Tap one for what it does, what hostile input reaches it, and what an upgrade has to re-prove. Versions are not on this page on purpose, because they change; they live in go.mod.

Connect the RPC layer

Does: serves the SandboxService over HTTP, decodes each request, and returns each result.

Hostile input reaches it: yes, first and always. Every request arrives through it before any plimsoll code runs.

An upgrade re-proves: that request bodies and decompressed messages are still independently capped, and that authentication still runs before the body is read.

protobuf the wire format

Does: encodes and decodes every request and response.

Hostile input reaches it: yes. A caller controls the bytes it parses.

An upgrade re-proves: that the generated code still matches the schema. The test gate (make audit) regenerates the code and fails on any difference, so a bump cannot quietly change the wire contract.

wazero the WebAssembly runtime

Does: runs the embedded JavaScript engine for the wasm provider, and applies the per-run memory cap in WebAssembly memory pages.

Hostile input reaches it: yes, the most directly of the four. It executes the guest, the code inside the sandbox.

An upgrade re-proves: that the memory cap is still enforced by the runtime rather than by a wrapper that reports it. This is the whole reason plimsoll drives wazero itself.

golang.org/x/net HTTP/2 plumbing

Does: the HTTP/2 paths the daemon serves and the broker, the part of plimsoll that makes API calls for sandboxed code, calls out through.

Hostile input reaches it: indirectly, on both sides of the broker.

An upgrade re-proves: nothing new, but this is the module whose security advisories most often decide which Go toolchain the official build is pinned to, fixed at one exact version. See the two numbers below.

golang.org/x/sys and golang.org/x/text come along indirectly. Adding a fifth direct module is a decision about what plimsoll trusts, not a convenience: anything that parses input or runs code has to earn its place.

The clients outside Go keep the same discipline. The Python client uses Python's standard library only. The TypeScript client has no runtime dependencies; its two add-ons import the agent framework they plug into (the SDK of Trigger.dev, a hosted job runner for TypeScript, with the ai package, or Mastra's core), and their tool's input schema imports zod; all of these are optional peer dependencies the application installs itself.

One module is banned by name

github.com/fastschema/qjs wraps the same JavaScript engine and offers a MemoryLimit setting. The setting does nothing. A run configured with a limit through that wrapper is a run with no limit, and it reports success either way, which is worse than an error. plimsoll drives wazero directly instead, so the cap is applied by the thing that owns the memory. Do not reintroduce it.

Needed at run time, per provider

WASM process tier

Needs outside the binary
Nothing. The engine is embedded in the executable.
Fails when
Effectively never for missing dependencies. It is the provider that still works on an air-gapped host, one with no network connection, with nothing installed.
The trade
An escape, a bug that lets code act outside its sandbox, from the engine lands inside plimsolld itself. So the process tier, the weakest of the four levels of wall strength plimsoll reports, is for development and for code you are not treating as hostile.

Docker container, or kernel with runsc

Needs outside the binary
A reachable container daemon, a runtime (runc, or runsc for a gVisor boundary: gVisor is Google's stand-in kernel, which answers a container's requests to the kernel itself), and the images named by SANDBOX_DOCKER_IMAGE and SANDBOX_DOCKER_PROJECT_IMAGE. Simulation runs (module runs, which run a compiled simulator once per row of a parameter table) add SANDBOX_DOCKER_MODULE_IMAGE.
Fails when
The daemon is down, an image is absent, or the runtime is not registered. All three are startup failures, not run failures: the daemon refuses to serve.
Worth knowing
Set SANDBOX_REQUIRE_PINNED_IMAGES and both images must be named by digest, the SHA-256 hash of the image's content. Every run then launches the exact image the startup check inspected, so pointing a tag at a different image later cannot slip it past that check.

E2B virtual machine

Needs outside the binary
The API of E2B, a hosted service that runs each sandbox in a small virtual machine; an E2B_API_KEY; and a template (the image E2B starts each sandbox from) named by E2B_TEMPLATE that already has the toolchain built in. Host-API grants, permission for a run's code to call listed routes of one API, need E2B_GUARD_URL as well.
Fails when
The network is down, the key is wrong, or the template is missing. Its startup configuration check cannot tell you any of that, which is why the behavioural smoke test creates one real sandbox before the service will serve.
Worth knowing
This is one of two providers whose dependency is somebody else's service, and one of the two that bill. That is also why no automated run in this project touches it.

Docker Cloud Sandboxes virtual machine

Needs outside the binary
Docker Cloud Sandboxes, Docker's hosted service that runs each sandbox as a small virtual machine on Docker's machines; a Docker personal access token with the "Cloud Sandboxes" scope in DOCKER_SBX_TOKEN and the account it belongs to in DOCKER_SBX_USERNAME; the management endpoint in SANDBOX_DOCKERCLOUD_API_URL, which is required because Docker documents no default; and the image each sandbox boots, SANDBOX_DOCKERCLOUD_IMAGE, which must already carry the toolchain. The account's cloud network policy must default to deny-all, blocking every connection that no rule allows (sbx --cloud policy init deny-all), because a create that carries its own policy fails.
Fails when
The network is down, the token is wrong or lacks a permission a run needs, or the account's policy is not deny-all. As with E2B, the startup configuration check cannot tell you the first two, so the smoke test creates one real sandbox first. The policy is read back on every run, and a sandbox that is not sealed (deny-all, plus, for a grant run, only the rule its grant needs) is refused.
Worth knowing
The sandbox API refuses the personal access token itself, so the daemon exchanges it at Docker Hub for a short-lived bearer token. A pinned image must name its linux/amd64 manifest digest, the hash of the file that describes that one platform's image, not of the multi-platform index that lists several, because the cloud reports the manifest it booted. With SANDBOX_CPUS and SANDBOX_MEMORY_MB unset, each sandbox is the Micro size (1 CPU, 2 GiB). Host-API grants need SANDBOX_DOCKERCLOUD_GUARD_URL: the grant run's one network rule, to the host of the guard (the one plimsoll address a run with a grant may reach) on port 443, is applied through a Docker call outside its published API and read back through the published one, and the guest holds its own run's short-lived guard credential. There are no module runs. Docker bills each sandbox per second, so, like E2B, no automated run touches it.

NVIDIA OpenShell container

Needs outside the binary
A gateway of OpenShell, NVIDIA's agent sandbox runtime; the gateway is its server that creates and deletes sandboxes on request. It must run the docker compute driver, the part that turns each sandbox into a docker container, and be reachable at SANDBOX_OPENSHELL_GATEWAY_URL over mutual TLS, in which both sides present a certificate: the certificate authority its certificate chains to in SANDBOX_OPENSHELL_CA_FILE, and a client identity the gateway accepts in SANDBOX_OPENSHELL_CERT_FILE and SANDBOX_OPENSHELL_KEY_FILE. The image each sandbox boots, SANDBOX_OPENSHELL_IMAGE, must carry node, sh and the project runner with its guard library, as the project image does, plus python3 for Python code in a session, one sandbox kept open for many calls.
Fails when
The gateway is unreachable, rejects the client identity or the call that reports its compute driver, or runs a driver other than docker. Each is a startup failure: the daemon asks the gateway for its driver, then proves the policy from inside one throwaway sandbox before it serves.
Worth knowing
plimsolld builds this provider itself rather than through the shared provider factory, because its generated protocol code registers the same names as OpenShell's own Go SDK, and a Go program that registers a name twice refuses to start. OpenShell's own default is no memory or CPU limit, so plimsoll always requests both (SANDBOX_MEMORY_MB, SANDBOX_CPUS). A process limit fails startup instead, because the gateway sets one for all of its sandboxes. A disk limit (SANDBOX_DISK_MB) makes /tmp, the only writable directory, a fixed-size filesystem held in memory, which the gateway allows only when its operator sets allow_driver_config. Each run gets a fresh sandbox, deleted when the run ends, and a session keeps one for many calls; module runs are not supported. The gateway is free, so its live suite costs nothing, but no automated run has a gateway to drive.

Guest tools are baked, never fetched

A project run compiles and lints inside a sandbox with no network. That works because the tools are already in the image: the project image carries Node and the TypeScript toolchain, a derived image carries Python with NumPy and SciPy (the image a session needs for Python code), a third carries the simulation worker with its compiled simulators, and a fourth, built on the third, carries a C compiler that produces WebAssembly. Third-party packages for projects are built into an image derived from plimsoll's, at build time, where Node's own lookup finds them.

The point is not convenience. A live install inside a run would mean fetching code you have not reviewed, at the moment you are least able to review it, into the process you trust least. Baking turns that into a reviewed ingredient of an image you built.

Pinned, because a tag is mutable

What the build fetches from outside the module is pinned to exact bytes, and each pin has a file that holds it. The guest images are the one exception, last in the list.

The embedded JavaScript engine
A QuickJS-ng build (QuickJS is a small JavaScript engine, and QuickJS-ng its maintained fork), pinned by SHA-256, with a test that fails if the embedded file stops matching. Recorded in sandbox/wasm/README.md.
The gVisor release
Installed by a script that pins a specific release and its checksum rather than tracking the latest one.
The three gate tools
The test gate runs tools that are not Go dependencies, so nothing in go.mod can pin them. gate-tools.versions does, and the gate's first step refuses to run against any other version. Otherwise a passing gate would only mean it passed against whatever happened to be installed.
The code generators
Pinned by go.mod tool directives and run through the Go toolchain, so regeneration needs no network at all.
The OpenShell protocol
OpenShell's protocol definitions, copied byte for byte from one tagged release, with each file's SHA-256 recorded in third_party/openshell/README.md. The client is generated from them, so moving to a newer OpenShell is a deliberate copy, not a drift.
Third-party CI actions
The reusable steps the GitHub workflows call. Pinned by commit, not by tag, for exactly the same reason as everything above.
The guest images, by version, not by bytes
The base image tag (node:22-alpine) moves with Alpine's releases, and the Python image pins Python, NumPy and SciPy to a minor version (3.14, 2.4, 1.17) rather than a patch: Alpine keeps only each package's current build, so a patch pin broke the image build the day a new patch shipped. What a run used stays exact: the daemon launches the image by the content ID its startup check inspected and states it on every run, and SANDBOX_REQUIRE_PINNED_IMAGES refuses an image not named by digest.

The toolchain is two numbers on purpose

go.mod carries a minimum Go version and a pinned toolchain, and they are deliberately different. The minimum is the oldest version that can build this module, kept low so a project that depends on plimsoll is not forced to upgrade. The pinned toolchain is the one the official build uses, kept current because the standard library of earlier versions had security advisories in the HTTP/2 and reverse-proxy paths this code actually calls. Build the daemon with the pinned one or newer.

Bumping any of it

An upgrade is not finished when the number changes. It is finished when the same checks that let the thing in the first time pass again, in this order.

  1. The local gate. Build, vet, race tests, lint, schema lint with a generated-code drift check, and a vulnerability scan.
    make audit
  2. The container suite. Adds the real isolation tests: a container comes up read-only, every writable mount is a sized no-execute temporary filesystem, the profile of allowed system calls (the requests a program makes to the kernel) loads, and the broker refuses what it should. The profile allows mknod only for the named pipes a session's interpreters use, never for a device or a regular file. Under this flag a missing daemon fails instead of skipping.
    make audit DOCKER=1
  3. The kernel-tier run. The same suite again under runsc, as its own run so a gVisor failure is attributable on its own.
  4. The live VM suites, by hand. They drive paid services (E2B and Docker Cloud Sandboxes), so they are never part of an automated run and never have a key or token in CI. A green automated gate therefore proves strictly less than a full local one, and the repository says so rather than implying coverage it does not have.
  5. The OpenShell suite, by hand. Free, but it needs a running gateway, which no automated run has. Like the VM suites, it fails rather than skips when its configuration is missing.
    make audit OPENSHELL=1

Unit tests alone cannot do this job. They cannot tell you that an artifact is the one you pinned, or that a runtime still enforces what its flag claims. Those are questions about the world, and only a run against the world answers them.

Keep reading