Same run, different sandboxes: one fingerprint on every provider

The results on this page were measured in real runs on 2026-09-28; the inputs the runs were given are named as inputs where they appear. The run is the one the physics oracle page publishes: the same accepted controller, the same runner, the same cart-pole plant, started 0.2 rad from upright for 20 seconds. Here it was sent to each plimsoll provider in turn, with nothing changed but the provider. Switching is one setting, SANDBOX_PROVIDER, and the code that calls plimsoll does not change.

The verdict, provider by provider

ProviderWhere it ranIsolation reportedNodeSteps tookTrajectory fingerprint
Docker, locked down (runc)this machinecontainerv22.23.3314 msIDENTICAL 63cfd676924a3896127d60c97333677f3b3c804dd764dba001121f6c234a937f
Docker with gVisor (runsc)this machinekernelv22.23.31393 msIDENTICAL 63cfd676924a3896127d60c97333677f3b3c804dd764dba001121f6c234a937f
NVIDIA OpenShell sandboxan OpenShell gateway (docker driver)containerv22.23.3290 msIDENTICAL 63cfd676924a3896127d60c97333677f3b3c804dd764dba001121f6c234a937f
E2B Firecracker microVME2B's cloudvmv20.9.0874 msIDENTICAL 63cfd676924a3896127d60c97333677f3b3c804dd764dba001121f6c234a937f
Docker Cloud SandboxesDocker's cloudvmv22.23.3973 msIDENTICAL 63cfd676924a3896127d60c97333677f3b3c804dd764dba001121f6c234a937f
WebAssembly (in-process QuickJS)inside the daemonNOT RUN snippets only: the process tier runs no multi-file projects, so it cannot host the runner

IDENTICAL means the fingerprint equals the one published on the oracle page, 63cfd676924a3896127d60c97333677f3b3c804dd764dba001121f6c234a937f, to the last bit. The plant sent with every run hashes to ada5376adfa2652ea19439be69f70cde1e284573d51683a9f9bd0168106bbeee.

The run every provider recorded

t = 0.00 s

One record replayed: every provider that ran produced the same bytes, so there is only one trajectory to show. The browser only replays it; no physics runs here.

Why the numbers agree

The plant is WebAssembly: it rounds once per operation, has no fused multiply-add, and carries its own sin and cos. The controller uses only + and * on doubles. So the bytes depend on the inputs and on nothing about the machine, the kernel, the hypervisor or the Node version, which is why a Firecracker microVM in one company's cloud and a container on a laptop agree to the last bit.

The isolation column is the other half: each provider reports the boundary it ran behind, and a caller can require one on every run, so the same code can run fast in a container while you develop and in a microVM when the code is untrusted.

How it works

  1. The same files everywhere. The controller, the runner (run.mjs) and the plant (cartpole.wasm, sent base64 encoded and decoded by the first step) are ordinary project files, so no provider needs a special image. Any image with Node 20 or later can host the run. One more file, a one-line package.json declaring ES modules, is what made that true: Node 22 recognizes the controller's import syntax on its own, Node 20 does not, and the first E2B attempt failed on exactly that.
  2. Built the way the daemon builds them. Each provider is built from its own settings by the constructor plimsolld uses (plimsoll's provider factory, or the OpenShell provider's own, since the daemon builds that one itself), must pass its readiness checks and startup smoke test before its run counts, and then runs three steps: print the Node version, decode the plant, start the runner.
  3. The same check. The runner prints the SHA-256 of the record it wrote; the program hashes the returned artifact itself, refuses a mismatch, and compares the result with the fingerprint the oracle page published.

What this does not show

Reproduce: make docker-images && go run ./examples/providers in the plimsoll repository, with docker, and with the OpenShell, E2B and Docker Cloud settings in the environment for those rows. The page is written from the run. plimsoll is an independent project, not affiliated with or endorsed by NVIDIA, E2B or Docker.