interactive field guide
Lessons on running code you do not trust.
Twelve compact trainers, each one page, with no build step and no network calls. They cover the walls plimsoll can put around code, called isolation tiers, and how the daemon proves each one; how to call it from Go, Python, TypeScript or an AI agent; and how sandboxed code can use an API without ever holding its key. That is the job of the broker, the part of plimsoll that makes API calls for the code, and four trainers go deeper into it: how it connects to each kind of sandbox, a worked example with a private API, the efficiency advisor that reads a run's API calls afterwards, and how it backs off from an overloaded API. Each new word is defined where it first appears and collected in the glossary (INTERNAL · trainer site →).
Learn the system
Plain English
Brand new? Start here. One idea is behind every page: how many walls stand between untrusted code and your stuff. This page names each wall in plain words, then gives its real name.
Concepts and providers
Why the project is named after the load line, the mark on a ship's hull that shows how deep it may sit: pick a provider (a backend that runs code), set the floor (the weakest tier a request accepts), and see whether the run goes ahead. Then what each provider proves at startup, what a grant (one run's permission to call listed API routes) is and is not, the three kinds of answer, how the daemon caps the runs it takes at once, how one sandbox can serve many calls with an interpreter that can keep its state, and a checklist for switching it on.
Architecture
Follow code, credentials, API calls, and replies across an interactive map. Click any component to explore its incoming and outgoing connections.
Dependencies
Everything plimsoll's safety depends on, on one page: the four Go modules it imports and what each one touches, what each provider needs outside the program, which versions are fixed by a hash of their content, and the checks an upgrade has to pass. A reference sheet, not a lesson.
Build with it
Integrations
Run plimsoll inside your own Go program, or call a remote daemon with the official client for Go, Python or TypeScript, and get the same request checks, permissions and isolation guarantees either way.
Agent products
Design a run_code tool for an AI agent over MCP, the standard way an AI application offers tools to a model: who is calling, strict limits on each request, quotas, failure messages the agent can act on, and a code tool that keeps one sandbox per conversation.
Discoverable tools
Stop clients guessing: describe each tool's output with a schema (outputSchema), label honestly what each tool does, give the sandboxed code type definitions for what it can call, serve that description at runtime, and add a test that fails when it drifts from the code.
Customer examples
Worked recipes: an AI agent in a SaaS product writing code against the product's API, analytics, automation that needs a person's approval, checking whole projects, and running plimsoll privately.
Broker deep dives
The API Broker
How agent code reaches your API without ever holding the key, how Docker, WASM (WebAssembly run inside plimsoll itself) and NVIDIA's OpenShell sandboxes reach the same broker by different routes, why E2B, a hosted service of small virtual machines, is the hard case, and what a grant means when one sandbox serves many calls.
Private API Flight Recorder
See which program calls which. Follow two requests from the language model through the MCP host, a Node gateway, a Go service, plimsoll, the sandboxed code and a private API, where their paths split, and back.
Efficiency Advisor
See how the advisor reads the list of API calls a run made (which routes, how many, how long; never the data), after the run has finished, and sends a suggested fix to the agent or to the operator without changing how anything ran.
The capacity endpoint
See how the broker stops sending calls to an API that says it is overloaded, how one cheap health request can tell it early that the API has recovered, and how to set that request up from an OpenAPI description.
See real runs
see one real run
An AI agent wrote code to balance a pole on a cart. The sandbox ran it against a simulator.Not a lesson: every number on that page came from one run in a locked-down docker container on 2026-09-18. The agent's controller, a program that decides every 10 ms of simulated time how to push the cart, was the only file the caller sent. The simulated cart-pole (a cart on a rail with a pole hinged on top) and a runner program were built into the sandbox image; the runner ran the controller as a separate process and recorded the state every 10 ms. The accepted controller balanced the pole for 20 s and produced the same record, bit for bit, twice. The agent's first draft dropped the pole at 5.96 s. Replay both in the browser.
Open the run report →see real runs · same run, different sandboxes
The same run on five sandboxes, local and in two clouds: identical results.Not a lesson: on 2026-09-28 the same accepted controller, runner and simulator ran on five sandboxes from four providers (the backends that run code), with nothing changed but the provider's settings. A locked-down local container reported the container tier, the same container under gVisor, a stand-in kernel that runs as an ordinary program, reported kernel, and a sandbox from NVIDIA's OpenShell, an agent sandbox runtime, reported container. E2B and Docker Cloud Sandboxes, two hosted services, each ran it in a microVM, a small virtual machine created for one run, and reported vm. The Node versions differed, and all five records hashed to the same value the first report published.
Open the run report →see real runs · controller in C
The same cart-pole, with a controller written in C and compiled inside the run.Not a lesson: the run's first step compiled the C to WebAssembly, a portable bytecode, with no network. A runner built into the image then ran it in four scenarios where it swings the pole up from hanging, and the page compares each record with the same controller in JavaScript, value by value. The records differ where the two languages' cos functions disagree in the last bit; the page shows exactly where, and that the motion itself never differed.
How to use these: read them here, or serve the repository's docs directory with any static file server (serving only docs/trainers breaks the links up a level, and the architecture map stays blank when a page is opened straight from disk). Nothing you click is saved. The lessons have no build step, analytics, or third-party JavaScript.