plimsoll trainer · 05 / AI agents
Code mode replaces a menu of read-only tools with one
run_code tool: the agent writes a short program against a small SDK, and the
program runs in a sandbox. The agent's host offers that tool to the model over
MCP (Model Context Protocol), the tool-calling standard most agent hosts
speak. It makes the agent's interface narrower, not its permissions wider. The way to check that is to follow one
tool call down through every layer and ask, at each one, what it can and cannot change.
The model writes the tool arguments. The gateway, your service between the agent and plimsoll, has already authenticated the user, and it adds everything that gives the run a permission. plimsolld checks the request and runs the code; its broker, the part of plimsoll that makes the code's API calls, attaches the credential to each one. The domain API, the customer's own API the code calls, sees a route and a token, never the model. Choose what the model sends and watch where it stops.
The four: the agent conversation; the tenant (the customer organisation) the gateway maps it to; the principal, the caller identity plimsolld authenticates; and the subject the domain API reads from the token plimsolld minted, meaning created fresh for that one run. The last two are built from the token the gateway holds, so the gateway's choice of token decides whether tenants stay apart downstream.
One response shape for every run, including the ones where the code failed, so the agent can read its own mistake and revise. A refusal is a different shape, and it never echoes the code.
The name of the provider, the backend that runs the code, is not a policy.
Decide the minimum wall for each kind of code, the request's floor, put it on
every request as minimum_isolation, and read the isolation tier,
the strength of the wall, that every result reports.
floor: none · process is fine
Code you basically trust. Your own team's formulas, a demo, the inner loop while building the product. The engine inside the daemon (the wasm provider) is instant and free; an escape from it, a bug that lets the code act outside the engine, lands in the daemon, which you have decided you can live with.
floor: kernel
Model-authored code over your API. Treat it as hostile unless a policy says otherwise. Docker with verified runsc, the runtime of gVisor (a stand-in kernel that runs as an ordinary program), earns this tier, and carries both snippet and project grants (permission for the code to call listed routes of one API) through the broker. A plain container is not a wall against hostile code, whatever its image is called.
floor: vm
Hostile code, or not on your hardware. One disposable microVM (a small virtual machine created for one run) per run, on E2B or Docker Cloud Sandboxes, two hosted services that bill per use. Grants need the provider's guard, the one address on the plimsoll daemon the VM may reach (E2B_GUARD_URL or SANDBOX_DOCKERCLOUD_GUARD_URL); on Docker Cloud Sandboxes the guest, the code in the VM, holds its own run's short-lived guard credential. The template or image must already carry every tool, because the VM has no way to fetch one.
A small SDK, the preamble (JavaScript loaded before the agent's code), can dress the
generic host.* client up as inventory.items.get(id). It runs
inside the sandbox with the hostile code, so it is a convenience and a backstop, never
the thing that enforces anything.
.d.ts beside it, so the model writes against types.plimsoll-specgen, so the three cannot drift.* matches one whole path segment.allowed_callers.The guest can skip the preamble. A crafted script calls host.get
directly with any path it likes. That is fine: the broker checks the route against the
profile exactly as it would have anyway, and refuses the same way. The grant
example in the repository does precisely this and is refused, then succeeds under a
separate grant that lists the route. Nothing in the guest is relied on for security.
A chat agent calls its code tool many times in one conversation, and the TypeScript
client (npm install @plimsollmark/client) ships that tool:
executeCode takes code, a language (Python or
JavaScript) and optional files, text files written into the working
directory before the code runs, so the model can hand the code another tool's output by
path instead of pasting it into the source. The value of the code's last expression is
printed, as in a notebook. CodeSandboxes, underneath it, keeps one sandbox
per conversation.
Each conversation holds one session, one sandbox kept open for many
calls (docker with a project image, openshell or
e2b; an e2b suspend ends the interpreters). Every
call is a cell, code run in a Python or Node.js interpreter the session
keeps alive, so variables, imports and data loaded by earlier calls can still be
there, and so can the files.
On dockercloud and wasm every call runs
in a fresh sandbox, as a small project whose runner prints the last expression the
same way, and nothing persists. wasm runs no projects, so there a call
is a JavaScript snippet; Python and files are refused before anything is sent.
stateKept and
filesPersist say whether this call's interpreter and its sandbox were
still alive when the call answered. They are not a promise about the next call: the
cleanup after the answer can still end the session (its files over the disk budget,
a sandbox it cannot prove clean), and the session's lifetime can run out between
calls. The next call's freshInterpreter and freshSandbox say
what actually survived: freshInterpreter that nothing earlier calls
defined exists (the first call in a language, or one after a deadline or a crash),
freshSandbox that no earlier file is there either (a new sandbox after
an idle close, a dispose, or an end such as the session's lifetime).Two add-ons wire it into agent frameworks. @plimsollmark/client/trigger
follows the code-sandbox recipe of Trigger.dev, a hosted job runner for
TypeScript whose chat agents can sleep between messages: the sandbox is warmed when a
turn starts and closed right before the run sleeps and when it completes.
@plimsollmark/client/mastra does the same for Mastra, a TypeScript agent
framework, keyed by its thread and resource, and closes a sandbox on
dispose or after 10 minutes without a call.
One successful demo proves the happy path. The launch is decided by what happens under overload, bad code, a weak backend, two tenants, one forbidden route and one downstream outage: each must produce a different, safe answer.