plimsoll home · All integrations

plimsoll in a Trigger.dev chat agent

plimsollCodeSandbox gives a Trigger.dev chat agent an executeCode tool whose sandbox lasts the whole run. Unlike the fresh-run pages (Agno, CrewAI, Google ADK), variables, loaded data and files survive between calls and between turns, until the run suspends: onChatSuspend closes the sandbox right before the run sleeps, and the next message gets a new one. The price: the sandbox is one trust domain for the whole conversation, and while the run waits for the next message its sandbox holds a daemon slot.

Trigger.dev: a platform for running long background jobs written in TypeScript; its chat.agent runs an AI chat as one run that can sleep between messages. The add-on follows Trigger.dev's own code-sandbox recipe. Trigger.dev (EXTERNAL · official docs ↗)

The Professor (a fictional narrator)

The conversation gets a desk while the run is awake, and the desk is cleared before the run goes to sleep. Nobody pays rent on an empty office.

Wire it up

npm install @plimsollmark/client @trigger.dev/sdk ai zod
import { chat } from "@trigger.dev/sdk/ai";
import { PlimsollClient } from "@plimsollmark/client";
import { plimsollCodeSandbox } from "@plimsollmark/client/trigger";

const sandbox = plimsollCodeSandbox({
  // Built on first use: indexing the task at deploy time needs no secrets.
  client: () => new PlimsollClient({ baseUrl: process.env.PLIMSOLL_URL!, token: process.env.PLIMSOLL_TOKEN }),
  minimumIsolation: "kernel", // the default, written out; "container" for local docker
});

export const codeChat = chat.agent({
  id: "code-chat",
  tools: { executeCode: sandbox.executeCode },
  onTurnStart: async ({ runId }) => sandbox.warm(runId),      // start it, do not wait
  onChatSuspend: async ({ runId }) => sandbox.dispose(runId), // close it before the run sleeps
  onComplete: async ({ ctx }) => sandbox.dispose(ctx.run.id), // or when the run ends
  run: async ({ messages, tools, signal, streamText }) =>
    streamText({ model, messages, tools, abortSignal: signal }),
});

Options, limits and errors: Trigger.dev guide (EXTERNAL · source repo ↗)

Follow one call

Replay a recorded scenario. The diagram lights each hop the call passes; a refused call stops where it is refused. These buttons send no execution requests. Each answer's source states how it was recorded, including any fixture used. Exception panels show what reached the application, rather than an invented code result.

The modelasks for executeCodeSTOPPED HEREMAY HAVE RUNThe language model writes the arguments: the code, its language and any text files. It never sees or chooses the run id, the sandbox or the floor.Trigger.dev agentstreamText + hooksSTOPPED HEREMAY HAVE RUNchat.agent runs each user message as a turn inside one run. onTurnStart warms the run's sandbox without waiting; streamText calls executeCode whenever the model asks; onChatSuspend and onComplete close the sandbox.executeCodeplimsollCodeSandboxSTOPPED HEREMAY HAVE RUNThe AI SDK tool from plimsollCodeSandbox. It reads the run id from chat.local, which holds nothing else, and hands the call to CodeSandboxes under that key.CodeSandboxeskeyed by run idSTOPPED HEREMAY HAVE RUNHolds one plimsoll session per run id in this worker's memory and sends each call to it as a cell. The session ID never leaves the process: it is a capability, and chat.local is serialized into subtask metadata. The client checks every answer's run record and isolation tier.plimsolldsession + floorSTOPPED HEREMAY HAVE RUNOpens the session only behind the floor the tool asked for, or refuses before anything runs; runs each call in the session and states a run record for it.Session sandboxkept for the runSTOPPED HEREMAY HAVE RUNOne locked-down container with no network, kept for the run: a Python and a JavaScript interpreter stay alive between calls, and files stay in the working directory. Closed when the run suspends or completes.The modelasks for executeCodeSTOP?The language model writes the arguments: the code, its language and any text files. It never sees or chooses the run id, the sandbox or the floor.Trigger.dev agentstreamText + hooksSTOP?chat.agent runs each user message as a turn inside one run. onTurnStart warms the run's sandbox without waiting; streamText calls executeCode whenever the model asks; onChatSuspend and onComplete close the sandbox.executeCodeplimsollCodeSandboxSTOP?The AI SDK tool from plimsollCodeSandbox. It reads the run id from chat.local, which holds nothing else, and hands the call to CodeSandboxes under that key.CodeSandboxeskeyed by run idSTOP?Holds one plimsoll session per run id in this worker's memory and sends each call to it as a cell. The session ID never leaves the process: it is a capability, and chat.local is serialized into subtask metadata. The client checks every answer's run record and isolation tier.plimsolldsession + floorSTOP?Opens the session only behind the floor the tool asked for, or refuses before anything runs; runs each call in the session and states a run record for it.Session sandboxkept for the runSTOP?One locked-down container with no network, kept for the run: a Python and a JavaScript interpreter stay alive between calls, and files stay in the working directory. Closed when the run suspends or completes.

The model asks for

The tool answers the model

The Professor (a fictional narrator)

Each step of the call

  1. The model: The language model writes the arguments: the code, its language and any text files. It never sees or chooses the run id, the sandbox or the floor.
  2. Trigger.dev agent: chat.agent runs each user message as a turn inside one run. onTurnStart warms the run's sandbox without waiting; streamText calls executeCode whenever the model asks; onChatSuspend and onComplete close the sandbox.
  3. executeCode: The AI SDK tool from plimsollCodeSandbox. It reads the run id from chat.local, which holds nothing else, and hands the call to CodeSandboxes under that key.
  4. CodeSandboxes: Holds one plimsoll session per run id in this worker's memory and sends each call to it as a cell. The session ID never leaves the process: it is a capability, and chat.local is serialized into subtask metadata. The client checks every answer's run record and isolation tier.
  5. plimsolld: Opens the session only behind the floor the tool asked for, or refuses before anything runs; runs each call in the session and states a run record for it.
  6. Session sandbox: One locked-down container with no network, kept for the run: a Python and a JavaScript interpreter stay alive between calls, and files stay in the working directory. Closed when the run suspends or completes.

How thick should the walls be?

The executor carries a floor; the daemon states its isolation tier. This picker is a teaching simulation: change either to see whether the tier meets the floor. It sends no request and does not check other capabilities.

The Professor (a fictional narrator)

plimsollCodeSandbox's floor defaults to kernel (gVisor): docker under runc is the container tier and refuses every call until it runs gVisor, or the tool is built with minimumIsolation container, which is for development on your own code only. Keep kernel, or set vm, for code you did not write. With vm, e2b keeps sessions (a suspend ends the interpreter) and dockercloud keeps none, so there every call runs fresh.

Surprises and limits

One sandbox per run, not per callEvery executeCode call in a run shares one sandbox: variables, imports and loaded data survive between calls and between turns, and files stay in the working directory. That makes the run one trust domain: code from an early call can leave variables, files or patched functions behind that later calls run with. The fresh-run integrations give every call a new sandbox instead.
Closed when the run suspendsonChatSuspend closes the sandbox right before the run sleeps (after idleTimeoutInSeconds without a message, 30 by default), and onComplete closes it when the run ends. A message after a suspend gets a new sandbox: the conversation survives, its variables and files do not, and freshSandbox and freshInterpreter tell the model to rebuild them.
chat.local holds only the run idThe add-on keeps its sandboxes in the worker's memory, keyed by run id, and stores only the run id in chat.local. Trigger.dev serializes chat.local into subtask metadata, and a plimsoll session ID is a capability: whoever holds it can run code in that sandbox. So it never goes there.
A waiting run holds a daemon slotWhile the run waits for the next message its sandbox stays open, and counts against the daemon's session limit, until the idle window passes. A run that dies without reaching onChatSuspend or onComplete leaves its sandbox open until the add-on's idleCloseMs (10 minutes by default) or, if the process is gone too, the daemon's own idle suspend and session lifetime. sandbox.sandboxes.disposeAll() closes every sandbox a worker holds, for shutdown. To cap each user's sandboxes across a deployment's workers, pass the run's user to the warm, sandbox.warm(runId, userId), as your server authenticated it (never clientData, which the browser sets: a user who could choose it could close another user's sandboxes), and set the daemon's SANDBOX_MAX_SESSIONS_PER_OWNER. Each open then names the user to the daemon as a digest keyed by the caller token, and at the cap the daemon closes that user's least recently used idle sandbox; its run's next call says freshSandbox. Without a user, no per-user cap applies, and the add-on's maxSessions counts one worker's sandboxes only.
Kernel floor unless you lower itplimsollCodeSandbox requires the kernel tier (docker under gVisor, as the daemon's startup checks verified) unless you pass minimumIsolation, so an ordinary docker daemon under runc, the container tier, refuses every call before anything runs. The example agent sets container for local development on your own code; code you did not write needs kernel or vm, and with vm, e2b keeps sessions (a suspend ends the interpreter) while dockercloud keeps none, so there every call runs fresh.
These buttons replay an offline harnessmockChatAgent ran the real example agent's turns and hooks in memory, with a scripted model and a local daemon on docker under runc; the idle window was 3 seconds instead of 30. The harness does not run task lifecycle hooks, so the recorder called the agent's own onComplete after the run ended, as Trigger.dev's runtime does. Scripted answers test the plumbing, not a model's judgment.