Vercel's AI SDK lets a model request tools with typed arguments. plimsollCodeTools supplies executeCode through the checked TypeScript client. Keep calls fresh, or bind a tool to an authenticated user and conversation when the daemon supports sessions.
Vercel AI SDK: A TypeScript library for calling models, streaming responses and letting models request tools. This integration does not require Trigger.dev or a Vercel deployment. Vercel AI SDK (EXTERNAL · official docs ↗)
The Professor (a fictional narrator)
The model gets a calculator; the application keeps the keys. Even a very clever guest should not choose the locks.
Real quote
Geoffrey Hinton: “I had a principle when selecting graduate students: 'If they're not smarter than me, what's the point?' And I've had quite a number of graduate students who were smarter than me.”
Replay a recorded scenario. The diagram lights each hop the call passes; a refused call stops where it is refused.
These buttons send no execution requests. Each answer's source states how it was recorded, including any fixture used.
Exception panels show what reached the application, rather than an invented code result.
The model asks for
Tool result, or exception returned to the application
The Professor (a fictional narrator)
Each step of the call
Model: The model requests executeCode with code, language and optional text files. It cannot select credentials, the image, the floor or a conversation identity.
AI SDK: The SDK validates the tool's argument schema and invokes its execute function. It passes the call's cancellation signal.
Tool manager: plimsollCodeTools uses CodeSandboxes to select a fresh run or a session bound to the application's user and conversation pair. It reserves its runner files; other unsafe paths are refused by the daemon.
Checked client: The client serializes the project request, sends its required floor to the daemon and checks returned isolation and the run record. It does not retry an execution whose outcome is unknown.
plimsolld: The daemon enforces the configured boundary and deadline, runs in the selected image and owns session lifecycle. Packages are installed in the image before runtime.
Python or JS: A fresh call runs in a new project. A supported session executes a cell in a retained interpreter. The result states whether the interpreter and files remained available.
How thick should the walls be?
The executor carries a floor; the daemon states its
isolation tier. This picker is a teaching simulation:
change either to see whether the tier meets the floor. It sends no request and does not check other capabilities.
The Professor (a fictional narrator)
The executor's floor defaults to kernel: docker under runc is the container tier and refuses every call until it runs gVisor, or the executor is built with an explicit container floor, which is for development on your own code only. Use the spelling shown in this page's wiring example.
Surprises and limits
Bind identities in your applicationCreate one tool manager per process. forConversation takes a user and conversation pair from your authenticated application. Those IDs are not model-generated tool arguments. The unscoped executeCode always runs fresh, even when sessions are available.
A limit for each user's retained sandboxesmaxSessionsPerOwner defaults to 3, keeping a few conversations warm while bounding one user's allocation in this manager. Opening a fourth closes that user's least recently used sandbox with no call in progress. If all three are busy, the open is refused with notDispatched: capacity. The evicted conversation's next call starts fresh, reports freshSandbox and loses its earlier variables and files. Set maxSessions for this process's share of the daemon pool; its default 0 adds no process-wide cap. Both manager caps count one process only.
Several workers need the daemon limitFor several application workers or serverless functions, set SANDBOX_MAX_SESSIONS_PER_OWNER on the daemon; its default 0 leaves this quota off. Workers using the same daemon and application token send the same user's keyed digest, so the daemon counts their sessions together. It replaces that user's least recently used idle session, or refuses capacity if all are busy. After a replacement refusal, the manager opens a fresh sandbox. The application must authenticate the user and keep the daemon token on its server: an owner hint is not independent user authentication. This quota counts one daemon; several daemon replicas still need shared application admission if you promise a global limit.
State follows capability, not the tool's nameThe current Docker, OpenShell and E2B providers support retained interpreters when sessions are enabled; an E2B suspend ends them, so the next call reports freshInterpreter. The current Docker Cloud adapter runs fresh projects. Inspect stateKept, filesPersist, freshInterpreter and freshSandbox instead of assuming the last call's state survived. The initial call in a new interpreter reports freshInterpreter even if state is then retained.
Switch endpoints, then check capabilitiesThe same tool contract can target your local daemon or a remote daemon. Language support, packages, isolation and sessions come from that daemon's configured provider and image. Changing endpoints does not migrate a held interpreter or its files. Cloud and local configurations still need the appropriate authentication and connectivity.
Lifecycle and cancellationdispose closes one conversation's sandbox and disposeAll closes every held sandbox. The AI SDK abort signal reaches the checked client. If a dispatched call loses its transport, its effect may be unknown; without a notDispatched refusal mark, do not automatically replay it.
Packages are prepared before the model callsSupply dependencies in the daemon's image. The tool does not install npm or pip packages at runtime. The kernel floor is the default; ordinary Docker with its default runc runtime is container tier and is refused. These pages do not prove every remote daemon's isolation, and checked records are consistency checks rather than independent attestation.
These buttons replay evidenceThe scenario buttons replay real local tool calls, not fresh requests to a model or sandbox. The floor picker is a teaching simulation. A separate smoke example drives the real AI SDK loop with a scripted model and local gVisor. Scripted answers test integration plumbing, not a model's reasoning ability.