13 requests became 1, and the sandbox is what said so
One run of plimsoll's efficiency advisor, generated 2026-10-03 by go run ./examples/advisor -report. Every number on this page was measured by the API or read from the daemon's own log, or is labelled as a model.
plimsoll is an open source sandbox service for running code that an AI agent wrote. Its broker authorizes every call that code makes to your API, keeps the credential outside the sandbox, and records only route templates, verbs, status codes, byte counts and timings. A route template is the granted route with its variable segment left as a *, so a call to /items/i-7 is recorded as /items/*: which route the code used, never which item it asked for. There is no field for a path, a body or a credential, so none can be recorded by accident.
The efficiency advisor reads that record after a run and reports where the call pattern cost the API more than the question needed. When the profile declares a batch route for the route the code looped over (its batch_of, the operator's statement that one request returns what the loop fetched) and grants it, the finding names that route and comes back to the caller, who can rewrite. Every other finding stays with the operator: a declared route the profile does not grant is one the agent cannot call, and a route found from its path alone may page, return fewer fields or cover another scope, so the operator checks it before declaring it. plimsoll never calls a model to do any of this. It is a work in progress, and it is off unless a profile asks for it: a profile that does not set advice has its traffic left unanalyzed, and nothing is computed.
This page is one real run of the repository's advisor example: a fake inventory API with twelve items, a plimsoll daemon on the WASM provider, and one grant profile with advice: caller that declares GET /items the batch form of GET /items/*. The same question is asked twice, first the way an agent tends to write it, then the way the advice suggests.
What the agent's code did
host.get("/items").then(function (list) {
return Promise.all(list.items.map(function (it) { return host.get("/items/" + it.id); }));
}).then(function (rows) {
var total = rows.reduce(function (a, r) { return a + r.onHand; }, 0);
console.log(JSON.stringify({ total: total, rows: rows.length }));
});
What came back
- guest output
{"total":132,"rows":12}- guest wall time
- 304 ms, engine start included
- advice on the result
- 1 finding(s), below
suggested: GET /items, declared and granted
11 calls beyond one (measured count minus one) · about 1 ms beyond one call (modelled, not wall time) · 490 bytes moved in total (gross, not a saving)
This run made 12 successful GET /items/* calls to one per-item route (the trace holds the template, not the item); a collection or batch read would cover them in one request if it returns the same items. The profile declares GET /items as this route's batch form and grants it; the declaration is the operator's, and plimsoll has not checked that it returns the same items.
severity ranks nothing but the size of the pattern, for display: low under 25 calls, medium from 25, high from 100. A low finding is the same mistake as a high one over a smaller collection, so it is worth the same rewrite. Severity never gates a run or changes a result.
What the API saw, and what plimsoll's trace holds
/items/i-7 collapses to /items/*: the record says which route the code called, never which item it asked for. The trace never leaves the daemon, so this view is reconstructed from the profile's route templates, and the audit line's host_calls=13 is what confirms the count.| # | method | path as the API saw it | request bytes | response bytes | |
|---|---|---|---|---|---|
| 0 | GET | /items | 0 | 502 | |
| 1 | GET | /items/i-1 | 0 | 40 | |
| 2 | GET | /items/i-2 | 0 | 41 | |
| 3 | GET | /items/i-3 | 0 | 41 | |
| 4 | GET | /items/i-4 | 0 | 40 | |
| 5 | GET | /items/i-5 | 0 | 41 | |
| 6 | GET | /items/i-6 | 0 | 41 | |
| 7 | GET | /items/i-7 | 0 | 40 | |
| 8 | GET | /items/i-8 | 0 | 41 | |
| 9 | GET | /items/i-9 | 0 | 41 | |
| 10 | GET | /items/i-10 | 0 | 41 | |
| 11 | GET | /items/i-11 | 0 | 41 | |
| 12 | GET | /items/i-12 | 0 | 42 | |
| 13 requests | 992 bytes | ||||
The operator's view: the daemon's audit line
detected fan_out on GET /items/* (severity low, remedy batch), and the finding was returned to the caller
A finding the agent could not act on (no declared, granted route to switch to) would stay here, operator only: a declared route to grant, a candidate route to check, or a prompt the API owner can paste into their own AI. This run produced none of those: the profile declares and grants the batch route, so its one finding went back to the caller.
What the agent's code did
host.get("/items").then(function (list) {
var total = list.items.reduce(function (a, r) { return a + r.onHand; }, 0);
console.log(JSON.stringify({ total: total, rows: list.items.length }));
});
What came back
- guest output
{"total":132,"rows":12}- guest wall time
- 302 ms, engine start included
- advice on the result
- none
What the API saw, and what plimsoll's trace holds
/items/i-7 collapses to /items/*: the record says which route the code called, never which item it asked for. The trace never leaves the daemon, so this view is reconstructed from the profile's route templates, and the audit line's host_calls=1 is what confirms the count.| # | method | path as the API saw it | request bytes | response bytes | |
|---|---|---|---|---|---|
| 0 | GET | /items | 0 | 502 | |
| 1 requests | 502 bytes | ||||
The operator's view: the daemon's audit line
A finding the agent could not act on (no declared, granted route to switch to) would stay here, operator only: a declared route to grant, a candidate route to check, or a prompt the API owner can paste into their own AI. This run produced no finding at all, so there is nothing here for either audience.
Predicted next to measured
The finding's numbers compare the measured pattern with an assumed ideal of one call. The ideal is never measured, so the second run is the measurement.
| number | the finding said | the API measured |
|---|---|---|
| calls | 11 beyond one on GET /items/* (measured count minus one) | the API served 13 requests, then 1: 12 fewer |
| latency | about 1ms beyond one call (a model, not wall time) | guest wall time 304ms, then 302ms (includes engine start; not the same quantity) |
| bytes | 490 moved in total by the flagged calls (gross) | the API moved 992 bytes, then 502: 490 fewer |
Bytes matched here only because the collection call was already being made in run 1; in general the replacement call moves bytes of its own, which is why the number is a gross total and not a saving.
Run it yourself
With Go installed, clone the repository and fetch its dependencies. The example run then needs no docker, credentials, model or outbound network access:
git clone https://github.com/plimsollmark/plimsoll.git && cd plimsoll go run ./examples/advisor # the run, in the terminal go run ./examples/advisor -report out.html # the same run as a page like this one
The program checks its own claims: it fails if the two answers differ, if the loop does not come back with a fan-out finding naming the declared, granted route, if the rewrite comes back with any finding at all, if the per-item loop is flagged as a repeated read, or if the audit line carries anything but that one finding with its 11 calls beyond one. Source: EXTERNAL · source repo ↗ examples/advisor/main.go. How the advisor fits the rest: EXTERNAL · source repo ↗ efficiency advisor docs.
What this page claims, and what it does not
- Advice never changes a run. It is computed after the result is final, over metadata the broker already held. A run with advice is byte-identical in execution to one without; advice cannot gate admission or change an exit code, an output byte or the isolation tier.
- The trace is metadata only. Route templates, verbs, status codes, byte counts, latency. The toggle above shows what that leaves out.
- No model is involved. Two deterministic detectors (a fan-out over a per-item route, and repeated reads of one fixed route), counting only calls the broker delivered with a 2xx status, and one lookup of the profile's declared batch routes (plus, for the operator only, a route the path suggests). The paste-ready prompt for an API-change finding is text the operator may choose to hand to their own AI.
- The numbers are what they say. The call count is measured. The latency and byte figures are a model of the pattern against an ideal of one call, and the table above puts them next to what the API measured.
- A finding claims only what the trace can support. In run 1, 8 of the 12 per-item responses were the same size, and the API's log shows they were 12 distinct items. The trace holds the route template and the size, not the item, so it cannot tell same-size from same-item, and the repeated-read detector does not read a wildcard route at all. It reads only a route without a wildcard, where the broker admits exactly one path and every call is the same request; an unchanged size there is evidence the data did not change, and the finding says evidence, not proof.
- This run used the WASM provider, which is process-tier isolation and fine for an example. The advisor is provider-independent: the same broker core records the trace under Docker, E2B and WASM, so the findings do not depend on the tier.
- The advisor is a work in progress, and off by default. A profile that does not set
advicehas its traffic left unanalyzed and nothing computed at all; this page had to opt in withadvice: callerto produce anything. Two detectors ship today and the rule set is not settled: two earlier ones were removed once it was clear the trace could not support them, because it holds no call start times (so it cannot tell serial calls from concurrent ones) and no guest content (so it cannot tell a client-side reduce from any other loop). Expect the detectors to change. - Status: pre-1.0, single author, no external users yet, no third-party security audit. The lesson INTERNAL · trainer site → API Efficiency Advisor walks the same ground with a 128-call example.