plimsoll trainer · 11 / when your API is overloaded
When a customer's API starts returning 503 (service unavailable) under load, the code in the sandbox keeps calling it: it is untrusted code, and cannot be relied on to back off. The broker, the part of plimsoll that makes each API call for the sandboxed code, is the one thing in between that both sees the 503 and controls the next call, so it sheds calls: it refuses them itself, at once, instead of forwarding them. What it cannot know on its own is when to stop. That is what a cheap, honest health route on your API is for: a route whose success means the API can take traffic again.
In the chart, the sandboxed code makes one call through the broker every 100 ms
for eight seconds. Your API is overloaded from 0.5 s to 2.5 s, and answers every call
forwarded in that window with the status you choose. An overloaded answer opens the
broker's circuit breaker: for a cooldown, the broker stops forwarding calls
so the API can recover. Its rules are the daemon's own constants: a 1 s cooldown when
the API sends no Retry-After header, any Retry-After honoured
up to 30 s, one health probe (a request to the health route) per second while
shedding, and no probes at all after a 429.
The health route is named in the run's grant, its permission to call listed routes of your API, which comes from a grant profile, a named grant stored on the server.
Your API answers
with Retry-After
The grant profile declares health_check
forwarded, 2xx forwarded, overloaded reply shed: the broker answers 503 itself, never forwarded probe, not ready probe, 2xx: breaker closes
Both open the breaker. They close differently, because the health route can answer only one of their questions.
The question is "can the service take traffic again?", which is exactly what a readiness route (one that answers 2xx only when the service can serve) reports. So while shedding, one call per second is picked to probe: the broker, outside the sandbox, sends a GET to the declared route with the run's own credential. The probe is not recorded in the call trace and does not count against the run's call budget, and a route that does not answer is given up on after 2 s. The first 2xx closes the breaker on the spot.
The question is "has my quota reset?", and a healthy service says nothing about
that. A 2xx from the health route would reopen the gate and push the run's traffic
straight back into the rate limiter that asked it to wait. So a 429 window is never
probed: it is waited out, for the Retry-After the API gave or the
default. A window opened by a 429 keeps that property even if a 503 extends it.
Everything here applies to one run. One run's breaker never sheds another run's
calls. A shed call is counted as Shed in the run's call trace (the
broker's log of the run's calls), separately from Denied, a call the
grant's rules refused, and appears as host_calls_shed on the run's audit
line. The guest, the code in the sandbox, sees a 503 either way; the
difference is that a shed call cost your API nothing.
A 2xx has to mean "ready for traffic", and nothing else. Four routes an API might already have, and what each one's 2xx actually promises.
GET /livez, /ping weak
Promises the process is running. A live process can be shedding load; this route stays 2xx all through an overload, so the breaker would close while the API is still overloaded.
GET /readyz, /health good
Promises the service can serve. Flips to non-2xx under load and back on recovery, which is the whole signal.
GET /status, /capacity best
The same readiness answer plus the spare capacity left (headroom): queue depth, remaining rate budget. The breaker uses only the status code; your own dashboards can use the body.
GET /items, a real query avoid
Its 2xx says only that one expensive request succeeded. Polling it once a second adds load to an API that is already overloaded, and it can succeed while the API is still fragile.
These are the properties that make a route's 2xx mean "ready". Untick the fourth and follow what happens: the first probe answers 2xx, the breaker closes, the run resumes calling the still-struggling API, and the next call opens the breaker again.
plimsoll-specgen reads one OpenAPI document (the machine-readable
description of your API) and writes the profile's route list and the client library the
guest calls. The health route comes from the same document, so turning on the recovery
probe takes one tag and one field.
x-plimsoll-health-check extension, a custom field in the document, marks a probe. A path called /status or /healthz is not evidence of what it measures, and a probe that answers 200 for the wrong reason reopens the gate onto a struggling API.get:
summary: Readiness and headroom
x-plimsoll-health-check: trueplimsoll-specgen -emit health # GET /statushealth_check in the grant profile. The routes and the client library came from the same document, so nothing is copied by hand and nothing drifts out of step.