"sync": true to the run request. StefanBrain holds the connection until the run ends or the sync window closes: about 90 seconds, counted from when StefanBrain receives the request, upload time included. A run that ends inside the window returns 200 with its final status, including output and structured_output. Check status, because failed and cancelled runs also return 200. Otherwise the response is the normal 202 envelope - no work is lost; finish by polling. A retry that reuses an Idempotency-Key answers at once without waiting.
Set your HTTP client’s timeout to 120 seconds for sync requests. The gateway in front of the API closes a connection that receives no response for 120 seconds. A closed connection, timeout, or 504 does not stop the run. It keeps running and is billed. If the start request carried an Idempotency-Key, retry with the same key on the same UTC day to get that run back instead of starting another.
Polling. The 202 envelope carries the run id and URLs. Poll GET /api/developers/v1/runs/{run_id} one time each 3–10 seconds until status is completed, failed, or cancelled. Polling does not count against rate limits. Calls that start work, file uploads, and transcript requests count.
Use sync mode for interactive requests and short tasks. Use polling for long work, batch pipelines, and environments where a held connection is fragile.
Choose what happens when a chat is busy
A chat runs one turn at a time. Onlyon_busy: "reject" (the default) is supported. A busy chat returns a refusal with retry guidance.
Poll the blocking run and retry after it finishes. Use separate chats or session labels for independent parallel work.
