invalid_request_errorauthentication_errorpermission_errorrate_limit_errorapi_errorservice_unavailable_error
Status codes
Rate-limit responses use
rate_limit_error as the error type and include a more specific code. Rate limits lists the request-limit codes. Other 429 codes include:
api_wallet_exhaustedapi_key_budget_exhaustedruntime_capacity_busywhen the service is near its limit and your account already holds its share of active runs. There is no fixed per-account cap: while capacity is free, your account can run more. Higher plans get a larger share when capacity is contended. The response always includesRetry-After: 5.monthly_usage_limit_reachedfor standard member-account MCP plan usage. Contracted partner accounts can have account-specific billing terms.feedback_rate_limitedwhen the account already sent 60 feedback reports in the last hour. The response includesRetry-After.
403 api_key_scope_forbidden.
Common 409 codes include agent_run_rejected (the Agent Run’s chat is busy; the body can include a top-level activeRunId, the UUID of the turn holding the chat), chat_project_mismatch, chat_id_retired, and, for tool submits only, idempotency_conflict. For a busy chat, wait for the blocking run to finish or address a separate chat.
A 503 brain_unavailable response means the brain is briefly unavailable. Retry later. It uses the service_unavailable_error type. A 503 workspace_busy response uses the rate_limit_error type and includes Retry-After: 5.
A 503 api_wallet_reconciliation_required response is a billing review hold for requests billed to the API wallet. A model receipt lacks billable token counts. Contact support; adding credit does not resolve the hold. Do not retry the request until support resolves the billing review.
When the agent runtime fails or times out, a request that reaches it, such as an Agent Run start or cancel, returns 503 with runtime_unavailable, runtime_timeout, or runtime_invalid_response. The work may have started. Check the chat or the run status, or retry a start with the same Idempotency-Key. A load balancer can still return 502 or 504 before the application answers.
Agent Run error codes
POST /api/developers/v1/runs can return these codes in addition to the shared codes above.
GET /api/developers/v1/runs/{run_id}, GET /api/developers/v1/runs/{run_id}/events, and POST /api/developers/v1/runs/{run_id}/cancel return 400 invalid_run_id for a malformed run id and 404 run_not_found when this account has no run with that id.
A run that fails after it starts reports why in last_error. budget_limited means the run reached the runtime’s limit of 1,500 model steps without finishing. Such a run is not billed.
Retry boundaries
Retry transient infrastructure failures. Do not retry malformed requests, authentication failures, permission failures, or attachment-size errors without changing the request or credentials. HonorRetry-After on 429 and 503. Use the returned effective limits and reset time to tune concurrency.
The TypeScript SDK retries requests that are safe to repeat: reads, cancellations, file uploads, event-stream reconnects, and submissions that carry an idempotency key (except research_shortform). It retries on network failures and 408, 429, 500, 502, 503, and 504, honors a Retry-After of up to 60 seconds, and makes at most two retries by default. It never creates an idempotency key for you.
Do not blindly retry a run or tool submission without an Idempotency-Key. Use the same key and address so the server can return the original accepted work.
- Agent Run and
ask_stefanbrainkeys deduplicate for the rest of the same UTC day. - Supported asynchronous submit tools deduplicate inside the resolved chat and family during the UTC-day window.
research_shortformis not idempotent.
429 api_wallet_exhausted. The server also validates the request body before it looks up the key. When the key matches, the retry answers at once. An Agent Run retry returns the 202 handle while the run is active, or the run’s final status after it ends. A tool submit retry returns the original job’s 202 envelope with the same job_id.
Run and job polling does not consume standard request limits.
When is_error is true, error.code classifies the tool failure and error.message gives the readable reason. The original content blocks remain in content.
For run failures, read output and last_error. For structured-output runs, structured_output is null unless the run completed and its final text is a JSON object.
For job failures, inspect the top-level status, is_error, and error, then the family-specific fields in data.
See Rate limits and Pricing for the controls behind 429 responses.
Stuck? Send feedback
When an error does not say how to fix the request, or the API cannot do what you need, report it withPOST /api/developers/v1/feedback. MCP clients call send_feedback. Include the endpoint or tool, the error code, and the run or chat id.
Reports are free, and the StefanBrain team reads every report and replies. See Agent feedback.
