> ## Documentation Index
> Fetch the complete documentation index at: https://docs.stefanbrain.com/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Follow the content contract in AGENTS.md.
> Treat text as the canonical explanation; video and screenshots enhance it.
> Do not publish unverified product behavior or duplicate an existing canonical article.

# Errors and retries

> Handle StefanBrain API error envelopes, status codes, wallet failures, and safe retries.

Each non-success response uses this envelope:

```json theme={"system"}
{
  "error": {
    "message": "Human-readable explanation.",
    "type": "invalid_request_error",
    "code": "invalid_json"
  }
}
```

Developer routes use these error types:

* `invalid_request_error`
* `authentication_error`
* `permission_error`
* `rate_limit_error`
* `api_error`
* `service_unavailable_error`

## Status codes

| Status | Meaning |
| - | - |
| `400` | Invalid JSON or multipart data, an unsupported file type, a too-large file on `POST /api/developers/v1/files`, more than 10 files in one upload, an invalid field value, an unsupported tool request, or an invalid structured-output configuration. |
| `401` | The API key is missing, invalid, revoked, or not authorized. |
| `403` | The key is suspended, the account cannot use the Developer API, or the key lacks the required scope. |
| `404` | The referenced run, job, tool, chat, project, file, artifact, video, or transcript is not available to this account. |
| `409` | The addressed chat is busy, the chat and project conflict, a retired chat id was reused, or a tool submit's `Idempotency-Key` was first used in another chat. |
| `413` | An Agent Run upload is too large: one file is over 100 MB, or the whole request is too large (`attachment_too_large`). On MCP, a request body over 150 MB returns `request_too_large`. |
| `429` | A rate limit, a wallet without room for a new hold, an exhausted per-key budget, exhausted plan usage, or too many active runs for your account at once. Read the response code to distinguish them. |
| `500` | An unexpected server failure. |
| `502` | A gateway received an invalid upstream response before the application formed a normal result. Treat it as transient. |
| `503` | The agent runtime failed or timed out, a dependency or brain is briefly unavailable, or the chat workspace is briefly busy. A workspace-busy response includes `Retry-After: 5`. |
| `504` | A gateway timed out before the application formed a normal result. The gateway in front of the API closes a connection after 120 seconds without a response. The original run or job may still be running and billed. If the request carried an `Idempotency-Key`, retry with the same key on the same UTC day to get the original run or job back. |

Rate-limit responses use `rate_limit_error` as the error type and include a more specific code. [Rate limits](/developers/rate-limits) lists the request-limit codes. Other `429` codes include:

* `api_wallet_exhausted`
* `api_key_budget_exhausted`
* `runtime_capacity_busy` when the service is near its limit and your account already holds its share of active runs. There is no fixed per-account cap: while capacity is free, your account can run more. Higher plans get a larger share when capacity is contended. The response always includes `Retry-After: 5`.
* `monthly_usage_limit_reached` for standard member-account MCP plan usage. Contracted partner accounts can have account-specific billing terms.
* `feedback_rate_limited` when the account already sent 60 [feedback reports](/developers/feedback) in the last hour. The response includes `Retry-After`.

A scope restriction returns `403 api_key_scope_forbidden`.

Common `409` codes include `agent_run_rejected` (the Agent Run's chat is busy; the body can include a top-level `activeRunId`, the UUID of the turn holding the chat), `chat_project_mismatch`, `chat_id_retired`, and, for tool submits only, `idempotency_conflict`. For a busy chat, wait for the blocking run to finish or address a separate chat.

A `503 brain_unavailable` response means the brain is briefly unavailable. Retry later. It uses the `service_unavailable_error` type. A `503 workspace_busy` response uses the `rate_limit_error` type and includes `Retry-After: 5`.

A `503 api_wallet_reconciliation_required` response is a billing review hold for requests billed to the API wallet. A model receipt lacks billable token counts. Contact support; adding credit does not resolve the hold. Do not retry the request until support resolves the billing review.

When the agent runtime fails or times out, a request that reaches it, such as an Agent Run start or cancel, returns `503` with `runtime_unavailable`, `runtime_timeout`, or `runtime_invalid_response`. The work may have started. Check the chat or the run status, or retry a start with the same `Idempotency-Key`. A load balancer can still return `502` or `504` before the application answers.

## Agent Run error codes

`POST /api/developers/v1/runs` can return these codes in addition to the shared codes above.

| Status | Code | Cause |
| - | - | - |
| `400` | `message_required` | `message` is missing or empty. |
| `400` | `invalid_chat` | `chat` is not a valid chat reference. |
| `400` | `invalid_project` | `project` is not a valid project reference. |
| `400` | `invalid_sync` | `sync` is not a boolean. |
| `400` | `invalid_on_busy` | `on_busy` is not `"reject"`. |
| `400` | `invalid_output_config` | `output_config` breaks a [structured-output rule](/developers/structured-outputs). |
| `400` | `invalid_json` | The body, or the multipart `payload` field, is not a JSON object. |
| `400` | `invalid_multipart` | The multipart body cannot be read. |
| `400` | `invalid_idempotency_key` | `Idempotency-Key` is not 1–128 characters of letters, digits, `_`, `.`, `:`, or `-`. |
| `400` | `too_many_attachments` | The request carries more than 10 files. |
| `400` | `unsupported_attachment` | A file type is not supported. Audio and video files are rejected. |
| `404` | `chat_not_found` | `chat` names a chat this account cannot access. |
| `404` | `project_not_found` | `project` names a project that does not exist or that this account cannot access. |
| `409` | `agent_run_rejected` | The chat already has an active turn. |
| `413` | `attachment_too_large` | A file is over 100 MB, or the whole request is too large. |

`GET /api/developers/v1/runs/{run_id}`, `GET /api/developers/v1/runs/{run_id}/events`, and `POST /api/developers/v1/runs/{run_id}/cancel` return `400 invalid_run_id` for a malformed run id and `404 run_not_found` when this account has no run with that id.

A run that fails after it starts reports why in `last_error`. `budget_limited` means the run reached the runtime's limit of 1,500 model steps without finishing. Such a run is not billed.

## Retry boundaries

Retry transient infrastructure failures. Do not retry malformed requests, authentication failures, permission failures, or attachment-size errors without changing the request or credentials.

Honor `Retry-After` on `429` and `503`. Use the returned effective limits and reset time to tune concurrency.

The TypeScript SDK retries requests that are safe to repeat: reads, cancellations, file uploads, event-stream reconnects, and submissions that carry an idempotency key (except `research_shortform`). It retries on network failures and `408`, `429`, `500`, `502`, `503`, and `504`, honors a `Retry-After` of up to 60 seconds, and makes at most two retries by default. It never creates an idempotency key for you.

Do not blindly retry a run or tool submission without an `Idempotency-Key`. Use the same key and address so the server can return the original accepted work.

* Agent Run and `ask_stefanbrain` keys deduplicate for the rest of the same UTC day.
* Supported asynchronous submit tools deduplicate inside the resolved chat and family during the UTC-day window.
* `research_shortform` is not idempotent.

A keyed retry is still a new request. It counts against rate limits and must pass the same wallet, key-budget, or plan-usage check as new work, so a wallet without room for a new hold can refuse it with `429 api_wallet_exhausted`. The server also validates the request body before it looks up the key. When the key matches, the retry answers at once. An Agent Run retry returns the `202` handle while the run is active, or the run's final status after it ends. A tool submit retry returns the original job's `202` envelope with the same `job_id`.

Run and job polling does not consume standard request limits.

<Warning>
  Tool and job operations can return HTTP `200` with `is_error: true`. Check the envelope instead of treating every `200` as a successful outcome.
</Warning>

When `is_error` is true, `error.code` classifies the tool failure and `error.message` gives the readable reason. The original content blocks remain in `content`.

For run failures, read `output` and `last_error`. For structured-output runs, `structured_output` is `null` unless the run completed and its final text is a JSON object.

For job failures, inspect the top-level `status`, `is_error`, and `error`, then the family-specific fields in `data`.

See [Rate limits](/developers/rate-limits) and [Pricing](/developers/pricing) for the controls behind `429` responses.

## Stuck? Send feedback

When an error does not say how to fix the request, or the API cannot do what you need, report it with `POST /api/developers/v1/feedback`. MCP clients call `send_feedback`. Include the endpoint or tool, the error code, and the run or chat id.

Reports are free, and the StefanBrain team reads every report and replies. See [Agent feedback](/developers/feedback).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.