Skip to main content
Request-count limits vary by plan.
  • When the active top-up balance is $100 or more, the effective limits move down one row in this table automatically.
  • The limits object also reports global_tokens_per_day and per_user_tokens_per_day, a 10-million-token daily guardrail by default. Both fields are informational; no Developer API surface enforces a daily token limit.
  • Per-user overrides are available for production integrations. Contact StefanBrain if you need more throughput.
Spend and throughput are separate limits:
  • The API wallet limits wallet-billed spend. A request returns 429 api_wallet_exhausted when the balance minus active holds is below one hold, so the wallet can refuse new work before $0. See Pricing.
  • A per-key budget returns 429 api_key_budget_exhausted when exhausted.
  • For standard member accounts, the plan’s monthly usage pool controls MCP usage and returns 429 monthly_usage_limit_reached when exhausted. Contracted partner accounts can have account-specific billing terms.

Rate-limit response

Each request-limit response includes the user’s effective limits, observed counts, and reset time. When reset_at is set, the response also includes a Retry-After header in seconds:
The code names the limit you reached: When a throttle or suspension has no end time, reset_at is null and the response has no Retry-After header. For request limits, the observed counts include the rejected call, so retries sent while you are limited still count. Retry after the returned delay. Run status, run events, job status, job results, and cancellation do not consume request limits. Other reads do not either: the tool catalog, file listings and downloads, job artifacts, transcript reads, and feedback reads over REST. Calls that start work, file uploads, transcript requests, and feedback reports count. Feedback reports also have their own limit of 60 per account in any rolling hour, which returns 429 feedback_rate_limited. See Agent feedback. See Errors and retries for retry boundaries.