Skip to main content
Agent Runs give your application the full StefanBrain agent. The agent plans the work, selects tools, executes them, and returns the finished response.

Start a run

The request accepts these fields: A run with project addresses a chat inside that project. Project instructions and summaries are not added to the run automatically. Use the list_projects tool to find project identifiers.

Understand the response

Starting a run normally returns 202 with the run handle: run_id, chat, project (when the chat is in a project), status: "queued", status_url, events_url, and cancel_url. A synchronous run that ends inside the wait window returns 200 with the run’s status body instead: run_id, chat, status, output, structured_output, last_error, cancel_requested, and timestamps. The 200 body has no URLs; the lifecycle paths are:
Poll the status until it is completed, failed, or cancelled.
  • output holds the latest assistant text. It is the final answer only once status is completed.
  • structured_output contains parsed JSON when the run requested a structured response.
  • last_error provides failure detail when the run fails.
Runs execute in the background. Closing the HTTP connection does not cancel the work; you can reconnect and poll again. POST .../cancel returns the run’s status with cancel_requested: true. A running run stops and reports cancelled shortly after. An admitted run that has not started can also be cancelled. Cancellation does not count against rate limits.

Control busy-chat behavior

A chat executes one run at a time. Only on_busy: "reject" (the default) is supported. A busy start returns 409 with the blocking run id when known. Poll the blocking run and retry after it finishes. Use a separate chat for independent parallel work.

Make a retry idempotent

Add an Idempotency-Key header when you may need to retry a start without knowing whether the first request succeeded:
A repeated key returns the original run: the 202 envelope while it is queued or running, or its final status (200) once it ends. It does not create another chat, stage attachments again, or start another billed run. The retry’s body is not compared with the original. An invalid key returns 400 invalid_idempotency_key. Send a key with every run start. A timeout, dropped connection, or gateway 504 does not stop a run, so a lost response is otherwise a second billed run on retry. Run keys are scoped to your account and remain effective for the rest of the same UTC day. Valid keys contain 1–128 URL-safe characters: letters, numbers, _, ., :, or -.

Wait synchronously or poll

Add "sync": true for short, interactive tasks. StefanBrain holds the connection for up to about 90 seconds, counted from when it receives the request, and answers as soon as the run ends. If the run takes longer, the request returns the ordinary 202 response and continues in the background. Set your HTTP client’s timeout to 120 seconds so it receives that response. Poll the returned status URL every 3–10 seconds until the run is terminal. Polling does not count against rate limits. Use synchronous mode for short interactive requests. Use polling for long work, batch pipelines, or environments where holding a connection is fragile.

Follow progress events

The events endpoint returns turn phases, tool activity, and text sections. Page through events with the cursor returned as next_after:
Each page returns up to 200 events. Pass next_after as after. When a page is empty, next_after is null; keep your previous cursor. is_done is true on the page that contains the terminal event. The same endpoint supports server-sent events. Send this header:
Each SSE id is an event identifier. Heartbeats arrive approximately every 15 seconds, and the stream closes after the terminal event. Reconnect with Last-Event-ID or ?after= to resume. Cursor polling and SSE contain the same event data; choose one transport for each client.

Attach files

Send multipart/form-data when StefanBrain must read files with the message.
  • Put the JSON request body in a payload field.
  • Add every upload as a files field.
  • A run accepts up to 10 files.
  • Each file can be up to 100 MB.
  • Supported inputs include images, PDFs, common Office documents, spreadsheets, and text files.
  • Audio and video files are rejected. Host the media and put its public URL in message instead.
Use uploaded files parts for run attachments. Internal attachment identifiers and referenced document identifiers are not part of this public run request.

Request structured output

Add output_config.format to constrain the final answer to a JSON schema. The raw text remains in output; parsed JSON returns in structured_output. Structured output uses strict mode:
  1. The root schema must use "type": "object".
  2. Every object must set "additionalProperties": false.
  3. Every object must list all property names in required. To make a value optional, include null in its type.
An invalid strict-mode schema returns 400 invalid_output_config and identifies the failing path. structured_output is null when the run is unfinished or its final text is not valid JSON, including refusals and hard failures.

Model behavior

There is no model selector. If a request includes a model field, StefanBrain ignores it. Agent Runs execute the model that currently operates the StefanBrain product. See the generated REST reference for exact request and response schemas. See Errors and retries before adding retry logic.