Skip to main content
Valar optimizes for throughput and cost on long-running agent work, not the latency of a single call. You pick an execution mode by how soon you need each result, then tune the cost-versus-latency trade-off with a completion window.

Completion windows at a glance

A completion window sets how much wall-clock time per turn you’re willing to trade for a lower rate. Valar runs four. Choose one with metadata.completion_window (or the X-Valar-Completion-Window header), or leave it off and get standard. Each is detailed in Completion windows below.

The three modes

Realtime

A normal synchronous request that returns the result immediately. You send the call without background and read the output from the response. This is the lowest-latency path, finishing in seconds, and it runs on the Now completion window. Realtime works across the Responses API (/v1/responses), Chat Completions (/v1/chat/completions). Use it for interactive chat, prototyping, and human-in-the-loop steps.

Tell Valar your client timeout

If your client gives up on a realtime call after a fixed time, tell Valar what that time is with the X-Valar-Client-Timeout header. It is read on every realtime endpoint: /v1/responses, /v1/chat/completions and /v1/messages. It changes what happens when a request produces nothing. Without it, Valar works to its own limits, which are longer than most clients wait: your client can time out while the request is still running, and all you see is a connection that stopped. With it, the request ends at your number and tells you why. The value is your timeout in seconds ("120"), or a duration ("2m", "90s"). It must be at least one second. Values longer than Valar’s own realtime limit are treated as that limit. Without the header Valar plans against its own limit, so set it to the timeout your SDK is actually configured with:
The budget covers the time to the first output. A stream that has started producing keeps running past it — it is working, and cutting it would throw away the answer you are already receiving. For a non-streaming call the first output is the whole answer, so there the budget covers the call. If the budget runs out with nothing produced, how the error reaches you depends on whether the endpoint has already committed to a response: Either way the request is recorded as failed, and it ends when you said you would stop waiting rather than later. The header does not change your prices, your rate limits, or which models you can call, and it is not a guarantee that an answer arrives inside the time you give — a model that needs longer than your budget will simply be reported as a timeout sooner. The header applies to realtime calls only. A request on one of the async windows is not waiting on a live connection, so it carries no client timeout.

Async

Set background=True on the Responses API. The create call returns a response id immediately, then you poll the retrieve endpoint or receive a webhook when the work finishes. On the Standard window a turn takes at most a minute and is usually much faster, so async clears high throughput at lower cost. Async jobs usually run on the Standard window (standard), the lower-cost default. Use async for agent loops, background jobs, and large fan-out. See Sending requests at scale.

Batch (Private Preview)

Batch lets you submit many requests at once as an uploaded JSONL file and retrieve a results file when the set completes - the lowest cost and highest throughput, with the longest turnaround (results land within the batch window, up to 12 hours). It follows the OpenAI Batch API, so existing client.batches.create() code works unchanged. For the end-to-end workflow, see Sending requests at scale.

Compare the modes

Completion windows

A completion window tells Valar how much wall-clock time per turn your workload can tolerate, and you pay less the more time you give it. There are four, from fastest to cheapest: Realtime uses the Now window; async work uses Priority for a firm ~10 s deadline, Standard for lower cost, or Flex for the cheapest background runs. Each model-and-window price pairing is on the Pricing page.

Set the window

Pass metadata.completion_window on the request:
Alternatively, set the X-Valar-Completion-Window header to the same value. This is useful when a client owns the request body on your behalf (for example, the Claude Agent SDK) and body metadata isn’t available. The body field takes precedence when both are set. Accepted values are "asap" (Now), "priority", "standard", and "flex".

How each tier behaves

Now runs immediately on the fastest available hardware in a latency-optimized setup, at the higher on-demand rate. Use it for realtime, interactive requests where a person or another system is waiting on the result. Priority is a durable async window that targets a tight ~10-second turn-time ceiling - a firm, low-latency completion target for agent loops - priced about 25% below the Now rate. Like the other async windows, request it explicitly with background=True. Standard is the default wherever a model supports it. It runs on Valar’s maximum-efficiency serving stack with a turn-time ceiling of one minute - and in practice most turns complete much faster than that. Most of Valar’s published prices reference this window, and it’s the right default for async agent loops and batch jobs. Flex is the lowest-cost tier, for background bulk work that can tolerate a little extra latency. It targets a turn-time ceiling of about five minutes and requires background=True - it isn’t available for realtime or synchronous calls. Flex is always explicit: Valar never selects it for you, so request it with metadata.completion_window: "flex" when you want it.

Default behavior

Leave completion_window off and the request defaults to standard when the model supports it; otherwise it falls back to the Now tier. Priority and Flex are never selected automatically - set either explicitly when you want it.
Explicitly choosing a window the model does not support fails with 400 invalid_request_error. The error message lists the model’s supported windows, and the Pricing page keeps a current support matrix.

Next steps

Quickstart

Send your first request with an OpenAI-compatible client.

Sending requests at scale

Fan out async and batch work across many requests.