Skip to main content
A loop that fires requests in sequence stalls the moment your workload grows past a few hundred calls. Valar gives you two ways to clear large volumes: fan out many concurrent Responses API calls in background mode, or hand a single bundle to the Batch API and let Valar work through it. This page covers both, starting with how to choose.

Which path fits your workload

Concurrent Responses API

You want results streaming back as each request finishes, need per-request control, or care about latency on individual items. You manage concurrency on your client.

Batch API

You have a fixed set of requests and would rather submit once and collect results later. Valar owns the queue; you poll for status.
The running example throughout is sentiment scoring for a backlog of product reviews, using zai-org/GLM-5.2.
Both paths below set max_output_tokens. On a reasoning model that cap covers hidden reasoning as well as the answer, so a large share of results settle incomplete with a perfectly good answer attached. Collect on incomplete_details.reason, not on status alone.

Fan out with the Responses API

Submit each request individually with background=True and run many in flight at once. For workloads of roughly 1,000 requests and up, four practices keep it stable:
  • Use AsyncOpenAI with DefaultAioHttpClient(). Under high concurrency the SDK’s aiohttp backend outperforms the default httpx transport.
  • Bound concurrency with an asyncio.Semaphore. A fixed limit (200 is a sound starting point) caps simultaneous connections so you never exhaust them.
  • Submit all requests behind the semaphore, collect the response IDs, then poll in a second pass. Separating submission from polling keeps both phases simple.
  • Attach an Idempotency-Key per request so a retry after a transient failure replays the reservation rather than paying for inference twice. See Idempotent Requests.
Install the dependencies first:
The script submits every review concurrently, gathers the response IDs, and then polls each one to completion before printing its score:
Treat a semaphore of 200 as the baseline. Lower it if connection errors or timeouts appear; raise it when you have headroom and want submissions to move through faster.

Hand off a bundle with the Batch API

The Batch API takes up to 10,000 requests per batch. You upload the requests as a JSONL file, create a batch that points at it, poll until it settles, and download the results file. It follows the OpenAI Batch API, so client.batches.create() code you already have works unchanged against Valar’s base URL.
1

Upload

Write one request per line to a .jsonl file and upload it with POST /files (purpose: batch). Each line carries your custom_id, "method": "POST", the target url, and the request body. A line can be up to 28 MiB - enough for a maximum-size 20 MB image, which takes about 26.7 MiB once base64-encoded.That cap is about transport, not the model. Each request also has to fit the context window of the model it names, and windows vary by more than an order of magnitude across the catalog - so a line well under 28 MiB can still be too long for the model you picked, which is easiest to hit when you pack base64 images. Read context_length from GET /models and size your requests against it: a request over the window is accepted with the file and fails when it runs, on that line alone.
2

Submit

POST /batches with the returned input_file_id, the endpoint the requests target, and completion_window: "24h" (the OpenAI-compatible value; Valar schedules against its own promised SLAs).
3

Poll

Call GET /batches/{batch_id} until status is terminal. Progress can arrive in large steps, so a quiet request_counts is normal - every batch settles within its window.
4

Retrieve

Download output_file_id with GET /files/{file_id}/content - one JSONL result line per request. To peek at a single result early, GET /batches/{batch_id}/{custom_id} works while the batch is still running (a Valar extension).
To review everything you’ve submitted previously, call GET /batches.
A line that fails validation - unparseable, missing a model, or over the 28 MiB line limit - is rejected on its own and counted in request_counts.failed; the rest of the batch keeps running. Batch is the lowest-cost, highest-latency lane. When latency matters, fan out through the Responses API instead. See Inference modes for the tradeoff.
This version scores the same review backlog as one batch: