Which path fits your workload
Concurrent Responses API
You want results streaming back as each request finishes, need per-request control, or care about latency on individual items. You manage concurrency on your client.
Batch API
You have a fixed set of requests and would rather submit once and collect results later. Valar owns the queue; you poll for status.
zai-org/GLM-5.2.
Both paths below set
max_output_tokens. On a reasoning model that cap covers hidden reasoning as well as the answer, so a large share of results settle incomplete with a perfectly good answer attached. Collect on incomplete_details.reason, not on status alone.Fan out with the Responses API
Submit each request individually withbackground=True and run many in flight at once. For workloads of roughly 1,000 requests and up, four practices keep it stable:
- Use
AsyncOpenAIwithDefaultAioHttpClient(). Under high concurrency the SDK’s aiohttp backend outperforms the defaulthttpxtransport. - Bound concurrency with an
asyncio.Semaphore. A fixed limit (200 is a sound starting point) caps simultaneous connections so you never exhaust them. - Submit all requests behind the semaphore, collect the response IDs, then poll in a second pass. Separating submission from polling keeps both phases simple.
- Attach an
Idempotency-Keyper request so a retry after a transient failure replays the reservation rather than paying for inference twice. See Idempotent Requests.
Hand off a bundle with the Batch API
The Batch API takes up to 10,000 requests per batch. You upload the requests as a JSONL file, create a batch that points at it, poll until it settles, and download the results file. It follows the OpenAI Batch API, soclient.batches.create() code you already have works unchanged against Valar’s base URL.
1
Upload
Write one request per line to a
.jsonl file and upload it with POST /files (purpose: batch). Each line carries your custom_id, "method": "POST", the target url, and the request body. A line can be up to 28 MiB - enough for a maximum-size 20 MB image, which takes about 26.7 MiB once base64-encoded.That cap is about transport, not the model. Each request also has to fit the context window of the model it names, and windows vary by more than an order of magnitude across the catalog - so a line well under 28 MiB can still be too long for the model you picked, which is easiest to hit when you pack base64 images. Read context_length from GET /models and size your requests against it: a request over the window is accepted with the file and fails when it runs, on that line alone.2
Submit
POST /batches with the returned input_file_id, the endpoint the requests target, and completion_window: "24h" (the OpenAI-compatible value; Valar schedules against its own promised SLAs).3
Poll
Call
GET /batches/{batch_id} until status is terminal. Progress can arrive in large steps, so a quiet request_counts is normal - every batch settles within its window.4
Retrieve
Download
output_file_id with GET /files/{file_id}/content - one JSONL result line per request. To peek at a single result early, GET /batches/{batch_id}/{custom_id} works while the batch is still running (a Valar extension).GET /batches.
This version scores the same review backlog as one batch: