Skip to main content
ModelSlugImageReasoning
DeepSeek V4 Pro
DeepSeek
Kimi-K3 Fast
Moonshot AI
moonshotai/Kimi-K3-fast
Nemotron Ultra
NVIDIA
GLM-5.3 Fast
Z.ai
zai-org/GLM-5.3-fast
DeepSeek-V4.1-Flash
DeepSeek
Qwen3.5 27B
Qwen
Qwen/Qwen3.5-27B
Kimi K2.7 Code
Moonshot AI
Kimi-K3
Moonshot AI
GLM-5.2
Z.ai
GLM-5.2 Fast
Z.ai
GLM-5.3
Z.ai
GLM-5.3 Flash
Z.ai
gpt-oss-120b
OpenAI
Qwen3.5-397B-A17B
Qwen
Qwen3.6 35B-A3B
Qwen
Qwen3.8-Max
Qwen
qwen/qwen3.8-max
Gemma 4 31B IT
Google
Gemma 4 26B A4B
Google
MiniMax M3
MiniMax

Reasoning models and output caps

max_output_tokens caps reasoning tokens plus the visible answer, not the answer alone. On a model marked Reasoning above, most of that budget goes to thinking you never see — on qwen/qwen3.7-plus, around 96% of it, roughly a thousand reasoning tokens behind a two-sentence answer. So a cap sized for the answer stops generation mid-thought, and the response settles incomplete:
That status is accurate - generation really did stop at the cap - but it does not mean the answer is missing. In a 500-request run at max_output_tokens: 1000, two thirds of the responses came back incomplete, and their answers matched the completed third in length and quality. The only difference was how long the model thought.
Gating on status alone throws those answers away:
Branch on incomplete_details.reason and read the output instead:

Size the cap for reasoning plus answer

Send a handful of representative requests, read usage.output_tokens_details.reasoning_tokens off the results, and set the cap to that plus the answer length you want. Omit max_output_tokens entirely when you don’t need a hard ceiling.
Don’t set a cap below the model’s typical reasoning volume. A cap of a few tokens leaves no room to finish thinking and start answering, and the request fails outright with a 503 rather than returning an empty incomplete.

Custom models

Beyond the catalog above, Valar can serve your own model weights. If you have a custom or fine-tuned open-weight model, we can host it on Valar’s inference stack and expose it through the same OpenAI-compatible API, completion windows, and billing as any catalog model. Reach out to your Valar contact to onboard a custom model. For adapter-based customization on top of a supported base model, register a PEFT-trained LoRA adapter and run it against an eligible base model such as zai-org/GLM-5.2.
To confirm what your API key can reach in a given environment at runtime, call GET /v1/models against that environment rather than relying on this table alone. Each model it returns carries its context window, capability flags, and the rate card your organization is billed at.