| Model | Slug | Image | Reasoning |
|---|---|---|---|
moonshotai/Kimi-K3-fast | |||
zai-org/GLM-5.3-fast | |||
Qwen/Qwen3.5-27B | |||
qwen/qwen3.8-max | |||
Reasoning models and output caps
max_output_tokens caps reasoning tokens plus the visible answer, not the answer alone. On a model marked Reasoning above, most of that budget goes to thinking you never see — on qwen/qwen3.7-plus, around 96% of it, roughly a thousand reasoning tokens behind a two-sentence answer.
So a cap sized for the answer stops generation mid-thought, and the response settles incomplete:
max_output_tokens: 1000, two thirds of the responses came back incomplete, and their answers matched the completed third in length and quality. The only difference was how long the model thought.
Branch on incomplete_details.reason and read the output instead:
Size the cap for reasoning plus answer
Send a handful of representative requests, readusage.output_tokens_details.reasoning_tokens off the results, and set the cap to that plus the answer length you want. Omit max_output_tokens entirely when you don’t need a hard ceiling.
Custom models
Beyond the catalog above, Valar can serve your own model weights. If you have a custom or fine-tuned open-weight model, we can host it on Valar’s inference stack and expose it through the same OpenAI-compatible API, completion windows, and billing as any catalog model. Reach out to your Valar contact to onboard a custom model. For adapter-based customization on top of a supported base model, register a PEFT-trained LoRA adapter and run it against an eligible base model such aszai-org/GLM-5.2.
To confirm what your API key can reach in a given environment at runtime, call
GET /v1/models against that environment rather than relying on this table alone. Each model it returns carries its context window, capability flags, and the rate card your organization is billed at.