Skip to main content

How billing works

You pay per token. Three rates apply to each request: input for the tokens you send, cached for input tokens served from a prefix cache, and output for the tokens the model generates. Claude and OpenAI models add a fourth, cache write, for input tokens written into the prompt cache; see Cache writes. Every figure in the table below is in USD per 1M tokens. Caching is automatic: Valar matches shared prompt prefixes for you and charges the lower cached rate on the tokens that hit. Prompt Caching explains how to raise your hit rate, but nothing is required to get the cached rate. The rate you pay also depends on the completion window you request. Faster scheduling carries a higher rate: the on-demand Now window costs the most, Priority about 25% less, Standard about 50% less, and Flex the least. The table prices out the windows available for each model; coverage varies, and you can mix windows per request. Completion Windows explains the trade-offs.
Window coverage differs by model, and we keep adding models and widening window support. If a model or window you want isn’t shown, get in touch.
You can also read these rates in code: GET /v1/models returns a pricing array on every model, one entry per window. Those figures are resolved for your organization, so if you’re on negotiated rates the API quotes yours rather than the table below.

Cache writes

When a request writes part of its prompt into the prompt cache, those tokens are counted separately from input and billed at the model’s cache-write rate: The written tokens are reported apart from input tokens, for example as cache_creation_input_tokens on the Anthropic Messages API and cache_write_tokens in the Usage API. A later request that reuses the cached prefix pays the lower cached rate on those tokens. Organizations on negotiated rates may have a different cache-write rate.

Use pricing from an agent

Read Pricing for agents for a compact Markdown table of the same public rates. Your agent can retrieve it through the docs MCP server, or you can open the Markdown export directly. No API key is required. Both tables are generated from the same public model catalog and update when the docs are published. They show list prices, not organization-specific rates.

Rate table

USDper 1M tokens
ModelWindowInputCachedOutput
DeepSeek V4 Pro
deepseek-ai/DeepSeek-V4-Pro
Now1.0560.03523.168
Priority0.7920.02642.376
Standard0.5280.01761.584
Flex0.36960.012321.1088
Nemotron Ultra
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
Now0.50.12.2
Priority0.3750.0751.65
Standard0.250.051.1
Flex0.1750.0350.77
DeepSeek-V4.1-Flash
deepseek-ai/DeepSeek-V4.1-Flash
Now0.2550.02551.02
Priority0.191250.0191250.765
Standard0.12750.012750.51
Flex0.089250.0089250.357
Kimi K2.7 Code
moonshotai/Kimi-K2.7
Now0.750.153.5
Priority0.60.122.8
Standard0.3750.0751.75
Flex0.26250.05251.225
Kimi-K3
moonshotai/Kimi-K3
Now2.250.22511.25
Priority1.68750.168758.4375
Standard1.1250.11255.625
Flex0.78750.078753.9375
Kimi-K3 Fast
moonshotai/Kimi-K3-fast
Now3.3750.22516.875
Priority2.531250.1687512.65625
Standard1.68750.11258.4375
Flex1.181250.078755.90625
GLM-5.2
zai-org/GLM-5.2
Now0.980.1123.4
Priority0.7840.08962.72
Standard0.490.0911.54
Flex0.3430.0641.078
GLM-5.2 Fast
zai-org/GLM-5.2-fast
Now1.680.1685.28
Priority1.260.1263.96
Standard0.840.0842.64
Flex0.5880.05881.848
GLM-5.3
zai-org/GLM-5.3
Now1.120.2083.52
Priority0.840.1562.64
Standard0.560.1041.76
Flex0.3920.07281.232
GLM-5.3 Fast
zai-org/GLM-5.3-fast
Now1.7850.33155.61
Priority1.338750.2486254.2075
Standard0.89250.165752.805
Flex0.624750.1160251.9635
GLM-5.3 Flash
zai-org/GLM-5.3-Flash
Now0.120.02320.4
Priority0.090.01740.3
Standard0.060.01160.2
Flex0.0420.008120.14
gpt-oss-120b
openai/gpt-oss-120b
Now0.060.030.4
Priority0.0450.0230.3
Standard0.040.020.3
Flex0.0240.0120.155
Qwen3.5 27B
Qwen/Qwen3.5-27B
Now0.270.0542.2
Priority0.2030.0411.65
Standard0.1350.0271.1
Flex0.09450.01890.77
Qwen3.5-397B-A17B
Qwen/Qwen3.5-397B-A17B
Now0.450.091.35
Priority0.3380.0681.013
Standard0.250.050.75
Flex0.1750.0350.525
Qwen3.6 35B-A3B
Qwen/Qwen3.6-35B-A3B
Now0.2250.0450.9
Priority0.1690.03380.675
Standard0.130.0260.52
Flex0.09250.01850.37
Qwen3.8-Max
qwen/qwen3.8-max
Now20.256
Priority20.256
Standard10.1253
Flex0.70.08752.1
Gemma 4 31B IT
google/gemma-4-31B-it
Now0.140.060.4
Priority0.1050.0450.3
Standard0.070.030.2
Flex0.0490.0210.14
Gemma 4 26B A4B
google/gemma-4-26B-A4B-it
Now0.0720.01440.216
Priority0.0540.01080.162
Standard0.0360.00720.108
Flex0.02520.005040.0756
MiniMax M3
MiniMaxAI/MiniMax-M3
Now0.20.061.2
Priority0.150.0450.9
Standard0.10.030.6
Flex0.070.0210.42
For each model’s capabilities (e.g., image input and reasoning support), see Models.