> ## Documentation Index
> Fetch the complete documentation index at: https://docs.valarhq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Pricing for agents

> Public model prices in a compact Markdown table for agents and exports

Use this table to read or compare Valar's public list prices without parsing the visual pricing page. All amounts are **USD per 1M tokens**. Model IDs are case-sensitive.

Retired IDs keep working as aliases of their replacement and are billed at the replacement's rates: requests for `deepseek-ai/DeepSeek-V4-Flash` are served as `deepseek-ai/DeepSeek-V4.1-Flash`.

This table contains the same models and rates as [Pricing](/pricing), covering **Now**, **Priority**, **Standard**, and **Flex**. Hidden and restricted models are excluded. Both tables are generated from the same public catalog and update when the docs are published; this is not a live quote or a feed of organization-specific rates.

## Read or export

No API key is required. [Open this page as Markdown](https://docs.valarhq.ai/pricing-for-agents.md), or connect to the [docs MCP server](/docs-mcp) and ask:

```text theme={"system"}
Read the full Pricing for agents page from the Valar docs MCP server.
Return each model ID and its Standard input, cached input, and output rates
as JSON. Keep the USD-per-1M-token units and decimal precision unchanged.
```

The Markdown table is the published export. Any JSON or CSV you ask an agent to produce is a conversion of that table, not a separate API response.

## Public rates

Each row identifies a model and a completion window. **Input** is uncached input; **cached input** is input served from the prompt cache; **output** includes generated reasoning tokens. See [Completion windows](/inference-modes#completion-windows) for scheduling behavior.

Claude and OpenAI models also bill input tokens written into the prompt cache at a cache-write rate: 1.25x the input rate, or 2x for a Claude write cached with the 1-hour TTL. Other models have no cache-write charge. See [Cache writes](/pricing#cache-writes).

| Model ID | Window | Input | Cached input | Output |
| - | - | -: | -: | -: |
| `deepseek-ai/DeepSeek-V4-Pro` | Now | 1.056 | 0.0352 | 3.168 |
| `deepseek-ai/DeepSeek-V4-Pro` | Priority | 0.792 | 0.0264 | 2.376 |
| `deepseek-ai/DeepSeek-V4-Pro` | Standard | 0.528 | 0.0176 | 1.584 |
| `deepseek-ai/DeepSeek-V4-Pro` | Flex | 0.3696 | 0.01232 | 1.1088 |
| `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` | Now | 0.5 | 0.1 | 2.2 |
| `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` | Priority | 0.375 | 0.075 | 1.65 |
| `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` | Standard | 0.25 | 0.05 | 1.1 |
| `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` | Flex | 0.175 | 0.035 | 0.77 |
| `deepseek-ai/DeepSeek-V4.1-Flash` | Now | 0.255 | 0.0255 | 1.02 |
| `deepseek-ai/DeepSeek-V4.1-Flash` | Priority | 0.19125 | 0.019125 | 0.765 |
| `deepseek-ai/DeepSeek-V4.1-Flash` | Standard | 0.1275 | 0.01275 | 0.51 |
| `deepseek-ai/DeepSeek-V4.1-Flash` | Flex | 0.08925 | 0.008925 | 0.357 |
| `moonshotai/Kimi-K2.7` | Now | 0.75 | 0.15 | 3.5 |
| `moonshotai/Kimi-K2.7` | Priority | 0.6 | 0.12 | 2.8 |
| `moonshotai/Kimi-K2.7` | Standard | 0.375 | 0.075 | 1.75 |
| `moonshotai/Kimi-K2.7` | Flex | 0.2625 | 0.0525 | 1.225 |
| `moonshotai/Kimi-K3` | Now | 2.25 | 0.225 | 11.25 |
| `moonshotai/Kimi-K3` | Priority | 1.6875 | 0.16875 | 8.4375 |
| `moonshotai/Kimi-K3` | Standard | 1.125 | 0.1125 | 5.625 |
| `moonshotai/Kimi-K3` | Flex | 0.7875 | 0.07875 | 3.9375 |
| `moonshotai/Kimi-K3-fast` | Now | 3.375 | 0.225 | 16.875 |
| `moonshotai/Kimi-K3-fast` | Priority | 2.53125 | 0.16875 | 12.65625 |
| `moonshotai/Kimi-K3-fast` | Standard | 1.6875 | 0.1125 | 8.4375 |
| `moonshotai/Kimi-K3-fast` | Flex | 1.18125 | 0.07875 | 5.90625 |
| `zai-org/GLM-5.2` | Now | 0.98 | 0.112 | 3.4 |
| `zai-org/GLM-5.2` | Priority | 0.784 | 0.0896 | 2.72 |
| `zai-org/GLM-5.2` | Standard | 0.49 | 0.091 | 1.54 |
| `zai-org/GLM-5.2` | Flex | 0.343 | 0.064 | 1.078 |
| `zai-org/GLM-5.2-fast` | Now | 1.68 | 0.168 | 5.28 |
| `zai-org/GLM-5.2-fast` | Priority | 1.26 | 0.126 | 3.96 |
| `zai-org/GLM-5.2-fast` | Standard | 0.84 | 0.084 | 2.64 |
| `zai-org/GLM-5.2-fast` | Flex | 0.588 | 0.0588 | 1.848 |
| `zai-org/GLM-5.3` | Now | 1.12 | 0.208 | 3.52 |
| `zai-org/GLM-5.3` | Priority | 0.84 | 0.156 | 2.64 |
| `zai-org/GLM-5.3` | Standard | 0.56 | 0.104 | 1.76 |
| `zai-org/GLM-5.3` | Flex | 0.392 | 0.0728 | 1.232 |
| `zai-org/GLM-5.3-fast` | Now | 1.785 | 0.3315 | 5.61 |
| `zai-org/GLM-5.3-fast` | Priority | 1.33875 | 0.248625 | 4.2075 |
| `zai-org/GLM-5.3-fast` | Standard | 0.8925 | 0.16575 | 2.805 |
| `zai-org/GLM-5.3-fast` | Flex | 0.62475 | 0.116025 | 1.9635 |
| `zai-org/GLM-5.3-Flash` | Now | 0.12 | 0.0232 | 0.4 |
| `zai-org/GLM-5.3-Flash` | Priority | 0.09 | 0.0174 | 0.3 |
| `zai-org/GLM-5.3-Flash` | Standard | 0.06 | 0.0116 | 0.2 |
| `zai-org/GLM-5.3-Flash` | Flex | 0.042 | 0.00812 | 0.14 |
| `openai/gpt-oss-120b` | Now | 0.06 | 0.03 | 0.4 |
| `openai/gpt-oss-120b` | Priority | 0.045 | 0.023 | 0.3 |
| `openai/gpt-oss-120b` | Standard | 0.04 | 0.02 | 0.3 |
| `openai/gpt-oss-120b` | Flex | 0.024 | 0.012 | 0.155 |
| `Qwen/Qwen3.5-27B` | Now | 0.27 | 0.054 | 2.2 |
| `Qwen/Qwen3.5-27B` | Priority | 0.203 | 0.041 | 1.65 |
| `Qwen/Qwen3.5-27B` | Standard | 0.135 | 0.027 | 1.1 |
| `Qwen/Qwen3.5-27B` | Flex | 0.0945 | 0.0189 | 0.77 |
| `Qwen/Qwen3.5-397B-A17B` | Now | 0.45 | 0.09 | 1.35 |
| `Qwen/Qwen3.5-397B-A17B` | Priority | 0.338 | 0.068 | 1.013 |
| `Qwen/Qwen3.5-397B-A17B` | Standard | 0.25 | 0.05 | 0.75 |
| `Qwen/Qwen3.5-397B-A17B` | Flex | 0.175 | 0.035 | 0.525 |
| `Qwen/Qwen3.6-35B-A3B` | Now | 0.225 | 0.045 | 0.9 |
| `Qwen/Qwen3.6-35B-A3B` | Priority | 0.169 | 0.0338 | 0.675 |
| `Qwen/Qwen3.6-35B-A3B` | Standard | 0.13 | 0.026 | 0.52 |
| `Qwen/Qwen3.6-35B-A3B` | Flex | 0.0925 | 0.0185 | 0.37 |
| `qwen/qwen3.8-max` | Now | 2 | 0.25 | 6 |
| `qwen/qwen3.8-max` | Priority | 2 | 0.25 | 6 |
| `qwen/qwen3.8-max` | Standard | 1 | 0.125 | 3 |
| `qwen/qwen3.8-max` | Flex | 0.7 | 0.0875 | 2.1 |
| `google/gemma-4-31B-it` | Now | 0.14 | 0.06 | 0.4 |
| `google/gemma-4-31B-it` | Priority | 0.105 | 0.045 | 0.3 |
| `google/gemma-4-31B-it` | Standard | 0.07 | 0.03 | 0.2 |
| `google/gemma-4-31B-it` | Flex | 0.049 | 0.021 | 0.14 |
| `google/gemma-4-26B-A4B-it` | Now | 0.072 | 0.0144 | 0.216 |
| `google/gemma-4-26B-A4B-it` | Priority | 0.054 | 0.0108 | 0.162 |
| `google/gemma-4-26B-A4B-it` | Standard | 0.036 | 0.0072 | 0.108 |
| `google/gemma-4-26B-A4B-it` | Flex | 0.0252 | 0.00504 | 0.0756 |
| `MiniMaxAI/MiniMax-M3` | Now | 0.2 | 0.06 | 1.2 |
| `MiniMaxAI/MiniMax-M3` | Priority | 0.15 | 0.045 | 0.9 |
| `MiniMaxAI/MiniMax-M3` | Standard | 0.1 | 0.03 | 0.6 |
| `MiniMaxAI/MiniMax-M3` | Flex | 0.07 | 0.021 | 0.42 |

For your organization's available models and negotiated rates, use the separate authenticated [`GET /v1/models` API](/api-reference/models-api/list-supported-models). The public docs MCP does not call that API.
