# Building a Tool-Calling Agent Source: https://docs.valarhq.ai/agents Wire external tools into a multi-turn Valar conversation and let the model orchestrate them An agent is a model that can reach outside the conversation: it asks to run a function you control, reads what comes back, and decides what to do next. The Responses API at `/v1/responses` exposes exactly the primitives you need for this - tool definitions on the way in, `function_call` items on the way out, and `function_call_output` items to feed results back in. This guide builds a small agent that books meeting rooms, walks through what each field does, and closes with the operational settings that matter once it's running for real. ## What you provide and what you get back You hand the model a list of `tools`. Each tool is a JSON Schema description of a function it may call. When the model decides a tool is needed, its `response.output` contains one or more `function_call` items instead of (or alongside) plain text. You execute those calls and report back. The key property to internalize: **the API is stateless across requests, so you carry the conversation yourself.** Every item in `response.output` is already a valid input item. You append those items to your running list verbatim - no reshaping - then append your tool results as `function_call_output` items and send the whole thing again. Because you replay the full history on each request, the model always sees its own earlier tool calls and their outputs. There is nothing to serialize or translate; output items go back in as-is. ## The conversation cycle Append the user message to your conversation list and `POST` it to `/v1/responses` along with the `tools` array. Scan `response.output` for items of type `function_call`. If there are none, the turn is finished and `output_text` holds the answer. Parse the `arguments` JSON on each call, execute the matching function on your side, and capture its return value. Push the model's `response.output` items onto your list, then append one `function_call_output` per call (matched by `call_id`). Send the updated conversation back and return to step 2. The loop ends on the first response that carries text with no `function_call` items. ## Example: a room-booking assistant The agent below answers scheduling questions with two tools: one that checks a room's availability and one that reserves it. On the first user turn it looks up availability; once you return the result it reserves the room and writes a confirmation. A second user turn then reuses everything already in context to book a follow-up. ```python theme={"system"} import json import time from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key="YOUR_VALAR_API_KEY", # or read OPENAI_API_KEY from the environment ) MODEL = "moonshotai/Kimi-K2.7" TOOLS = [ { "type": "function", "name": "check_room", "description": "Check whether a meeting room is free for a time slot.", "parameters": { "type": "object", "properties": { "room": {"type": "string", "description": "Room name, e.g. Birch"}, "slot": {"type": "string", "description": "Time slot, e.g. 2026-06-08 14:00"}, }, "required": ["room", "slot"], "additionalProperties": False, }, "strict": True, }, { "type": "function", "name": "reserve_room", "description": "Reserve a meeting room for a time slot.", "parameters": { "type": "object", "properties": { "room": {"type": "string", "description": "Room name, e.g. Birch"}, "slot": {"type": "string", "description": "Time slot, e.g. 2026-06-08 14:00"}, }, "required": ["room", "slot"], "additionalProperties": False, }, "strict": True, }, ] TOOL_DISPATCH = { "check_room": lambda room, slot: json.dumps({"room": room, "slot": slot, "free": True}), "reserve_room": lambda room, slot: json.dumps({"room": room, "slot": slot, "confirmation": "RSV-4417"}), } def await_completion(response, timeout=300): """Poll a background response until it settles.""" deadline = time.time() + timeout while response.status in {"queued", "in_progress"}: if time.time() > deadline: raise TimeoutError(f"{response.id} did not complete within {timeout}s") time.sleep(2) response = client.responses.retrieve(response.id) if response.status != "completed": raise RuntimeError(f"{response.id} status: {response.status}") return response def run_turn(conversation, user_message): """Append a user message and drive the tool loop to a final text answer.""" conversation.append({"role": "user", "content": user_message}) while True: response = client.responses.create( model=MODEL, input=conversation, tools=TOOLS, max_output_tokens=4096, background=True, ) response = await_completion(response) # Output items are valid input items: append them unchanged. conversation.extend(response.output) tool_calls = [ item for item in (response.output or []) if getattr(item, "type", None) == "function_call" ] if not tool_calls: return response for call in tool_calls: args = json.loads(call.arguments) result = TOOL_DISPATCH[call.name](**args) conversation.append( {"type": "function_call_output", "call_id": call.call_id, "output": result} ) conversation = [] # First turn: the model checks availability, then reserves and confirms. first = run_turn(conversation, "Is the Birch room free at 2pm on June 8? If so, book it.") print("Turn 1:", first.output_text) # Second turn: the model reuses the booking already in context. second = run_turn(conversation, "Great - also grab the Cedar room right after, at 3pm the same day.") print("Turn 2:", second.output_text) ``` On turn one the model emits a `check_room` call. You return `free: true`, resend the conversation, and the model follows up with a `reserve_room` call. After you return the confirmation number, the next response is plain text so the loop exits. On turn two you resend the entire history, including turn one's calls and their outputs. The model already knows June 8 from context, so it only needs to call `check_room` and `reserve_room` for Cedar before confirming. ## Running it in production Set `strict: true` on each tool's parameters. This enables structured-output guarantees, so the `arguments` JSON the model returns always conforms to your schema and `json.loads` never trips over a malformed payload. * **Prefer** `background=True `**for anything long-running.** Valar optimizes for throughput, so an individual request can run longer than a latency-first API would. Background mode avoids HTTP timeouts and lets you poll for completion instead. * **Match the completion window to your loop.** The default `standard` window balances cost against trajectory time for most agents. Switch to the **Now** (`asap`) window when a single turn is latency-sensitive and a person or system is waiting on it. See [Completion windows](/inference-modes#completion-windows) for response-time and pricing details. * **Replay the complete conversation every request.** Every prior message, all `response.output` items, and every `function_call_output` must be present. Output items drop back in without conversion. * **Handle parallel calls.** A single response can contain several `function_call` items at once; execute them all and return one `function_call_output` per `call_id`. # Anthropic SDK & Claude Agent SDK Source: https://docs.valarhq.ai/anthropic-sdk Use the Anthropic Python SDK and the Claude Agent SDK against Valar's /v1/messages surface, and pick a completion window with a header Valar's `POST /v1/messages` endpoint speaks the Anthropic Messages API, so the official Anthropic Python SDK and the Claude Agent SDK work against Valar with a base URL and key change. The same [models](/models) and [completion windows](/inference-modes#completion-windows) work as on the Responses and Chat Completions surfaces; see the [API support matrix](/support) for the full feature list. ## Anthropic Python SDK ```bash theme={"system"} pip install anthropic ``` Point the SDK at Valar. Use `auth_token` (sends `Authorization: Bearer`) — or `api_key` (sends `x-api-key`); Valar accepts both. ```python theme={"system"} from anthropic import Anthropic client = Anthropic( auth_token="YOUR_VALAR_API_KEY", base_url="https://api.valarhq.ai", # no /v1 — the SDK appends it ) message = client.messages.create( model="zai-org/GLM-5.2", max_tokens=2048, messages=[{"role": "user", "content": "Explain transformers in one sentence."}], ) # Models may emit thinking blocks before the text response — extract the text block. text = next(block.text for block in message.content if block.type == "text") print(text) print(message.stop_reason, message.usage.input_tokens, message.usage.output_tokens) ``` The full Anthropic surface is supported — `system` prompts, `tools`, `thinking`, `stop_sequences`, `top_k`, `stream: true`, and [structured outputs](/structured-outputs) via `output_config.format`. System prompts go in the top-level `system` field, not as a `role: "system"` message. Reasoning-capable models emit `thinking` blocks by default; control the depth with `thinking` or `output_config.effort` (`low` / `medium` / `high` / `xhigh` / `max`). Use `max_tokens` generously (thinking counts toward the output budget) and iterate `content` for the `text` block rather than indexing `content[0]`. ### Streaming ```python theme={"system"} with client.messages.stream( model="zai-org/GLM-5.2", max_tokens=2048, messages=[{"role": "user", "content": "Explain transformers."}], ) as stream: for text in stream.text_stream: print(text, end="", flush=True) final = stream.get_final_message() ``` ### Choose a completion window Set the window either in the request body or through a header. The body field takes precedence when both are set. ```python theme={"system"} # Body — the natural path for the Anthropic SDK message = client.messages.create( model="zai-org/GLM-5.2", max_tokens=2048, messages=[{"role": "user", "content": "hi"}], metadata={"completion_window": "standard"}, ) ``` ```python theme={"system"} # Header — set once on the client, or per request with extra_headers client = Anthropic( auth_token="YOUR_VALAR_API_KEY", base_url="https://api.valarhq.ai", default_headers={"X-Valar-Completion-Window": "standard"}, ) ``` Accepted values are `"asap"` (Now), `"priority"`, `"standard"`, and `"flex"`. See [Inference modes](/inference-modes#completion-windows). `"flex"` targets a \~5-minute turn time and is designed for background use. The Messages API has no `background` parameter, so a synchronous flex call blocks for up to 5 minutes and may 504 at the gateway's wait ceiling. For flex work, use the [Responses API](/inference-modes#completion-windows) with `background=true`, which returns a queued response you poll for later. ## Claude Agent SDK The Claude Agent SDK runs the Claude Code agent loop as a library. Because the SDK owns the request body, set the completion window through the `X-Valar-Completion-Window` header via `ANTHROPIC_CUSTOM_HEADERS` rather than body metadata. ```bash theme={"system"} pip install claude-agent-sdk ``` ```python theme={"system"} import asyncio from claude_agent_sdk import ClaudeAgentOptions, query async def main(): options = ClaudeAgentOptions( model="zai-org/GLM-5.2", max_turns=5, allowed_tools=[], env={ "ANTHROPIC_BASE_URL": "https://api.valarhq.ai", "ANTHROPIC_API_KEY": "YOUR_VALAR_API_KEY", "ANTHROPIC_CUSTOM_HEADERS": "X-Valar-Completion-Window: standard", # No effect until you attach MCP servers; see below. "ENABLE_TOOL_SEARCH": "true", }, ) async for message in query(prompt="Summarize what this repo does in two sentences.", options=options): for block in getattr(message, "content", []) or []: if getattr(block, "text", None): print(block.text, end="") asyncio.run(main()) ``` `ANTHROPIC_API_KEY` sends `x-api-key`, which Valar accepts. `allowed_tools` enables the SDK's built-in tools (`Read`, `Grep`, `Bash`, …) or your own MCP servers — leave it empty for a plain completion. ### Keep MCP tool definitions out of the prompt Set `ENABLE_TOOL_SEARCH` to `true` whenever you connect MCP servers. The SDK withholds tool definitions by default and loads only the ones a turn needs, but it turns that off when `ANTHROPIC_BASE_URL` points anywhere other than Anthropic. Pointing the SDK at Valar therefore sends every tool definition on every request. With a few hundred MCP tools connected, those definitions can account for more than half of every prompt. The cost compounds. Tool definitions are rendered ahead of your system prompt and messages, so any change to the tool list invalidates the prompt cache for the whole request. MCP servers connect asynchronously, which means the first turns of a session often carry a partial tool list that grows as servers finish connecting, and each change re-prefills everything behind it. With `ENABLE_TOOL_SEARCH` set, the agent receives a short summary of the available tools and loads full definitions on demand, at the cost of one extra round-trip per search. ## Next steps Realtime, async, and completion windows. What /v1/messages accepts and returns. # Check batch status Source: https://docs.valarhq.ai/api-reference/batches-api/check-batch-status /openapi.json get /batches/{batch_id} See where a batch currently stands. Poll this until `status` is terminal (`completed`, `failed`, or `expired`), then download `output_file_id`. Progress can arrive in large steps - a batch can move from quiet to completed at once, often well before its window ends. # Create a batch Source: https://docs.valarhq.ai/api-reference/batches-api/create-a-batch /openapi.json post /batches Start processing a previously uploaded JSONL file of requests. Upload the file first with [Upload a file](/api-reference/files-api/upload-a-file) using `purpose: batch`. Each line of the file is one request: ```json {"custom_id": "req-1", "method": "POST", "url": "/v1/chat/completions", "body": {"model": "zai-org/GLM-5.2", "messages": [{"role": "user", "content": "Hello!"}]}} ``` A batch holds up to 10,000 requests, and each line can be up to 28 MiB — enough for a maximum-size 20 MB image, which takes about 26.7 MiB once base64-encoded. The line size is a transport limit; each request must also fit the context window of the model it names, which is a separate and often much tighter bound — read `context_length` from [List supported models](/api-reference/models-api/list-supported-models) and size your requests against it. A request over the window is accepted with the file and fails when it runs. A file with more than 10,000 requests fails the whole batch after creation (`status: failed`). A line that fails validation is rejected on its own and counted in `request_counts.failed`; the rest of the batch keeps running. # Get a single batch request result Source: https://docs.valarhq.ai/api-reference/batches-api/get-a-single-batch-request-result /openapi.json get /batches/{batch_id}/{custom_id} Retrieve one request's result by its `custom_id` without waiting for the whole batch — a Valar extension to the OpenAI Batch API. Returns the request's Response object; while the batch is running the response reports `in_progress`. A line that was rejected at validation never became a request, so it returns 404. # Get batch request result Source: https://docs.valarhq.ai/api-reference/batches-api/get-batch-request-result GET /batches/{batch_id}/{custom_id} Retrieve the result of a specific request within a batch by its custom_id. # Get batch status Source: https://docs.valarhq.ai/api-reference/batches-api/get-batch-status GET /batches/{batch_id} Retrieve the current status of a batch. # List batches Source: https://docs.valarhq.ai/api-reference/batches-api/list-batches /openapi.json get /batches Retrieve your batches, paging through them when you need to. # Create a chat completion Source: https://docs.valarhq.ai/api-reference/chat-completions-api/create-a-chat-completion /openapi.json post /chat/completions A Chat Completions endpoint that follows the OpenAI interface. Pass stream: true to stream the reply back as Server-Sent Events made up of chat.completion.chunk objects. # Delete a file Source: https://docs.valarhq.ai/api-reference/files-api/delete-a-file /openapi.json delete /files/{file_id} # Get file metadata Source: https://docs.valarhq.ai/api-reference/files-api/get-file GET /files/{file_id} Retrieve a file's metadata by ID. # Get file content Source: https://docs.valarhq.ai/api-reference/files-api/get-file-content /openapi.json get /files/{file_id}/content Download a file's raw bytes. For a completed batch, download `output_file_id` to get one JSONL result line per request. # Get file metadata Source: https://docs.valarhq.ai/api-reference/files-api/get-file-metadata /openapi.json get /files/{file_id} # List files Source: https://docs.valarhq.ai/api-reference/files-api/list-files /openapi.json get /files Retrieve your uploaded files, newest first. # Upload a file Source: https://docs.valarhq.ai/api-reference/files-api/upload-a-file /openapi.json post /files Upload a file as `multipart/form-data`. For the Batch API, upload your JSONL request file with `purpose: batch`, then pass the returned file ID to [Create a batch](/api-reference/batches-api/create-a-batch). Batch input files can be up to 500 MB. # List supported models Source: https://docs.valarhq.ai/api-reference/models-api/list-supported-models GET /models # List the models you can call Source: https://docs.valarhq.ai/api-reference/models-api/list-the-models-you-can-call /openapi.json get /models # Create a response Source: https://docs.valarhq.ai/api-reference/responses-api/create-a-response /openapi.json post /responses Starts a Responses API task in the OpenAI style. You get back 202 when you pass background=true; otherwise the call returns 200 once the work finishes. # Fetch a response Source: https://docs.valarhq.ai/api-reference/responses-api/fetch-a-response /openapi.json get /responses/{response_id} # Retrieve a response Source: https://docs.valarhq.ai/api-reference/responses-api/retrieve-a-response GET /responses/{response_id} # OpenAI API compatibility Source: https://docs.valarhq.ai/compatibility What to set for an agent loop, and which request fields Valar ignores Valar speaks the OpenAI **Responses** and **Chat Completions** APIs, so your SDK works unchanged. Two things to know. ## Set `store: false` for an agent loop ```ts Vercel AI SDK {13} theme={"system"} import { createOpenAI } from "@ai-sdk/openai"; import { generateText } from "ai"; const valar = createOpenAI({ baseURL: "https://api.valarhq.ai/v1", apiKey: process.env.VALAR_API_KEY, }); const result = await generateText({ model: valar.responses("deepseek-ai/DeepSeek-V4.1-Flash"), messages, tools, providerOptions: { openai: { store: false } }, }); ``` ```ts OpenAI SDK {11} theme={"system"} import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.valarhq.ai/v1", apiKey: process.env.VALAR_API_KEY, }); const response = await client.responses.create({ model: "deepseek-ai/DeepSeek-V4.1-Flash", input: conversation, store: false, }); ``` ```python Python {13} theme={"system"} import os from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", api_key=os.environ["VALAR_API_KEY"], ) response = client.responses.create( model="deepseek-ai/DeepSeek-V4.1-Flash", input=conversation, store=False, ) ``` Reasoning then comes back with `encrypted_content`, your SDK replays it on the next turn, and the model keeps the thinking it already did. Leave `store` unset and that thinking is lost from the second step onward — the usual cause of "my agent forgets what it was doing". Pass `encrypted_content` back unchanged and do not parse it. `store` is not a retention control. ## Fields Valar ignores Every request is answered from the content you send in `input` — there is no server-side thread to point at. These fields are removed rather than rejected, so your request still runs; it just does not see the turn they name: | Field | Effect | | - | - | | `{"type": "item_reference", "id": "rs_…"}` input item | The item it names is not added to the prompt | | `previous_response_id` | The previous turn is not prepended | | `conversation` | No stored conversation is loaded | Send the conversation on every turn instead — every OpenAI-compatible SDK does this by default. # Use the docs MCP server Source: https://docs.valarhq.ai/docs-mcp Connect an AI agent to Valar documentation and public pricing tables The Valar docs MCP server lets an AI agent search the published documentation and read complete pages, including a compact public pricing table. ```text theme={"system"} https://docs.valarhq.ai/mcp ``` The docs MCP is public. You do not need a Valar API key to read documentation or pricing through it. ## Connect your client ### Claude Code Run this command once: ```bash theme={"system"} claude mcp add --transport http valar-docs https://docs.valarhq.ai/mcp ``` Restart Claude Code if the server does not appear in the current session. ### Cursor Add the server to `.cursor/mcp.json` in your project, or to your global Cursor MCP settings: ```json theme={"system"} { "mcpServers": { "valar-docs": { "url": "https://docs.valarhq.ai/mcp" } } } ``` Reload Cursor after saving the file. ### Other MCP clients Add a remote MCP server with these settings: | Setting | Value | | - | - | | Name | `valar-docs` | | Transport | HTTP | | URL | `https://docs.valarhq.ai/mcp` | Use the HTTP transport supported by your client. You do not need to run a local MCP process. ## What the server provides Your agent can discover the server's tools, search for a page, and use the documentation filesystem tool to read its full content. Search snippets help find a page; read the complete page before comparing all prices. The docs MCP reads published content. It does not run inference or call the authenticated model-listing API. ## Read pricing through MCP [Pricing for agents](/pricing-for-agents) is a compact Markdown table generated from the same public catalog as the visual [Pricing](/pricing) page. It lists model IDs and input, cached input, and output rates for Now, Priority, Standard, and Flex in USD per 1M tokens. Ask your agent: ```text theme={"system"} Use the Valar docs MCP server to find Pricing for agents. Read the full page, then show the input, cached input, and output price for each listed window on moonshotai/Kimi-K3-fast. Include the units. ``` If you already know the page path, ask the agent to use the documentation filesystem tool to read `/pricing-for-agents.mdx`. Tool names are supplied by the server when the client connects. These are public list prices, not live quotes. They update when the docs are published and exclude hidden and restricted models. For your organization's available models and negotiated rates, use the separate authenticated [`GET /v1/models` API](/api-reference/models-api/list-supported-models). ## Example prompts ```text theme={"system"} Search the Valar docs for the difference between Now, Priority, Standard, and Flex. ``` ```text theme={"system"} Find the Valar model-listing API and explain its pricing fields. ``` ```text theme={"system"} Read the full Pricing for agents page from the Valar docs MCP server. Convert the Standard rates for all listed models to JSON with model ID, input, cached input, and output fields. Preserve the decimal precision and label the prices as USD per 1M tokens. ``` ```text theme={"system"} Using only the Valar docs MCP server, show me how to migrate this project to Valar. ``` ## Troubleshooting ### The server does not connect Confirm that the URL ends in `/mcp` and that your client is configured for HTTP. Then reload the client or reconnect the server. ### The pricing page does not appear in search Ask the documentation filesystem tool to read `/pricing-for-agents.mdx` directly. If the page was just published, allow time for indexing. You can also [read the Markdown export](https://docs.valarhq.ai/pricing-for-agents.md) without MCP. ### The client asks for an API key Public documentation and pricing do not require a key. Check that you connected to `https://docs.valarhq.ai/mcp`, not an inference API endpoint. Do not paste a Valar API key into your prompt. ### The answer is incomplete or has different prices Ask the agent to read the full [Pricing for agents](/pricing-for-agents) page again instead of using search snippets or remembered prices. The compact and visual tables share the same source. Rates returned by the authenticated API may differ because of negotiated pricing or changes not yet published in the docs. ## Machine-readable alternatives If your client does not support MCP, [open the pricing Markdown export](https://docs.valarhq.ai/pricing-for-agents.md) directly. You can also use [`/llms.txt`](https://docs.valarhq.ai/llms.txt) for the documentation index or append `.md` to a documentation page URL. The export is Markdown. An agent can convert it to JSON or CSV on request, but that conversion is not a separate Valar API response. # Data Processing Addendum Source: https://docs.valarhq.ai/dpa How Valar processes and protects personal data on your behalf: roles, permitted use, zero data retention, sub-processors, security, incidents, and international transfers **Binding text:** the [Valar Customer Data Processing Addendum](https://valarhq.ai/dpa) is incorporated by reference into the Valar Terms of Use and prevails over this page. This page summarizes it so you can find the terms you need quickly. \ **Applies to:** every customer using the Services under the Agreement \ **Contact:** [privacy@valarhq.ai](mailto:privacy@valarhq.ai) for questions or a countersigned copy The DPA governs how Valar Ltd. (Hasolelim Street 17, Tel Aviv, Israel) processes personal data solely on your behalf when you use the Services. By using the Services you accept it. Where it conflicts with the Agreement, the DPA prevails for the processing of personal data; its Schedules prevail over its main body for the matters they cover. ## Definitions * **Prompt** - the content you submit to the Services for execution against a Model. * **Output** - the content a Model generates in response to a Prompt and returns to you. * **Model** - a third-party AI model made available through the Services and selected by you for a Prompt. * **Completion Window** - the maximum period within which Valar undertakes to return the first token, chosen per request through the API's completion window parameter, from real-time execution up to twelve (12) hours. * **Zero Data Retention** - the standing configuration under which Prompts and Outputs are not retained after the request they relate to has been executed. * **Self-Managed Deployment** - a deployment of the Services inside your own infrastructure or a cloud environment you control, where Prompts and Outputs are processed within that environment. * **Customer Personal Data** - personal data Valar processes solely on your behalf under the DPA and the Agreement. * **Data Incident** - accidental or unlawful destruction, loss, alteration, unauthorized disclosure of, or access to Customer Personal Data processed by Valar on your behalf. * **Security Documentation** - the technical and organizational measures for the Services, detailed in the [Valar Trust Center](https://trust.valarhq.ai/). Valar grants access on request. * **Data Protection Laws** - the GDPR, UK GDPR, CCPA, Israel's PPL (including Amendment 13, in force from 14 August 2025) and the Swiss FADP, as applicable to the processing. ## 1. Roles You are the **controller** of Customer Personal Data and Valar is the **processor**. Where you act as a processor for your own customers, Valar acts as your sub-processor. You are responsible for complying with Data Protection Laws in your use of the Services and for the lawful bases, notices and consents needed to send personal data to Valar. ## 2. What Valar may do with your data Valar processes Customer Personal Data only for the **Permitted Purposes**: processing in accordance with the Agreement and the DPA, providing the Services, following your reasonable documented instructions consistent with the Agreement, and complying with applicable law under a court order or a competent authority (with notice to you unless legally prohibited). If Valar believes an instruction infringes Data Protection Laws, it tells you without undue delay and may temporarily cease the affected processing. Valar will **not**, and will not permit any sub-processor to, use Customer Personal Data to train, fine-tune, or otherwise develop or improve artificial intelligence or machine learning systems. ## 3. Zero data retention Prompts and Outputs are not retained after the request they relate to has been executed: * Prompts submitted for **real-time execution** are not written to persistent storage. * Prompts submitted with a **deferred Completion Window** are held in queue storage only until the request has been executed, and in no case for more than twelve (12) hours from receipt, after which they are deleted. * Outputs are returned to you and are not retained after delivery. Because of this, no personal data from Prompts or Outputs remains in Valar's possession on termination, and no deletion request is needed for it. For a **Self-Managed Deployment**, Valar does not receive, store or process Prompts or Outputs at all. The DPA then applies only to account administration, support, and the operational telemetry described in Schedule 1. ## 4. Data subject requests If Valar receives a request from a data subject about Customer Personal Data, it notifies you or refers the data subject to you, to the extent legally permitted. Valar assists you in responding, insofar as possible and reasonable, limited to the information and means reasonably available to it. ## 5. Confidentiality Personnel, contractors and advisors who process Customer Personal Data are bound by confidentiality obligations at least as protective as the DPA and the Agreement, or by a statutory duty of confidentiality, and access is granted on a need-to-know basis. ## 6. Sub-processors * You give Valar general written authorization to engage sub-processors, subject to the conditions below. * The current **Sub-processor List**, with identities, locations and the service each provides, is published at [trust.valarhq.ai](https://trust.valarhq.ai/). * Valar gives at least **fourteen (14) days'** notice before adding or replacing a sub-processor, by updating the list and by email. * You may object in writing, for reasons relating to the protection of Customer Personal Data, within **thirty (30) days** of the notice. Silence is deemed acceptance. If you object, Valar uses reasonable efforts to offer a change that avoids the sub-processor within thirty (30) days; if it cannot, you may terminate the affected Services and pay only for Services provided. * Every sub-processor is bound by a written agreement with data-protection obligations that are the same or materially similar to the DPA. Valar remains responsible for its sub-processors, and none may use Customer Personal Data for training or model development or for any purpose beyond providing the Services. ## 7. Security and audits Valar maintains appropriate technical and organizational measures against unauthorized or unlawful processing and against accidental or unlawful destruction, loss, alteration, disclosure or access, taking into account the state of the art, implementation cost, and the nature, scope, context, purposes and risks of the processing. The measures are described in the Security Documentation and may be updated as long as the level of security is not lowered. On reasonable request and at your cost, Valar assists with your obligations under Articles 32 to 36 of the GDPR and the equivalent UK GDPR provisions. **Audits.** On fourteen (14) days' prior written request, no more than once every twelve (12) months (except after a Data Incident or where a supervisory authority requires it), and subject to reasonable confidentiality undertakings, Valar makes available at its own cost the information needed to demonstrate compliance, to you or to an independent, reputable third-party auditor who is not a competitor. Valar may first offer recent third-party certifications, attestations and audit reports; you may reasonably require on-site access where those do not adequately demonstrate compliance. Audit results are used solely to assess compliance with the DPA or to meet your own contractual obligations. Where the Standard Contractual Clauses apply, their audit rights prevail. ## 8. Data incidents * Valar keeps documented incident management policies and notifies you **without undue delay, and in any event within seventy-two (72) hours**, after becoming aware of a Data Incident, to the extent Data Protection Laws require it. This does not cover incidents caused by your own acts or omissions. * The notification includes, as it becomes known: the nature of the incident, including where possible the categories and approximate number of data subjects and records; the likely consequences; the measures taken or proposed; and a contact for more information. * You do not publish findings, notices or reports about a Data Incident that identify Valar without Valar's prior written approval, except where Data Protection Laws or a mandatory regulatory requirement compel it, in which case you give reasonable prior notice and limit the disclosure to the minimum required. Disclosure to your advisors under confidentiality is allowed. * Valar promptly reimburses your reasonable costs from a Data Incident caused by Valar's breach of the DPA, including notices to authorities and data subjects, audits and security testing, and data subject claims. ## 9. Return and deletion * **Prompts and Outputs** are deleted in the ordinary course under Zero Data Retention. Nothing remains to return or delete at termination. * **Other Customer Personal Data** (account, administrative and support data) is, at your written choice, deleted or returned within **thirty (30) days** after termination or expiry, and existing copies are deleted unless Data Protection Laws require retention. Valar may retain data for evidential purposes, to establish, exercise or defend legal claims, or to comply with law; retained data stays subject to the DPA. ## 10. International transfers * Transfers from the EEA, Switzerland and the UK to countries covered by an adequacy decision need no additional safeguards. * For other destinations, the transfer mechanism is the **EU Standard Contractual Clauses** (Commission Implementing Decision (EU) 2021/914) for EEA transfers, the **UK Addendum** (ICO template B.1.0) for UK transfers, and the EU SCCs as adjusted for the FADP for Swiss transfers. The full terms are Schedule 2 of the [binding addendum](https://valarhq.ai/dpa); the points below are the ones that decide how it applies to you. * Module Two applies when you are the controller, Module Three when you are a processor, and Module Four when Valar transfers to you as a controller outside the GDPR. Clause 9 uses general written authorization with the fourteen-day notice above. The EU SCCs are governed by the laws of Ireland, with disputes before Irish courts, and prevail over the DPA on any conflict. * **Additional safeguards.** Valar keeps encryption in transit and at rest and network protection per good industry practice, makes commercially reasonable efforts to resist bulk surveillance requests including under FISA section 702, and on a government access request, unless legally prohibited, informs the authority that Valar is a processor and directs it to you, and challenges the demand through reasonable legal mechanisms. No more than once every twelve (12) months, on your written request, Valar reports the types of binding legal demands it has received, to the extent the law permits. ## 11. Authorized affiliates You enter the DPA for yourself and for your Authorized Affiliates that use the Services under your Agreement without their own contract. They are bound by your obligations, and you coordinate all communication with Valar on their behalf. ## 12. Other provisions * **Impact assessments.** On reasonable request and at your cost, Valar cooperates with your data protection impact assessments and prior consultations with supervisory authorities, to the extent the information is available to Valar and not to you. * **Changes to the DPA.** Either party may request variations on at least **forty-five (45) calendar days'** written notice where a change in Data Protection Laws or a decision of a competent authority requires it. If no agreement is reached within thirty (30) days, either party may terminate the affected Services, paying only for Services provided. ## Schedule 1: details of the processing | | | | - | - | | **Subject matter** | Performance of the Services under the Agreement | | **Nature and purpose** | Providing the Services; performing the Agreement and the DPA; acting on your instructions consistent with the Agreement; sharing data with third parties on your instructions or through your use of the Services; complying with law | | **Duration** | Prompts and Outputs: continuously and on demand for the term of the Agreement, retained under Zero Data Retention. Account, administrative and support data: for the term, then deleted or returned per Section 9. Usage and log data: for as long as needed to provide the Services | | **Types of personal data** | Personal data you or your users include in a Prompt or a Model generates in an Output; account and administrative data (names, business email addresses, authentication identifiers and credentials); usage and log data (API request metadata, IP addresses, logs, metrics) | | **Data subjects** | Your administrative users, developers and other personnel who use the Services; natural persons whose personal data appears in a Prompt or an Output, depending on your use case | | **Sensitive data** | Valar does not require, request or intentionally process special categories of personal data under Article 9 of the GDPR. You decide whether to include such data in a Prompt and are responsible for the lawful basis and any notices or consents. Where sector-specific requirements apply, notify Valar in advance so any required further agreement is in place before the data is submitted | ## Schedule 3: CCPA terms Where you are a Business under the CCPA and Valar processes Personal Information subject to it, Valar acts as your **Service Provider**. Valar processes Personal Information solely for the Permitted Purposes; does not receive it as consideration for the Services; does not sell or share it, and does not retain, use or disclose it outside the Permitted Purposes or the direct business relationship; does not combine it with other data in a way the CCPA prohibits for Service Providers; and notifies you if it can no longer meet these obligations. Sections 4 to 9 and 12 of this summary apply with CCPA terminology. ## Schedule 4: Israel PPL supplement Where Valar processes personal data subject to Israel's Protection of Privacy Law on your behalf, you are the **Database Controller** and Valar is the **Holder**, with sub-processors as sub-Holders. Valar maintains the measures the Information Security Regulations require for the relevant database security level: a data security procedure, up-to-date documentation of the database structure and systems, periodic risk assessments and penetration tests with findings remediated, monitoring records, controls on portable devices and remote access, and restorable backups. Transfers out of Israel comply with the Transfer of Data to Databases Abroad regulations. Valar assists with inspection, rectification and erasure requests, performs no direct mailing under section 17C of the PPL without your written instruction, and on request reports at least annually on its performance under the supplement and supports your notifications to the Privacy Protection Authority. ## Questions For questions about the DPA or a countersigned copy, contact [**privacy@valarhq.ai**](mailto:privacy@valarhq.ai). # GCP Private Service Connect Source: https://docs.valarhq.ai/gcp-private-service-connect Reach the Valar API from a Google Cloud VPC with no internet egress Reach `valar-psc-gateway.gcp.api.valarhq.ai` over Private Service Connect instead of the public internet. Your VPC needs no Cloud NAT, no external IPs, no egress firewall rules for Valar, and no VPN or peering. Your IP ranges can overlap ours. The connection is one-way: you connect to us, and nothing on our side can reach into your VPC. ## Send us your project We need two things: 1. The Google Cloud **project ID** your workloads run in. 2. The **region** they run in. We allowlist the project. We need no access to your environment. ## Create the endpoint ```hcl Terraform theme={"system"} resource "google_compute_address" "valar" { name = "valar-psc-gateway" region = var.region subnetwork = var.subnet address_type = "INTERNAL" } resource "google_service_directory_namespace" "valar" { namespace_id = "valar-psc" location = var.region } resource "google_compute_forwarding_rule" "valar" { name = "valar-psc-gateway" region = var.region network = var.network ip_address = google_compute_address.valar.id target = "projects/valar-cp-prod/regions/us-east4/serviceAttachments/valar-prod-gateway" load_balancing_scheme = "" service_directory_registrations { namespace = google_service_directory_namespace.valar.namespace_id } } ``` ```bash gcloud theme={"system"} gcloud compute addresses create valar-psc-gateway \ --region=REGION --subnet=SUBNET gcloud service-directory namespaces create valar-psc --location=REGION gcloud compute forwarding-rules create valar-psc-gateway \ --region=REGION \ --network=NETWORK \ --address=valar-psc-gateway \ --target-service-attachment=projects/valar-cp-prod/regions/us-east4/serviceAttachments/valar-prod-gateway \ --service-directory-registration=projects/PROJECT/locations/REGION/namespaces/valar-psc ``` `load_balancing_scheme` must be the empty string. Omitting it builds an ordinary load balancer, and the error will not mention Private Service Connect. Keep the name `valar-psc-gateway` — it becomes your hostname. Keep the Service Directory registration; it is what creates your DNS record. ## Point your client at it DNS is created for you inside your VPC. Set the base URL: ```python Python theme={"system"} client = OpenAI( base_url="https://valar-psc-gateway.gcp.api.valarhq.ai/v1", api_key=os.environ["VALAR_API_KEY"], ) ``` ```typescript TypeScript theme={"system"} const client = new OpenAI({ baseURL: "https://valar-psc-gateway.gcp.api.valarhq.ai/v1", apiKey: process.env.VALAR_API_KEY, }); ``` ```bash curl theme={"system"} curl https://valar-psc-gateway.gcp.api.valarhq.ai/v1/models \ -H "Authorization: Bearer $VALAR_API_KEY" ``` ## Troubleshooting ```bash theme={"system"} gcloud compute forwarding-rules describe valar-psc-gateway \ --region=REGION --format='value(pscConnectionStatus)' ``` `PENDING` means we have not allowlisted your project. Send us the project ID. Check the forwarding rule has a Service Directory registration — without it the endpoint works but no DNS is created. Query from inside the VPC the endpoint lives in, and confirm the rule is named `valar-psc-gateway`. Add `allow_psc_global_access = true` to the forwarding rule (`--allow-psc-global-access` with `gcloud`). Traffic crosses regions on Google's backbone, which adds latency. Tell us your region and we can place an attachment closer. # Geo Control Source: https://docs.valarhq.ai/geo-control Pin your traffic to a regional endpoint instead of the geo-steered global one ## The global endpoint and the regional ones The base URL used throughout these docs, `https://api.valarhq.ai/v1`, is the **global** endpoint: DNS geo-steering resolves it to the nearest Valar region automatically, so a client in Frankfurt enters through the EU edge and a client in Chicago through the US edge, with failover between regions if one is unhealthy. If you want to be explicit about which region your traffic enters through — a fixed egress path for firewall rules, reproducible latency, or simply no dependence on where your resolver sits — point at a regional endpoint directly: | Endpoint | Region | | - | - | | `https://us.api.valarhq.ai/v1` | United States | | `https://eu.api.valarhq.ai/v1` | European Union (Ireland) | Everything else stays identical: use exactly the same API keys and model slugs from the [dashboard](https://app.valarhq.ai), the same request shapes, and the same [API surface](/api-reference/responses-api/create-a-response). The base URL is the only thing that changes. ## Configure it ```python Python theme={"system"} import os from openai import OpenAI client = OpenAI( base_url="https://eu.api.valarhq.ai/v1", # or https://us.api.valarhq.ai/v1 api_key=os.environ["VALAR_API_KEY"], ) ``` ```ts TypeScript theme={"system"} import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://eu.api.valarhq.ai/v1", // or https://us.api.valarhq.ai/v1 apiKey: process.env.VALAR_API_KEY, }); ``` ```bash cURL theme={"system"} curl https://eu.api.valarhq.ai/v1/responses \ -H "Authorization: Bearer $VALAR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "zai-org/GLM-5.2", "input": "Hello from the EU edge." }' ``` No key or workspace changes are needed to switch endpoints — the same key works against the global endpoint and both regional ones, and usage from all three lands in the same reporting. # Idempotent Requests Source: https://docs.valarhq.ai/idempotency How an idempotency key lets you retry inference calls without paying for the same work twice ## The problem retries create A `POST` that times out or returns a `5xx` leaves you guessing. The server may have received the request and run the inference, or it may not have. Retrying blindly risks running and billing the same job twice. Not retrying risks dropping a job that never actually completed. An idempotency key removes the guesswork. You attach a key to the first attempt, Valar stores the response under that key, and every later attempt with the same key returns that stored response instead of running inference again. Retrying becomes safe across flaky networks, timeouts, and ambiguous `5xx` responses. ## How it works Send the key in the `Idempotency-Key` request header. Use one key per logical request (a UUID or any unique string up to 255 characters) and send the same key on the first attempt and every retry. ```python theme={"system"} from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key="YOUR_VALAR_API_KEY", # or read OPENAI_API_KEY from the environment ) response = client.responses.create( model="zai-org/GLM-5.2", input="Summarize this document.", extra_headers={"Idempotency-Key": "order-9f8e7d6c"}, ) ``` The first call reserves the key and runs the job. Each subsequent call with that key lands on the same reservation and returns the stored response without re-running inference. The key only earns its keep once you actually retry. ## Rules that govern a reservation A reservation is identified by all three together. The same key under a different API key is a separate reservation. Rotating API keys therefore won't break replay, but the same idempotency key won't dedupe across two different API keys. Valar fingerprints the request body. Reusing a key with a meaningfully different body returns `400 idempotency_error` rather than the earlier response, which prevents a retry from silently returning the wrong answer after the client changed the request. Any `4xx` raised before the reservation is written leaves the key unreserved. You can fix the body and retry under the same key. A replay returns the current state of the underlying response record. For background requests the `status` tracks the task's latest transition, such as `queued` → `in_progress` → `completed`. ## Deciding when to retry Reach for the same idempotency key whenever the outcome is ambiguous. Don't retry when the work genuinely failed. | Signal | What it means | What to do | | - | - | - | | Network error / timeout | Ambiguous - the server may or may not have received the request | Retry with the same idempotency key | | `5xx` on a POST | Transient server-side failure | Retry with the same idempotency key | | `429` + `Retry-After` | Rate limit | Wait the `Retry-After` value, then retry | | `body.status: "failed"` | Inference genuinely failed | Investigate the cause; do not blind-retry | ## See also * [Sending Requests at Scale](/requests_at_scale) - the batch and background workflows where idempotent retries pay off most. * [Webhooks](/webhooks) - pair idempotent retries with webhook delivery so a client can re-drive submission without triggering another inference run. # Images Source: https://docs.valarhq.ai/images Send images alongside text to multimodal models on the Responses and Chat Completions APIs Multimodal models on Valar read images mixed in with your text. You attach each image one of two ways - a public URL that Valar fetches, or a base64 data URI you embed directly - and the model treats it as part of the prompt. Two things decide whether a request works: the model has to be multimodal, and each image has to fall inside the size and format limits below. ## Check the model supports images Vision is a per-model capability. Look for a check in the **Image** column on the [Models](/models) page, or call `GET /v1/models` to confirm at runtime. Sending image content to a text-only model fails fast: A non-multimodal model returns `400` with `this model does not support image input`. No tokens are charged. ## Attach an image The content block is shaped to match whichever API you're already calling. Each tab shows the URL form first, then the inline base64 form. Add an `input_image` part to the message `content`. Its `image_url` takes either a public URL or a `data:` URI, and the optional `detail` field accepts `"auto"`, `"low"`, or `"high"`. ```python theme={"system"} from openai import OpenAI client = OpenAI(base_url="https://api.valarhq.ai/v1", api_key="YOUR_VALAR_API_KEY") response = client.responses.create( model="moonshotai/Kimi-K2.7", input=[ { "role": "user", "content": [ {"type": "input_text", "text": "What's on the sign in this photo?"}, {"type": "input_image", "image_url": "https://example.com/storefront.jpg"}, ], } ], ) print(response.output_text) ``` To send the bytes yourself, base64-encode the file into a data URI and pass it in the same field: ```python theme={"system"} import base64, pathlib b64 = base64.b64encode(pathlib.Path("storefront.jpg").read_bytes()).decode() image = {"type": "input_image", "image_url": f"data:image/jpeg;base64,{b64}"} ``` Use the OpenAI `image_url` content part, where `url` holds the public URL or data URI: ```python theme={"system"} response = client.chat.completions.create( model="moonshotai/Kimi-K2.7", messages=[ { "role": "user", "content": [ {"type": "text", "text": "What's on the sign in this photo?"}, {"type": "image_url", "image_url": {"url": "https://example.com/storefront.jpg", "detail": "auto"}}, ], } ], ) ``` The same `url` field accepts a data URI - `{"url": f"data:image/jpeg;base64,{b64}"}` - for inline bytes. ## Limits | Limit | Value | | - | - | | Images per request | 20 (GLM-5.3-Flash: 30). Every image in the request counts, including those in earlier turns | | Size per image | 20 MB, measured on the decoded bytes (not the base64 string) | | Formats | JPEG, PNG, WebP, GIF | | Pixel dimensions | No cap - the 20 MB ceiling is the real bound; oversized images may be resized or tiled before decoding | | URL fetch | `http`/`https` only, reachable on the public internet; Valar gives up after 10 s | ## When an image is rejected Most problems come back as a `400` with a JSON error whose `message` names the failing check: | Status | Message | Cause | | - | - | - | | `400` | `this model does not support image input` | The model isn't multimodal. | | `400` | `a request may include at most N images` | More images in one request than the model accepts; `N` is that model's limit. | | `400` | `each image must be at most 20 MB` | An image exceeded 20 MB once decoded. | | `400` | `image format must be JPEG, PNG, WebP, or GIF` | The format isn't one Valar accepts. | | `400` | `image_url must be an https URL or a base64 data URI` | The `image_url` wasn't an `https` URL or a `data:` URI. | Two fetch failures happen **below** the JSON layer, so branch on the HTTP status rather than the body: * **`403 Forbidden`** - a referenced URL blocked at the edge comes back as an HTML page from the WAF, not a JSON error. * **Upstream fetch error** - when Valar can't retrieve a referenced image (unreachable host, timeout, or a non-200 response), the request fails with a generic provider error, sometimes a `503`. ## URL or base64? Either works - pick based on where the bytes already are: * **Base64** when you already hold the file (uploads, generated images). It skips a fetch and avoids exposing a URL, at the cost of a larger request body. * **URL** when the image is already hosted somewhere public. Smaller payload, but Valar has to reach it within 10 seconds. Either way, image bytes are **never cached between requests** - every call re-sends or re-fetches its images, so reusing the same image across turns pays the transfer each time. # Introduction Source: https://docs.valarhq.ai/index The inference provider tuned for high-throughput agentic work Most inference APIs were built for chat: one prompt, one response, optimized for the latency of a single call. But a growing share of AI work is asynchronous. Agents, batch jobs, labeling, data extraction, document processing, evaluation, and background workflows often run across hundreds or thousands of model calls. For these workloads, the goal is not the fastest first token - it is finishing the full job reliably, efficiently, and at the lowest possible cost. ## How requests work Valar speaks the [OpenAI Responses API](https://developers.openai.com/api/reference/resources/responses/methods/create) at `/v1/responses`, reachable under the base URL `https://api.valarhq.ai/v1`. Any OpenAI-compatible client works once you repoint its base URL and key. Because agent work is rarely interactive, the API leans on a `background` mode. Send `background: true` and the request returns a response id straight away instead of blocking; you then retrieve that id until the job reports `completed`. Pair that with [completion windows](/inference-modes#completion-windows) to trade latency for price, and you can keep large batches in flight without holding open a connection per request. Authenticate with `Authorization: Bearer `. Generate keys from the dashboard at [app.valarhq.ai](https://app.valarhq.ai), and store them in the `OPENAI_API_KEY` environment variable so the SDK and OpenAI clients pick them up automatically. ## Choose your execution mode How soon you need each result decides how you call Valar. [Inference modes](/inference-modes) covers this in depth; in short: | Mode | How you call it | Latency | Cost | Best for | | - | - | - | - | - | | Realtime | Synchronous request on the **Now** window | Seconds | Highest | Interactive chat, prototyping | | Async | `background=True`, then poll or wait on a [webhook](/webhooks), on the **Standard** window | Minutes | Lower | Agent loops, background jobs | | Batch | The [Batches API](/requests_at_scale) on the **Standard** window | Up to hours | Lowest | Large datasets, evals, offline jobs | ## Where to go next Make your first API request with Valar Move an OpenAI-compatible app to Valar with a base URL, key, and model change Read these docs by machine: MCP server, llms.txt, and Markdown views Route Claude Code and other coding agents through Valar and cut the bill # Inference modes Source: https://docs.valarhq.ai/inference-modes Pick how to run inference on Valar by how soon you need each result, and tune cost with completion windows Valar optimizes for throughput and cost on long-running agent work, not the latency of a single call. You pick an execution mode by how soon you need each result, then tune the cost-versus-latency trade-off with a completion window. ## Completion windows at a glance A completion window sets how much wall-clock time per turn you're willing to trade for a lower rate. Valar runs four. Choose one with `metadata.completion_window` (or the `X-Valar-Completion-Window` header), or leave it off and get `standard`. Each is detailed in [Completion windows](#completion-windows) below. | Window | Turnaround | Best for | Relative cost | | - | - | - | - | | Now (`asap`) | Immediate | Interactive and human-in-the-loop requests | Baseline | | Priority (`priority`) | Under \~10 seconds | Latency-sensitive agent loops that need a firm deadline | \~25% lower | | Standard (`standard`) | Under a minute, usually much faster | Everyday agent loops where cost matters | \~50% lower | | Flex (`flex`) | Up to 5 minutes, background only | Bulk jobs, evals, and offline runs | \~50% lower | ## The three modes ### Realtime A normal synchronous request that returns the result immediately. You send the call without `background` and read the output from the response. This is the lowest-latency path, finishing in seconds, and it runs on the **Now** completion window. Realtime works across the [Responses API](/api-reference/responses-api/create-a-response) (`/v1/responses`), Chat Completions (`/v1/chat/completions`). Use it for interactive chat, prototyping, and human-in-the-loop steps. ```python theme={"system"} import os from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key="YOUR_VALAR_API_KEY", # or read OPENAI_API_KEY from the environment ) response = client.responses.create( model="moonshotai/Kimi-K2.7", input="Summarize the latest support ticket in one sentence.", ) print(response.output_text) ``` #### Tell Valar your client timeout If your client gives up on a realtime call after a fixed time, tell Valar what that time is with the `X-Valar-Client-Timeout` header. It is read on every realtime endpoint: `/v1/responses`, `/v1/chat/completions` and `/v1/messages`. It changes what happens when a request produces nothing. Without it, Valar works to its own limits, which are longer than most clients wait: your client can time out while the request is still running, and all you see is a connection that stopped. With it, the request ends at *your* number and tells you why. The value is your timeout in seconds (`"120"`), or a duration (`"2m"`, `"90s"`). It must be at least one second. Values longer than Valar's own realtime limit are treated as that limit. Without the header Valar plans against its own limit, so set it to the timeout your SDK is actually configured with: ```python theme={"system"} client = OpenAI( base_url="https://api.valarhq.ai/v1", api_key="YOUR_VALAR_API_KEY", timeout=120, default_headers={"X-Valar-Client-Timeout": "120"}, ) ``` The budget covers the time to the **first output**. A stream that has started producing keeps running past it — it is working, and cutting it would throw away the answer you are already receiving. For a non-streaming call the first output is the whole answer, so there the budget covers the call. If the budget runs out with nothing produced, how the error reaches you depends on whether the endpoint has already committed to a response: | Endpoint | What you get | | - | - | | `/v1/chat/completions` | HTTP `504` with an error of type `timeout_error` | | `/v1/responses`, non-streaming | HTTP `200` with the response's `status` set to `failed` and the same `timeout_error` on its `error` field — the response object already exists, so the failure is reported on it | | `/v1/responses`, streaming | HTTP `200`; the stream opens with `response.created` as usual and ends with a `response.failed` event carrying the same error | | `/v1/messages` | as `/v1/responses` | Either way the request is recorded as failed, and it ends when you said you would stop waiting rather than later. The header does not change your prices, your rate limits, or which models you can call, and it is not a guarantee that an answer arrives inside the time you give — a model that needs longer than your budget will simply be reported as a timeout sooner. The header applies to realtime calls only. A request on one of the async windows is not waiting on a live connection, so it carries no client timeout. ### Async Set `background=True` on the Responses API. The create call returns a response id immediately, then you poll the [retrieve endpoint](/api-reference/responses-api/retrieve-a-response) or receive a [webhook](/webhooks) when the work finishes. On the Standard window a turn takes at most a minute and is usually much faster, so async clears high throughput at lower cost. Async jobs usually run on the **Standard** window (`standard`), the lower-cost default. Use async for agent loops, background jobs, and large fan-out. See [Sending requests at scale](/requests_at_scale). ```python theme={"system"} started = client.responses.create( model="moonshotai/Kimi-K2.7", input="Classify this ticket and draft a reply.", background=True, # returns a response id right away metadata={"completion_window": "standard"}, ) print("Queued:", started.id) ``` ### Batch (Private Preview) Batch lets you submit many requests at once as an uploaded JSONL file and retrieve a results file when the set completes - the lowest cost and highest throughput, with the longest turnaround (results land within the batch window, up to 12 hours). It follows the OpenAI Batch API, so existing `client.batches.create()` code works unchanged. For the end-to-end workflow, see [Sending requests at scale](/requests_at_scale). ## Compare the modes | Mode | How you call it | Typical latency | Cost | Best for | | - | - | - | - | - | | Realtime | Synchronous request, no `background` | Seconds | Highest | Interactive chat, prototyping, human-in-the-loop | | Async | Responses API with `background=True` | Minutes | Lower | Agent loops, background jobs, large fan-out | | Batch | Batches API, retrieve on completion | Up to hours | Lowest | Large datasets, evals, offline transforms | ## Completion windows A completion window tells Valar how much wall-clock time per turn your workload can tolerate, and you pay less the more time you give it. There are four, from fastest to cheapest: | Window | Avg. turn time | Price | Best for | | - | - | - | - | | Now (`asap`) | Immediate | Highest | Realtime calls, interactive UIs, human-in-the-loop | | Priority (`priority`) | Up to \~10 s | \~25% below Now | Latency-sensitive agents that need a firm deadline | | Standard (`standard`) | Up to 1 min, usually much faster | Lower | Cost-optimized agents, async loops | | Flex (`flex`) | Up to 5 min | Lowest | Background bulk work that can wait | Realtime uses the **Now** window; async work uses **Priority** for a firm \~10 s deadline, **Standard** for lower cost, or **Flex** for the cheapest background runs. Each model-and-window price pairing is on the [Pricing](/pricing) page. ### Set the window Pass `metadata.completion_window` on the request: ```python theme={"system"} response = client.responses.create( model="zai-org/GLM-5.2", input="Explain the key ideas behind transformers.", background=True, metadata={"completion_window": "standard"}, ) ``` Alternatively, set the `X-Valar-Completion-Window` header to the same value. This is useful when a client owns the request body on your behalf (for example, the Claude Agent SDK) and body metadata isn't available. The body field takes precedence when both are set. Accepted values are `"asap"` (Now), `"priority"`, `"standard"`, and `"flex"`. ### How each tier behaves **Now** runs immediately on the fastest available hardware in a latency-optimized setup, at the higher on-demand rate. Use it for realtime, interactive requests where a person or another system is waiting on the result. **Priority** is a durable async window that targets a tight \~10-second turn-time ceiling - a firm, low-latency completion target for agent loops - priced about 25% below the Now rate. Like the other async windows, request it explicitly with `background=True`. **Standard** is the default wherever a model supports it. It runs on Valar's maximum-efficiency serving stack with a turn-time ceiling of one minute - and in practice most turns complete much faster than that. Most of Valar's published prices reference this window, and it's the right default for async agent loops and batch jobs. **Flex** is the lowest-cost tier, for background bulk work that can tolerate a little extra latency. It targets a turn-time ceiling of about five minutes and **requires `background=True`** - it isn't available for realtime or synchronous calls. Flex is always explicit: Valar never selects it for you, so request it with `metadata.completion_window: "flex"` when you want it. ### Default behavior Leave `completion_window` off and the request defaults to `standard` when the model supports it; otherwise it falls back to the **Now** tier. **Priority** and **Flex** are never selected automatically - set either explicitly when you want it. Explicitly choosing a window the model does not support fails with `400 invalid_request_error`. The error message lists the model's supported windows, and the [Pricing](/pricing) page keeps a current support matrix. ## Next steps Send your first request with an OpenAI-compatible client. Fan out async and batch work across many requests. # LLM & agent access Source: https://docs.valarhq.ai/llm-access Read these docs by machine: MCP server, llms.txt, Markdown views, and a copy-paste migration prompt Everything on this site is designed to be read by an LLM as easily as by a person. If you are wiring up a coding agent — or you *are* the coding agent — use the surfaces below instead of scraping HTML. ## MCP server The docs run a hosted [MCP](https://modelcontextprotocol.io) server at: ```text theme={"system"} https://docs.valarhq.ai/mcp ``` It exposes search and retrieval tools over the full published site. Connect it to your client: ```bash Claude Code theme={"system"} claude mcp add --transport http valar-docs https://docs.valarhq.ai/mcp ``` ```json Cursor (mcp.json) theme={"system"} { "mcpServers": { "valar-docs": { "url": "https://docs.valarhq.ai/mcp" } } } ``` You can also use the contextual menu on any page (the button next to the page title) to connect the MCP server, copy the page as Markdown, or open the page in ChatGPT, Claude, or Cursor. [Use the docs MCP server](/docs-mcp) for connection steps, pricing examples, and troubleshooting. [Pricing for agents](/pricing-for-agents) provides a compact public rate table that you can retrieve through MCP or read as Markdown, with no API key. ## llms.txt and Markdown views * [`/llms.txt`](https://docs.valarhq.ai/llms.txt) — an index of every page, one URL per line, for agents that want to pick what to read. * [`/llms-full.txt`](https://docs.valarhq.ai/llms-full.txt) — the entire documentation in one file, for one-shot context loading. * **Any page as Markdown** — append `.md` to a page URL, e.g. [`/quickstart.md`](https://docs.valarhq.ai/quickstart.md). * [`/openapi.json`](https://docs.valarhq.ai/openapi.json) — the machine-readable API specification behind the API reference tab. ## Migrate an app with one prompt Valar speaks the OpenAI Responses and Chat Completions APIs, so migrating is a base-URL, key, and model-id change. Paste this into a coding agent pointed at your repo: ```text theme={"system"} Migrate this project from OpenAI to Valar (https://docs.valarhq.ai). - Change the OpenAI client base_url to https://api.valarhq.ai/v1 and read the key from VALAR_API_KEY. - Replace OpenAI model ids with a Valar model from /models (e.g. moonshotai/Kimi-K2.7). - Keep request and response shapes the same; flag any OpenAI-only params Valar doesn't support (see /support). ``` [Migrating to Valar](/migrate) walks through the same change by hand, with the full compatibility table. ## Route your coding agent through Valar Reading the docs is one thing; the agent itself can run on Valar. [ValarCode](/valarcode/overview) connects Claude Code, Cursor, Codex, and other harnesses to Valar with one command and cuts coding-model spend by more than half: ```bash theme={"system"} valar claude on ``` See [Set up ValarCode](/valarcode/setup). # LoRA adapters Source: https://docs.valarhq.ai/loras Register PEFT-trained LoRA adapters and run them on supported base models LoRA adapters let you customize a supported base model with your own PEFT-trained weights. You upload the adapter, register it against one or more eligible base models, and Valar validates it before it becomes usable. LoRA is rolling out in phases. Today you can upload, register, and manage adapters, and Valar validates them per base model. Running a registered adapter at request time is coming soon, and this page will be updated with the request syntax when it ships. ## Entitlement LoRA is enabled per organization. If your org is not entitled, the LoRA and file-upload endpoints return `403` with `lora entitlement required`. Contact your Valar representative to turn it on. ## Supported base models An adapter can only target a base model that Valar has marked LoRA-eligible. Each eligible model declares a maximum adapter rank and the attention and MLP modules a LoRA may target. | Base model | Max rank | Target modules | | - | - | - | | `zai-org/GLM-5.2` | 32 | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` | | `Qwen/Qwen3.5-397B-A17B` | 32 | same | | `Qwen/Qwen3.6-35B-A3B` | 32 | same | | `MiniMaxAI/MiniMax-M3` | 32 | same | The eligible set can change over time. Call [`GET /v1/models`](/api-reference/models-api/list-supported-models) to confirm what your key can reach, and register only against models from the list above. ## Adapter requirements Train with [PEFT](https://huggingface.co/docs/peft) and export the two standard adapter files: * `adapter_config.json` with `peft_type` set to `"LORA"`, `task_type` set to `"CAUSAL_LM"`, `base_model_name_or_path` matching the base model you register against, an adapter rank `r` no greater than the base model's max rank, and `target_modules` within the base model's allowed set. * `adapter_model.safetensors` with the adapter weights. Each file may be up to 5 GiB. The config file name must end in `.json` and the weights file name must end in `.safetensors`. ## Register an adapter Upload both files to Valar with `purpose` set to `lora`. The Files endpoint is OpenAI-compatible, so the OpenAI SDK works directly. ```python theme={"system"} from openai import OpenAI client = OpenAI(base_url="https://api.valarhq.ai/v1", api_key="YOUR_VALAR_KEY") with open("adapter_config.json", "rb") as f: cfg = client.files.create(file=f, purpose="lora") with open("adapter_model.safetensors", "rb") as f: wts = client.files.create(file=f, purpose="lora") ``` Each upload returns a file object with an `id` you pass to the next step. Register the adapter against one or more supported base models. `name` must be 2 to 64 characters, lowercase alphanumeric or dashes, start and end with an alphanumeric character, and be unique within your organization. ```bash theme={"system"} curl https://api.valarhq.ai/v1/loras \ -H "Authorization: Bearer YOUR_VALAR_KEY" \ -H "Content-Type: application/json" \ -d '{ "name": "support-tone-v1", "supported_models": ["zai-org/GLM-5.2"], "config_file_id": "file-abc123", "weights_file_id": "file-def456", "display_name": "Support tone v1", "description": "House support voice, trained on resolved tickets" }' ``` The response is a LoRA object created in the `verifying` status: ```json theme={"system"} { "id": "lora-9f2c1a7b", "object": "lora", "name": "support-tone-v1", "display_name": "Support tone v1", "description": "House support voice, trained on resolved tickets", "supported_models": ["zai-org/GLM-5.2"], "config_file_id": "file-abc123", "weights_file_id": "file-def456", "status": "verifying", "created_at": 1731000000, "updated_at": 1731000000 } ``` ## Validation and status Every adapter is reviewed before it can be used. A LoRA moves through three states: | Status | Meaning | | - | - | | `verifying` | Registered and awaiting validation. | | `deployed` | Validated and ready. | | `failed` | Validation failed. `status_details` carries the reason. | ## Manage adapters List every adapter in your organization, fetch one by name or id, or delete one. ```bash theme={"system"} # List curl https://api.valarhq.ai/v1/loras -H "Authorization: Bearer YOUR_VALAR_KEY" # Fetch by name or by id curl https://api.valarhq.ai/v1/loras/support-tone-v1 -H "Authorization: Bearer YOUR_VALAR_KEY" # Delete curl -X DELETE https://api.valarhq.ai/v1/loras/support-tone-v1 -H "Authorization: Bearer YOUR_VALAR_KEY" ``` `GET /v1/loras` returns `{ "object": "list", "data": [ ... ] }`. A delete returns `{ "id": "...", "object": "lora.deleted", "deleted": true }`. ## Using a LoRA (coming soon) Once an adapter is `deployed`, you will be able to run it by referencing it on a request to a supported base model. This is not yet available. When it ships, this section will document the exact request field and any completion-window constraints. # Migrating to Valar Source: https://docs.valarhq.ai/migrate Move an OpenAI-compatible app to Valar with a base URL, key, and model change Valar speaks the OpenAI **Responses** and **Chat Completions** APIs, so moving an app from OpenAI - or any OpenAI-compatible provider - is mostly three changes: the base URL, the API key, and the model id. Your request and response shapes stay the same. Sign in at the [Valar dashboard](https://app.valarhq.ai) and create a key. Store it as `VALAR_API_KEY` so the OpenAI SDK and other clients pick it up. ```bash theme={"system"} export VALAR_API_KEY=sk_... ``` Keep your existing OpenAI client. Change the base URL to `https://api.valarhq.ai/v1` and pass your Valar key - Valar authenticates with `Authorization: Bearer `. ```python Python theme={"system"} from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", api_key="YOUR_VALAR_API_KEY", ) ``` ```ts TypeScript theme={"system"} import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.valarhq.ai/v1", apiKey: process.env.VALAR_API_KEY, }); ``` ```bash cURL theme={"system"} curl https://api.valarhq.ai/v1/responses \ -H "Authorization: Bearer $VALAR_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "moonshotai/Kimi-K2.7", "input": "Hello"}' ``` Valar serves open-weight models, so update the `model` field - OpenAI names like `gpt-4o` won't resolve. Pick one from the [Models](/models) page, such as `moonshotai/Kimi-K2.7` or `zai-org/GLM-5.2`, and confirm availability at runtime with `GET /v1/models`. For work that doesn't need an instant answer, set `background=True` and choose a [completion window](/inference-modes#completion-windows) - `standard` or `flex` - to trade a little latency for a lower rate. See [Inference modes](/inference-modes). ## What changes | | OpenAI | Valar | | - | - | - | | Base URL | `https://api.openai.com/v1` | `https://api.valarhq.ai/v1` | | Auth header | `Authorization: Bearer ` | `Authorization: Bearer ` | | Key env var | `OPENAI_API_KEY` | `VALAR_API_KEY` | | Model id | `gpt-4o`, `o3`, … | open-weight ids from [Models](/models) | | Surfaces | Responses, Chat Completions | Responses, Chat Completions | A few behaviors differ from OpenAI - worth checking before you ship: * **Streaming is available on Chat Completions and Messages.** The Responses API rejects `stream: true`; for long jobs use `background: true` and poll or wait on [webhooks](/webhooks). [Inference modes](/inference-modes) covers the realtime / async / batch split. * **Completion windows replace latency tuning.** Use `metadata.completion_window` (`asap`, `standard`, or `flex`) instead of `service_tier`. See [Pricing](/pricing). * **Structured outputs work the same way** - `text.format` on Responses, or `response_format` on Chat Completions, with a JSON schema. See [Structured outputs](/structured-outputs). * **Some OpenAI-only parameters are ignored or rejected** (server-side tools, conversation chaining, sampling penalties, and others). The [API support matrix](/support) lists exactly what each surface accepts. Coming from **Anthropic**? The OpenAI SDK works against Valar's Responses and Chat Completions APIs, and the Anthropic SDK works against Valar's Messages API — see [Anthropic SDK & Claude Agent SDK](/anthropic-sdk). ## Let your coding agent do it Valar's docs are built to be read by agents: every page has a Markdown view, the full index lives at [`/llms.txt`](https://docs.valarhq.ai/llms.txt), and there's an MCP server at `https://docs.valarhq.ai/mcp` — see [LLM & agent access](/llm-access) for all of it. Point your coding agent (Cursor, Claude Code, and the like) at the docs and hand it a prompt such as: ```text theme={"system"} Migrate this project from OpenAI to Valar (https://docs.valarhq.ai). - Change the OpenAI client base_url to https://api.valarhq.ai/v1 and read the key from VALAR_API_KEY. - Replace OpenAI model ids with a Valar model from /models (e.g. moonshotai/Kimi-K2.7). - Keep request and response shapes the same; flag any OpenAI-only params Valar doesn't support (see /support). ``` ## Next steps Make your first Valar request. Pick a model and copy its id. Realtime, async, batch, and completion windows. Get schema-constrained JSON back. # Models Source: https://docs.valarhq.ai/models All models currently served by Valar
Model Slug Image Reasoning
DeepSeek V4 Pro
DeepSeek
Kimi-K3 Fast
Moonshot AI
moonshotai/Kimi-K3-fast
Nemotron Ultra
NVIDIA
GLM-5.3 Fast
Z.ai
zai-org/GLM-5.3-fast
DeepSeek-V4.1-Flash
DeepSeek
Qwen3.5 27B
Qwen
Qwen/Qwen3.5-27B
Kimi K2.7 Code
Moonshot AI
Kimi-K3
Moonshot AI
GLM-5.2
Z.ai
GLM-5.2 Fast
Z.ai
GLM-5.3
Z.ai
GLM-5.3 Flash
Z.ai
gpt-oss-120b
OpenAI
Qwen3.5-397B-A17B
Qwen
Qwen3.6 35B-A3B
Qwen
Qwen3.8-Max
Qwen
qwen/qwen3.8-max
Gemma 4 31B IT
Google
Gemma 4 26B A4B
Google
MiniMax M3
MiniMax
## Reasoning models and output caps `max_output_tokens` caps **reasoning tokens plus the visible answer**, not the answer alone. On a model marked **Reasoning** above, most of that budget goes to thinking you never see — on `qwen/qwen3.7-plus`, around 96% of it, roughly a thousand reasoning tokens behind a two-sentence answer. So a cap sized for the answer stops generation mid-thought, and the response settles `incomplete`: ```json theme={"system"} { "status": "incomplete", "incomplete_details": { "reason": "max_output_tokens" }, "max_output_tokens": 1000 } ``` That status is accurate - generation really did stop at the cap - but it does not mean the answer is missing. In a 500-request run at `max_output_tokens: 1000`, two thirds of the responses came back `incomplete`, and their answers matched the `completed` third in length and quality. The only difference was how long the model thought. Gating on `status` alone throws those answers away: ```python theme={"system"} if response.status == "completed": # discards two thirds of good answers keep(response) ``` Branch on `incomplete_details.reason` and read the output instead: ```python theme={"system"} text = (response.output_text or "").strip() if response.status == "completed": keep(text) elif response.incomplete_details and response.incomplete_details.reason == "max_output_tokens": # Stopped at the cap, usually on a finished answer. Keep it if there is text; # an empty one means the model never got past reasoning, so raise the cap. if text: keep(text) else: retry(response, cap=response.max_output_tokens * 2) else: handle_failure(response) ``` ### Size the cap for reasoning plus answer Send a handful of representative requests, read `usage.output_tokens_details.reasoning_tokens` off the results, and set the cap to that plus the answer length you want. Omit `max_output_tokens` entirely when you don't need a hard ceiling. Don't set a cap below the model's typical reasoning volume. A cap of a few tokens leaves no room to finish thinking and start answering, and the request fails outright with a `503` rather than returning an empty `incomplete`. ## Custom models Beyond the catalog above, Valar can serve your own model weights. If you have a custom or fine-tuned open-weight model, we can host it on Valar's inference stack and expose it through the same OpenAI-compatible API, completion windows, and billing as any catalog model. Reach out to your Valar contact to onboard a custom model. For adapter-based customization on top of a supported base model, register a PEFT-trained [LoRA adapter](/loras) and run it against an eligible base model such as `zai-org/GLM-5.2`. To confirm what your API key can reach in a given environment at runtime, call [`GET /v1/models`](/api-reference/models-api/list-supported-models) against that environment rather than relying on this table alone. Each model it returns carries its context window, capability flags, and the [rate card](/pricing) your organization is billed at. # Pricing Source: https://docs.valarhq.ai/pricing How Valar bills inference and the per-token rate for every model ## How billing works You pay per token. Three rates apply to each request: **input** for the tokens you send, **cached** for input tokens served from a prefix cache, and **output** for the tokens the model generates. Claude and OpenAI models add a fourth, **cache write**, for input tokens written into the prompt cache; see [Cache writes](#cache-writes). Every figure in the table below is in **USD per 1M tokens**. Caching is automatic: Valar matches shared prompt prefixes for you and charges the lower cached rate on the tokens that hit. [Prompt Caching](/prompt-caching) explains how to raise your hit rate, but nothing is required to get the cached rate. The rate you pay also depends on the completion window you request. Faster scheduling carries a higher rate: the on-demand **Now** window costs the most, **Priority** about 25% less, **Standard** about 50% less, and **Flex** the least. The table prices out the windows available for each model; coverage varies, and you can mix windows per request. [Completion Windows](/inference-modes#completion-windows) explains the trade-offs. Window coverage differs by model, and we keep adding models and widening window support. If a model or window you want isn't shown, get in touch. You can also read these rates in code: [`GET /v1/models`](/api-reference/models-api/list-supported-models) returns a `pricing` array on every model, one entry per window. Those figures are resolved for your organization, so if you're on negotiated rates the API quotes yours rather than the table below. ## Cache writes When a request writes part of its prompt into the prompt cache, those tokens are counted separately from input and billed at the model's cache-write rate: | Models | Cache-write rate | | - | - | | Claude | 1.25x the input rate for the default 5-minute cache, 2x the input rate for a write cached with the 1-hour TTL | | OpenAI | 1.25x the input rate | | All other models | No cache-write charge | The written tokens are reported apart from input tokens, for example as `cache_creation_input_tokens` on the Anthropic Messages API and `cache_write_tokens` in the [Usage API](/usage-endpoints). A later request that reuses the cached prefix pays the lower cached rate on those tokens. Organizations on negotiated rates may have a different cache-write rate. ## Use pricing from an agent Read [Pricing for agents](/pricing-for-agents) for a compact Markdown table of the same public rates. Your agent can retrieve it through the [docs MCP server](/docs-mcp), or you can [open the Markdown export](https://docs.valarhq.ai/pricing-for-agents.md) directly. No API key is required. Both tables are generated from the same public model catalog and update when the docs are published. They show list prices, not organization-specific rates. ## Rate table
USD · per 1M tokens
Model Window Input Cached Output
DeepSeek V4 Pro
deepseek-ai/DeepSeek-V4-Pro
Now \$ 1.056 \$ 0.0352 \$ 3.168
Priority \$ 0.792 \$ 0.0264 \$ 2.376
Standard \$ 0.528 \$ 0.0176 \$ 1.584
Flex \$ 0.3696 \$ 0.01232 \$ 1.1088
Nemotron Ultra
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4
Now \$ 0.5 \$ 0.1 \$ 2.2
Priority \$ 0.375 \$ 0.075 \$ 1.65
Standard \$ 0.25 \$ 0.05 \$ 1.1
Flex \$ 0.175 \$ 0.035 \$ 0.77
DeepSeek-V4.1-Flash
deepseek-ai/DeepSeek-V4.1-Flash
Now \$ 0.255 \$ 0.0255 \$ 1.02
Priority \$ 0.19125 \$ 0.019125 \$ 0.765
Standard \$ 0.1275 \$ 0.01275 \$ 0.51
Flex \$ 0.08925 \$ 0.008925 \$ 0.357
Kimi K2.7 Code
moonshotai/Kimi-K2.7
Now \$ 0.75 \$ 0.15 \$ 3.5
Priority \$ 0.6 \$ 0.12 \$ 2.8
Standard \$ 0.375 \$ 0.075 \$ 1.75
Flex \$ 0.2625 \$ 0.0525 \$ 1.225
Kimi-K3
moonshotai/Kimi-K3
Now \$ 2.25 \$ 0.225 \$ 11.25
Priority \$ 1.6875 \$ 0.16875 \$ 8.4375
Standard \$ 1.125 \$ 0.1125 \$ 5.625
Flex \$ 0.7875 \$ 0.07875 \$ 3.9375
Kimi-K3 Fast
moonshotai/Kimi-K3-fast
Now \$ 3.375 \$ 0.225 \$ 16.875
Priority \$ 2.53125 \$ 0.16875 \$ 12.65625
Standard \$ 1.6875 \$ 0.1125 \$ 8.4375
Flex \$ 1.18125 \$ 0.07875 \$ 5.90625
GLM-5.2
zai-org/GLM-5.2
Now \$ 0.98 \$ 0.112 \$ 3.4
Priority \$ 0.784 \$ 0.0896 \$ 2.72
Standard \$ 0.49 \$ 0.091 \$ 1.54
Flex \$ 0.343 \$ 0.064 \$ 1.078
GLM-5.2 Fast
zai-org/GLM-5.2-fast
Now \$ 1.68 \$ 0.168 \$ 5.28
Priority \$ 1.26 \$ 0.126 \$ 3.96
Standard \$ 0.84 \$ 0.084 \$ 2.64
Flex \$ 0.588 \$ 0.0588 \$ 1.848
GLM-5.3
zai-org/GLM-5.3
Now \$ 1.12 \$ 0.208 \$ 3.52
Priority \$ 0.84 \$ 0.156 \$ 2.64
Standard \$ 0.56 \$ 0.104 \$ 1.76
Flex \$ 0.392 \$ 0.0728 \$ 1.232
GLM-5.3 Fast
zai-org/GLM-5.3-fast
Now \$ 1.785 \$ 0.3315 \$ 5.61
Priority \$ 1.33875 \$ 0.248625 \$ 4.2075
Standard \$ 0.8925 \$ 0.16575 \$ 2.805
Flex \$ 0.62475 \$ 0.116025 \$ 1.9635
GLM-5.3 Flash
zai-org/GLM-5.3-Flash
Now \$ 0.12 \$ 0.0232 \$ 0.4
Priority \$ 0.09 \$ 0.0174 \$ 0.3
Standard \$ 0.06 \$ 0.0116 \$ 0.2
Flex \$ 0.042 \$ 0.00812 \$ 0.14
gpt-oss-120b
openai/gpt-oss-120b
Now \$ 0.06 \$ 0.03 \$ 0.4
Priority \$ 0.045 \$ 0.023 \$ 0.3
Standard \$ 0.04 \$ 0.02 \$ 0.3
Flex \$ 0.024 \$ 0.012 \$ 0.155
Qwen3.5 27B
Qwen/Qwen3.5-27B
Now \$ 0.27 \$ 0.054 \$ 2.2
Priority \$ 0.203 \$ 0.041 \$ 1.65
Standard \$ 0.135 \$ 0.027 \$ 1.1
Flex \$ 0.0945 \$ 0.0189 \$ 0.77
Qwen3.5-397B-A17B
Qwen/Qwen3.5-397B-A17B
Now \$ 0.45 \$ 0.09 \$ 1.35
Priority \$ 0.338 \$ 0.068 \$ 1.013
Standard \$ 0.25 \$ 0.05 \$ 0.75
Flex \$ 0.175 \$ 0.035 \$ 0.525
Qwen3.6 35B-A3B
Qwen/Qwen3.6-35B-A3B
Now \$ 0.225 \$ 0.045 \$ 0.9
Priority \$ 0.169 \$ 0.0338 \$ 0.675
Standard \$ 0.13 \$ 0.026 \$ 0.52
Flex \$ 0.0925 \$ 0.0185 \$ 0.37
Qwen3.8-Max
qwen/qwen3.8-max
Now \$ 2 \$ 0.25 \$ 6
Priority \$ 2 \$ 0.25 \$ 6
Standard \$ 1 \$ 0.125 \$ 3
Flex \$ 0.7 \$ 0.0875 \$ 2.1
Gemma 4 31B IT
google/gemma-4-31B-it
Now \$ 0.14 \$ 0.06 \$ 0.4
Priority \$ 0.105 \$ 0.045 \$ 0.3
Standard \$ 0.07 \$ 0.03 \$ 0.2
Flex \$ 0.049 \$ 0.021 \$ 0.14
Gemma 4 26B A4B
google/gemma-4-26B-A4B-it
Now \$ 0.072 \$ 0.0144 \$ 0.216
Priority \$ 0.054 \$ 0.0108 \$ 0.162
Standard \$ 0.036 \$ 0.0072 \$ 0.108
Flex \$ 0.0252 \$ 0.00504 \$ 0.0756
MiniMax M3
MiniMaxAI/MiniMax-M3
Now \$ 0.2 \$ 0.06 \$ 1.2
Priority \$ 0.15 \$ 0.045 \$ 0.9
Standard \$ 0.1 \$ 0.03 \$ 0.6
Flex \$ 0.07 \$ 0.021 \$ 0.42
For each model's capabilities (e.g., image input and reasoning support), see [Models](/models). # Pricing for agents Source: https://docs.valarhq.ai/pricing-for-agents Public model prices in a compact Markdown table for agents and exports Use this table to read or compare Valar's public list prices without parsing the visual pricing page. All amounts are **USD per 1M tokens**. Model IDs are case-sensitive. Retired IDs keep working as aliases of their replacement and are billed at the replacement's rates: requests for `deepseek-ai/DeepSeek-V4-Flash` are served as `deepseek-ai/DeepSeek-V4.1-Flash`. This table contains the same models and rates as [Pricing](/pricing), covering **Now**, **Priority**, **Standard**, and **Flex**. Hidden and restricted models are excluded. Both tables are generated from the same public catalog and update when the docs are published; this is not a live quote or a feed of organization-specific rates. ## Read or export No API key is required. [Open this page as Markdown](https://docs.valarhq.ai/pricing-for-agents.md), or connect to the [docs MCP server](/docs-mcp) and ask: ```text theme={"system"} Read the full Pricing for agents page from the Valar docs MCP server. Return each model ID and its Standard input, cached input, and output rates as JSON. Keep the USD-per-1M-token units and decimal precision unchanged. ``` The Markdown table is the published export. Any JSON or CSV you ask an agent to produce is a conversion of that table, not a separate API response. ## Public rates Each row identifies a model and a completion window. **Input** is uncached input; **cached input** is input served from the prompt cache; **output** includes generated reasoning tokens. See [Completion windows](/inference-modes#completion-windows) for scheduling behavior. Claude and OpenAI models also bill input tokens written into the prompt cache at a cache-write rate: 1.25x the input rate, or 2x for a Claude write cached with the 1-hour TTL. Other models have no cache-write charge. See [Cache writes](/pricing#cache-writes). | Model ID | Window | Input | Cached input | Output | | - | - | -: | -: | -: | | `deepseek-ai/DeepSeek-V4-Pro` | Now | 1.056 | 0.0352 | 3.168 | | `deepseek-ai/DeepSeek-V4-Pro` | Priority | 0.792 | 0.0264 | 2.376 | | `deepseek-ai/DeepSeek-V4-Pro` | Standard | 0.528 | 0.0176 | 1.584 | | `deepseek-ai/DeepSeek-V4-Pro` | Flex | 0.3696 | 0.01232 | 1.1088 | | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` | Now | 0.5 | 0.1 | 2.2 | | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` | Priority | 0.375 | 0.075 | 1.65 | | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` | Standard | 0.25 | 0.05 | 1.1 | | `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` | Flex | 0.175 | 0.035 | 0.77 | | `deepseek-ai/DeepSeek-V4.1-Flash` | Now | 0.255 | 0.0255 | 1.02 | | `deepseek-ai/DeepSeek-V4.1-Flash` | Priority | 0.19125 | 0.019125 | 0.765 | | `deepseek-ai/DeepSeek-V4.1-Flash` | Standard | 0.1275 | 0.01275 | 0.51 | | `deepseek-ai/DeepSeek-V4.1-Flash` | Flex | 0.08925 | 0.008925 | 0.357 | | `moonshotai/Kimi-K2.7` | Now | 0.75 | 0.15 | 3.5 | | `moonshotai/Kimi-K2.7` | Priority | 0.6 | 0.12 | 2.8 | | `moonshotai/Kimi-K2.7` | Standard | 0.375 | 0.075 | 1.75 | | `moonshotai/Kimi-K2.7` | Flex | 0.2625 | 0.0525 | 1.225 | | `moonshotai/Kimi-K3` | Now | 2.25 | 0.225 | 11.25 | | `moonshotai/Kimi-K3` | Priority | 1.6875 | 0.16875 | 8.4375 | | `moonshotai/Kimi-K3` | Standard | 1.125 | 0.1125 | 5.625 | | `moonshotai/Kimi-K3` | Flex | 0.7875 | 0.07875 | 3.9375 | | `moonshotai/Kimi-K3-fast` | Now | 3.375 | 0.225 | 16.875 | | `moonshotai/Kimi-K3-fast` | Priority | 2.53125 | 0.16875 | 12.65625 | | `moonshotai/Kimi-K3-fast` | Standard | 1.6875 | 0.1125 | 8.4375 | | `moonshotai/Kimi-K3-fast` | Flex | 1.18125 | 0.07875 | 5.90625 | | `zai-org/GLM-5.2` | Now | 0.98 | 0.112 | 3.4 | | `zai-org/GLM-5.2` | Priority | 0.784 | 0.0896 | 2.72 | | `zai-org/GLM-5.2` | Standard | 0.49 | 0.091 | 1.54 | | `zai-org/GLM-5.2` | Flex | 0.343 | 0.064 | 1.078 | | `zai-org/GLM-5.2-fast` | Now | 1.68 | 0.168 | 5.28 | | `zai-org/GLM-5.2-fast` | Priority | 1.26 | 0.126 | 3.96 | | `zai-org/GLM-5.2-fast` | Standard | 0.84 | 0.084 | 2.64 | | `zai-org/GLM-5.2-fast` | Flex | 0.588 | 0.0588 | 1.848 | | `zai-org/GLM-5.3` | Now | 1.12 | 0.208 | 3.52 | | `zai-org/GLM-5.3` | Priority | 0.84 | 0.156 | 2.64 | | `zai-org/GLM-5.3` | Standard | 0.56 | 0.104 | 1.76 | | `zai-org/GLM-5.3` | Flex | 0.392 | 0.0728 | 1.232 | | `zai-org/GLM-5.3-fast` | Now | 1.785 | 0.3315 | 5.61 | | `zai-org/GLM-5.3-fast` | Priority | 1.33875 | 0.248625 | 4.2075 | | `zai-org/GLM-5.3-fast` | Standard | 0.8925 | 0.16575 | 2.805 | | `zai-org/GLM-5.3-fast` | Flex | 0.62475 | 0.116025 | 1.9635 | | `zai-org/GLM-5.3-Flash` | Now | 0.12 | 0.0232 | 0.4 | | `zai-org/GLM-5.3-Flash` | Priority | 0.09 | 0.0174 | 0.3 | | `zai-org/GLM-5.3-Flash` | Standard | 0.06 | 0.0116 | 0.2 | | `zai-org/GLM-5.3-Flash` | Flex | 0.042 | 0.00812 | 0.14 | | `openai/gpt-oss-120b` | Now | 0.06 | 0.03 | 0.4 | | `openai/gpt-oss-120b` | Priority | 0.045 | 0.023 | 0.3 | | `openai/gpt-oss-120b` | Standard | 0.04 | 0.02 | 0.3 | | `openai/gpt-oss-120b` | Flex | 0.024 | 0.012 | 0.155 | | `Qwen/Qwen3.5-27B` | Now | 0.27 | 0.054 | 2.2 | | `Qwen/Qwen3.5-27B` | Priority | 0.203 | 0.041 | 1.65 | | `Qwen/Qwen3.5-27B` | Standard | 0.135 | 0.027 | 1.1 | | `Qwen/Qwen3.5-27B` | Flex | 0.0945 | 0.0189 | 0.77 | | `Qwen/Qwen3.5-397B-A17B` | Now | 0.45 | 0.09 | 1.35 | | `Qwen/Qwen3.5-397B-A17B` | Priority | 0.338 | 0.068 | 1.013 | | `Qwen/Qwen3.5-397B-A17B` | Standard | 0.25 | 0.05 | 0.75 | | `Qwen/Qwen3.5-397B-A17B` | Flex | 0.175 | 0.035 | 0.525 | | `Qwen/Qwen3.6-35B-A3B` | Now | 0.225 | 0.045 | 0.9 | | `Qwen/Qwen3.6-35B-A3B` | Priority | 0.169 | 0.0338 | 0.675 | | `Qwen/Qwen3.6-35B-A3B` | Standard | 0.13 | 0.026 | 0.52 | | `Qwen/Qwen3.6-35B-A3B` | Flex | 0.0925 | 0.0185 | 0.37 | | `qwen/qwen3.8-max` | Now | 2 | 0.25 | 6 | | `qwen/qwen3.8-max` | Priority | 2 | 0.25 | 6 | | `qwen/qwen3.8-max` | Standard | 1 | 0.125 | 3 | | `qwen/qwen3.8-max` | Flex | 0.7 | 0.0875 | 2.1 | | `google/gemma-4-31B-it` | Now | 0.14 | 0.06 | 0.4 | | `google/gemma-4-31B-it` | Priority | 0.105 | 0.045 | 0.3 | | `google/gemma-4-31B-it` | Standard | 0.07 | 0.03 | 0.2 | | `google/gemma-4-31B-it` | Flex | 0.049 | 0.021 | 0.14 | | `google/gemma-4-26B-A4B-it` | Now | 0.072 | 0.0144 | 0.216 | | `google/gemma-4-26B-A4B-it` | Priority | 0.054 | 0.0108 | 0.162 | | `google/gemma-4-26B-A4B-it` | Standard | 0.036 | 0.0072 | 0.108 | | `google/gemma-4-26B-A4B-it` | Flex | 0.0252 | 0.00504 | 0.0756 | | `MiniMaxAI/MiniMax-M3` | Now | 0.2 | 0.06 | 1.2 | | `MiniMaxAI/MiniMax-M3` | Priority | 0.15 | 0.045 | 0.9 | | `MiniMaxAI/MiniMax-M3` | Standard | 0.1 | 0.03 | 0.6 | | `MiniMaxAI/MiniMax-M3` | Flex | 0.07 | 0.021 | 0.42 | For your organization's available models and negotiated rates, use the separate authenticated [`GET /v1/models` API](/api-reference/models-api/list-supported-models). The public docs MCP does not call that API. # Prompt Caching Source: https://docs.valarhq.ai/prompt-caching How Valar reuses shared prompt prefixes, and what you can do to hit the cache more often ## How it works Requests that share an opening are common: an agent that sends the same system prompt and tool definitions on every run, a chat that resends its history on every turn, a document a user asks several questions about. Paying full price to reprocess that shared text every time is waste, and it is usually the largest part of the bill. Prompt caching removes it. Valar matches the shared prefix for you and bills those tokens at the **cached** rate, a fraction of the input rate on every model that offers one — so the repeated part of your prompt gets dramatically cheaper the more you send it. Responses also start sooner, since the model skips work it has already done. Caching is on by default and there is nothing to configure. The rest of this page is about hitting it more often. Two things decide whether a request hits: 1. **A prefix matches byte for byte, from the very start of the prompt.** Matching stops at the first byte that differs. 2. **A cache is local to the machine that served the request.** A later request hits only if it lands on that same machine. You control the first completely. The second is what [session affinity](#routing-to-a-warm-cache) is for. ## Keep your prompt prefix stable Because matching starts at the beginning of the prompt and stops at the first difference, the order of your prompt decides how much of it can be reused. Put what stays the same at the front — system prompt, tool definitions, few-shot examples, a long document — and what varies at the end. Usually only the user's question changes, so it belongs last. Dynamic content at the *start* of a prompt kills your hit rate. A timestamp, a request id, or a session id in the opening line changes the first bytes on every request, so nothing behind it can be reused — including a system prompt that never changed at all. ```python ❌ Don't — the timestamp changes the first bytes, so nothing after it is reused theme={"system"} response = client.responses.create( model="zai-org/GLM-5.3", instructions=f"Current time: {datetime.now()}\n\n{SYSTEM_PROMPT}", input=question, ) ``` ```python ✅ Do — the prefix is identical every time; only the tail varies theme={"system"} response = client.responses.create( model="zai-org/GLM-5.3", instructions=SYSTEM_PROMPT, input=f"{question}\n\nCurrent time: {datetime.now()}", ) ``` The same mistake wears other clothes: * A request id or trace id interpolated into the system prompt. * Tool definitions serialized in a different order between runs, or with object keys emitted in a different order. * Retrieved documents or a conversation summary placed before the system prompt instead of after it. * Few-shot examples sampled at random per request. Caching pays in proportion to how long and how repeated your prefix is. A short prompt may not cache at all. If a prompt you believe is identical reports no cached tokens on its second call, diff two real request bodies before changing anything else. The difference is almost always in the first few lines. ## Routing to a warm cache The second property is the one you cannot fix by editing your prompt. A cache belongs to the machine that built it, so two requests share a cache only if the same machine serves them. Turns of one conversation usually land together. Separate runs of the same agent usually do not — which is why an agent with a long shared system prompt can cache well *within* a run and hardly at all *across* runs. Send an `X-Session-Affinity` header to tell us which requests belong together. We use it to steer them toward a machine that may already have their prefix hot, instead of picking one blind. It is a hint, not a pin: capacity, load, and eviction all still get a vote, and a busy machine may hand your request to another one. ```python OpenAI SDK theme={"system"} from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key="YOUR_VALAR_API_KEY", # or read OPENAI_API_KEY from the environment ) response = client.responses.create( model="zai-org/GLM-5.3", instructions=SYSTEM_PROMPT, input="Review pull request #4127.", extra_headers={"X-Session-Affinity": "pr-review-agent"}, ) ``` ```python Anthropic SDK theme={"system"} from anthropic import Anthropic client = Anthropic( auth_token="YOUR_VALAR_API_KEY", base_url="https://api.valarhq.ai", # no /v1 — the SDK appends it ) message = client.messages.create( model="zai-org/GLM-5.3", max_tokens=2048, system=SYSTEM_PROMPT, messages=[{"role": "user", "content": "Review pull request #4127."}], extra_headers={"X-Session-Affinity": "pr-review-agent"}, ) ``` ```bash cURL theme={"system"} curl https://api.valarhq.ai/v1/responses \ -H "Authorization: Bearer $VALAR_API_KEY" \ -H "Content-Type: application/json" \ -H "X-Session-Affinity: pr-review-agent" \ -d '{ "model": "zai-org/GLM-5.3", "instructions": "...your shared system prompt...", "input": "Review pull request #4127." }' ``` Available on `/v1/responses`, `/v1/chat/completions`, and `/v1/messages`. There is nothing to pre-register and no value to look up: the string is yours to pick, up to 1024 bytes. Values are scoped to your organization, so yours never collide with another customer's. The header is optional. Omit it and nothing changes: requests are routed as they are today, and automatic prefix caching still applies. ### Choosing a value Use **one value per group of requests that share a prefix**. For an agent, that is usually one value for the agent itself — every run of a PR reviewer that opens with the same system prompt and tools sends `pr-review-agent`. For a chat product it is usually one value per conversation or per user, because there the repeated prefix is the history rather than the system prompt. Group by whatever the requests actually share. Do not use a per-request value. A fresh UUID on every call puts every request in its own group, so nothing shares a cache and you end up worse off than sending no header at all. The same goes for routing traffic randomly. ### Rules Affinity only influences where a request goes. It cannot create a hit that would not otherwise exist — if the front of your prompt varies, it misses wherever it lands, so fix the prompt first. And even with a stable prefix, a hit is never certain: caches evict their oldest entries, machines restart, and a machine under load may not be the one that serves you. Neither caching nor affinity changes model selection, output, or the price of a token. They change only how many of your input tokens qualify for the cached rate. ## Measuring your hit rate Every response reports its cached input tokens, in the shape its dialect uses: | Endpoint | Field | | - | - | | `/v1/responses` | `usage.input_tokens_details.cached_tokens` | | `/v1/chat/completions` | `usage.prompt_tokens_details.cached_tokens` | | `/v1/messages` | `usage.cache_read_input_tokens` | They also appear in the [usage endpoints](/usage-endpoints), which is the easier place to look at a hit rate across many requests. Measure before and after a change rather than assuming. Cached tokens bill at the cached input rate on the [pricing page](/pricing). ## Batch requests Batch requests take no affinity hint — a batch carries no per-request header. Prefix stability still applies to every item, so the ordering advice above is worth following there too. # Quickstart Source: https://docs.valarhq.ai/quickstart Start using Valar with OpenAI clients in Python, TypeScript, or cURL The endpoint is the [OpenAI Responses API](/api-reference/responses-api/create-a-response) at `/v1/responses`, so any OpenAI-compatible client works after you change two settings: the base URL and the key. This walkthrough runs one realistic task end to end: classifying an inbound support ticket and drafting a reply. You send it as a background job, then retrieve the result once Valar finishes. The same pattern scales from this single call to the thousands of concurrent requests an agent fans out at runtime. **For AI agents** — if you are an agent migrating an existing app to Valar, follow the condensed instructions in [Migrating to Valar](/migrate#let-your-coding-agent-do-it). These docs are machine-readable: see [LLM & agent access](/llm-access). Sign up at the [Valar Dashboard](https://app.valarhq.ai) — new workspaces are approved before activation — then create an API key from the dashboard. Keep your existing OpenAI client. Change two settings: 1. Set the base URL to `https://api.valarhq.ai/v1` 2. Pass the API key as a bearer token. `https://api.valarhq.ai/v1` is the global endpoint, which geo-steers to the nearest region automatically. To pin a specific region, use a regional endpoint instead — see [Geo Control](/geo-control). ```python Python theme={"system"} import os from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key=os.environ["VALAR_API_KEY"], ) ``` ```ts TypeScript theme={"system"} import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.valarhq.ai/v1", // or read OPENAI_BASE_URL from the environment apiKey: process.env.VALAR_API_KEY, }); ``` ```bash cURL theme={"system"} export VALAR_API_KEY=sk_... # every call below is a plain HTTPS request with a bearer token: # -H "Authorization: Bearer $VALAR_API_KEY" ``` Setting `background` returns a response id immediately rather than holding the connection open. For one ticket this is convenient; across a queue of them it is what lets the work run concurrently. Use a model from the [Models](/models) page - here, `zai-org/GLM-5.2`. ```python Python theme={"system"} ticket = ( "Subject: Charged twice this month\n" "I see two identical $49 charges on the 3rd. Can you refund one and " "tell me why it happened?" ) started = client.responses.create( model="zai-org/GLM-5.2", instructions=( "You are a support triage agent. Classify the ticket as one of " "billing, technical, or account, then draft a short reply." ), input=ticket, background=True, # returns a response id right away to poll ) print("Queued:", started.id) ``` ```ts TypeScript theme={"system"} const ticket = "Subject: Charged twice this month\n" + "I see two identical $49 charges on the 3rd. Can you refund one and " + "tell me why it happened?"; const started = await client.responses.create({ model: "zai-org/GLM-5.2", instructions: "You are a support triage agent. Classify the ticket as one of " + "billing, technical, or account, then draft a short reply.", input: ticket, background: true, // returns a response id right away to poll }); console.log("Queued:", started.id); ``` ```bash cURL theme={"system"} curl https://api.valarhq.ai/v1/responses \ -H "Authorization: Bearer $VALAR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "zai-org/GLM-5.2", "instructions": "You are a support triage agent. Classify the ticket as one of billing, technical, or account, then draft a short reply.", "input": "Subject: Charged twice this month\nI see two identical $49 charges on the 3rd. Can you refund one and tell me why it happened?", "background": true }' # → {"id": "resp_...", "status": "queued", ...} ``` The create call hands back a response id and a status of `queued` or `in_progress`. Retrieve that id until it reaches `completed`, then read `output_text`. In production you can replace this poll loop with a [webhook](/webhooks) so you aren't holding a thread per job. ```python Python theme={"system"} import time response = started while response.status in {"queued", "in_progress"}: time.sleep(2) response = client.responses.retrieve(response.id) if response.status != "completed": raise RuntimeError(f"Task ended as {response.status}") print(response.output_text) ``` ```ts TypeScript theme={"system"} let response = started; while (response.status === "queued" || response.status === "in_progress") { await new Promise((resolve) => setTimeout(resolve, 2000)); response = await client.responses.retrieve(started.id); } if (response.status !== "completed") { throw new Error(`Task ended as ${response.status}`); } console.log(response.output_text); ``` ```bash cURL theme={"system"} # poll until "status" is "completed", then read output_text curl https://api.valarhq.ai/v1/responses/resp_YOUR_RESPONSE_ID \ -H "Authorization: Bearer $VALAR_API_KEY" ``` ## Going further A single triaged ticket is the unit; an agent is many of them in a loop. From here: * See [**Models**](/models) for the full list of supported models, and [**Pricing**](/pricing) for per-token rates. * Turn this into a tool-using agent that looks up the customer's billing record before replying - see [Building a tool-calling agent](/agents). * Run the same task over a backlog of tickets at once with [Requests at scale](/requests_at_scale), choosing a [completion window](/inference-modes#completion-windows) per the latency you can tolerate. Questions about a specific workload can go to [support@valarhq.ai](mailto:support@valarhq.ai). # Sending Requests at Scale Source: https://docs.valarhq.ai/requests_at_scale Move tens of thousands of requests through Valar without serializing them one by one A loop that fires requests in sequence stalls the moment your workload grows past a few hundred calls. Valar gives you two ways to clear large volumes: fan out many concurrent Responses API calls in background mode, or hand a single bundle to the Batch API and let Valar work through it. This page covers both, starting with how to choose. ## Which path fits your workload You want results streaming back as each request finishes, need per-request control, or care about latency on individual items. You manage concurrency on your client. You have a fixed set of requests and would rather submit once and collect results later. Valar owns the queue; you poll for status. The running example throughout is sentiment scoring for a backlog of product reviews, using `zai-org/GLM-5.2`. Both paths below set `max_output_tokens`. On a [reasoning model](/models#reasoning-models-and-output-caps) that cap covers hidden reasoning as well as the answer, so a large share of results settle `incomplete` with a perfectly good answer attached. Collect on `incomplete_details.reason`, not on `status` alone. ## Fan out with the Responses API Submit each request individually with `background=True` and run many in flight at once. For workloads of roughly 1,000 requests and up, four practices keep it stable: * **Use** `AsyncOpenAI `**with** `DefaultAioHttpClient()`**.** Under high concurrency the SDK's aiohttp backend outperforms the default `httpx` transport. * **Bound concurrency with an** `asyncio.Semaphore`**.** A fixed limit (200 is a sound starting point) caps simultaneous connections so you never exhaust them. * **Submit all requests behind the semaphore, collect the response IDs, then poll in a second pass.** Separating submission from polling keeps both phases simple. * **Attach an** `Idempotency-Key `**per request** so a retry after a transient failure replays the reservation rather than paying for inference twice. See [Idempotent Requests](/idempotency). Install the dependencies first: ```bash theme={"system"} pip install tqdm 'openai[aiohttp]' ``` The script submits every review concurrently, gathers the response IDs, and then polls each one to completion before printing its score: ```python theme={"system"} import argparse import asyncio from tqdm import tqdm from openai import AsyncOpenAI, DefaultAioHttpClient REVIEW_TEMPLATE = ( "Classify the sentiment of this product review as positive, negative, or " "neutral. Reply with one word.\n\nReview {i}: {text}" ) async def main(num_requests: int, max_output_tokens: int, model: str): async with AsyncOpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key="YOUR_VALAR_API_KEY", # or read OPENAI_API_KEY from the environment http_client=DefaultAioHttpClient(), ) as client: submit_sem = asyncio.Semaphore(200) submit_bar = tqdm(total=num_requests, desc="Submitting") async def submit(i): content = REVIEW_TEMPLATE.format( i=i, text=f"Sample review body number {i} goes here." ) async with submit_sem: response = await client.responses.create( model=model, input=[{"role": "user", "content": content}], max_output_tokens=max_output_tokens, background=True, ) submit_bar.update(1) return response.id response_ids = list(await asyncio.gather(*[submit(i) for i in range(num_requests)])) submit_bar.close() print(f"Submitted {len(response_ids)} requests") # Second pass: poll each background response until it completes. completed = {} poll_sem = asyncio.Semaphore(200) poll_bar = tqdm(total=len(response_ids), desc="Polling") async def poll(response_id): for _ in range(3600): # up to ~1 hour async with poll_sem: response = await client.responses.retrieve(response_id) if response.status in {"queued", "in_progress"}: await asyncio.sleep(1) continue completed[response_id] = response poll_bar.update(1) return print(f"Timed out waiting for {response_id}") await asyncio.gather(*[poll(rid) for rid in response_ids]) poll_bar.close() for response_id in response_ids: resp = completed[response_id] text = "".join( c.text for item in resp.output for c in getattr(item, "content", []) or [] if getattr(c, "type", None) == "output_text" ) print(f"{response_id}: {text.strip()}") if __name__ == "__main__": parser = argparse.ArgumentParser() parser.add_argument("--num-requests", type=int, default=20000) parser.add_argument("--max-output-tokens", type=int, default=16) parser.add_argument("--model", type=str, default="zai-org/GLM-5.2") args = parser.parse_args() asyncio.run( main( num_requests=args.num_requests, max_output_tokens=args.max_output_tokens, model=args.model, ) ) ``` Treat a semaphore of 200 as the baseline. Lower it if connection errors or timeouts appear; raise it when you have headroom and want submissions to move through faster. ## Hand off a bundle with the Batch API The Batch API takes up to 10,000 requests per batch. You upload the requests as a JSONL file, create a batch that points at it, poll until it settles, and download the results file. It follows the OpenAI Batch API, so `client.batches.create()` code you already have works unchanged against Valar's base URL. Write one request per line to a `.jsonl` file and upload it with `POST /files` (`purpose: batch`). Each line carries your `custom_id`, `"method": "POST"`, the target `url`, and the request `body`. A line can be up to 28 MiB - enough for a maximum-size 20 MB image, which takes about 26.7 MiB once base64-encoded. That cap is about transport, not the model. Each request also has to fit the context window of the model it names, and windows vary by more than an order of magnitude across the catalog - so a line well under 28 MiB can still be too long for the model you picked, which is easiest to hit when you pack base64 images. Read `context_length` from [GET /models](/api-reference/models-api/list-supported-models) and size your requests against it: a request over the window is accepted with the file and fails when it runs, on that line alone. `POST /batches` with the returned `input_file_id`, the `endpoint` the requests target, and `completion_window: "24h"` (the OpenAI-compatible value; Valar schedules against its own promised SLAs). Call `GET /batches/{batch_id}` until `status` is terminal. Progress can arrive in large steps, so a quiet `request_counts` is normal - every batch settles within its window. Download `output_file_id` with `GET /files/{file_id}/content` - one JSONL result line per request. To peek at a single result early, `GET /batches/{batch_id}/{custom_id}` works while the batch is still running (a Valar extension). To review everything you've submitted previously, call `GET /batches`. A line that fails validation - unparseable, missing a `model`, or over the 28 MiB line limit - is rejected on its own and counted in `request_counts.failed`; the rest of the batch keeps running. Batch is the lowest-cost, highest-latency lane. When latency matters, fan out through the [Responses API](#fan-out-with-the-responses-api) instead. See [Inference modes](/inference-modes) for the tradeoff. This version scores the same review backlog as one batch: ```python theme={"system"} import json import time from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", api_key="YOUR_VALAR_API_KEY", ) # 1. Write the requests file - one request per line with open("requests.jsonl", "w") as f: for i in range(100): f.write(json.dumps({ "custom_id": f"review-{i}", "method": "POST", "url": "/v1/chat/completions", "body": { "model": "zai-org/GLM-5.2", "max_tokens": 16, "messages": [{ "role": "user", "content": ( "Classify the sentiment of this review as positive, " f"negative, or neutral. Review {i}: Sample review body {i}." ), }], }, }) + "\n") # 2. Upload it and create the batch batch_input = client.files.create(file=open("requests.jsonl", "rb"), purpose="batch") batch = client.batches.create( input_file_id=batch_input.id, endpoint="/v1/chat/completions", completion_window="24h", ) print(f"Created batch {batch.id}") # 3. Poll until the batch settles while batch.status not in ("completed", "failed", "expired"): time.sleep(30) batch = client.batches.retrieve(batch.id) counts = batch.request_counts print(f"{batch.status}: {counts.completed + counts.failed}/{counts.total}") # 4. Download the results - one JSONL line per request for line in client.files.content(batch.output_file_id).text.splitlines(): result = json.loads(line) body = result["response"]["body"] text = body["choices"][0]["message"]["content"] print(f"{result['custom_id']}: {text.strip()}") ``` # Structured outputs Source: https://docs.valarhq.ai/structured-outputs Constrain a model's response to a JSON schema so you get parseable, predictable JSON ## What structured outputs are A structured output is a model response that is forced to match a [JSON Schema](https://json-schema.org/) you supply. Instead of asking for JSON in the prompt and hoping the model complies, you hand Valar the shape you want and the response comes back as JSON that conforms to it. Reach for structured outputs when the response feeds code rather than a human: * **Reliable parsing**: every response is valid JSON with the fields you declared, so `json.loads` never trips over prose, code fences, or trailing commentary. * **Extraction**: pull typed fields out of unstructured text, such as turning a support email into `{ category, priority, summary }`. * **Downstream automation**: route, store, or act on the result without a brittle post-processing step. Structured outputs are supported on the Responses, Chat Completions, and Messages APIs. The field you set differs per API, but the JSON Schema you pass is the same. ## Defining a schema per API Each tab points at the same base URL, `https://api.valarhq.ai/v1`, and authenticates with `Authorization: Bearer $VALAR_API_KEY`. The example extracts a support ticket into a fixed shape. On the [Responses API](/api-reference/responses-api/create-a-response), set `text.format` to a `json_schema` object. The schema lives directly under `schema`, alongside a `name` and `strict`. `text.format.type: "json_object"` is **not** supported on the Responses API. Use `json_schema` to constrain the output. ```python theme={"system"} import json from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key="YOUR_VALAR_API_KEY", # or read OPENAI_API_KEY from the environment ) ticket_schema = { "type": "object", "properties": { "category": { "type": "string", "enum": ["billing", "bug", "feature_request", "account", "other"], }, "priority": {"type": "string", "enum": ["low", "medium", "high", "urgent"]}, "summary": {"type": "string"}, }, "required": ["category", "priority", "summary"], "additionalProperties": False, } response = client.responses.create( model="moonshotai/Kimi-K2.7", input="Triage this ticket: 'I was charged twice for my May invoice and need a refund ASAP.'", text={ "format": { "type": "json_schema", "name": "support_ticket", "schema": ticket_schema, "strict": True, } }, ) ticket = json.loads(response.output_text) print(ticket["category"], ticket["priority"]) ``` On [Chat Completions](/api-reference/chat-completions-api/create-a-chat-completion), set `response_format` to a `json_schema` object. The schema is nested one level deeper, under `json_schema.schema`. This API also accepts `{ "type": "json_object" }` for free-form JSON without a schema. ```python theme={"system"} import json from openai import OpenAI client = OpenAI( api_key="YOUR_VALAR_API_KEY", base_url="https://api.valarhq.ai/v1", ) ticket_schema = { "type": "object", "properties": { "category": { "type": "string", "enum": ["billing", "bug", "feature_request", "account", "other"], }, "priority": {"type": "string", "enum": ["low", "medium", "high", "urgent"]}, "summary": {"type": "string"}, }, "required": ["category", "priority", "summary"], "additionalProperties": False, } completion = client.chat.completions.create( model="moonshotai/Kimi-K2.7", messages=[ { "role": "user", "content": "Triage this ticket: 'I was charged twice for my May invoice and need a refund ASAP.'", } ], response_format={ "type": "json_schema", "json_schema": { "name": "support_ticket", "schema": ticket_schema, "strict": True, }, }, ) ticket = json.loads(completion.choices[0].message.content) print(ticket["category"], ticket["priority"]) ``` On the [Messages API](/anthropic-sdk) (`/v1/messages`), set `output_config.format` to a `json_schema` object. The schema lives directly under `schema`, alongside `name` and `strict` — the same flat shape the Responses API uses. The Anthropic Python SDK doesn't type this parameter, so pass it through `extra_body`. ```python theme={"system"} import json from anthropic import Anthropic client = Anthropic( auth_token="YOUR_VALAR_API_KEY", base_url="https://api.valarhq.ai", # no /v1 — the SDK appends it ) ticket_schema = { "type": "object", "properties": { "category": { "type": "string", "enum": ["billing", "bug", "feature_request", "account", "other"], }, "priority": {"type": "string", "enum": ["low", "medium", "high", "urgent"]}, "summary": {"type": "string"}, }, "required": ["category", "priority", "summary"], "additionalProperties": False, } message = client.messages.create( model="moonshotai/Kimi-K2.7", max_tokens=2048, messages=[ { "role": "user", "content": "Triage this ticket: 'I was charged twice for my May invoice and need a refund ASAP.'", } ], extra_body={"output_config": {"format": { "type": "json_schema", "name": "support_ticket", "schema": ticket_schema, "strict": True, }}}, ) # The model may emit thinking blocks first — extract the text block and parse it. text = next(block.text for block in message.content if block.type == "text") ticket = json.loads(text) print(ticket["category"], ticket["priority"]) ``` ## Enforcing the schema with strict Setting `strict: true` makes Valar enforce the schema during decoding, so the response is guaranteed to match the structure you declared - required fields are present, types line up, and `enum` values stay within the allowed set. Without it, the schema is treated as guidance and the model may drift. For strict mode to hold, your schema must be one Valar can enforce. A schema that is malformed or uses an unsupported construct returns `400 invalid_request_error` rather than running the request, so validate the shape before you ship it. You rarely need to hand-write the schema. Generate it from a [Pydantic](https://docs.pydantic.dev/) model with `Model.model_json_schema()`, or from a [Zod](https://zod.dev/) schema with a JSON Schema converter, then drop the result into the `schema` field. ```python theme={"system"} from pydantic import BaseModel from typing import Literal class SupportTicket(BaseModel): category: Literal["billing", "bug", "feature_request", "account", "other"] priority: Literal["low", "medium", "high", "urgent"] summary: str ticket_schema = SupportTicket.model_json_schema() ``` ## See also * [Create a response](/api-reference/responses-api/create-a-response) - full `text.format` reference. * [Create a chat completion](/api-reference/chat-completions-api/create-a-chat-completion) - full `response_format` reference. * [Anthropic SDK & Claude Agent SDK](/anthropic-sdk) - the `/v1/messages` surface and `output_config.format` usage. # API Support Matrix Source: https://docs.valarhq.ai/support What each Valar inference API accepts today. Valar gives you four endpoints. Three are inference surfaces shaped after APIs you already know, and one batches work asynchronously: | API | Endpoint | Maturity | | - | - | - | | OpenAI **Responses** | `POST /v1/responses` | Stable | | OpenAI **Chat Completions** | `POST /v1/chat/completions` | Stable | | Anthropic **Messages** | `POST /v1/messages` | Stable | | **Batch** | `POST /v1/batches` | Private Preview | The same [models](/models) and [completion windows](/completion-windows) work across all three inference surfaces. The Batch API layers on top: it wraps a large set of [Responses API](#responses-api) calls into one asynchronous job. ## Behavior shared across every API Before the per-API detail, a few rules hold no matter which surface you call: * **Streaming is available on Chat Completions and Messages.** Pass `stream: true` and you get Server-Sent Events (`chat.completion.chunk` on Chat Completions, Anthropic-style events on Messages); add `stream_options.include_usage` on Chat Completions for a closing usage chunk. The Responses API rejects `stream: true`. For long jobs, use `background: true` on the Responses API and poll or wait on [webhooks](/webhooks). * **Completion windows steer scheduling and price.** Set `metadata.completion_window` (or the `X-Valar-Completion-Window` header) to `"asap"` (the **Now** tier), `"priority"`, `"standard"`, or `"flex"` (background only). See [Completion windows](/completion-windows) and [Pricing](/pricing). * **Webhooks fire on completion.** Set `metadata.completion_webhook` to receive a POST when processing finishes. See [Webhooks](/webhooks). * **Responses are always stored.** `store: false` is unsupported. ## Inference APIs Each accordion below lists what the API accepts and what it rejects. Open the one that matches the SDK you're using.
Recommended OpenAI Responses format OpenAI SDK compatible API reference →
This is the surface we recommend reaching for first. **Supported** | Feature | Details | | - | - | | **Core parameters** | `model`, `input` (string or message array), `max_output_tokens`, `temperature`, `top_p`, `user`, `prompt_cache_key` | | **Structured outputs** | `text.format` with `type: "text"` or `type: "json_schema"` | | **Reasoning** | `reasoning.effort` (`none` / `minimal` / `low` / `medium` / `high` / `xhigh`), `reasoning.generate_summary` (`auto` / `concise` / `detailed`) | | **Function tools** | `tools` with `type: "function"` - client-side function calling with `name`, `description`, `parameters`, `strict` | | **Custom tools** | `tools` with `type: "custom"` | | **Tool choice** | `tool_choice`: `"none"`, `"auto"`, `"required"`, or a specific function/custom tool | | **Background mode** | `background: true` returns `202` immediately; poll with `GET /v1/responses/{id}` | | **Prompt cache routing** | `prompt_cache_key` is an optional routing hint for requests that share a large prompt prefix | | **Image input** | `input_image` content blocks on [multimodal models](/models). Non-multimodal models accept text only. | | **Output logprobs** | `include: ["message.output_text.logprobs"]` returns one logprob per output token (best effort; omitted for models served via a proxy that does not return logprobs). | **Not yet supported** | Feature | Notes | | - | - | | **Streaming** | `stream: true` is rejected. Every response comes back as one JSON object. | | **Instructions** | `instructions` is unsupported. Put system messages straight into `input`. | | **Conversation chaining** | `previous_response_id` and `conversation` are unsupported. Resend the full input on each call. | | **Prompt templates** | The `prompt` parameter is unsupported. | | **Server-side tools** | `web_search`, `file_search`, `code_interpreter`, `computer_use`, `mcp`, `image_generation`, `shell`, `apply_patch` are unsupported. | | **Multimodal input** | Audio and file input blocks are unsupported. Image input works on multimodal models (see above). | | **Include** | Accepted for compatibility when passed as an array of strings. A request is rejected if it includes `reasoning.encrypted_content`, `web_search_call.action.sources`, `code_interpreter_call.outputs`, `computer_call_output.output.image_url`, or `file_search_call.results`. | | **Truncation** | `"disabled"` is the only accepted value; custom truncation strategies are unsupported. | | **Parallel tool calls** | `parallel_tool_calls` is unsupported. | | **json\_object format** | `text.format.type: "json_object"` is unsupported. Reach for `"json_schema"` instead. | | **Service tier** | `"auto"` is the only accepted value. Use `metadata.completion_window` to govern response timing instead. | | **Delete / cancel** | `DELETE /v1/responses/{id}` and the cancel endpoints are not implemented. |
OpenAI Chat Completions format OpenAI SDK compatible API reference →
This surface streams `chat.completion.chunk` events. **Supported** | Feature | Details | | - | - | | **Core parameters** | `model`, `messages`, `max_completion_tokens`, `temperature`, `top_p`, `user` | | **Message roles** | `system`, `user`, `assistant`, `tool`, `function` (deprecated), `developer` | | **Structured outputs** | `response_format` with `type: "text"`, `"json_object"`, or `"json_schema"` | | **Reasoning** | `reasoning_effort` (`none` / `minimal` / `low` / `medium` / `high` / `xhigh`) | | **Function tools** | `tools` with `type: "function"` - standard `{type, function: {name, description, parameters, strict}}` format | | **Custom tools** | `tools` with `type: "custom"` | | **Tool choice** | `tool_choice`: `"none"`, `"auto"`, `"required"`, or a specific function/custom tool | | **Parallel tool calls** | `parallel_tool_calls` is passed through | | **Metadata** | `metadata` with string key-value pairs, including `completion_window` and `completion_webhook` | | **Streaming** | `stream: true` returns Server-Sent Events (`chat.completion.chunk`); `stream_options.include_usage` adds a final usage chunk. Reasoning is streamed as `reasoning_content` deltas and tool calls are emitted atomically. | | **Image input** | `image_url` content parts on [multimodal models](/models). Non-multimodal models accept text only. | **Not yet supported** | Feature | Notes | | - | - | | **Multiple choices** | `n` is required to be `1`. | | **Multimodal content** | Audio (`input_audio`) content parts are unsupported. Image (`image_url`) input works on multimodal models (see above). | | **Sampling controls** | `frequency_penalty`, `presence_penalty`, `logit_bias`, `stop`, `seed`, `top_logprobs`, `logprobs`, `verbosity` are unsupported. | | **Audio modality** | Neither `audio` nor `modalities: ["audio"]` is supported. | | **Predicted output** | `prediction` is unsupported. | | **Web search** | `web_search_options` is unsupported. | | **Service tier** | `"auto"` is the only accepted value. | | **CRUD endpoints** | The `GET`, `POST`, and `DELETE` operations on stored completions are not implemented. | | **Deprecated fields** | `max_tokens`, `functions`, and `function_call` are rejected; switch to their modern replacements. | **What comes back** * A response always carries exactly one choice (`n=1`). * `finish_reason` is either `"stop"` or `"tool_calls"` - values such as `"length"` and `"content_filter"` are never returned. * Neither `system_fingerprint` nor `service_tier` appears in responses. * `logprobs` is always `null`.
Anthropic Messages format Anthropic SDK compatible API reference →
Valar's Anthropic-compatible surface. The [Anthropic Python SDK](/anthropic-sdk) and the Claude Agent SDK work against it with a base URL and key change. **Supported** | Feature | Details | | - | - | | **Core parameters** | `model`, `max_tokens` (required), `messages`, `system` (string or content blocks) | | **Sampling** | `temperature` (0–1), `top_p` (0–1), `top_k` | | **Streaming** | `stream: true` returns Anthropic SSE events (`message_start`, `content_block_delta`, `message_delta`, `message_stop`) | | **Extended thinking** | `thinking` with `type: "enabled"` and `budget_tokens`, or `output_config.effort` (`low` / `medium` / `high` / `xhigh` / `max`) | | **Structured outputs** | `output_config.format` with `type: "json_schema"` (translated to the Responses API's `text.format`); see [Structured outputs](/structured-outputs) | | **Function tools** | `tools` with `name`, `description`, `input_schema`; `tool_choice` (`auto` / `any` / `{type:"tool",name}`) | | **Tool-result blocks** | `tool_result` content blocks in subsequent `user` messages | | **Stop sequences** | `stop_sequences` | | **Image input** | `image` content blocks on [multimodal models](/models). Non-multimodal models accept text only. | | **Metadata** | `metadata` with string key-value pairs, including `completion_window` and `completion_webhook` | | **Completion window header** | `X-Valar-Completion-Window` (`asap` / `priority` / `standard` / `flex`) as a fallback when `metadata.completion_window` is absent | | **Count tokens** | `POST /v1/messages/count_tokens` returns an estimated token count | **Not yet supported** | Feature | Notes | | - | - | | **Multimodal content** | Document (`document`) content blocks are unsupported. Image input works on multimodal models (see above). | | **Service tier** | `service_tier` is unsupported. Use `metadata.completion_window` or the `X-Valar-Completion-Window` header. | | **Inference geo** | `inference_geo` is unsupported. | | **Batches** | `POST /v1/messages/batches` and its related endpoints are not implemented. | **What comes back** * `stop_reason` is `end_turn`, `tool_use`, or `max_tokens`. (A run halted on a `stop_sequences` match comes back as `end_turn`, not `stop_sequence`.) * The response content is one or more blocks: `text`, `thinking`, and `tool_use`. * `usage` includes `input_tokens`, `output_tokens`, `cache_read_input_tokens`, and `cache_creation_input_tokens`. **Authenticating with the Anthropic SDK** Valar accepts both `Authorization: Bearer ` (the OpenAI convention) and the Anthropic `x-api-key: ` header. With the Anthropic SDK, either `auth_token` or `api_key` works: ```python theme={"system"} from anthropic import Anthropic client = Anthropic( auth_token="YOUR_VALAR_API_KEY", # sends Authorization: Bearer base_url="https://api.valarhq.ai", # no /v1 — the SDK appends it ) ``` The `anthropic-version` header is neither required nor inspected, and errors come back in the OpenAI-style error envelope format. See [Anthropic SDK & Claude Agent SDK](/anthropic-sdk) for end-to-end examples.
## Batch API The Batch API runs large volumes of requests asynchronously, following the OpenAI Batch API: upload a JSONL request file, create a batch that points at it, poll for status, and download the results file. Batches target `/v1/chat/completions` or `/v1/responses` - you can't batch `/v1/messages` today. One batch takes up to 10,000 requests; single results are also fetchable early by `custom_id`, a Valar extension. For the end-to-end workflow, see [Sending Requests at Scale](/requests_at_scale); for the request and response schemas, see the [Batch API reference](/api-reference/batches-api/create-a-batch). # Overview Source: https://docs.valarhq.ai/usage Query spend, balance, and operational metrics for your account programmatically Need your Valar usage figures inside a script, CLI tool, or custom dashboard? The Usage API exposes them directly. It authenticates with the very same key that powers your calls to `api.valarhq.ai`. ## Base URL ``` https://api.valarhq.ai ``` ## Authentication Pass your API key as a Bearer token in the `Authorization` header on every request. The key here is identical to the one used by the inference API at `api.valarhq.ai`. Generate and manage keys from the [Valar dashboard](https://app.valarhq.ai). ```python theme={null} theme={"system"} import requests headers = {"Authorization": "Bearer YOUR_VALAR_API_KEY"} resp = requests.get("https://api.valarhq.ai/v1/usage", headers=headers) print(resp.json()) ``` ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ https://api.valarhq.ai/v1/usage | jq ``` The complete reference for each `GET /v1/usage` route, covering parameters, response shapes, error formats, and headers. ## Organization and workspace costs By default, `/v1/usage` and `/v1/usage/breakdown` include **all workspaces in your organization**, not just the workspace that owns your API key. * Add `workspace_id=ws_...` to either endpoint to filter costs to one workspace. * Add `group_by=workspace` to `/v1/usage/breakdown` to get workspace IDs, names, totals, and per-model costs in one response. ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/breakdown?range=7d&group_by=workspace" \ | jq '.workspaces[] | {workspace_id, workspace_name, dollars: (.total / 100)}' ``` To compare with the dashboard's **This workspace** view, filter the API request to the same workspace and UTC date range. For an unfiltered API request, compare with **All workspaces** instead. All API monetary values are integer **cents**; divide by 100 to get dollars. See [workspace filtering and breakdowns](/usage-endpoints#workspace-filtering-and-breakdowns) for examples and response fields. These options apply to the two spend endpoints, not the activity or coding endpoints. Expect a small lag in usage data - it is not reported in real time. # Endpoints Source: https://docs.valarhq.ai/usage-endpoints Routes for spend, balance, activity, time-series usage metrics, and per-key ValarCode analytics ## GET /v1/usage Summarizes spending and balance over the time range you request. By default, spend includes all workspaces in your organization, regardless of which workspace owns your API key. | Parameter | Values | Default | Description | | - | - | - | - | | `range` | `24h`, `7d`, `30d`, `period` | `30d` | Time window. `period` = the current billing period (a rolling 30 days today). | | `workspace_id` | Workspace ID | All organization workspaces | Filters spend and token metrics, including `prior_period`, to one workspace. | ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage?range=7d" | jq ``` ```json theme={null} theme={"system"} { "object": "usage.summary", "range": "7d", "period_spend": 5432, "burn_rate": 776, "balance": 0, "balance_unavailable": true, "plan_type": "self_serve", "model_count": 3, "tokens": { "total": 1000000, "input": 600000, "output": 350000, "cached": 50000 }, "avg_cost_per_day": 776, "sla_mix": { "asap": 0.3, "standard": 0.7 }, "prior_period": { "period_spend": 4000, "model_count": 2, "tokens": { "total": 800000, "input": 500000, "output": 280000, "cached": 20000 }, "avg_cost_per_day": 571 } } ``` Every monetary value is expressed in **cents**. The `prior_period` block mirrors the same-length window directly preceding your requested range, so you can compare the two. When you supply `workspace_id`, the response echoes it as a top-level field. `plan_type` still describes your organization, not a separate workspace plan. Workspace grouping is available on `/v1/usage/breakdown`, not this endpoint. The keys in `sla_mix` are [completion window](/inference-modes#completion-windows) names: `asap` (the **Now** tier), `standard`, and `flex`. `balance`, `balance_unavailable`, and `days_remaining` describe a prepaid-credit balance. Until prepaid credits are available, `balance_unavailable` is always `true`, `balance` is `0`, and `days_remaining` is omitted - spend, burn rate, tokens, and the SLA mix are always reported. *** ## GET /v1/usage/breakdown Provides per-model rankings alongside time-series spend data - ideal for building charts. | Parameter | Values | Default | Description | | - | - | - | - | | `range` | `24h`, `7d`, `30d`, `period`, `day` | `30d` | Time window. | | `date` | `YYYY-MM-DD` | - | Required when `range=day`. Drills into hourly data for that date. | | `workspace_id` | Workspace ID | All organization workspaces | Filters the entire response to one workspace. | | `group_by` | `workspace` | Omitted | Adds a `workspaces` array with each workspace's spend, tokens, time series, and model rankings. | ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/breakdown?range=7d" | jq ``` ```json theme={null} theme={"system"} { "object": "usage.breakdown", "range": "7d", "granularity": "day", "data": [ { "timestamp": "2025-01-15", "total": 800, "models": { "zai-org/GLM-5.2": { "total": 500, "tokens": 50000, "input_tokens": 30000, "output_tokens": 15000, "cached_tokens": 5000 } }, "slas": { "standard": 800 } } ], "models": [ { "model": "zai-org/GLM-5.2", "total": 3500, "tokens": 350000, "input_tokens": 210000, "output_tokens": 105000, "cached_tokens": 35000, "slas": { "standard": 3500 }, "percentage": 0.65 } ] } ``` For 7d, 30d, and period ranges `granularity` is `"day"`; for a 24h range or a day drill-down it is `"hour"`. ### Workspace filtering and breakdowns To retrieve one workspace's costs for a UTC day: ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/breakdown?range=day&date=2026-09-07&workspace_id=ws_prod" \ | jq '.models[] | {model, dollars: (.total / 100)}' ``` Use the same `workspace_id` parameter on `/v1/usage` for a filtered summary, including its prior-period comparison: ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage?range=7d&workspace_id=ws_prod" | jq ``` To separate costs by workspace in a single response, use `group_by=workspace`. This also gives you the workspace IDs to use in filtered requests: ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/breakdown?range=day&date=2026-09-07&group_by=workspace" \ | jq '[.workspaces[] | { workspace_id, workspace_name, dollars: (.total / 100), models: [.models[] | {model, dollars: (.total / 100)}] }]' ``` For example, the command above can produce: ```json theme={null} theme={"system"} [ { "workspace_id": "ws_prod", "workspace_name": "Production", "dollars": 125, "models": [ { "model": "zai-org/GLM-5.3", "dollars": 100 }, { "model": "claude-opus-5", "dollars": 25 } ] }, { "workspace_id": "ws_dev", "workspace_name": "Development", "dollars": 10, "models": [ { "model": "zai-org/GLM-5.3", "dollars": 10 } ] } ] ``` The API response keeps the existing top-level `data` and `models` fields and adds `group_by: "workspace"` and `workspaces`. Each workspace contains: | Field | Description | | - | - | | `workspace_id` | Stable workspace ID. Use this for filters and joins, even if the workspace is renamed. | | `workspace_name` | Current workspace display name. | | `total` | Total spend over the requested range, in cents. | | `tokens` | Token counts: `total`, fresh `input`, `output`, and `cached`. | | `data` | The same time-bucket structure as top-level `data`, limited to this workspace. | | `models` | The same model-ranking structure as top-level `models`. Each `percentage` is relative to this workspace's spend. | Workspaces are sorted by `total` descending, then by `workspace_id`. Only workspaces with usage records in the selected range are included; a workspace with tokens but zero spend is still included. An empty result returns `workspaces: []`. BYOK usage remains excluded, as it is from the ungrouped spend response. You can combine `workspace_id` and `group_by=workspace`. The top-level totals and the workspace array then include only that workspace. The response echoes the filter in a top-level `workspace_id` field. Without either parameter, the response shape and organization-wide scope stay unchanged. A workspace filter can only narrow your authenticated organization's usage. An unknown workspace ID, a workspace outside your organization, or a workspace without usage returns an empty breakdown and zero summary usage metrics. An omitted or empty `workspace_id` includes all of your organization's workspaces. All monetary fields in the API response are integer cents. The `dollars` fields above are calculated by `jq`, not returned by the API. Values are rounded after aggregation, so summing separately rounded workspace or bucket totals can differ slightly from a combined total. Match the dashboard's workspace scope and UTC date range when comparing costs; usage data can also lag slightly. *** ## GET /v1/usage/activity Surfaces operational metrics - request counts, token throughput, latency, and a list of recent requests. | Parameter | Values | Default | Description | | - | - | - | - | | `limit` | `1`–`100` | `10` | Number of recent requests to return. | ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ https://api.valarhq.ai/v1/usage/activity | jq ``` ```json theme={null} theme={"system"} { "object": "usage.activity", "available": true, "requests": { "last_1m": 5, "last_1h": 120, "last_24h": 2000, "last_7d": 14000 }, "tokens": { "last_1m": 500, "last_1h": 12000, "last_24h": 200000, "last_7d": 1400000, "token_breakdown_1h": { "input": 8000, "output": 3500, "cached": 500 } }, "latency": { "avg_1m_ms": 1200, "avg_1h_ms": 1500 }, "recent_requests": [ { "response_id": "resp_abc123", "model": "zai-org/GLM-5.2", "sla": "standard", "status": "completed", "created_at": "2025-01-15T10:00:00Z", "updated_at": "2025-01-15T10:01:00Z" } ], "has_more": false } ``` A `has_more` value of `true` means more recent requests exist beyond the `limit` you asked for. Like the other endpoints, activity data lags slightly and is not delivered as a real-time stream. *** ## GET /v1/usage/activity/timeseries Returns bucketed series for either requests or tokens - use it to draw throughput charts. | Parameter | Values | Default | Description | | - | - | - | - | | `type` | `requests`, `tokens` | `requests` | What to chart. | | `range` | `1h`, `6h`, `24h` | `24h` | Time window. Bucket size: 1min / 5min / 1 hour respectively. | **Requests by model:** ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/activity/timeseries?type=requests&range=1h" | jq ``` ```json theme={null} theme={"system"} { "object": "usage.activity.timeseries", "type": "requests", "range": "1h", "available": true, "series": [ { "time_bucket": "2025-01-15T10:00:00Z", "model": "zai-org/GLM-5.2", "count": 5 }, { "time_bucket": "2025-01-15T10:01:00Z", "model": "zai-org/GLM-5.2", "count": 3 } ] } ``` **Token breakdown:** ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/activity/timeseries?type=tokens&range=24h" | jq ``` ```json theme={null} theme={"system"} { "object": "usage.activity.timeseries", "type": "tokens", "range": "24h", "available": true, "series": [ { "time_bucket": "2025-01-15T10:00:00Z", "total_tokens": 1000, "input_tokens": 600, "output_tokens": 350, "cached_tokens": 50 } ] } ``` *** ## ValarCode analytics The `/v1/usage/coding/*` endpoints report the traffic your [ValarCode](/valarcode/overview) keys send through the coding lane: which key, which engineer, which model was served, and what it cost. They accept the same `Authorization: Bearer` API key as the endpoints above (a coding key or a standard key) and always cover the whole organization. Unlike the endpoints above, they also accept the [organization automation credential](/valarcode/key-automation) as the bearer, so a provisioning service that manages keys can read their usage with the token it already holds. All three share the same window and filter parameters. ### Window Pick one of the two forms. Combining them returns `400`. | Parameter | Values | Default | Description | | - | - | - | - | | `range` | `1h`, `24h`, `7d`, `30d`, `period`, `mtd` | `30d` | Preset window ending now. `7d`, `30d`, and `period` are whole UTC days ending today; `mtd` is the current UTC calendar month to date; `1h` is the last rolling hour; `24h` is the last rolling 24 hours, not the current UTC day. | | `start`, `end` | `YYYY-MM-DD` | - | An explicit window of whole UTC days, both bounds inclusive. Both are required together; `start` must not be after `end`, `end` must not be after today, and the span is at most 366 days. The response echoes `"range": "custom"`. | ### Filters Every filter narrows the rows before they are aggregated, so the filters compose with each other and with any window. | Parameter | Description | | - | - | | `key_id` | One coding key. Use the `id` from the [key list](/valarcode/key-automation#list-provisioned-keys) or the dashboard. | | `cohort_id` | One routing cohort. | | `model` | One served model id, as it appears in `served_model`. | | `byok` | `true` restricts the view to traffic served on your own upstream credentials. Omitted, the view excludes BYOK traffic. | ### GET /v1/usage/coding/keys One row per coding key, ranked by spend. ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/coding/keys?start=2026-08-01&end=2026-08-31" | jq ``` ```json theme={null} theme={"system"} { "object": "usage.coding.keys", "range": "custom", "keys": [ { "key_id": "key_01J9X4M2K7Q3", "key_name": "alice@example.com", "requests": 1840, "engineers": 1, "tokens": 60850000, "input_tokens": 59700000, "output_tokens": 1150000, "cached_tokens": 55600000, "cache_write_tokens": 400000, "successes": 1831, "failures": 9, "spend": 18420, "baseline_spend": 61300, "savings": 42880, "savings_pct": 0.6995 } ] } ``` | Field | Description | | - | - | | `key_id`, `key_name` | The key's stable id and its current display name. For an automation-provisioned key the name is the engineer's email. | | `requests` | Requests the coding lane routed for this key in the window. | | `engineers` | Distinct engineer ids seen on the key. | | `tokens` | `input_tokens + output_tokens`. | | `input_tokens` | All input tokens, including the cached ones. `input_tokens - cached_tokens` is the fresh input. | | `output_tokens` | Generated tokens. | | `cached_tokens` | The part of `input_tokens` served from the prompt cache. | | `cache_write_tokens` | Input tokens written to the prompt cache. Counted separately from `input_tokens` and priced at the cache-write rate. | | `successes`, `failures` | Requests that completed vs. failed. | | `spend` | What Valar billed for the key's traffic, in **cents**. | | `baseline_spend`, `savings`, `savings_pct` | What the same traffic would have cost at Anthropic list prices, the difference, and the difference as a fraction of the baseline. | ### GET /v1/usage/coding/models One row per served model, ranked by spend. Add `key_id` to get one key's model mix. ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/coding/models?start=2026-08-01&end=2026-08-31&key_id=key_01J9X4M2K7Q3" | jq ``` ```json theme={null} theme={"system"} { "object": "usage.coding.models", "range": "custom", "models": [ { "served_model": "zai-org/GLM-5.2", "requests": 1420, "tokens": 47800000, "input_tokens": 46900000, "output_tokens": 900000, "cached_tokens": 43900000, "cache_write_tokens": 300000, "spend": 9210, "share": 0.5 }, { "served_model": "claude-opus-5", "requests": 420, "tokens": 13050000, "input_tokens": 12800000, "output_tokens": 250000, "cached_tokens": 11700000, "cache_write_tokens": 100000, "spend": 9210, "share": 0.5 } ] } ``` `served_model` is the model that actually answered, which under Auto routing can differ from the model the coding tool asked for. `share` is the model's fraction of the spend in the response (`0`–`1`), so with a `key_id` filter it is the share of that key's spend. ### GET /v1/usage/coding/requests The last requests one engineer made, newest first. This is the per-request view: everything else on this page is an aggregate, so this is where you attribute a single turn to the model that answered it. `engineer_id` is required. `limit` defaults to `10` and accepts `1`-`100`. The feed covers the last 7 days and spans every coding harness, so a Claude Code turn, a Cursor completion and a Codex run all appear in one list, each tagged by its `harness`. ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/coding/requests?engineer_id=amara&limit=5" | jq ``` ```json theme={null} theme={"system"} { "object": "usage.coding.requests", "engineer_id": "amara", "data": [ { "response_id": "resp_01J9X4M2K7Q3", "created_at": "2026-09-15T10:04:00Z", "harness": "claude", "class": "medium", "served_model": "zai-org/GLM-5.3-fast", "status": "completed" }, { "response_id": "resp_01J9X4M2K7Q4", "created_at": "2026-09-15T10:03:00Z", "harness": "claude", "class": "high", "served_model": "claude-opus-5", "status": "failed", "failure_code": "rate_limit" } ] } ``` `class` is the routing tier the request resolved into, one of `low`, `medium`, `high` or `max`, and `served_model` is what answered. Under Auto routing they differ by design. The tiers correspond to the Claude families: `low` to Haiku, `medium` to Sonnet, `high` to Opus and `max` to Fable. The feed reports the tier rather than the family because a request need not arrive as a Claude alias at all: an explicit pick of an open-weight model resolves to a tier too. The feed does not carry the exact model id the tool sent. `class` is the tier a request **resolved into**, not the one it asked for. Under Auto routing the router can re-class a request, so a Sonnet call sent to a low-tier model is recorded as `low`. Read the field as the routing decision. `status` is the request's terminal state (`completed`, `incomplete`, `failed`, `cancelled`) and `failure_code` is present only when Valar itself refused the request, for example `rate_limit` for a quota shed. A request the upstream rejected shows `failed` with no `failure_code`. `response_id` is Valar's own id for the request. It is not the `id` your client sees on the response, so use it the way this feed hands it to you: find the row by time, then quote its `response_id` to support. `valar usage` prints this feed at the bottom of its output, so an engineer can read it without writing a request. Use `-n` to ask for more rows. ### GET /v1/usage/coding/timeseries A bucketed request series for charts. This endpoint takes a `window` (`6h`, `24h`, `7d`, `30d`, default `24h`) instead of `range` or `start`/`end`, because its buckets are aligned to the current clock and the final bucket is still filling. It accepts the same filters, so `?window=7d&key_id=…` charts one key. ### Putting it together Per-key cost and token usage for a calendar month, then the model mix behind the top key: ```bash theme={null} theme={"system"} curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/coding/keys?start=2026-08-01&end=2026-08-31" \ | jq '.keys[] | {key_name, dollars: (.spend / 100), input_tokens, output_tokens}' TOP_KEY=$(curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/coding/keys?start=2026-08-01&end=2026-08-31" \ | jq -r '.keys[0].key_id') curl -s -H "Authorization: Bearer $VALAR_API_KEY" \ "https://api.valarhq.ai/v1/usage/coding/models?start=2026-08-01&end=2026-08-31&key_id=$TOP_KEY" \ | jq '.models[] | {served_model, dollars: (.spend / 100), share}' ``` Spend is integer cents, like the rest of the Usage API. The rollup behind these endpoints is per UTC day and lags live traffic by a few minutes; `range=1h` reads the live ledger instead and is the only sub-day view. *** ## Errors Errors share the format used throughout the inference API: ```json theme={null} theme={"system"} { "error": { "message": "Missing or invalid Authorization header", "type": "authentication_error", "param": null, "code": null } } ``` | HTTP Status | `type` | When | | - | - | - | | 400 | `invalid_request_error` | Bad query parameter, or a missing required one | | 401 | `authentication_error` | Missing, invalid, or expired API key | | 404 | `invalid_request_error` | Unknown route or unsupported method | | 500 | `api_error` | Internal server error | ## Response headers Each response carries an `X-Request-ID: `. Quote it in [support requests](mailto:support@valarhq.ai) to speed up troubleshooting. # Analytics & savings Source: https://docs.valarhq.ai/valarcode/analytics Per-engineer usage, operational metrics, and how ValarCode works out savings against a Claude-only baseline The **ValarCode** dashboard shows two things: how your coding traffic is doing (who is connected, what is flowing, how it performs) and how much you are saving, with every request priced against what it would have cost on Claude. The numbers are live. If the usage source is briefly unreachable, the dashboard shows zeros and an "unavailable" badge instead of guessing. ## Key numbers The setup view shows a 30-day, org-wide strip: * **Engineers connected**: distinct client ids seen. By default the client id is the engineer's OS username, so the leaderboard reads as real names; teams that prefer opaque ids can connect with `--user-id RANDOM` (see [per-engineer attribution](/valarcode/setup#per-engineer-attribution)). * **Requests routed**: total coding requests. * **Tokens**: input, cached, and output. * **Saved**: dollars saved against a Claude-only baseline (never negative). ## Savings and spend A 30-day card breaks down the money: * **Saved this period**, with a "% vs Claude-only" figure. * **Actual spend**, what you paid across all served models. * **Claude-only cost**, what the same traffic would have cost at Claude list prices. * **Spend vs. Claude-only**, a per-day chart of actual against baseline. The gap between the two lines is your savings. ## Operational metrics With a window selector (1h, 6h, 24h, 7d): * **Inference volume**: requests by status class (2xx, 4xx, 5xx). * **Response time**: end-to-end p50, p90, and p99 latency. * **Rate-limited requests**: requests shed by plan limits. * **Token usage**: throughput in tokens per second. ## Where your requests went A served-model mix table shows, for each model, its share of traffic, its request count, and its average savings against the Claude baseline. Rows served by a frontier Claude model show "n/a", since they are the baseline. Models you have configured but that saw no traffic still show up, zero-filled, so nothing is missing from the picture. ## Your own recent requests The dashboard aggregates. To attribute a single turn, run `valar usage` on the machine that made it. Under the spend summary it lists your last requests, newest first, with the routing tier each request resolved into beside the model that actually answered. ```text theme={"system"} Recent requests (last 2): TIME HARNESS CLASS SERVED STATUS REQUEST ID 2026-09-15 13:04:21 claude medium zai-org/GLM-5.3-fast completed resp_01J9X4M2K7Q3 2026-09-15 13:03:58 claude high claude-opus-5 failed (rate_limit) resp_01J9X4M2K7Q4 ``` It prints 10 rows by default. Pass `-n` for up to 100, or `--no-requests` to skip the table and see the spend summary alone. The feed covers the last 7 days, is scoped to your own client id, and spans every harness you use, so a Cursor or Codex request shows up beside a Claude Code one. That is what lets it answer "which model produced that bad turn" for a tool that keeps no local log. The **class** column is the routing tier: `low`, `medium`, `high` or `max`, corresponding to Haiku, Sonnet, Opus and Fable. It is the tier the request resolved into, not the one it asked for. Under Auto routing the router can re-class a request, so a Sonnet call sent to a low-tier model reads as `low`. Read it as the routing decision rather than as what you typed. The request id is Valar's own, not the id your tool displays. Find the row by time, then quote that id to us. ## How savings are worked out Savings compares two numbers for the same traffic, each priced on its own: * **Actual spend**: the real per-token rate of the model that served each request. Open-weight models use their [catalog price](/pricing), and the frontier Claude tiers use Claude list prices. * **Claude-only baseline**: what the same traffic would have cost on Claude at list prices. Requests already served by Claude use that tier's own list price. For an open-weight model, the tier it is priced against is the one that model stands in for — a top-tier open model is compared with a top-tier Claude, not with the same Claude for every model. **Savings = Claude-only baseline - Actual spend**, summed across served models. Claude-served traffic differences to about \$0, because its baseline and its actual cost are the same model at the same price. Claude-served traffic showing about \$0 in savings is working as intended. Savings show up only when traffic routes to an open-weight model. The baseline uses Anthropic list prices throughout. Which Claude tier a given open model is compared against depends on the tier that model stands in for, so two models can show different savings rates on identical traffic. ## Next steps Adjust the split and watch savings respond. Connect more engineers to grow the sample. # Use your own provider key Source: https://docs.valarhq.ai/valarcode/byok Serve your Claude or OpenAI coding traffic on your organization's own Anthropic, AWS Bedrock or OpenAI API key, billed by that provider instead of Valar If your organization already buys model capacity directly - from Anthropic, from AWS through Amazon Bedrock, or from OpenAI - you can serve the matching ValarCode traffic on your own key. Valar routes the request, meters the tokens, and charges you nothing for the inference. Your provider bills your account for it. You can store one key of each kind. **A key covers the models of its own provider, and nothing else.** An Anthropic or AWS Bedrock key serves your Claude-tier traffic; an OpenAI key serves the OpenAI models. A Bedrock key also serves OpenAI models when you enable **Also serve OpenAI models on this key**. Traffic to a provider you have stored no key for keeps running on Valar-managed providers and is billed by Valar, exactly as it was before you stored anything. This is often called BYOK — bring your own key. ## Who it is for Use your own key when: * You hold commitments, credits, or negotiated rates with Anthropic, AWS or OpenAI that you want your coding traffic to draw down. * Your procurement or compliance process requires frontier spend to sit on your own provider account. Stay on Valar-managed providers when: * You want a single bill, or you have no account of your own with that provider. * You want Valar's spend caps and failover to keep protecting your team. Both stop applying to requests served on your key — see [What you give up](#what-you-give-up), which you should read before you turn this on. Using your own key is turned on per organization, by arrangement with Valar. Talk to your Valar contact to have it enabled for yours. ## What you give up Two guardrails stop applying to any request served on your key. Both are deliberate, and neither is recoverable by retrying. **No fallback onto Valar's keys. If your provider cannot serve the request, the request fails.** Normally a request has more than one way to reach the model, on Valar's own accounts. With one key of yours, the chain collapses to that single hop. With two keys for the same provider, both hops are **yours**, never Valar's. This holds whichever Bedrock [endpoint](#endpoint-and-region) your key is set to: if that endpoint fails the request, it moves to your Anthropic key. If your provider is down, or your key is rate-limited, revoked, or wrong, the request returns an error to the harness — Valar does **not** quietly retry it on a Valar-paid provider key. Retrying would move the cost back onto us and put your traffic on a credential you did not choose, so we fail the request instead. Plan for it the way you would plan for calling the provider directly: keep enough rate limit headroom on your own account for your whole team's coding traffic, and know that a provider incident is now visible to your engineers. **No spend cap. Your Valar spend limits no longer apply to this traffic.** Valar's per-engineer monthly quotas and your organization's spend limits and credit checks exist to stop a runaway agent from spending your money with us. They cannot govern spend that never reaches us: on your key the tokens are billed directly to your own provider account, with no Valar-side ceiling in front of them. Valar's per-organization request rate limit is skipped for the same reason, so your provider account's own rate limits are the only ones in front of this traffic. Set your budget where the spend now lands - use your provider's own spend limits and usage alerts on the account that owns the key. Per-engineer **pause** still works, so you can still stop an individual engineer. ## Add, rotate, or remove your key Once your organization is enabled, a box per provider appears under **Bring your own key** on the **Settings - ValarCode settings** page: **Anthropic API key**, **AWS Bedrock API key** and **OpenAI API key**. Each box carries a badge for the traffic that key serves - Claude, OpenAI, or both. You need the organization **admin** role to change a key; any member can see whether one is set. **Add a key** 1. Create the key on the account you want the traffic billed to - an API key in the [Anthropic console](https://console.anthropic.com) (starts with `sk-ant-`), a **long-term** Bedrock API key in the AWS console under **Bedrock -> API keys** (starts with `ABSK`), or an API key in the [OpenAI platform console](https://platform.openai.com/api-keys) (starts with `sk-`). Give it enough rate limit for your whole team. 2. Expand the box, paste the key, and click **Check & save**. Valar checks the key with the provider before storing anything: a key the provider rejects is **not stored**, and the box says why. A key that passes is stored encrypted and switched on. 3. The box now shows the verdict — **Verified just now: Anthropic accepted this key.** — and a locked **API key** field with a redacted fingerprint (`sk-ant-api0…`, `ABSK…bz0=`) so you can tell which key is in place. The verdict stays: after a reload it reads **Verified 3 minutes ago**, or **Check failed 2 days ago: …** with the reason. It takes effect on the next request. There is nothing to change on any engineer's machine — no new connect command, no config edit. **Check a key later** Click **Check key** at any time to re-test the stored key (expired, revoked, model access changed). The result is remembered, so the box tells the next person what the last check found. **Rotate a key** Click **Rotate key**, paste the new key, and click **Check & save**. The new key is tested and then replaces the old one in place. Rotate on your own schedule the same way you would rotate any provider key: create the new one, save it here, then revoke the old one at the provider. **Turn a key off without deleting it** Use the switch next to the box title. Off means the key stays saved but is not used - the traffic that key serves uses your other enabled key for the same models, if you have one. Otherwise it runs on Valar-managed providers with normal charges. **Remove a key** Click **Remove key** and confirm. From the next request, the affected models use your other enabled key for those models, if you have one. Otherwise they return to Valar-managed providers with normal Valar inference charges and spend caps. The key is stored encrypted and is never shown again — not in the dashboard, not in the API, not in a log. Only the redacted fingerprint is ever displayed. ## OpenAI key An OpenAI API key serves OpenAI models through the Messages API and Codex Responses API. It does not affect your Claude traffic: that keeps running wherever it ran before - on your own Anthropic or AWS key if you store one, on Valar-managed providers otherwise. When you save or check an OpenAI key, Valar asks OpenAI which models the key can reach and lists any OpenAI model it cannot. That usually means the key's OpenAI **project** has not been granted access to the model; grant it in the OpenAI console and click **Check key** again. The key is still stored - it authenticated, and the models it can reach are served on it. This check lists available models without generating tokens. It does not verify your credit balance or permission to generate responses. A request can still fail if either is missing. Any OpenAI key shape works (`sk-`, `sk-proj-`, `sk-svcacct-`). Use one whose project has the spend limits and rate limits you want for your whole team. ## AWS Bedrock key A Bedrock API key works from any AWS region. Which AWS endpoint Valar calls with it, and where in the world the model runs, is set by the [Endpoint and region](#endpoint-and-region) controls in the key box; the default lets AWS decide. ### Also serve OpenAI models on your Bedrock key Bedrock supports OpenAI models as well as Claude models; your AWS account must have access to them. The Bedrock box has a switch - **Also serve OpenAI models on this key** - that extends the key you already stored to that traffic too. It is **off** by default and stays off until you turn it on, so a key you stored for Claude keeps serving Claude only, and your OpenAI-model spend does not move onto your AWS bill without you asking for it. Turning it on re-checks the key. With it on, AWS bills you for OpenAI traffic served on this key and Valar charges \$0. With it off, OpenAI traffic uses your enabled OpenAI API key, if present, or Valar-managed providers with normal charges. If both keys are enabled for OpenAI, requests can use either of your accounts. The Bedrock check verifies Claude access; it does not verify OpenAI model access or change that access in AWS. When you save or check a Bedrock key, Valar tests it against **every Claude model** it serves through Bedrock, on the endpoint you selected, and the box lists any model the key authenticated for but cannot use, with the reason. Two kinds of line: * **Transient** — AWS was overloaded, timed out, or is still enabling the model on your account (the first use of each model creates a Marketplace subscription, which takes about a minute). The key is fine; click **Check key** again. If an *Overloaded* line persists for one model while the others work, your account's on-demand throughput quota for that model is probably zero — new accounts start there for Opus-class models. Request an increase in the AWS console (Service Quotas → Bedrock), then check again. * **Not usable on your AWS account** — the account itself blocks the model. The line says what to change; the common cases are under [What the key check can tell you](#what-the-key-check-can-tell-you). ### Endpoint and region AWS offers two endpoints that accept the same Bedrock API key for Claude. The Bedrock key box lets you pick which one your traffic uses: | Setting | Options | What it does | | - | - | - | | **Endpoint** | **Mantle (default)** or **bedrock-runtime** | Which AWS endpoint Valar calls with your key. | | **Region** | Disabled on Mantle ("AWS decides routing"); **US** on bedrock-runtime | Where the model runs. Applies to bedrock-runtime only — on Mantle, AWS decides the routing, which may include regions outside the US. More regions are added as Valar verifies them. | Changing a select does not apply anything by itself — the row reads **Not applied yet** until you commit it with a button: * **Key already stored**: click **Save & check**. Valar saves the setting, then re-tests your stored key on the endpoint you picked. **Cancel** puts the selects back and saves nothing. * **Adding or rotating a key**: the selects are part of the form and are saved together with the key when you click **Check & save**. The new key is tested on the endpoint you picked. **When to use bedrock-runtime + US.** Choose it when your AWS organization blocks global routing, or requires that your Claude traffic stays in US regions. On Mantle, AWS resolves the model itself and may serve it from any region; a policy that denies that will fail your Fable 5 requests (and may let other models quietly leave the US). On bedrock-runtime with **US**, Valar asks for the model's `us.` inference profile, so the request is served only from US regions. Note that AWS prices US-region (geographic) inference about 10% higher than global inference — see [Amazon Bedrock pricing](https://aws.amazon.com/bedrock/pricing/). **The same key, different permissions.** The two endpoints are gated by different IAM actions — `bedrock:InvokeModel` and `bedrock:CallWithBearerToken` for bedrock-runtime, `bedrock-mantle:CreateInference` and `bedrock-mantle:CallWithBearerToken` for Mantle (see AWS: [API keys → Control who can generate and use API keys](https://docs.aws.amazon.com/bedrock/latest/userguide/api-keys.html) for the two `CallWithBearerToken` actions, and [Bedrock Mantle → Prerequisites → Permissions](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-mantle.html) for `bedrock-mantle:CreateInference` and `bedrock:InvokeModel`) — so a key that passes on one can fail on the other, and a policy in your AWS organization can allow one and deny the other. That is why the key check always runs on the endpoint you selected. If it fails after switching, ask your AWS administrator to allow the actions for that endpoint on the key's IAM user. **Failover is unchanged.** With both keys, a request that fails on your selected Bedrock endpoint falls back to your Anthropic key. With a Bedrock key only, it returns a clear error to the harness instead — it is never served, and never billed, on Valar's account. ### Claude Fable 5 needs an opt-in Anthropic treats Claude Fable 5 (and Mythos 5) as *covered models*: AWS only serves them if your account has agreed that prompts and completions may be **shared with Anthropic and kept for up to 30 days** for safety review. Older Claude models do not require this. AWS enforces it through an account setting called the *data retention mode*; a new account starts on `default`, which means "not agreed", so Bedrock refuses Fable 5 until you set it to `provider_data_share`. **AWS stores this setting separately for each endpoint.** Opting in through Mantle's API does **not** cover bedrock-runtime, and the other way round. Opt in on the endpoint your key is set to; if you switch endpoints later, opt in again on the new one. **On Mantle (the default endpoint)** — one API call, made with your Bedrock key: ```bash theme={"system"} curl -X PUT https://bedrock-mantle.us-east-1.api.aws/v1/data_retention \ -H "x-api-key: $BEDROCK_API_KEY" \ -H "content-type: application/json" \ -d '{"mode":"provider_data_share"}' ``` To limit the sharing to one Bedrock project instead of the whole account, set it on the project the key belongs to: ```bash theme={"system"} curl -X POST https://bedrock-mantle.us-east-1.api.aws/v1/organization/projects/ \ -H "x-api-key: $BEDROCK_API_KEY" \ -H "content-type: application/json" \ -d '{"data_retention":{"mode":"provider_data_share"}}' ``` **On bedrock-runtime + US** — the setting lives on the Bedrock control plane, per region. Under cross-region inference AWS stores retained data in the region that served the request, so set it in **each** US region the `us.` profile can serve from (`us-east-1`, `us-east-2`, `us-west-2`): ```bash theme={"system"} for r in us-east-1 us-east-2 us-west-2; do aws bedrock put-account-data-retention --mode provider_data_share --region "$r" done ``` Or, without the AWS CLI: `PUT https://bedrock..amazonaws.com/data-retention` with the body `{"mode":"provider_data_share"}`, once per region, signed with AWS credentials. Both forms are the same control-plane API — see AWS: [Data retention → Configuring data retention → Set account-wide data retention → Bedrock Control Plane](https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html) (the `PUT /data-retention` call; the CLI command is its `aws bedrock` surface). Two things to know about this call: * A Bedrock API key **cannot** make it. It needs an IAM identity (a user or role) with the `bedrock:PutAccountDataRetention` permission — typically an administrator signed in to the AWS account (the same AWS page's "IAM actions reference" maps `PUT /data-retention` to `bedrock:PutAccountDataRetention` and `GET /data-retention` to `bedrock:GetAccountDataRetention`). * It can take **up to about an hour** to propagate. Confirm it with `aws bedrock get-account-data-retention --region ` in each region (AWS: "Check your current configuration → Bedrock Control Plane"); when all three report `provider_data_share`, click **Check key**. Then click **Check key**; the Fable 5 line disappears. Until you opt in, the other Claude models work on your key and a Fable 5 request on it fails. Details: [Amazon Bedrock — Data retention](https://docs.aws.amazon.com/bedrock/latest/userguide/data-retention.html). Claude Fable 5.1, the current Fable fallback, is not on Bedrock yet, so it is served only through an Anthropic key. A Bedrock-only key cannot serve a Fable-class request that falls back to it; add an Anthropic key alongside, or pin the Fable slot to Claude Fable 5 in the routing editor. ### What the key check can tell you Each line in the box is one situation and its action. The common ones: | The box says | What it means | What to do | | - | - | - | | **AWS doesn't recognize this key.** | The key is not a Bedrock API key AWS knows — usually a partial paste, a deleted key, or an expired temporary key. | Copy the whole key again from **Bedrock → API keys**, or create a new long-term key, and save it. | | **AWS won't let this key use Claude.** | AWS authenticated the key but its IAM user may not invoke Claude on the endpoint you selected, or the Claude models are not enabled for your account. | Ask your AWS administrator to allow the IAM actions for that endpoint (see [Endpoint and region](#endpoint-and-region)) and to enable the Claude models in Bedrock, then click **Check key**. | | **Fable 5 — your AWS account hasn't opted in to sharing prompts with Anthropic … or your AWS organization may be restricting Bedrock to US regions** (on Mantle) | AWS gives the same answer for two different causes: no opt-in yet, or an organization policy that denies global routing. | If you have not opted in, do so on Mantle as above. If you already have, set **Endpoint** to **bedrock-runtime** and click **Save & check**. | | **Fable 5 — your AWS account hasn't opted in to sharing prompts with Anthropic for US-region Bedrock** (on bedrock-runtime + US) | The opt-in is missing on the Bedrock control plane in one or more US regions. The Mantle opt-in does not count here. | Run the per-region opt-in above, wait for it to propagate, then click **Check key**. | | A model line ending **Fix this in your AWS Bedrock account** with AWS's own words, such as the model not existing or not being available | The model is not available to your account on the endpoint and region you selected — it is not enabled there, or AWS is not serving it there right now. | Check the model is enabled for your account in the US regions. For Claude Haiku 4.5, see the [known limitation](#known-limitation-claude-haiku-45-on-bedrock-runtime) below. | ### Known limitation: Claude Haiku 4.5 on bedrock-runtime AWS intermittently reports Claude Haiku 4.5 (inference profile `us.anthropic.claude-haiku-4-5-20251001-v1:0`) as unavailable in some US regions on bedrock-runtime. This is on AWS's side and not something your key or account can fix. * If you also have an Anthropic key, Haiku requests that fail on Bedrock fall back to it — you will not notice beyond a line in the key check. * If your organization uses **only** a Bedrock key on bedrock-runtime + US, Haiku 4.5 requests may fail until AWS resolves it. Other Claude models are not affected. **Temporary keys expire.** A 12-hour Bedrock API key (starts with `bedrock-api-key-`), or a long-term key with an expiry set in AWS, stops your Claude traffic when it lapses — and **Check key** cannot see an expiry date. Use a long-term key, and rotate it here before it expires. The box warns when you paste a temporary key, but does not stop you. ## What it costs | | On your key | On Valar-managed providers | | - | - | - | | Claude and OpenAI tokens | Billed by your provider to your account | Included in what Valar charges you | | Valar inference charge | **\$0** | Valar's per-token rate for the model | | Tokens metered by Valar | Yes — input, cached, and output | Yes | Valar still counts every token so the traffic is measurable. It is simply priced at zero, because your provider bills your account for it. Your open-weight traffic is unchanged. Models such as GLM-5.2 and Kimi-K3 still run on Valar and are billed by Valar. A routing policy that mixes open-weight, Claude, and OpenAI models can produce both provider charges and Valar charges. ## Exactly which requests use your key On the supported APIs below, requests served by a model for which you have an enabled, eligible key use that key. This includes Auto routing and capability escalations to those models. Claude fallback requests on the Messages API also use your eligible Claude key. Requests served by an open-weight model remain billed by Valar. BYOK covers these coding requests: | Request | Customer keys used | | - | - | | Claude through the Messages API (`/v1/messages`) | Your enabled Anthropic or AWS Bedrock key | | OpenAI through the Messages API or Codex Responses API (`/v1/responses`) | Your enabled OpenAI API key or a Bedrock key with **Also serve OpenAI models on this key** enabled | | Chat Completions API (`/v1/chat/completions`), including Cursor | Stored provider keys do not apply; normal Valar charges apply | Claude requests made through Cursor or Codex do not use your stored provider keys. The OpenAI Responses support above does not extend Claude BYOK to those tools. ## Where this traffic shows up Requests served on your key appear in **BYOK usage** on the ValarCode **Metrics** tab once the selected time window contains own-key requests. Valar records the tokens and charges \$0 for inference. The regular usage and savings panels exclude this traffic. For the actual cost, use your provider's usage dashboard: Anthropic, OpenAI, or AWS Cost Explorer for Bedrock. That provider bills your account. ## Next steps Decide which requests reach Claude in the first place. See what the panels measure, and why own-key traffic sits outside them. # Claude Code Source: https://docs.valarhq.ai/valarcode/claude-code Connect Claude Code to Valar, choose proxy or direct routing, and manage running sessions Use Claude Code with Valar and choose models with `/model`. Install the [Valar CLI and configure your coding key](/valarcode/setup) first. ## Choose a routing lane | | Proxy lane | Direct lane | | - | - | - | | Available on | macOS and Windows 11 with the on-device proxy | macOS, Linux, and Windows | | Setup | `valar proxy enable`, then `valar claude on` | `valar claude on --direct` | | Routing | Claude keeps its Anthropic endpoint; the local proxy handles routing | Claude sends inference to the Valar endpoint | | Claude sign-in | Required; Claude Code keeps its own sign-in | Not required | | Claude subscription usage | Optional quota-first | Not available; inference uses Valar | | Remote Control | Retains first-party eligibility, subject to your Claude account and policy | Not available with the custom endpoint | | Request diagnostics | `valar proxy logs` | No local proxy log | `valar claude on` uses the proxy when usable and otherwise selects direct routing. It explains the selected lane. Add `--direct` to choose direct routing explicitly. ## Enable routing For the proxy lane on macOS or Windows: ```bash theme={"system"} valar proxy enable valar claude on ``` See [On-device proxy](/valarcode/proxy) for certificate trust, administrator access, and corporate networks. To deploy with MDM, see [Roll out with MDM](/valarcode/mdm). For direct routing: ```bash theme={"system"} valar claude on --direct ``` Check the result: ```bash theme={"system"} valar claude status ``` Claude Code and Claude Desktop have independent routing controls. This command changes the CLI in your terminal; use `valar claude-desktop on` for the app. Valar leaves your Claude Code version unchanged unless you use [the version-hold options](#stable-claude-code). ## Model selection Valar updates the `/model` picker with the models available to your coding key. An explicit `valar claude on` selects the Valar default, normally **Valar Auto**. You can choose another model; that choice persists until the next explicit `on`. With quota-first enabled, eligible Claude-model requests can use your subscription. See [Quota-first](#quota-first-use-your-own-subscription-first) for model and account behavior. ## Stable Claude Code To enable routing and hold Claude Code at Valar's stable release: ```bash theme={"system"} valar claude on --stable ``` This requires Claude Code from Anthropic's native installer. Valar installs the selected release with `claude install `, then sets `DISABLE_AUTOUPDATER=1` in Claude Code's settings to stop automatic updates. You can still run `claude update` manually. If installation fails, Valar does not add the hold or change routing. The hold survives later plain `valar claude on` commands and Valar upgrades. When Valar's stable release moves forward, the next `valar claude on` installs it and keeps the hold. It never moves Claude Code back to an older release. To remove Valar's hold and update Claude Code now: ```bash theme={"system"} valar claude on --latest ``` On a native installation, this removes Valar's hold and runs `claude update`. For npm or Homebrew installations, it removes Valar's hold; update Claude Code through your package manager. `--stable` does not support package-manager installations. `valar claude off` also removes Valar's hold. Both OFF and `--latest` preserve a `DISABLE_AUTOUPDATER` setting you had before enabling Valar, so automatic updates can remain disabled. Neither flag changes your `autoUpdatesChannel` setting. Use one version option at a time. Neither works with `--single-session`, because both affect the installed Claude Code. `valar claude status` reports a Valar-managed hold. `valar configure --stable` and `valar configure --latest` do the same: they save your configuration, then run `valar claude on --stable` or `valar claude on --latest`, which also turns routing on. `valar upgrade` updates Valar itself. ## Switch lanes and running sessions Switch directly between lanes without turning routing off first: ```bash theme={"system"} valar claude on --direct valar claude on ``` When the change requires refreshing running sessions, Valar asks before interrupting them: ```text theme={"system"} Save your work. Valar will try to restart running sessions. You may need to reopen sessions and resume interrupted tasks. Continue? [y/N] ``` * **No:** the requested routing change and session restart are cancelled. * **Yes:** Valar updates routing, attempts to recover background agents, and closes affected interactive sessions. Reopen saved conversations with `claude --resume` or reopen `claude agents`. * **`--force`:** approve the same interruption without a prompt. It does not make an unnecessary restart happen during an ordinary proxy toggle. Saved conversations remain available. Interrupted requests and tools are not replayed automatically. A session running the Valar command itself is left open; restart it from an external terminal if instructed. Recovery is best effort. If a session cannot refresh, Valar reports it and applies the routing configuration. That process can keep its previous routing until you restart it. Once the proxy arrangement is established, ordinary proxy ON/OFF toggles need no restart. Changing the lane, certificate, or proxy connection settings can require one again. ## Quota-first: use your own subscription first Quota-first requires the proxy lane and a Claude login with included subscription allowance. It is unavailable on direct routing, including Linux. ```bash theme={"system"} valar configure --quota-first=true valar claude status ``` The preference is shared with Claude Desktop and read live by the proxy. Each request uses the account carried by its login. * Eligible Claude-model requests use included subscription allowance while available. Requests that cannot use it, including requests that would incur extra usage, use Valar. * Enterprise accounts and accounts without included allowance use Valar. * Valar Auto and explicit open-model selections always use Valar. To send every routed request through Valar: ```bash theme={"system"} valar configure --quota-first=false ``` To clear temporary subscription holds and retry on the next eligible request: ```bash theme={"system"} valar claude reset ``` `valar claude-desktop reset` is equivalent. Reset applies to the shared proxy; it does not erase counters or reset the subscription allowance. ### Check what is serving Send a request, then inspect `valar claude status`. The account may not be known until a request carries your login. If Claude Code and Desktop are signed into different accounts, their status can show different limits. `Checked` shows when quota information was confirmed. `Served` counts subscription and Valar requests, split into main and background calls. The same account used by both apps shares counters; different accounts do not. These counters are temporary diagnostics and can reset, including when the proxy restarts. They are not billing totals. On the proxy lane, `valar claude status --verify` checks successful Valar forwarding. A request served by your subscription alone does not satisfy that check; select Valar Auto or temporarily disable quota-first for the test. ## Route a single session ```bash theme={"system"} valar claude on --single-session ``` This launches one Claude Code process on the **direct lane**, using temporary settings. It does not rewrite your main Claude settings, save a new routing preference, or change other sessions. Single-session routing does not support quota-first. When the process exits, its routing ends. ## Turn off routing ```bash theme={"system"} valar claude off ``` On the proxy lane, new connections pass through to Anthropic without decryption. Proxy pointers remain installed for the next ON, and the services keep running. Remove them with [the shared proxy teardown](/valarcode/proxy#remove-the-proxy). On the direct lane, Valar removes its endpoint and credential settings and restores the settings it previously replaced. Running sessions may need recovery; follow the prompt or use `--force` when interruption is acceptable. If recovery is incomplete, existing direct sessions may still send requests to Valar until restarted. When Valar can safely merge readable settings, OFF preserves unrelated changes such as plugins, theme, permissions, and hooks. ## What gets written Both lanes use Claude Code's settings file, normally `~/.claude/settings.json`. Valar does not install a launcher wrapper. The optional [version hold](#stable-claude-code) uses Claude Code's own installer and updater. | Setting | Proxy lane | Direct lane | | - | - | - | | Inference endpoint | Keeps the native endpoint | Sets `ANTHROPIC_BASE_URL` to Valar | | Proxy connection | Points Claude at its local listener | Removes Valar's proxy pointer and restores displaced values | | Valar credential | Read by the local proxy from Valar configuration | Written in custom request headers; a token is also configured when no Claude login is available | | Model picker | Updated for your coding key | Updated for your coding key | Valar saves the original settings before its first change and keeps credentials in user-only files. If you log into Claude after enabling direct routing without a login, rerun `valar claude on` so the configuration can reflect the new login. ## Upgrading an older installation Run `valar upgrade` from your own terminal session. The upgrade removes the retired socket routing and preserves your routing choice where it can safely determine it. If no usable on-device proxy is installed, Claude Code moves to the direct lane. Your quota-first preference stays saved, but requests use Valar until you set up the proxy and enable its lane. On macOS, that setup may require administrator access. Read any session-recovery warnings. See [Upgrade and recovery](/valarcode/setup#keeping-the-cli-current) before retrying a failed migration. ## Next steps Shared proxy setup and corporate networks. Configure the app separately from the CLI. # Claude Desktop Source: https://docs.valarhq.ai/valarcode/claude-desktop Connect Claude Desktop's Code and local Cowork tabs to Valar using proxy or direct routing Valar routes Claude Desktop's Code and local Cowork tabs on macOS, and its Code tab on Windows 11. Local Cowork on Windows is not supported yet. The app and the [Claude Code CLI](/valarcode/claude-code) are configured and reported separately. You need Claude Desktop, the [Valar CLI, and a coding key](/valarcode/setup). ## Choose a routing method | | Proxy lane | Direct lane | | - | - | - | | Configuration | Uses the shared on-device proxy | Adds a Valar provider profile | | Claude account and connectors | Keeps the app's native account and connections | Some native account features are unavailable | | Code and local Cowork | Supported (Code only on Windows) | Supported (Code only on Windows) | | Quota-first | Supported | Not available | | Chat | Outside this proxy integration | Optional with `--with-chat`; separate history | | Administrator access | Required for initial machine setup on macOS; not required on Windows | Not required for the provider profile | `valar claude-desktop on` selects the proxy lane when the proxy is provisioned; otherwise it uses the direct lane. Pass `--direct` to explicitly choose the provider profile. ## On-device proxy Provision the shared component once, then enable Desktop: ```bash theme={"system"} valar proxy enable valar claude-desktop on ``` See [On-device proxy](/valarcode/proxy) for setup, certificates, network policies, logs, and removal. To deploy with MDM, see [Roll out with MDM](/valarcode/mdm). Turning Desktop OFF does not turn Claude Code OFF. Once Desktop's proxy arrangement is established, ordinary ON/OFF toggles do not restart the app. ## Direct lane ```bash theme={"system"} valar claude-desktop on --direct ``` Direct routing sends inference to Valar without using your Claude subscription. The app reads its provider profile at startup, so changing to or from this lane may require a restart. To include the Chat tab: ```bash theme={"system"} valar claude-desktop on --direct --with-chat ``` Pass `--with-chat` on each `on` that should include Chat. The direct profile has its own chat history; it does not merge with your claude.ai history when you switch lanes. ## Switch lanes and running sessions You can switch directly between lanes: ```bash theme={"system"} valar claude-desktop on --direct valar claude-desktop on ``` When a running app needs to restart, Valar explains the change and asks: ```text theme={"system"} Save your work. Valar will try to restart running sessions. You may need to reopen sessions and resume interrupted tasks. Continue? [y/N] ``` * **No:** routing and the running app remain unchanged. * **Yes:** Valar quits the app, applies the change, and reopens it. On Windows, Claude Desktop stays in the notification area after its windows close, so Valar ends the app a few seconds after asking it to close. * **`--force`:** approve that restart without a prompt, for example `valar claude-desktop on --direct --force`. If the app cannot quit, Valar leaves its configuration unchanged. If it cannot reopen, Valar asks you to open it manually. Saved history remains in its original mode; active requests and tasks may be interrupted. First-time proxy activation and changes to proxy connection settings or certificates may require a restart. Established proxy ON/OFF toggles do not, including when `--force` is supplied. ## History and MCP servers on the direct lane On the proxy lane, Claude Desktop keeps its normal history and MCP servers unchanged. The direct lane runs the app in a separate provider profile. When you switch lanes, Valar copies your local Code sessions (and local Cowork sessions on macOS) and scheduled tasks between the two, and never deletes a session. It also adds your local MCP servers to the direct-lane profile each time you turn the direct lane on. Chat history is stored by Claude and is not affected. To distribute a fixed list of servers for the direct lane, place a JSON array at `~/.valar/claude/desktop-mcp.json` (`%USERPROFILE%\.valar\claude\desktop-mcp.json` on Windows), in the schema Claude Desktop uses for managed MCP servers (`name`, `transport`, then `command`/`args`/`env` or `url`/`headers`). An entry there replaces any server with the same name. The file is read whenever Valar writes the profile; a file that is not valid JSON stops the command with an error rather than being skipped. ## Check routing ```bash theme={"system"} valar claude-desktop status ``` Status reports OFF or the active lane. On the proxy lane it also reports service health, certificate information, and forwarding activity. To check Valar forwarding, send a Code or local Cowork request that uses Valar, then run: ```bash theme={"system"} valar claude-desktop status --verify ``` The check exits nonzero until the proxy observes a Desktop request successfully forwarded to Valar after routing was enabled. A subscription-served request does not satisfy it. See [Status and logs](/valarcode/proxy#status-and-logs). ## Quota-first: use your own subscription first With Desktop on the proxy lane and signed into an eligible Claude account: ```bash theme={"system"} valar configure --quota-first=true valar claude-desktop status ``` Quota-first uses included subscription allowance before Valar. The preference is shared with Claude Code, but each app's requests identify the account they use. See [the shared Claude quota-first controls](/valarcode/claude-code#quota-first-use-your-own-subscription-first) for limits, counters, reset, and opt-out. The direct lane cannot use quota-first. Desktop Chat and cloud Cowork sessions are outside the proxy integration's coverage. ## Turn off routing ```bash theme={"system"} valar claude-desktop off ``` On the proxy lane, new connections pass through to Anthropic without decryption. The shared proxy remains installed. On the direct lane, Valar removes its profile or restores the configuration it replaced, with a restart when required. To uninstall the shared services and their machine settings, first turn off every harness using them, then follow [Remove the proxy](/valarcode/proxy#remove-the-proxy). ## Provider profile details The direct lane writes a Valar profile to Claude Desktop's provider-profile folder: | OS | Folder | | - | - | | macOS | `~/Library/Application Support/Claude-3p/configLibrary/` | | Windows | `%LOCALAPPDATA%\Claude-3p\configLibrary\` | It records the Valar endpoint, coding key, model list, and request-attribution headers. The pre-enable profile configuration is backed up at `~/.valar/claude/desktop-backup.json` (`%USERPROFILE%\.valar\claude\desktop-backup.json` on Windows). OFF restores that configuration; it removes the configuration tree if Valar created it. Open a separate direct-routed instance: ```bash theme={"system"} valar claude-desktop on --direct --single-session ``` It uses `~/.valar/claude/desktop-session-3p` (`%USERPROFILE%\.valar\claude\desktop-session-windows\Claude-3p` on Windows) and keeps its history there. Your main app's configuration, login, and reported routing stay unchanged. Only one such instance runs at a time. Add `--force` to replace an existing instance. Explicit `--direct` allows this mode even when the shared proxy is provisioned. ## Security and machine changes Certificate trust, shared proxy services, and corporate network controls are documented under [On-device proxy](/valarcode/proxy#certificates-and-machine-changes). ## Collect CLI logs for support Inspect `~/.valar/cli.log` for command/setup failures and `valar proxy logs` for proxy traffic diagnostics. See [Support logs](/valarcode/proxy#support-logs) for collection, redaction and rotated history. Review logs before sharing them. ## Next steps Shared setup, certificates, and network policies. Configure the terminal CLI separately. # Codex Source: https://docs.valarhq.ai/valarcode/codex Route the OpenAI Codex CLI through Valar, and what the CLI writes to ~/.codex/config.toml ValarCode routes the OpenAI Codex CLI through Valar by editing its `config.toml`. Codex keeps speaking the OpenAI Responses API and you keep picking your model; the gateway maps whatever it sends to the target your routing resolves to. See [Set up ValarCode](/valarcode/setup) for the shared setup and the client-id model. ## Prerequisites * The OpenAI Codex CLI installed, with a `~/.codex` config directory. Run Codex once so the directory exists. * A Valar coding key (`vlrcode_…`) and the `valar` CLI ([install](/valarcode/setup#install-the-cli)). ## Enable routing Run once per machine: ```bash theme={"system"} valar codex on --api-key ``` The routing lives entirely in `config.toml`, with no env var, shell-rc edit, or new-terminal step. Then restart Codex (or the ChatGPT app, if that is where you run Codex) and **start a new conversation**. Check the state any time: ```bash theme={"system"} valar codex status ``` ## What gets written `valar codex on` edits `~/.codex/config.toml`. It sets the root provider to Valar and adds a provider block that points at the gateway and carries your credentials in request headers: ```toml theme={"system"} model_provider = "valar" [model_providers.valar] name = "Valar" base_url = "https://api.valarhq.ai/v1" wire_api = "responses" experimental_bearer_token = "vlrcode_…" http_headers = { "X-Valar-Client-Id" = "jdoe", "X-Valar-Harness" = "codex", "X-Valar-Cli-Version" = "1.4.2" } ``` In detail, enabling: * Sets the root `model_provider` to `valar` so Codex uses the Valar provider block. * Writes `base_url` with the `/v1` surface. Codex speaks the OpenAI Responses API (`wire_api = "responses"`) and appends `/responses`, so requests land on `https://api.valarhq.ai/v1/responses`. * Puts the coding key on `experimental_bearer_token`, which Codex renders as `Authorization: Bearer` on the wire. It must **not** be an `Authorization` entry in `http_headers`. Codex treats that name as reserved and drops it (confirmed from codex-cli 0.145), which sends requests with no credential at all. * Adds the `X-Valar-Client-Id`, `X-Valar-Harness` and `X-Valar-Cli-Version` request headers. The client-id header is omitted entirely when there is no client id, rather than written blank. Your comments, layout, and any other tables pass through byte-for-byte. The file is written user-only (`0600`), since it now holds a bearer token, and the original `config.toml` is snapshotted under `~/.valar/codex/` before the first change so `off` restores it exactly. The client id defaults to your OS username, so the header reads `X-Valar-Client-Id: jdoe`. Use `--user-id RANDOM` for an opaque `vc_…` id instead. See [per-engineer attribution](/valarcode/setup#per-engineer-attribution). ## Route a single session To try Valar without changing anything in your Codex setup, or to run one routed session beside it, open a single Valar-routed session instead of enabling globally: ```bash theme={"system"} valar codex on --single-session ``` This starts one `codex` pinned to its own data dir (`~/.valar/codex/session-home`), routed through Valar. Your real `~/.codex` config, and every other Codex session, stay untouched, and `valar codex status` stays `off`. Session history persists in that data dir across single-session runs. ## Turn off routing ```bash theme={"system"} valar codex off ``` This restores the pre-enable `config.toml` byte-for-byte. Restart Codex and start a new conversation. ## Known limitations ### Existing conversations require a fork between valar on and off Moving a conversation across a `valar codex on` or `off` boundary is not supported. Start a new conversation, or `/fork`. Conversations opened before `valar codex on` stay with OpenAI. Run `/fork` in one and the fork carries the context through Valar. Running `valar codex off` under a conversation that is still open breaks that conversation. Routing reverts with the file, but the model stays the one you had selected, and OpenAI does not serve it: ```json theme={"system"} {"type":"error","status":400,"error":{"type":"invalid_request_error","message":"The 'valar-auto' model is not supported when using Codex with a ChatGPT account."}} ``` Finish or close a routed conversation before you turn routing off. Resuming across a switch is the quieter trap. `codex resume` reads whatever `config.toml` says at that moment and drops both the provider and the model the session recorded. The only hint is about the model: ```text theme={"system"} warning: This session was recorded with model `valar-auto` but is resuming with `gpt-6-astra`. ``` Nothing mentions the endpoint. Because `off` restores the file without the Valar block, a session recorded on Valar resumes straight against OpenAI on OpenAI's default model. Forcing the provider back does not rescue it, since the block it needs is the thing `off` removed: ```text theme={"system"} $ codex resume -c model_provider="valar" Error: Model provider `valar` not found ``` ### Voice is off while routing is on Codex voice, realtime sessions and in-app dictation both, terminates at OpenAI. The audio leg runs against OpenAI's Realtime API under your ChatGPT sign-in, and `model_provider` does not redirect it. Valar routing authenticates with your coding key instead, so the voice path is unavailable and Codex stays text-only. Turn routing off, or keep an unrouted Codex beside it, if you need voice. ## Next steps How cohorts and the split decide which model serves each request. The open-weight targets and the frontier Claude tiers. # Cursor Source: https://docs.valarhq.ai/valarcode/cursor Route Cursor through Valar, side by side with Cursor's own models or Valar only ValarCode puts Valar's models in Cursor's model picker. On macOS with the on-device proxy, it also answers Cursor's **Auto** with Valar Auto. See [Set up ValarCode](/valarcode/setup) for the shared setup and the client-id model. `valar cursor` runs on macOS and Linux. It is not available on Windows. ## Prerequisites * Cursor installed and signed in, on a paid plan (Pro, Business, or Enterprise). `valar cursor on` fails on the free plan or when Cursor is not signed in. * For the on-device proxy, Cursor **3.21.16 or newer**. With an older Cursor and the proxy set up, `valar cursor on` warns in side by side, and Valar only stops until Cursor is updated. Without the proxy, any Cursor version works as before. * A Valar coding key (`vlrcode_…`) and the `valar` CLI ([install](/valarcode/setup#install-the-cli)). * On macOS, the [on-device proxy](/valarcode/proxy) is recommended: run `valar proxy enable` once. Without it, Cursor's own models such as Composer don't work next to Valar's, and Valar only can be worked around. If you set up the proxy before Cursor support, run it again so its certificate covers Cursor. If your organization supplies the proxy certificate, it must also cover Cursor's servers; see [Organization-managed certificates](/valarcode/proxy#organization-managed-certificates). * If your Cursor team settings control custom API keys, keep them allowed. Valar's key is saved in Cursor as one. ## Two modes | | Side by side | Valar only | | - | - | - | | Command | `valar cursor on` | `valar cursor on --valar-only` | | Picker | Valar's models next to Cursor's | Valar's models and Auto only | | Cursor's **Auto** on **This Mac** | Answered by Valar Auto | Answered by Valar Auto | | Cursor's own models (Composer, …) | Answered by Cursor, as before | Not offered; a stray pick is answered by Valar Auto | | A message that uses a personal API key saved in Cursor | Blocked | Blocked | | Cloud runs | Cursor's models run as usual; a Valar model is blocked | Blocked | `valar cursor on` also turns off the personal API keys saved in Cursor, and `valar cursor off` turns them back on. If your admin turned on **Show serving model name in the response body** ([details](/valarcode/routing#under-each-answer)), an answer that Valar Auto picked ends with a short `[served by ]` line. This table describes macOS with the on-device proxy. Without it, see [the proxy lane and the direct lane](#the-proxy-lane-and-the-direct-lane). ## Enable routing **Quit Cursor first**, or let valar do it for you. Cursor keeps its settings in memory and writes them on exit, so a change made while it runs would be undone. Run once for each user account: ```bash theme={"system"} valar proxy enable # macOS, once per machine valar cursor on --api-key # side by side valar cursor on --valar-only --api-key # or: Valar only ``` If Cursor is running, valar offers to quit it, apply the change, and reopen it. Pass `--force` to pre-answer that prompt; scripts and MDM need it, since a non-interactive run fails rather than ask. Valar only is sticky: a later `valar cursor on` keeps it. Run `valar cursor on --valar-only=false` to go back to side by side. `valar cursor off` also clears it. ### The proxy lane and the direct lane On macOS, Cursor's chat goes through the on-device proxy (the **proxy lane**). The proxy is what lets Cursor's own models keep working side by side, and what makes Valar only hold: it checks each message, switches a pick of one of Cursor's models to Valar Auto, and blocks what it cannot keep on Valar. Without the proxy, Cursor uses the **direct lane**: Valar is Cursor's OpenAI endpoint, and nothing checks messages on the way. In Valar only, a check inside Cursor blocks a send on one of Cursor's own models with `'' is not routed through Valar. Select a valar- model from the picker.` The direct lane is the only lane on Linux. | Situation | What `valar cursor on` does | | - | - | | macOS, proxy ready | Proxy lane | | macOS, proxy missing or not ready, side by side | Direct lane, and prints the fix (such as `valar proxy enable`). On the direct lane some of Cursor's own models, such as Composer, don't work next to Valar's. | | macOS, no proxy installed, Valar only | Direct lane with Valar only, as in earlier releases. A developer can get around it; `valar proxy enable`, then `valar cursor on --valar-only`, makes it strict. | | macOS, proxy installed but not ready, Valar only | Stops without changing Cursor, prints the fix (such as `valar proxy restart`), and names `--direct` for the weaker setup | | macOS, Cursor already set to another proxy, or a company VPN proxy the Valar proxy doesn't go through | Side by side: direct lane. Valar only: stops. Both print the fix. | | `--direct` | Direct lane, on purpose. On this lane a developer can get around Valar only. | | Linux | Direct lane | Check the lane and state any time: ```bash theme={"system"} valar cursor status ``` `status` shows `on, proxy lane` or `on, direct lane`, whether Valar only is on, and a fix when something needs one. `partial` means one of Valar's settings was turned off inside Cursor. Signing out of Cursor does this. Run the `valar cursor on` command that `status` prints. ## When Cursor says "Message not sent" On the proxy lane, a blocked message shows Cursor's own **Message not sent** box with a short message saying why and what to do: | Message | What to do | | - | - | | Valar blocked this message because it uses a personal API key. | Pick a Valar model, or (side by side) one of Cursor's own models. Personal keys saved in Cursor are not sent. | | Valar models only work on This Mac. | Switch the environment to **This Mac**, or pick a Cursor model. | | Cloud runs aren't available with Valar. | Valar only: switch the environment to **This Mac**. | | Valar's key is turned off in Cursor, so this message was not sent. | Run `valar cursor on` (with `--valar-only` if you use it). | | Valar couldn't check this message, so it was not sent. | Usually a Cursor update Valar doesn't read yet. Run `valar upgrade`, then try again. | | Too many chats are open through Valar right now, so this one was not started. | Close some chats, or try again in a few minutes. | If the proxy is stopped, Cursor's chat stops too ("Reconnecting…"). `valar cursor status` names the fix, usually `valar proxy restart`. ## What gets written Cursor stores its settings in a SQLite database, `state.vscdb`, and in `settings.json`. `valar cursor on` sets: | Where | Key | Value | | - | - | - | | `state.vscdb` | `openAIBaseUrl` | `https://api.valarhq.ai/v1` | | `state.vscdb` | `useOpenAIKey` | `true` | | `state.vscdb` | *(the OpenAI key row)* | `vlrcode_…~jdoe` | | `state.vscdb` | `userAddedModels` | adds your Valar models (and `valar-auto` on the direct lane) | | `state.vscdb` | `modelOverrideEnabled`, `modelOverrideDisabled` | turns Valar's rows on; in Valar only, turns Cursor's own models off | | `settings.json` | `http.proxy` | proxy lane only: `http://127.0.0.1:18083` (the proxy's Cursor address by default) | | `settings.json` | `http.proxySupport` | proxy lane only: `override` | | `settings.json` | `cursor.general.disableHttp2` | proxy lane only: `true` (Cursor's HTTP compatibility mode) | On the proxy lane, Cursor's **Auto** is selected; on the direct lane, `valar-auto` is. Existing chats are switched too. In Valar only on the direct lane, `valar cursor on` also adds a check to `~/.cursor/hooks.json`. Before the first change, the touched rows and `settings.json` are snapshotted under `~/.valar/cursor/` (keyed by a hash of each file's path, so separate installs never collide). `off` restores them. Changes you made to `settings.json` after that are kept. Cursor's settings stored in `state.vscdb`, including model toggles and models you added, go back to how they were before `valar cursor on`. Cursor cannot set custom request headers, so the client id rides as a **token suffix** (`~`) rather than an `X-Valar-Client-Id` header. The gateway splits on the last `~`, authenticates the base key, and routes on the id. By default the id is your OS username; use `--user-id RANDOM` for an opaque one. See [per-engineer attribution](/valarcode/setup#per-engineer-attribution). Cursor keeps `state.vscdb` values as SQLite **TEXT** and silently discards any row whose value is a BLOB on launch. The CLI writes the values as TEXT for this reason; a raw byte write would be wiped on the next restart. ## Limits * Valar only covers chat. Cursor's Agent Review (Bugbot) still runs on Cursor's models; turn it off in your Cursor team settings if you need everything on Valar. Chat titles are also named by Cursor's own model, and Tab completions always run on Cursor's models. * Claude models inside Cursor are not served by Valar. In Valar only they are not offered. * `valar cursor` routes the Cursor **editor** (the desktop app). The Cursor CLI (`cursor-agent`) is not supported. * Behind a company VPN or proxy, set up the Valar proxy to go through it: `valar proxy enable --upstream-proxy
`. See [Corporate networks](/valarcode/proxy#corporate-networks). ## Turn off routing Quit Cursor first (or pass `--force`), then: ```bash theme={"system"} valar cursor off ``` This puts back what `valar cursor on` changed and turns Valar only off. Start Cursor again to pick up the change. The proxy stays installed for other tools; `valar proxy disable` removes it once nothing uses it. ## Next steps How cohorts and the split decide which model serves each request. The open-weight targets and the frontier Claude tiers. # Automate coding key provisioning Source: https://docs.valarhq.ai/valarcode/key-automation Use an organization OAuth credential to create, list, suspend, and revoke ValarCode keys for engineers Private Preview Organization admins can give an internal provisioning service an OAuth credential for managing ValarCode keys. The credential is limited to the organization that created it and can manage only keys created through this automation flow. Client secrets and generated ValarCode keys stay valid until you revoke them. OAuth access tokens expire after one hour; request another token with the same client credentials when needed. ## Create a credential 1. Open **Settings → ValarCode settings → Automation credentials** in the Valar dashboard. 2. Click **Create credential**. 3. Store the client secret immediately. It is shown only once. 4. Copy the client ID and token endpoint shown beside it. You can create additional client secrets for rotation. Deploy the new secret before revoking the old one. ## Provision a key with Python Install Requests: ```bash theme={"system"} python -m pip install requests ``` Provide the credential through your secret manager or environment. Do not commit the client secret. ```bash theme={"system"} export VALAR_M2M_CLIENT_ID="client_..." export VALAR_M2M_CLIENT_SECRET="..." export VALAR_M2M_TOKEN_ENDPOINT="https://signin.valarhq.ai/oauth2/token" export VALAR_ENGINEER_EMAIL="engineer@example.com" # Optional: copy the routing policy from an existing coding key. export VALAR_ROUTING_SOURCE_KEY_ID="key_..." ``` To find the source ID, open **ValarCode → Setup**, select the configured key under **Model routing**, and copy the **Key ID for the automation API** shown below the selector. The `vlrcode_...` value is only the key's display prefix and cannot be used as `source_key_id`. Exchange the credential for an access token, then create the engineer's key: ```python theme={"system"} import os import requests API_BASE_URL = os.getenv("VALAR_API_BASE_URL", "https://api.valarhq.ai") CLIENT_ID = os.environ["VALAR_M2M_CLIENT_ID"] CLIENT_SECRET = os.environ["VALAR_M2M_CLIENT_SECRET"] TOKEN_ENDPOINT = os.environ["VALAR_M2M_TOKEN_ENDPOINT"] ENGINEER_EMAIL = os.environ["VALAR_ENGINEER_EMAIL"] token_response = requests.post( TOKEN_ENDPOINT, data={ "grant_type": "client_credentials", "client_id": CLIENT_ID, "client_secret": CLIENT_SECRET, }, timeout=15, ) token_response.raise_for_status() access_token = token_response.json()["access_token"] headers = {"Authorization": f"Bearer {access_token}"} workspace_id = os.getenv("VALAR_WORKSPACE_ID") if workspace_id: headers["X-Workspace-Id"] = workspace_id payload = {"email": ENGINEER_EMAIL} source_key_id = os.getenv("VALAR_ROUTING_SOURCE_KEY_ID") if source_key_id: payload["routing_policy"] = {"source_key_id": source_key_id} response = requests.post( f"{API_BASE_URL}/v1/organization/keys", headers=headers, json=payload, timeout=15, ) response.raise_for_status() created = response.json() print(f"Created {created['id']} for {created['email']}") print(f"ValarCode key (shown once): {created['token']}") ``` The routing source key must belong to the selected workspace and already have a routing policy. If you omit `VALAR_WORKSPACE_ID`, Valar uses the organization's default workspace. Store the returned `token` immediately. Later list calls return the key ID, prefix, state, and creation time, but never the full key again. ## Update an engineer's routing Keep one coding key per routing policy as a template, for example `template: open_weights` on Auto routing and `template: frontier_and_open` with a manual mapping that includes Claude models, and copy a template onto an engineer's live key whenever their policy changes. The engineer keeps the same key and nothing is redistributed; the gateway serves the new routing within about a minute. ```python theme={"system"} response = requests.put( f"{API_BASE_URL}/v1/organization/keys", headers=headers, params={"email": ENGINEER_EMAIL}, json={"routing_policy": {"source_key_id": os.environ["VALAR_FRONTIER_TEMPLATE_KEY_ID"]}}, timeout=15, ) response.raise_for_status() ``` The response is the key record without its token, including `routing.version` and `routing.changed_at` for the routing it just wrote. Record the version: a later list showing a higher version means the routing was changed again, for example by hand in the dashboard. `404 Not Found` means the email has no live automation-managed key in the selected workspace. The request is idempotent: repeating it with the same template returns `200 OK` and changes nothing. It honors `X-Workspace-Id` like the other key endpoints, and the template must live in the same workspace as the engineer's key. Copies are snapshots. Editing a template later does not move the engineers already on it; send the same request again to roll the change out. The gateway refreshes routing about every 30 seconds per instance, so for up to half a minute after a change requests can still see either policy. Every copy appears in the routing change history under **Model routing → Change history** with the template's name as the reason and "Automation" as the actor. To move the engineer back, copy the Auto template the same way. A manual mapping applies to requests that name a Claude model (`claude-opus-5-5`, `claude-opus-5`, `claude-sonnet-5-5`, `claude-sonnet-5`, `claude-haiku-4-5`, `claude-fable-5-1`). A request for the Auto model id (`valar-auto`) is always served by Auto routing, whatever the key's mapping, because that id means "let Valar pick". To move an engineer onto a manual mapping, their coding tool must be set to a Claude model, not to Auto. ## List provisioned keys Use the same `headers` from the example above: ```python theme={"system"} response = requests.get( f"{API_BASE_URL}/v1/organization/keys", headers=headers, params={"email": ENGINEER_EMAIL}, # Omit params to list every automation-managed key. timeout=15, ) response.raise_for_status() for key in response.json(): routing = key.get("routing") # Absent until a routing policy has been saved for the key. print(key["id"], key["email"], "revoked" if key["revoked"] else "active", routing and f"routing v{routing['version']} since {routing['changed_at']}") ``` Add `include=templates` to also list the workspace's live console-made keys, so your automation can resolve a template by name instead of storing its id: ```python theme={"system"} response = requests.get( f"{API_BASE_URL}/v1/organization/keys", headers=headers, params={"include": "templates"}, timeout=15, ) response.raise_for_status() templates = {k["name"]: k["id"] for k in response.json() if k["provisioning_source"] == "dashboard"} frontier_template_id = templates["template: frontier_and_open"] ``` Template rows carry `name`, `id`, `prefix`, and `routing`, never a token. Key ids never change: updating a key's routing, from the API or the dashboard, writes a new routing version under the same id. Only revoking and recreating a key produces a new id. ## Read a key's usage The `id` of a provisioned key is the `key_id` the [ValarCode analytics endpoints](/usage-endpoints#valarcode-analytics) report on and filter by. Those endpoints accept the same automation credential (the `/v1/usage/coding/*` routes only; the organization-wide billing endpoints still need a Valar API key), so the provisioning service can read cost, input and output tokens, and the served-model mix per engineer with the `headers` it already has: ```python theme={"system"} usage = requests.get( f"{API_BASE_URL}/v1/usage/coding/models", headers=headers, params={"key_id": created["id"], "start": "2026-08-01", "end": "2026-08-31"}, timeout=15, ) usage.raise_for_status() for model in usage.json()["models"]: print(model["served_model"], model["input_tokens"], model["output_tokens"], model["spend"] / 100) ``` ## Suspend and restore an engineer's key Disabling a key is the reversible alternative to revoking it. The key keeps its id, its routing, and the token the engineer already has; it just refuses every request until you enable it again. Use it for leave, an offboarding hold, or a spend investigation, where revoke-and-remint would force a new token onto the engineer's machine. ```python theme={"system"} response = requests.patch( f"{API_BASE_URL}/v1/organization/keys", headers=headers, params={"email": ENGINEER_EMAIL}, json={"enabled": False}, # True to restore timeout=15, ) response.raise_for_status() print(response.json()["enabled"]) # False ``` The response is the key record without its token. Every list row also carries `enabled`, so a periodic reconciliation can tell suspended engineers from active ones without keeping state of its own. `404 Not Found` means the email has no live automation-managed key in the selected workspace. The request is idempotent: disabling a disabled key, or enabling an enabled one, returns `200 OK` and changes nothing. While a key is disabled, the engineer's coding tool receives `403 Forbidden` with the message `this API key is disabled — ask your organization admin to re-enable it`, not the `401` an unknown key gets, so they know who to ask. The gateway checks the switch on every request; there is no propagation delay. Usage recorded before the suspension stays attributed to the key, and the dashboard shows the key as **Disabled** in both key lists. A revoked key cannot be disabled or re-enabled: revocation is permanent, and a `PATCH` on an email whose only key is revoked returns `404`. ## Revoke an engineer's key Revocation is explicit and immediate: ```python theme={"system"} response = requests.delete( f"{API_BASE_URL}/v1/organization/keys", headers=headers, params={"email": ENGINEER_EMAIL}, timeout=15, ) response.raise_for_status() ``` Create requests return `409 Conflict` when that email already has a live automation-managed key in the selected workspace. Revoke the existing key before creating its replacement. # LibreChat Source: https://docs.valarhq.ai/valarcode/librechat Route a self-hosted LibreChat deployment through Valar, with usage attributed per user and per app [LibreChat](https://www.librechat.ai) is a self-hosted chat front-end your whole organisation shares. Unlike the other tools on this page it runs as a **server you deploy**, not as a CLI on an engineer's machine — so there is no `valar librechat on` command, and everything below is configuration you make once, in LibreChat's own config, rather than per person. Because it is shared, attribution is the part worth getting right: without it every request in the deployment arrives as one anonymous stream. LibreChat knows who its users are, and it can pass that through. ## What Valar needs Three headers, set once in your LibreChat endpoint config. | Header | Value | What it buys | | - | - | - | | `X-Valar-Client-Id` | `{{LIBRECHAT_USER_EMAIL}}` | attributes usage to the individual user | | `X-Valar-Harness` | `librechat` | attributes usage to the app | | `X-Session-Id` | `{{LIBRECHAT_BODY_CONVERSATIONID}}` | groups a conversation's requests together | `{{LIBRECHAT_USER_EMAIL}}` and `{{LIBRECHAT_BODY_CONVERSATIONID}}` are LibreChat's own template variables, substituted per request. **Send `librechat`, not your deployment's name.** `X-Valar-Harness` accepts a fixed set of tool names — `claude`, `codex`, `cursor`, `gemini`, `opencode`, `pi`, `copilot`, `librechat`. Anything else is ignored rather than rejected, and unidentified traffic on this path is recorded as **Claude Code** — which also means it is served on Claude Code's routing configuration. If you call your install something else internally, that name belongs in your own dashboards, not in this header. ### Use the email, not the username `X-Valar-Client-Id` is normalised to `[a-z0-9._@+-]`, and characters outside that set — **including spaces** — are dropped rather than replaced. `{{LIBRECHAT_USER_NAME}}` = "Jane Doe" becomes `janedoe`, which is hard to reconcile with a real person and collides with anyone else whose name flattens the same way. `{{LIBRECHAT_USER_EMAIL}}` survives intact (`@` and `+` are both kept) and is already unique. Prefer it. ## Connecting LibreChat reaches Valar through the Anthropic Messages API. Point it at an AI gateway you already run — the setup is the gateway's, and both of ours carry the same header contract: * **[Portkey](/valarcode/portkey)** — add Valar as an Anthropic provider with a custom host, [provision every Valar model](/valarcode/portkey#provision-every-valar-model), then name the three headers in the Config's `forward_headers`. LibreChat can select the auto-router with `@valar/valar-auto` or a direct model such as `@valar/zai-org/GLM-5.3`. Portkey does not forward client headers unless you list them, and omitting one fails silently. If the Portkey key behind LibreChat also fronts other providers, read [what a pinned Config does to them](/valarcode/portkey#if-one-portkey-key-serves-other-backends) first. A Config naming a provider routes every request through it and discards the `@slug/` prefix LibreChat sent. * **[LiteLLM](/valarcode/litellm)** — enable `forward_client_headers_to_llm_api` and every `x-`-prefixed header passes through together. Whichever you use, the headers go in LibreChat's endpoint config alongside your gateway's own: ```yaml theme={"system"} headers: x-valar-client-id: "{{LIBRECHAT_USER_EMAIL}}" x-valar-harness: "librechat" x-session-id: "{{LIBRECHAT_BODY_CONVERSATIONID}}" ``` **Only the Anthropic Messages path is supported.** LibreChat can also be configured as an OpenAI-compatible endpoint, which reaches a different Valar ingress. That path is untested for ValarCode and Valar's automatic model selection may pick a target it cannot serve. If you need it, [tell us](mailto:support@valarhq.ai) rather than working around it. ## Verify it worked In the Valar analytics dashboard, after a few requests: * Users appear individually, under their email address. * Their harness list shows **LibreChat** — not Claude Code. If everything lands under Claude Code, `X-Valar-Harness` is not arriving: either your gateway is not forwarding it, or it is carrying a value outside the supported set. If usage appears but is not attributed to anyone, `X-Valar-Client-Id` is not arriving. Both fail silently — a request with neither header still returns a correct answer — so the dashboard is the only place the problem shows up. ## What you get Once both headers land, [analytics](/valarcode/analytics) breaks the deployment down per user: spend, tokens, which models served them, and savings against the Claude tier each request stood in for. LibreChat appears as its own tool in [routing](/valarcode/routing), separate from your engineers' coding tools, and gets its own **Auto or Manual** choice: leave it on Auto and Valar picks a model per request, or switch it to Manual and map each Claude size to the model you want it served by. Changing it affects LibreChat only — your coding tools keep whatever they are set to. # LiteLLM integration Source: https://docs.valarhq.ai/valarcode/litellm Run ValarCode through a LiteLLM proxy so Claude Code's per-engineer client id and routing still reach Valar If you already run a [LiteLLM](https://docs.litellm.ai) proxy in front of your models, you can keep it and still use ValarCode. Claude Code points at LiteLLM, LiteLLM forwards to Valar, and Valar's per-engineer routing and attribution work as long as the client id makes it through. This page covers Claude Code. **Pi** works the same way: it also sends the client id in an `X-Valar-Client-Id` header, so the LiteLLM config below applies unchanged. Just connect Pi with `valar pi on --endpoint --api-key ` and restart Pi. Instructions for connecting Cursor through LiteLLM are coming soon (Cursor carries its client id as a token suffix rather than a header). ## How it fits together ```mermaid theme={"system"} flowchart LR CC["Claude Code"] -->|"X-Valar-* headers"| L["LiteLLM proxy"] L -->|"forwards x- headers"| V["Valar gateway"] V --> RT["Route request
Opus / Sonnet / Haiku / Fable to target model"] RT --> MS["Valar model serving"] ``` `valar claude on` writes three Valar headers into `~/.claude/settings.json`: `X-Valar-Client-Id` (attributes usage to an engineer, defaulting to their username), `X-Valar-Harness` (names the tool, from a fixed set of supported values) and `X-Valar-Cli-Version`. All are `x-`-prefixed, so LiteLLM forwards them together — getting that hop right is the one thing that matters. It matters more through a proxy than it does directly. LiteLLM makes its own request upstream, so the user agent Valar sees is LiteLLM's, not your harness's — the headers are the only signal left. Without `X-Valar-Harness`, everything on this path is recorded as Claude Code. An **unrecognised** value does the same thing. `X-Valar-Harness` accepts a fixed set of tool names — `claude`, `codex`, `cursor`, `gemini`, `opencode`, `pi`, `copilot`, `librechat` — and anything outside that set is ignored rather than rejected, so putting your own application or deployment name here also records the traffic as Claude Code, and serves it on Claude Code's routing configuration. Send the name of the tool actually making the request. Running [LibreChat](/valarcode/librechat) behind this proxy is a supported case with a page of its own. ## Configure LiteLLM Valar speaks the Anthropic Messages API at `https://api.valarhq.ai`, so LiteLLM routes to it as an `anthropic/*` model. Two settings matter: 1. A **wildcard model entry** so any model string Claude Code requests (Opus, Sonnet, Haiku, or whatever it names across versions) is forwarded to Valar in Anthropic Messages format. Token and cost accounting still works, because every request maps to a known Anthropic model. 2. **Header forwarding**, so LiteLLM passes Claude Code's `x-`-prefixed headers (the `X-Valar-*` set) on to Valar. Without this, they are dropped at the proxy and per-engineer attribution is lost. Add this to your LiteLLM config: ```yaml theme={"system"} model_list: # Wildcard: any model string Claude Code requests is routed to Valar using the # Anthropic Messages format. Handles Claude Code changing its model names across # versions. Token/cost accounting works because these map to known Anthropic models. - model_name: "*" litellm_params: model: "anthropic/*" api_base: https://api.valarhq.ai api_key: os.environ/VALAR_API_KEY general_settings: # Forward Claude Code's x- headers (the X-Valar-* set) to Valar. # NOTE: verify on first request that the header actually lands at Valar -- only # x-prefixed (and anthropic-beta) headers are forwarded; keep the "x-valar-" name. forward_client_headers_to_llm_api: true ``` Set `VALAR_API_KEY` to your **Valar coding key**, the `vlrcode_…` key you generated in the [Valar Dashboard](https://app.valarhq.ai). LiteLLM uses it to authenticate to Valar upstream. The `X-Valar-Client-Id` header is `x-`-prefixed, which is exactly what LiteLLM's `forward_client_headers_to_llm_api` forwards. The header name matters: keep the `x-valar-` prefix. It is worth confirming on your first request (in Valar's request logs or the analytics dashboard) that the header actually arrives, since per-engineer attribution only works if it does. ## Create a virtual key in LiteLLM LiteLLM sits between Claude Code and Valar, so Claude Code no longer talks to Valar directly. Instead it authenticates to LiteLLM with a **virtual key** you create there. In the LiteLLM admin UI (or via the LiteLLM API), create a virtual key that is allowed to reach the Anthropic models you just configured, i.e. the wildcard `*` entry that forwards to Valar. Copy the key; Claude Code will use it as its bearer token. What flows through after this: * Claude Code authenticates to LiteLLM with the **virtual key**. * LiteLLM authenticates to Valar with your **Valar coding key** (`VALAR_API_KEY`). * Claude Code's `X-Valar-*` headers are forwarded to Valar, so routing and attribution behave exactly as in the direct path. ## Connect Claude Code Now point Claude Code at the LiteLLM proxy instead of Valar directly. Run `valar claude on` with `--endpoint` set to your LiteLLM base URL and `--api-key` set to the virtual key you just created: ```bash theme={"system"} valar claude on --endpoint --api-key ``` For example, if LiteLLM is running at `https://litellm.example.com`: ```bash theme={"system"} valar claude on --endpoint https://litellm.example.com --api-key sk-litellm-virtual-key ``` The CLI writes LiteLLM's base URL to `ANTHROPIC_BASE_URL`, the virtual key to `ANTHROPIC_AUTH_TOKEN`, and the Valar headers to `ANTHROPIC_CUSTOM_HEADERS` in `~/.claude/settings.json`, the same as the direct path but pointed at your proxy. Claude Code appends `/v1/messages` to the base URL, so pass the host without a `/v1` suffix, the same way you would for the default Valar endpoint. If your LiteLLM proxy serves the Anthropic surface under a path prefix, include that prefix in `--endpoint`. ## Verify it works After the first few requests, check: * **LiteLLM logs**: requests are arriving from Claude Code and being forwarded upstream. * **Valar analytics dashboard**: the engineer shows up under their client id (their username by default) with attributed usage. If usage appears but is not attributed to an engineer, the `X-Valar-Client-Id` header is not making it through; re-check `forward_client_headers_to_llm_api` and the header name. To stop routing through LiteLLM and restore Claude Code's previous settings: ```bash theme={"system"} valar claude off ``` ## Next steps The direct connect path and what the CLI writes to `settings.json`. The same setup for a Portkey gateway. How the client id drives cohort assignment and the split. # Roll out with MDM Source: https://docs.valarhq.ai/valarcode/mdm Deploy and run valar from Jamf, Kandji, Intune, or other MDM tools on macOS and Windows Use your MDM to install `valar` on engineers' Macs and Windows 11 PCs, set it up, and keep it current without anyone typing a command. This page explains how `valar` runs under MDM and in what order to deploy it. The sections before [Windows](#windows) describe macOS. ## How valar runs under MDM MDM scripts run as root, in the background, with no terminal. `valar` is built for a signed-in engineer, so three rules follow. **Every command acts for one user.** `valar` changes that user's configuration in `~/.valar`, each app's settings in the user's home, and the proxy services that run in the user's login session. A command run as root with no user selected is refused (`V1A-1004`). A command acts for the user in one of two ways: * **Proxy setup runs as root, for a selected user.** `valar proxy enable` needs administrator access: it sets the network proxy and installs the proxy services. Run it from a root policy with `SUDO_USER` set to the target account. `valar` does the user's part of the setup as that user. * **Everything else runs as the user.** `valar configure`, `valar upgrade`, and `valar on|off` need no administrator access. Run them as the signed-in user, either through your MDM's option to run as the signed-in user or by switching to that user from a root script (see the [example](#example-run-as-the-signed-in-user)). Run `valar upgrade` this way, not through `SUDO_USER`. **The user must be signed in.** The proxy services run in the user's login session, so the account and its home directory must exist and a session must be open. **Nothing can ask for confirmation.** `valar` asks before it restarts a running app. Under MDM it cannot ask, so pass `--force` (see [Restarts](#restarts)). ## Choose a lane The on-device proxy is optional. Without it, `valar on` uses the direct lane. The proxy lane adds [quota-first](/valarcode/claude-code#quota-first-use-your-own-subscription-first) and Remote Control. For [Cursor](/valarcode/cursor) on macOS, it keeps Cursor's own models working next to Valar's and makes Valar only strict. It needs two things on every Mac: * **A root certificate your MDM trusts.** See [Set up the proxy from a root policy](#set-up-the-proxy-from-a-root-policy). * **Each engineer signed in to Claude Code with their Claude account.** Without it, Claude Code stops at "Not logged in". For the direct lane, skip step 4 below. ## Deploy 1. **Install the CLI** at a path the user can run, such as `/usr/local/bin/valar`. See [Install the CLI](/valarcode/setup#install-the-cli). 2. **Give each engineer a coding key.** Create keys in the dashboard or with [key automation](/valarcode/key-automation). Store each key as that user with `valar configure --api-key `, or provide it through a protected `VALAR_API_KEY` environment variable. Keep keys out of policy arguments and logs. 3. **Upgrade existing installs, as the user:** `valar upgrade --force`. Finish this before setting up the proxy. 4. **Set up the proxy, from a root policy** (proxy lane only, see [below](#set-up-the-proxy-from-a-root-policy)). 5. **Turn on each harness, as the user:** ```bash theme={"system"} valar claude-desktop on --force valar claude on --force valar cursor on --force # macOS and Linux; add --valar-only for Valar only ``` 6. **Verify the proxy lane** with `valar status --verify` after the user sends a request through that harness. On the direct lane, use `valar status`. 7. **Keep it current.** Whenever your MDM installs a new `valar` binary, run `valar upgrade --force` as the user. Until it runs, most commands refuse (`V1A-1903`). ## Set up the proxy from a root policy Under MDM, bring your organization's certificate. macOS trusts a new root only from an MDM profile or a person at the screen, so the certificate `valar proxy enable` creates on its own fails from a policy (`V1A-1113`). Before the policy runs: 1. **Deliver your root CA as an MDM configuration profile** with a Certificate payload. Until it arrives, `valar proxy enable` stops with `V1A-1106`. 2. **Put a leaf certificate for `api.anthropic.com` and its key on the Mac**, readable by root. See [Create an organization-managed certificate](/valarcode/proxy#create-an-organization-managed-certificate). Then run: ```bash theme={"system"} #!/bin/bash set -euo pipefail TARGET_USER="jdoe" VALAR_BIN="/usr/local/bin/valar" if [ "$(id -u)" -ne 0 ]; then echo "Run this policy as root." >&2 exit 1 fi if ! TARGET_UID="$(id -u "$TARGET_USER" 2>/dev/null)" || [ "$TARGET_UID" -eq 0 ]; then echo "Select an existing, non-root user." >&2 exit 1 fi SUDO_USER="$TARGET_USER" "$VALAR_BIN" proxy enable \ --cert /absolute/path/leaf-bundle.pem --key /absolute/path/leaf.key ``` A root policy can read root-only certificate files; the installed copies are private to the user. Use absolute paths. If IT manages traffic steering, add `--ingress external`; if your network requires an outbound proxy, add `--upstream-proxy`. See [Corporate networks](/valarcode/proxy#corporate-networks). Setting up the proxy does not turn on any harness. Running it again keeps each harness's on/off choice but restarts the proxy, interrupting requests in flight, so run it on first setup and after renewing the certificate. If you use Cursor, also run it once after upgrading to a release with Cursor support, so the certificate covers Cursor (reissue an organization-managed certificate with Cursor's names first). ## Restarts Some changes apply only after a running app restarts. With a terminal, `valar` asks first. Under MDM it cannot ask, so pass `--force` to every `valar upgrade` and `valar on|off`. Without it, a command that needs a restart changes nothing and exits with an error (`1` for `on`/`off`, `20` for `upgrade`). `--force` restarts what the change needs: * **Claude Desktop** quits and reopens in the background. * **Cursor** quits and reopens. * **Claude Code** closes interactive sessions and restarts background agents. Users reopen conversations with `claude --resume`. On `valar upgrade`, `--force` also accepts any breaking changes the upgrade lists. Restarts interrupt work in progress, so schedule forced runs at login or outside working hours. `--force` does not bypass certificate, ownership, or network checks. ## Example: run as the signed-in user If your MDM can run a script as the signed-in user, run the `valar` commands directly. Otherwise, switch to that user from a root script, as this macOS example does. Adapt it to your fleet, for example by listing only the harnesses your engineers use. ```bash theme={"system"} #!/bin/bash set -e user=$(stat -f %Su /dev/console) if [ "$user" = root ]; then echo "No user is signed in; retry later." >&2 exit 1 fi uid=$(id -u "$user") as_user() { launchctl asuser "$uid" sudo -u "$user" -H /usr/local/bin/valar "$@" } as_user upgrade --force as_user claude-desktop on --force as_user claude on --force ``` ## Windows On Windows, no `valar` step needs administrator access, and every command runs as the signed-in user, including `valar proxy enable`. `valar` sets up the account that runs it, so a script that runs as SYSTEM does not set up the engineer's account. In Intune, set **Run this script using the logged on credentials** to **Yes**; with other tools, use their equivalent option. Windows asks the user to confirm before a certificate is added to their own store, and an unattended script cannot answer. An unattended `valar proxy enable` that would create its own certificate therefore stops before changing anything (`V1A-1128`). Deploy an organization certificate instead: 1. **Trust your organization's root** on each device: an Intune trusted certificate profile or Group Policy that installs it in the computer's **Trusted Root Certification Authorities** store. See [Trust and distribute](/valarcode/proxy#trust-and-distribute). 2. **Deliver each device's leaf bundle and key** to a location only the signed-in user can read. Setup copies both into its own protected folder. 3. **Install the CLI.** The [installer](/valarcode/setup#install-the-cli) places it at `%USERPROFILE%\.local\bin\valar.exe`. 4. **Run the setup as the signed-in user:** ```powershell theme={"system"} $ErrorActionPreference = "Stop" $valar = Join-Path $env:USERPROFILE ".local\bin\valar.exe" function Invoke-Valar { & $valar @args if ($LASTEXITCODE -ne 0) { exit $LASTEXITCODE } } Invoke-Valar upgrade --force Invoke-Valar proxy enable --cert C:\path\to\leaf-bundle.pem --key C:\path\to\leaf.key Invoke-Valar claude-desktop on --force Invoke-Valar claude on --force ``` Store the coding key first with `valar configure --api-key ` as the same user, or provide `VALAR_API_KEY` through a protected channel. Keep keys out of script arguments and logs. If your organization applies one proxy setting to every user of the machine, `proxy enable` stops (`V1A-1125`): add `--ingress external` and deploy the Valar rule in your PAC. If your network requires an outbound proxy, add `--upstream-proxy`. See [Corporate networks](/valarcode/proxy#corporate-networks). **Keep it current.** When your tool installs a new `valar.exe`, run `valar upgrade --force` and then `valar proxy restart` as the user. The restart moves the proxy services onto the new version. **Verify** with `valar status --verify` after the user sends a request through that harness. To check the device without Valar, see [Deployment checks and diagnostics](/valarcode/proxy#deployment-checks-and-diagnostics). ## Update older policies * **Selecting the user:** root policies that relied only on `HOME` must set `SUDO_USER` or run as the user. * **Old command names:** `valar claude-desktop proxy` still works as an alias for `valar proxy`. On `proxy enable`, `--force`, `--no-restart`, and `--off` are accepted and ignored with a warning. Put `--force` on the harness command instead, and use the harness's `off` to stop routing. # ValarCode models Source: https://docs.valarhq.ai/valarcode/models Available coding models, dashboard model selection, and how each harness is served ValarCode supports open-weight coding models, Claude tiers, and OpenAI models. Choose the supported model for each key and tool in the dashboard. ## Open-weight coding models These are the cheaper targets you route traffic to. They are regular catalog models, and ValarCode prices each request at the model's **Now** tier rate, which you can see on the [Pricing](/pricing) page. | Model | Model id | Maker | | - | - | - | | GLM-5.3 Fast | `zai-org/GLM-5.3-fast` | Z.ai | | GLM-5.3 | `zai-org/GLM-5.3` | Z.ai | | GLM-5.3 Flash | `zai-org/GLM-5.3-Flash` | Z.ai | | GLM-5.2 | `zai-org/GLM-5.2` | Z.ai | | GLM-5.2 Fast | `zai-org/GLM-5.2-fast` | Z.ai | | Kimi-K3 | `moonshotai/Kimi-K3` | Moonshot AI | | Kimi-K2.7 | `moonshotai/Kimi-K2.7` | Moonshot AI | | DeepSeek-V4-Pro | `deepseek-ai/DeepSeek-V4-Pro` | DeepSeek AI | ## Frontier Claude tiers These stand in for the Opus, Sonnet, Haiku, and Fable classes your harness asks for, and they are the baseline that savings are measured against. They are also the fallback when a key's routing cannot be read. | Alias | Fallback target | | - | - | | Opus | Claude Opus 5 | | Sonnet | Claude Sonnet 5.5 | | Haiku | Claude Haiku 4.5 | | Fable | Claude Fable 5.1 | Claude Sonnet 5 stays selectable in the routing picker. Claude Fable 5, Opus 4.8 and Sonnet 4.6 are also selectable, though nothing falls back to them. The frontier Claude tiers are specific to ValarCode routing. They do not appear on the public [Models](/models) page and are not callable as standalone models from the Responses or Chat Completions APIs. They exist so a cohort can hold a Claude baseline and so savings compare against real Claude list prices. ## OpenAI models You can select these models in the ValarCode dashboard for Claude Code, Codex, and OpenCode when enabled for your organization. | Model | Model id | | - | - | | GPT-6 Astra | `gpt-6-astra` | | GPT-5.6 Sol | `gpt-5.6-sol` | | GPT-5.6 Terra | `gpt-5.6-terra` | | GPT-5.6 Luna | `gpt-5.6-luna` | Astra currently supports up to 272,000 tokens of context on Valar and requires reasoning to stay enabled. Claude and OpenAI models also bill input tokens written into the prompt cache, at 1.25x the input rate (2x for a Claude write cached with the 1-hour TTL). Other models have no cache-write charge. See [Cache writes](/pricing#cache-writes). On the **Codex** and **OpenCode** tabs, choose **Manual** and select a model for all requests from that key and tool. Claude Code keeps its model-by-model mapping. Each picker shows the models supported by that tool. Choose **US only 🇺🇸** under **Auto-routing** to use models from US vendors for the selected key and tool. ## How each harness is served ValarCode meets each tool on its own API: | Harness | API | Endpoint | | - | - | - | | Claude Code | Anthropic Messages | `/v1/messages` | | Claude Desktop | Anthropic Messages | `/v1/messages` | | Pi | Anthropic Messages | `/v1/messages` | | Oh My Pi | Anthropic Messages | `/v1/messages` | | VS Code (Copilot Chat) | Anthropic Messages | `/v1/messages` | | OpenCode | Anthropic Messages | `/v1/messages` | | Codex | OpenAI Responses | `/v1/responses` | | Cursor | OpenAI Chat Completions | `/v1/chat/completions` | Every endpoint requires a coding key. Requests are forwarded to the resolved target model, and the response is rewritten so the model your harness sees is the one it asked for. That keeps the harness working across restarts, whatever target actually served the request. Whatever the target, Valar serves it on throughput-optimized inference tuned for agentic workloads. That highly efficient serving is what makes the open-weight models a practical everyday default rather than something you reach for only on cheap background work. ## Next steps Point each alias at one of these models per cohort. See the served-model mix and per-model savings. # Oh My Pi Source: https://docs.valarhq.ai/valarcode/oh-my-pi Route Oh My Pi (omp) through Valar, and what the CLI writes to its ~/.omp/agent config files [Oh My Pi](https://omp.sh) is a fork of Pi, and ValarCode routes it the same way: the CLI registers its own `valar` provider with a pinned model list, so the picker offers exactly Valar's models (Auto plus the open-weight models). omp reports to Valar as Pi, so its routing and usage attribution are Pi's. See [Set up ValarCode](/valarcode/setup) for the shared setup and the client-id model. ## Prerequisites * Oh My Pi installed and launched at least once (so `~/.omp/agent` exists). * A Valar coding key (`vlrcode_...`) and the `valar` CLI ([install](/valarcode/setup#install-the-cli)). ## Enable routing Run once per machine: ```bash theme={"system"} valar ohmypi on --api-key ``` omp reads its config at startup, so **restart omp** for the change to take effect. Check the state any time: ```bash theme={"system"} valar ohmypi status ``` `omp models` lists the Valar catalog; pick a model with `/model` in a session or `--model valar/` on the command line. Auto (`valar-auto`) is the default. ## What gets written omp keeps its config in YAML under `~/.omp/agent/`. `valar ohmypi on` edits two files, preserving other providers, roles and unrelated keys: **`models.yml`** declares the `valar` provider against the gateway, with the key, the model list and the headers that attribute your traffic: ```yaml theme={"system"} providers: valar: baseUrl: https://api.valarhq.ai apiKey: vlrcode_... api: anthropic-messages auth: apiKey headers: X-Valar-Client-Id: jdoe X-Valar-Harness: pi X-Valar-Cli-Version: 1.4.2 models: - id: valar-auto name: Valar Auto ... ``` **`config.yml`** points the default model role at it: ```yaml theme={"system"} modelRoles: default: valar/valar-auto ``` A default already on a Valar model is kept when Valar still offers it, along with any thinking selector you appended to it (`valar/zai-org/GLM-5.2:high`). `auth: apiKey` is required. Without a declared auth mode, omp treats a custom `anthropic-messages` provider as a Claude Code proxy and changes the request: it places Claude Code's system prompt ahead of yours and caps `max_tokens` at 64000. Declaring the mode turns that off, so your models get omp's own prompt and the full output length. If you run omp under a profile (`OMP_PROFILE`, `PI_PROFILE`) or with a relocated config (`PI_CONFIG_DIR`, `PI_CODING_AGENT_DIR`), the CLI follows it and writes that profile's files, so run `valar ohmypi on` with the same environment your omp sessions use. Both files are written user-only (`0600`), since they hold the key. Before the first change, the originals are snapshotted to `~/.valar/ohmypi/backup.json`, so `off` can restore them. The files are rewritten as plain YAML, so comments in them do not survive `on`. `off` restores the originals byte-for-byte when nothing else changed while routed; otherwise it removes only the `valar` provider and default role and keeps your other edits. ## Turn off routing ```bash theme={"system"} valar ohmypi off ``` Restart omp for the change to take effect. ## Next steps How cohorts and the split decide which model serves each request. The original Pi harness, which omp mirrors. # opencode Source: https://docs.valarhq.ai/valarcode/opencode Route the opencode CLI, desktop app, and mobile clients through Valar by editing ~/.opencode/opencode.jsonc ValarCode routes [opencode](https://opencode.ai) through Valar by editing its `opencode.jsonc`. opencode keeps speaking the Anthropic Messages API and you keep picking your model; the gateway maps whatever it sends to the target your routing resolves to. See [Set up ValarCode](/valarcode/setup) for the shared setup and the client-id model. One command covers every way opencode runs — the terminal UI, the desktop app, and a phone driving a session over `opencode serve` — because all three read the same config file. There is nothing extra to enable per surface. ## Prerequisites * opencode installed, by any method — the install script, Homebrew (`brew install opencode`), or the desktop app (`brew install --cask opencode-desktop`). * A Valar coding key (`vlrcode_…`) and the `valar` CLI ([install](/valarcode/setup#install-the-cli)). ## Enable routing Run once per machine: ```bash theme={"system"} valar opencode on --api-key ``` The routing lives entirely in `opencode.jsonc`, so a new opencode session picks it up on start with no env var, shell-rc edit, or new-terminal step. Check the state any time: ```bash theme={"system"} valar opencode status ``` **Restart opencode to apply a change.** opencode reads its config once when it starts and caches it for the life of the process. If you run `valar opencode on` or `off` while the terminal UI, the desktop app, or a `serve` process is already open, quit and reopen it to pick up the change. ## What gets written `valar opencode on` edits `~/.opencode/opencode.jsonc`. It adds a Valar provider that points at the gateway and carries your credentials, and pins both model slots to the Valar auto-router: ```jsonc theme={"system"} { "provider": { "valar": { "name": "Valar", "options": { "baseURL": "https://api.valarhq.ai/v1", "apiKey": "vlrcode_…", "headers": { "X-Valar-Client-Id": "jdoe", "X-Valar-Harness": "opencode", "X-Valar-Cli-Version": "0.9.32" } }, "models": { "…": "the Valar catalog" } } }, "model": "valar/valar-auto", "small_model": "valar/valar-auto" } ``` In detail, enabling: * Adds a `valar` provider whose `baseURL` is the `/v1` surface. opencode's `@ai-sdk/anthropic` dialect appends `/messages`, so requests land on `https://api.valarhq.ai/v1/messages`. * Puts the coding key on `options.apiKey`. In config it wins over opencode's own auth store, so there is no login conflict. * Adds the `X-Valar-Client-Id`, `X-Valar-Harness` and `X-Valar-Cli-Version` request headers. The client-id header is omitted entirely when there is no client id, rather than written blank. * Writes the full Valar model catalog under the provider's `models`. opencode has no public entry for these models, so what Valar writes here **is** the model picker you see in the app. * Pins **both** `model` and `small_model` to `valar/valar-auto`. The small slot matters: left unset, opencode picks its own model for the quick "name this session" call, which would route outside Valar. Your comments, layout, trailing commas, and any other providers pass through byte-for-byte. The original `opencode.jsonc` is snapshotted under `~/.valar/opencode/` before the first change, so `off` restores it exactly — including any edits you made while routing was on. The client id defaults to your OS username, so the header reads `X-Valar-Client-Id: jdoe`. Use `--user-id RANDOM` for an opaque `vc_…` id instead. See [per-engineer attribution](/valarcode/setup#per-engineer-attribution). ## The desktop app and mobile The opencode desktop app and the phone clients (which drive a session over `opencode serve`) run the **same engine** as the terminal, and read the same `~/.opencode/opencode.jsonc`. So `valar opencode on` routes them too — there is no separate desktop or mobile command. * **Desktop app** — enable routing, then open (or restart) the app. It shows the Valar models in its picker with **Valar Auto** selected. * **Remote / mobile** — start the server with `opencode serve`, connect your phone client to it on your network, and the server's requests route through Valar. The phone never contacts Valar directly; your machine's engine does, so usage attributes to `opencode` exactly as the terminal does. For any real remote use, set `OPENCODE_SERVER_PASSWORD` — an unsecured `serve` is open to anyone on your network. ## Turn off routing ```bash theme={"system"} valar opencode off ``` This restores the pre-enable `opencode.jsonc`. Start a new opencode session — or restart the app — to pick up the change. ## Next steps How cohorts and the split decide which model serves each request. The open-weight targets and the frontier Claude tiers. # ValarCode overview Source: https://docs.valarhq.ai/valarcode/overview Route your coding agents through Valar to cut spend, with per-engineer routing and measured savings ValarCode connects your coding tools to Valar, with model selection, routing controls, and usage attributed to each engineer. It supports Claude Code, Claude Desktop, Cursor, Codex, Pi, Oh My Pi, opencode, and VS Code through Copilot Chat. On macOS and Windows 11, Claude Code and Claude Desktop can use a shared [on-device proxy](/valarcode/proxy) while keeping their Claude login; on macOS, Cursor can use it too. Direct integrations configure the tool to use the Valar endpoint. You control routing with your coding key and review usage and savings in the dashboard. ## Why teams use it * Choose frontier and open models from supported coding tools. * Manage model routing without changing each engineer's setup for every policy change. * Use eligible Claude subscription allowance first through the optional on-device proxy. * Track usage and savings per engineer. ## How it works 1. Configure a coding key with the Valar CLI. 2. Connect each harness. Claude Code and Desktop can use the shared proxy or direct routing; other harnesses use their own integrations. 3. Choose models in the harness. Your key's policy controls requests served through Valar. With quota-first enabled on the proxy, eligible Claude requests can use your subscription instead. 4. Review Valar-served usage and savings in the dashboard. See [Model routing](/valarcode/routing) for routing controls and [On-device proxy](/valarcode/proxy) for shared setup and subscription-first routing. To deploy with MDM, see [Roll out with MDM](/valarcode/mdm). ## Get started Install the CLI, create a coding key, and connect a harness. Choose how your coding key routes model requests. Shared proxy setup and lifecycle. The open-weight targets and the frontier Claude tiers. Per-engineer usage and how savings are worked out. ## Supported harnesses Each harness has its own page covering how to connect it and what the CLI writes behind the scenes: `valar claude on` `valar claude-desktop on` `valar cursor on` `valar codex on` `valar opencode on` `valar pi on` `valar ohmypi on` `valar copilot on` ## LiteLLM If you run a LiteLLM proxy, you can keep it and still use ValarCode. Claude Code (or Pi) points at LiteLLM, which forwards requests to Valar. For per-engineer routing and attribution, LiteLLM must preserve the client ID header when forwarding requests. Route Claude Code through a LiteLLM proxy without losing per-engineer attribution. # Pi Source: https://docs.valarhq.ai/valarcode/pi Route Pi through Valar, and what the CLI writes to Pi's ~/.pi/agent config files ValarCode routes Pi through Valar by registering its own `valar` provider in Pi's config. Pi speaks the Anthropic Messages API, so it works the same way as Claude Code: the gateway maps what Pi sends to the target your routing resolves to. See [Set up ValarCode](/valarcode/setup) for the shared setup and the client-id model. ## Prerequisites * Pi installed, with a `~/.pi/agent` config directory. * A Valar coding key (`vlrcode_…`) and the `valar` CLI ([install](/valarcode/setup#install-the-cli)). ## Enable routing Run once per machine: ```bash theme={"system"} valar pi on --api-key ``` Pi reads its config at startup, so **restart Pi** for the change to take effect. Check the state any time: ```bash theme={"system"} valar pi status ``` ## What gets written Pi keeps its config in three JSON files under `~/.pi/agent/`. `valar pi on` edits all three, preserving other providers and unrelated keys: **`settings.json`** makes the `valar` provider active and pins a model from Valar's list: ```json theme={"system"} { "defaultProvider": "valar", "defaultModel": "claude-opus-4-8" } ``` Your current selection is kept when Valar offers that model; otherwise it moves to Valar's default. **`auth.json`** stores the bare coding key under a Valar-managed entry: ```json theme={"system"} { "valar": { "type": "api_key", "key": "vlrcode_…", "managedBy": "valar" } } ``` **`models.json`** declares the `valar` provider against the gateway, with its model list and the headers that attribute your traffic: ```json theme={"system"} { "providers": { "valar": { "name": "Valar", "baseUrl": "https://api.valarhq.ai", "api": "anthropic-messages", "compat": { "sendSessionAffinityHeaders": true }, "headers": { "X-Valar-Client-Id": "jdoe", "X-Valar-Harness": "pi", "X-Valar-Cli-Version": "1.4.2" }, "models": [ … ] } } } ``` Each entry in `models` carries one model's id, display name, context window, max output tokens, cost, and whether it accepts images. The array is the whole catalog for this provider — it is exactly what Pi's model picker offers. Your own builtin `anthropic` provider is left alone, still pointed at `api.anthropic.com`. Pi's Anthropic SDK appends `/v1/messages`, so a trailing `/v1` on the base URL is stripped before it is written. All three files are written user-only (`0600`), since they hold the key and the client-id header. Before the first change, the originals are snapshotted to `~/.valar/pi/backup.json`, so `off` restores them byte-for-byte. The client id defaults to your OS username, so the header reads `X-Valar-Client-Id: jdoe`. Use `--user-id RANDOM` for an opaque `vc_…` id instead. See [per-engineer attribution](/valarcode/setup#per-engineer-attribution). ## Turn off routing ```bash theme={"system"} valar pi off ``` This restores the three pre-enable files. Restart Pi for the change to take effect. ## Next steps How cohorts and the split decide which model serves each request. Pi works through a LiteLLM proxy the same way Claude Code does. # Portkey integration Source: https://docs.valarhq.ai/valarcode/portkey Run ValarCode through a Portkey gateway so Claude Code's per-engineer client id and routing still reach Valar If you already run [Portkey](https://portkey.ai/docs) as your AI gateway, you can keep it and still use ValarCode. Claude Code points at Portkey, Portkey forwards to Valar, and Valar's per-engineer routing and attribution work as long as the client id makes it through. The same Portkey provider can also expose every Valar model for direct selection by clients such as LibreChat. This page covers Claude Code. **Pi** works the same way: it also sends the client id in an `X-Valar-Client-Id` header, so the Portkey setup below applies unchanged. Just connect Pi with `valar pi on --endpoint https://api.portkey.ai --api-key ` and restart Pi. Instructions for connecting Cursor through Portkey are coming soon (Cursor carries its client id as a token suffix rather than a header). ## How it fits together ```mermaid theme={"system"} flowchart LR CC["Claude Code"] -->|"X-Valar-* headers"| P["Portkey gateway"] P -->|"saved forward_headers allowlist"| V["Valar gateway"] V --> RT["Route request
Opus / Sonnet / Haiku / Fable to target model"] RT --> MS["Valar model serving"] ``` `valar claude on` writes three Valar headers into `~/.claude/settings.json`: `X-Valar-Client-Id` (attributes usage to an engineer, defaulting to their username), `X-Valar-Harness` (names the tool, from a fixed set of supported values) and `X-Valar-Cli-Version`. Getting them across the Portkey hop is the one thing that matters. It matters more through a proxy than it does directly. Portkey makes its own request upstream, so the user agent Valar sees is Portkey's, not your harness's. The headers are the only reliable signal left. Without `X-Valar-Harness`, everything on this path is recorded as Claude Code. An **unrecognised** value does the same thing. `X-Valar-Harness` accepts a fixed set of tool names — `claude`, `codex`, `cursor`, `gemini`, `opencode`, `pi`, `copilot`, `librechat` — and anything outside that set is ignored rather than rejected, so putting your own application or deployment name here also records the traffic as Claude Code. That affects routing as well as reporting: traffic recorded as Claude Code is served on Claude Code's cohort configuration. Send the name of the tool actually making the request, and [talk to us](mailto:support@valarhq.ai) if yours isn't on the list. Portkey does **not** forward client headers by default. You must name the headers that may pass through. The recommended setup stores that allowlist in a Portkey Config. Some workspaces reject the equivalent per-request `x-portkey-forward-headers` setting as inline configuration. If the allowlist is missing or misspelled, requests can still succeed while the usage arrives at Valar unattributed. ## Add Valar to the Portkey Model Catalog Valar speaks the Anthropic Messages API at `https://api.valarhq.ai`, so Portkey treats it as an `anthropic` provider pointed at a custom host. In the Portkey dashboard, go to **Model Catalog → Add Provider** and create a provider with: | Field | Value | | - | - | | Provider | `Anthropic` | | Custom Host | `https://api.valarhq.ai/v1` | | API Key | your Valar coding key (`vlrcode_…`) | | Slug | `valar` | Set the API key to the `vlrcode_…` key you generated in the [Valar Dashboard](https://app.valarhq.ai). Portkey stores it and authenticates to Valar upstream, so the key never lands on an engineer's machine. This is the same split LiteLLM gives you with a virtual key. The Custom Host **must** include the `/v1` path segment. Portkey appends the endpoint path (`/messages`) to whatever you give it, so `https://api.valarhq.ai` without the suffix produces a 404 on every request. This is the opposite of the convention everywhere else in these docs, where you pass the Valar root and the client adds `/v1` itself. ## Provision every Valar model A new Anthropic integration starts with Anthropic's own model catalog. Add the Valar models to the integration before you create the Portkey API key. In Portkey, open the integration you created, then go to **Model Provisioning → Add Model**. Leave **Model Type** on **Custom Model** and add: 1. `valar-auto`, the Valar auto-router. 2. Every model id returned by Valar's Models API for your coding key, except the OpenAI-only ids below. ```bash theme={"system"} curl https://api.valarhq.ai/v1/models \ -H "Authorization: Bearer " ``` Use each returned `data[].id` as the Portkey **Model Slug**, preserving its spelling and case. You can also browse the current catalog on the [Models](/models) page. Skip the ids served only on Valar's OpenAI-compatible APIs. This integration speaks the Anthropic Messages API, and they have no route on it, so selecting one through Portkey returns an error instead of an answer. The Models API does not mark them; currently they are: `google/gemma-4-26B-A4B-it` · `google/gemma-4-31B-it` · `MiniMaxAI/MiniMax-M3` · `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` · `openai/gpt-oss-120b` · `Qwen/Qwen3.5-27B` · `Qwen/Qwen3.5-397B-A17B` · `Qwen/Qwen3.6-35B-A3B` Portkey combines the provider slug and model id in the request's `model` field: ```text theme={"system"} @valar/valar-auto @valar/zai-org/GLM-5.3 @valar/moonshotai/Kimi-K3 @valar/deepseek-ai/DeepSeek-V4-Pro ``` The first value lets Valar choose a model per request. The others select that exact Valar model. The model id can contain another `/`; Portkey treats only the first segment, `@valar`, as the provider selector. Provisioning a model makes it available through Portkey. It does not route every request to that model. The caller still chooses with `@valar/`, unless a Portkey Config fixes the provider or overrides the model. ## Generate a Portkey API key Portkey sits between Claude Code and Valar, so Claude Code no longer talks to Valar directly. Instead it authenticates to Portkey with a **Portkey API key** you create there. In the Portkey dashboard, go to **API Keys → Generate**, and scope it to the `valar` provider you just created. Copy the key; Claude Code will use it as its bearer token. What flows through after this: * Claude Code authenticates to Portkey with the **Portkey API key**. * Portkey authenticates to Valar with the **Valar coding key** stored in the Model Catalog. * Claude Code's `X-Valar-*` headers are forwarded to Valar, so routing and attribution behave exactly as in the direct path. ## Create a Portkey Config Store the forwarding allowlist in a Portkey Config. This works in workspaces that block inline gateway configuration. In the Portkey dashboard, go to **Configs → Create** and enter: ```json theme={"system"} { "provider": "@valar", "forward_headers": [ "x-valar-client-id", "x-valar-harness", "x-valar-cli-version", "anthropic-beta" ] } ``` Save the Config and copy its ID. Portkey Config IDs start with `pc-`, for example `pc-valar-a1bbcf`. The Config selects the `@valar` Model Catalog provider and tells Portkey which request headers it may pass to Valar. `anthropic-beta` is included because Claude Code uses Anthropic beta features that fail closed when the header is dropped. Use this pinned Config for clients such as Claude Code that send bare model names. Clients that send `@valar/` can select Valar without it. **This Config sends every request that carries it to Valar.** `provider` names one fixed target and there is no `strategy`, so Portkey makes no routing decision. A request that named its own provider does not get it: an `@slug/model` string like `@anthropic/claude-sonnet-5`, or an `x-portkey-provider` header, is discarded. Portkey resolves the provider from the Config unless the Config opts out, which is what [`passthrough`](#if-one-portkey-key-serves-other-backends) is for. That is what you want here. Claude Code sends bare `claude-*` model strings with no provider prefix, so something has to pick the provider, and the Config is where that happens. It is not what you want on a Config other traffic sees. Select this one per request, with the `x-portkey-config` header below, and **do not set it as the default Config on a Portkey API key that other applications share**. [If one key has to serve both](#if-one-portkey-key-serves-other-backends), drop the pin. ## Connect Claude Code Point Claude Code at Portkey with `valar claude on`, using your Portkey API key: ```bash theme={"system"} valar claude on --endpoint https://api.portkey.ai --api-key ``` `ANTHROPIC_AUTH_TOKEN` authenticates Claude Code to Portkey. Add the saved Config ID to the `ANTHROPIC_CUSTOM_HEADERS` line in `~/.claude/settings.json`: ```json theme={"system"} { "env": { "ANTHROPIC_BASE_URL": "https://api.portkey.ai", "ANTHROPIC_AUTH_TOKEN": "", "ANTHROPIC_CUSTOM_HEADERS": "X-Valar-Client-Id: jdoe\nX-Valar-Harness: claude\nX-Valar-Cli-Version: 1.4.2\nx-portkey-config: " } } ``` One header goes on each line as `Name: value`. This is the newline-separated format `valar claude on` already writes. Order does not matter, and neither does which you do first. `valar claude on` rewrites only its own three `X-Valar-*` lines and preserves every other header you set, so the Portkey lines survive re-enables, client-id changes and CLI upgrades. You can add them before or after the first `valar claude on`. Three things to keep straight in that block: * **`x-portkey-config` selects the saved Config.** Do not also set `x-portkey-provider`; the Config already selects `@valar`. The one exception is the unpinned Config [below](#if-one-portkey-key-serves-other-backends), which deliberately selects nothing. * **`ANTHROPIC_AUTH_TOKEN` is the Portkey credential.** You do not need to repeat it in an `x-portkey-api-key` custom header on the Model Catalog path. * **Do not set `ANTHROPIC_DEFAULT_*_MODEL`.** Portkey's generic Claude Code guide tells you to pin model IDs per provider; that advice does not apply here. Valar maps whatever model string Claude Code sends to the target model your cohort resolves to, and `valar claude on` deliberately strips those pins so the gateway's server-side routing stays authoritative. Keep choosing tiers with `/model`. Claude Code appends `/v1/messages` to the base URL, so pass `https://api.portkey.ai` without a `/v1` suffix. `valar claude on` strips a trailing `/v1` if you include one. ### If your workspace allows inline configuration If your Portkey workspace permits inline configuration, you can skip the saved Config and use these two lines instead of `x-portkey-config`: ```text theme={"system"} x-portkey-provider: @valar x-portkey-forward-headers: x-valar-client-id,x-valar-harness,x-valar-cli-version,anthropic-beta ``` The `@` prefix on `@valar` is required. A bare `valar` value is interpreted as a built-in provider name and fails. The forwarding list is comma-separated, and its header names are lowercase. If Portkey returns this error, your workspace blocks the inline form: ```text theme={"system"} Inline forward headers are not allowed when block_inline_config is enabled. Configure forward_headers on a saved provider instead. ``` Use the saved Config setup above. Replace `x-portkey-provider` and `x-portkey-forward-headers` with `x-portkey-config: `. ## If one Portkey key serves other backends One Portkey key in front of several providers — a LibreChat deployment offering a few backends, say — must not carry the pinned Config as its default. Every request on that key is force-routed to Valar, whatever model it asked for, and the `@slug/` prefix the caller sent is discarded. **Nothing errors when that happens.** A Valar coding key maps a model string it does not recognise onto the tier your [cohort](/valarcode/routing) resolves to rather than rejecting it, so a request for `@anthropic/eu.anthropic.claude-sonnet-5` or `@z-ai/zai.glm-5` comes back `200`, with a sensible answer, from a Valar-routed model — and on Valar's bill. The symptom is a provider that went quiet, not a failure. Two ways round it. **Give ValarCode its own Portkey API key**, with the pinned Config as its default. The shared key keeps its own Config and the two never meet. Take this one unless you have a reason not to. **Or drop the pin and let each request choose.** A Config can carry the forwarding allowlist and no provider at all: ```json theme={"system"} { "passthrough": true, "forward_headers": [ "x-valar-client-id", "x-valar-harness", "x-valar-cli-version", "anthropic-beta" ] } ``` `passthrough` tells Portkey to resolve the provider from the request rather than from the Config. `@valar/valar-auto` uses Valar's auto-router, `@valar/zai-org/GLM-5.3` selects that model directly, and `@anthropic/…` still reaches Anthropic. The Valar headers are forwarded on whichever hop carries them. A caller that names its provider in the model string, such as LibreChat, needs nothing further. Claude Code is the exception. It sends `claude-sonnet-…`, with no prefix for Portkey to read, so it has to name the provider itself: add `x-portkey-provider: @valar` to `ANTHROPIC_CUSTOM_HEADERS` alongside `x-portkey-config`. That is inline configuration, which some workspaces block; where it is blocked, give Claude Code its own key. The unpinned form comes from Portkey's [Config schema](https://portkey.ai/docs/api-reference/inference-api/config-object), and we have not run it end to end ourselves. The pinned Config above is the path we document and test. [Tell us](mailto:support@valarhq.ai) if the unpinned one misbehaves. ## Without the Model Catalog If you would rather not store the Valar key in Portkey, you can pass the upstream host and credential inline instead. This option requires a workspace that allows inline configuration. Drop the `x-portkey-provider: @valar` line and use: ``` x-portkey-api-key: x-portkey-provider: anthropic x-portkey-custom-host: https://api.valarhq.ai/v1 x-portkey-forward-headers: Authorization,x-valar-client-id,x-valar-harness,x-valar-cli-version,anthropic-beta ``` Then run `valar claude on --endpoint https://api.portkey.ai --api-key ` so `ANTHROPIC_AUTH_TOKEN` holds the `vlrcode_…` key. `Authorization` is in the forward list, so Portkey passes that credential through to Valar untouched rather than processing it; `x-portkey-api-key` is what authenticates to Portkey. The Valar coding key now sits in every engineer's `settings.json` instead of in Portkey. Prefer the Model Catalog path unless you have a reason not to. If `block_inline_config` is enabled, use the Model Catalog and saved Config path above. ## Verify it works After the first few requests, check: * **Portkey logs**: the request shows the saved Config ID and provider `valar`, and the response is successful. * **Valar analytics dashboard**: the engineer appears under their client id (their username by default), and their harness list includes the tool you configured — `claude` for Claude Code. Portkey's **User** field identifies the Portkey API-key owner. It does not prove that Valar received `X-Valar-Client-Id`. Check Valar analytics for the attribution result. If usage appears but is not attributed to an engineer, `X-Valar-Client-Id` is not making it through. Check the saved Config's `forward_headers` list first. A missing or misspelled entry is the common cause. This failure is silent. Portkey's Anthropic transform re-issues the upstream credential as `x-api-key` and adds `anthropic-version` independently of the forwarding allowlist. Authentication can still succeed when the allowlist is wrong. The request returns `200`, the answer is correct, and the only symptom is unattributed Valar usage. To stop routing through Portkey and restore Claude Code's previous settings: ```bash theme={"system"} valar claude off ``` `off` restores the whole `settings.json` from its backup, so the Portkey header lines go with it. ## Next steps The direct connect path and what the CLI writes to `settings.json`. The same setup for a LiteLLM proxy. How the client id drives cohort assignment and the split. Where per-engineer attribution shows up. # On-device proxy Source: https://docs.valarhq.ai/valarcode/proxy Set up, manage, and deploy the shared local proxy for supported coding harnesses The on-device proxy is a shared local service for supported coding harnesses. Install it once, then control Valar routing separately for each harness. The proxy runs on macOS and Windows 11. It currently supports [Claude Code](/valarcode/claude-code) (`claude`) and [Claude Desktop](/valarcode/claude-desktop) (`claude-desktop`), and on macOS [Cursor](/valarcode/cursor) (`cursor`). Other harnesses use separate routing integrations. ## Set up the proxy You need the [Valar CLI and a coding key](/valarcode/setup) and a supported harness. On macOS you also need administrator access for certificate trust and proxy settings. On Windows, setup needs no administrator access. Run these commands from the account that will use Valar: ```bash theme={"system"} valar configure --api-key valar proxy enable ``` Setup explains its machine changes before requesting administrator access on macOS, or before Windows asks you to confirm the local certificate. It installs the services without enabling harness routing or restarting applications. Existing ON/OFF choices stay unchanged. Then enable each harness you want to route. Replace `` with a supported harness name: ```text theme={"system"} valar on ``` First-time activation may require restarting running sessions. Follow the harness command's prompt and recovery instructions. ## Routing controls | Command | Effect | | - | - | | `valar proxy enable` | Provision or refresh the shared proxy. Preserve each harness's ON/OFF choice. | | `valar on` / `off` | Change routing for that harness only. | | `valar proxy restart` | Restart the installed proxy services. Active connections may be interrupted. | | `valar proxy disable` | Remove the proxy after every harness has left the proxy lane. | On an established proxy setup, ordinary ON/OFF toggles do not restart sessions. OFF stops inspection and Valar routing; new connections pass through the local proxy to their original destination without decryption. Proxy settings stay installed so the next ON needs no restart. Changes to the certificate, connection settings, or routing lane may require one. The services remain available when all harnesses are OFF. To remove them, follow [Remove the proxy](#remove-the-proxy). ## Subscription-first routing The proxy can apply subscription-first routing for integrations that support it. Today, [quota-first for Claude](/valarcode/claude-code#quota-first-use-your-own-subscription-first) is shared by Claude Code and Claude Desktop. It uses eligible included allowance before Valar; other harness subscriptions are not supported. ```bash theme={"system"} valar configure --quota-first=true ``` The running proxy reads this preference automatically. Use `--quota-first=false` to send all routed inference through Valar. Turning the preference off does not turn a harness off. See the harness documentation for supported accounts, model selection, counters, and reset commands. ## Status and logs ```bash theme={"system"} valar proxy status valar proxy logs valar proxy logs -f ``` `status` reports the services, listeners, certificate, and enabled harnesses. `logs` prints recent proxy activity; `-f` keeps following it. To verify Valar forwarding, send a request that uses Valar, then check the harness: ```text theme={"system"} valar status --verify ``` Verification requires a healthy proxy setup and a request successfully forwarded to Valar from that harness since routing was enabled. A subscription-served request does not satisfy the check. Restore subscription-first routing afterward if you disabled it for the test. ## Corporate networks Valar does not silently replace an existing corporate proxy or automatic-proxy policy. Configure the two directions separately: | Direction | Configuration | | - | - | | Harness → local Valar proxy | Your corporate PAC or the harness's proxy settings | | Local Valar proxy → network | `--upstream-proxy` with the corporate proxy's concrete URL | ```mermaid theme={"system"} flowchart LR H[Supported harness] -->|PAC or app proxy setting| V[Local Valar proxy] V -->|Configured upstream proxy| C[Corporate proxy] C --> D[Valar or original destination] ``` **Corporate PAC and an upstream proxy work together.** Provision without replacing the corporate PAC, and explicitly configure the onward connection: ```bash theme={"system"} valar proxy enable \ --ingress external \ --upstream-proxy http://proxy.example.com:8080 ``` Add `--cert` and `--key` when using organization-managed certificates. The upstream setting applies to every harness using this proxy, including subscription requests, other forwarded traffic, and undecrypted passthrough. Direct-lane sessions do not use this setting. ### Outbound corporate proxy If your shell already sets an outbound proxy, explicitly approve the same URL: ```bash theme={"system"} valar proxy enable --upstream-proxy http://proxy.example.com:8080 ``` The proxy uses the recorded URL for outbound connections. For HTTPS it asks the corporate proxy to open a tunnel to the destination with HTTP `CONNECT`: ```http theme={"system"} CONNECT api.valarhq.ai:443 HTTP/1.1 Host: api.valarhq.ai:443 ``` After the corporate proxy accepts the tunnel, Valar's local proxy establishes TLS to the destination through it and sends the request. Subscription/original-service traffic uses that service's host instead, such as `CONNECT api.anthropic.com:443`. Valar selects the destination; the corporate proxy applies its access policy and carries the connection. It must allow both destinations. The hostname changes if you configured another Valar endpoint. Supply an HTTP or HTTPS **proxy endpoint**, not the URL of a `.pac` file. Valar does not evaluate PAC JavaScript or inherit its destination-specific proxy choices. If your PAC selects several corporate proxies, ask IT for one endpoint that can reach all required destinations. The corporate proxy can be unauthenticated or use Basic authentication: `http://user:password@proxy.example.com:8080`, with reserved characters percent-encoded. NTLM, Kerberos and browser-based sign-in are not supported, which rules out most Windows proxies that use integrated Windows authentication. Ask IT for an endpoint that accepts Basic or no authentication for the destinations below. The proxy credential is separate from your Valar coding key. If the corporate proxy inspects TLS, its certificate chain must be trusted on the machine. Do not disable TLS verification. Test both a Valar-served request and an original-service request after setup. The URL is saved in a user-only local file and its credentials are redacted in status output. The command-line argument itself can still appear in shell history or process listings; use your organization's approved secret-handling process. A later `proxy enable` preserves the saved upstream if you omit the flag. Use `--upstream-proxy ` to replace it. `--upstream-proxy=""` clears it only when no conflicting outbound proxy remains in the shell. Never point the upstream at a Valar loopback listener. ### How this relates to HTTPS\_PROXY Programs launched from your shell can use its `HTTPS_PROXY` value if they support that variable. It does not configure the Valar daemon. The daemon reads only its saved upstream setting, which prevents it from accidentally proxying back to itself. * If `HTTPS_PROXY` already points to a corporate proxy, `proxy enable` requires the same URL to be explicitly approved by `--upstream-proxy`, or already recorded from an earlier setup. * Claude Code's own proxy settings point to its local Valar listener after `on`. Its requests then travel through Valar and onward through the configured corporate proxy. * A managed proxy setting takes precedence over user settings. IT must point it to the correct local listener; configuring an upstream does not override managed policy. * The `valar` commands themselves use standard proxy environment settings for their HTTP calls. On a PAC-only network that blocks direct access, provide the corporate `HTTPS_PROXY` value to CLI setup/upgrade commands as well. The saved daemon upstream is not a proxy setting for every program on the machine. ### A proxy set in network settings A fixed proxy in the machine's network settings (on Windows, **Use a proxy server** in **Settings > Network & internet > Proxy**) is treated the same way as `HTTPS_PROXY`. Apps that read the automatic-proxy document try it before the fixed proxy, so Valar's document would otherwise move those apps off your corporate proxy. `proxy enable` therefore stops and names the proxy until you approve that exact proxy: ```bash theme={"system"} valar proxy enable --upstream-proxy http://proxy.example.com:8080 ``` The approval must name the same scheme, host and port as the network setting. It may add a username and password that the network setting does not carry. Once it is approved, Valar's automatic-proxy document keeps your proxy for every other app: * Hosts Valar routes try the local Valar proxy first, then your corporate proxy. * Hosts on the proxy's bypass list and local addresses connect directly, as before. On macOS, plain host names do too. * Every other host goes to your corporate proxy. The document names only the proxy's host and port, never its credentials. `proxy enable` also stops when network services name different proxies or different bypass lists, or when a bypass entry cannot be reproduced exactly. It names the entry so you can rewrite it. `valar proxy status` warns if the network setting later changes to a proxy that is not approved. A machine set up with an earlier release over a fixed proxy keeps working as it did. `valar proxy status` warns that other apps are bypassing that proxy and shows the `--upstream-proxy` command that fixes it. ### Automatic proxy discovery (WPAD) Automatic proxy discovery does not stop `proxy enable`. When discovery is on (on Windows, **Automatically detect settings**, which is on by default), `proxy enable` checks whether the current network publishes a proxy script. Apps such as Claude Desktop use a discovered script ahead of Valar's document, so on a network that publishes one, Claude Desktop may not be routed through Valar. `proxy enable` and `valar proxy status` warn only in that case. To route Claude Desktop there, IT can add the Valar rule to the published script and provision with `--ingress external`, as described below. ### Administrator-managed traffic steering If IT owns the automatic-proxy policy or managed app settings, provision without replacing them: ```bash theme={"system"} valar proxy enable --ingress external ``` IT must then direct each harness to its own local listener. This option does not redirect traffic by itself. Current default addresses are listed below. Each harness must use its own listener: | Purpose | Address | | - | - | | Claude Desktop proxy | `127.0.0.1:18080` | | Claude Code proxy | `127.0.0.1:18082` | | Cursor proxy (macOS) | `127.0.0.1:18083` | | Automatic-proxy document | `http://127.0.0.1:18081/proxy.pac` | Use `valar proxy status` for the installed addresses. Set custom ports with `--port =` and `--pac-port `. Existing corporate outbound proxy requirements still apply with external steering; combine `--ingress external` with `--upstream-proxy` when both are needed. ### Add Valar to an existing corporate PAC Keep the corporate PAC URL. Add an exact-host rule for supported traffic before the existing catch-all rules. For the current PAC-based integration, Claude Desktop, use its listener (`18080` by default): ```javascript theme={"system"} function FindProxyForURL(url, host) { var normalizedHost = host.toLowerCase(); if (normalizedHost === "api.anthropic.com" || normalizedHost === "api.anthropic.com.") { return "PROXY 127.0.0.1:18080; PROXY proxy.example.com:8080"; } return CorporateProxyForURL(url, host); } function CorporateProxyForURL(url, host) { // Keep your existing FindProxyForURL body here, including its exclusions. // This minimal example sends other traffic through the corporate proxy. return "PROXY proxy.example.com:8080"; } ``` The same PAC works on macOS and Windows. On Windows, deliver its URL through the **Use setup script** proxy setting, with Intune or Group Policy. Rename your existing `FindProxyForURL` to `CorporateProxyForURL` without changing its body, then add the wrapper. The selected host uses the new local-proxy rule; all other hosts retain their existing corporate rules. Review that override with IT. The final function above is only a minimal example. Deploy this PAC change only to devices where Valar is provisioned. Keep local proxy/PAC addresses reachable without sending them back through another proxy. The semicolon-separated return value is a **fallback list, not a chain**. The app tries the local proxy first and may try the corporate proxy directly if the local proxy is unavailable. To make a healthy local Valar proxy send onward through the corporate proxy, you still need `--upstream-proxy`. Choose the failure policy with IT: | PAC result for the selected host | If the local proxy is unavailable | | - | - | | `PROXY 127.0.0.1:18080; PROXY proxy.example.com:8080` | The app may connect through the corporate proxy without Valar. | | `PROXY 127.0.0.1:18080` | No alternate route is supplied; affected requests can fail. | Do not append `DIRECT` if your organization prohibits direct internet access. PAC fallback is client-dependent and does not guarantee retrying an active request or bypassing HTTP errors. An upstream refusal returned by Valar is not the same as an unreachable local proxy. A PAC function receives a URL and host, **not the calling application's identity**. This host rule also affects other PAC-aware apps requesting that host. They reach the Desktop listener and follow its routing switch. Review that scope with IT; do not send multiple harnesses to one listener when you need independent controls. ### Configure harnesses that do not use system PAC Cursor does not use the system PAC either; `valar cursor on` points Cursor's own `http.proxy` setting at its listener. Claude Code does not use the system PAC for its proxy connection. Run `valar claude on` to configure its listener, or merge equivalent proxy values into IT-managed Claude settings: ```json theme={"system"} { "env": { "HTTPS_PROXY": "http://127.0.0.1:18082", "https_proxy": "http://127.0.0.1:18082", "HTTP_PROXY": "http://127.0.0.1:18082", "http_proxy": "http://127.0.0.1:18082", "NO_PROXY": "localhost,127.0.0.1,::1", "no_proxy": "localhost,127.0.0.1,::1" } } ``` Claude Code reads IT-managed settings from `/Library/Application Support/ClaudeCode/managed-settings.json` on macOS and `C:\Program Files\ClaudeCode\managed-settings.json` on Windows. Merge the proxy settings into the existing policy without replacing the whole file. For `NO_PROXY` and `no_proxy`, add the listed loopback addresses to your required internal bypasses rather than replacing the existing lists. Neither list should exclude `api.anthropic.com` or use `*`, which would bypass Valar. A conflicting lowercase proxy value can take precedence over uppercase, so keep both forms consistent. Use the installed listener address if you changed the default port. Even when IT delivers the connection settings, run the harness's `on` command to enable Valar routing. Settings alone do not turn its routing switch on. ### Roll out and verify corporate networking 1. Provision with `--ingress external`, the corporate upstream, and trusted certificates. 2. Deploy the PAC rule and any app-specific proxy settings to the selected devices. 3. Activate each supported harness and follow any restart prompt. A PAC update can remain cached in a running app. 4. Check `valar proxy status` for `external` ingress and the redacted upstream. Check each harness's status and send a Valar-served request. 5. Run each harness's `status --verify`. Test subscription/native traffic separately; successful Valar verification confirms only the Valar path. 6. In a pilot device, test your chosen local-proxy failure policy and confirm the corporate proxy sees the intended destinations. When retiring the deployment, turn the harnesses OFF, remove IT-managed local proxy pointers/PAC rules, and reload affected apps before removing the shared proxy. Valar cannot remove rules inside your corporate PAC or managed settings. ## Certificates and machine changes By default, setup: 1. Creates a local root certificate named **Valar Local Root CA** and trusts its public certificate: in the System keychain on macOS, and in your own user certificate store (`CurrentUser\Root`) on Windows. 2. Uses its private key once to sign a leaf certificate covering the supported harness destinations, then discards the root private key without saving it to disk. 3. Stores the leaf certificate and private key under `~/.valar/proxy/` with user-only permissions. 4. Installs two per-user background services: the proxy and the automatic-proxy document server. 5. Configures the automatic-proxy URL, unless you selected external steering: on each enabled network service on macOS, and in your Windows proxy setting (**Use setup script**) on Windows. The generated root is valid for 1,095 days and the leaf for 730 days. On macOS, administrator privileges handle certificate trust and system proxy changes, and can read protected certificate inputs supplied during root-run provisioning. The services and user configuration belong to the selected user. One user owns the proxy installation on a machine; on Windows, setup refuses while another account's proxy is running (`V1A-1015`), because the listening ports are shared by every account. ### On Windows Setup changes only your own Windows account and never asks for administrator access. Windows asks you once to confirm the Valar root certificate; choose **Yes**. If you choose **No**, setup stops without changes (`V1A-1123`); run `valar proxy enable` again and accept. If your organization's policy does not trust certificates that users install, setup stops without changes (`V1A-1124`); deploy an [organization-managed certificate](#organization-managed-certificates) instead. Setup stops before making changes when: * A policy makes Windows use one proxy setting for every user of the machine (`V1A-1125`). Your own setting would be ignored, so have IT add the Valar rule to the machine's script and provision with `--ingress external`. See [Administrator-managed traffic steering](#administrator-managed-traffic-steering). * A port it needs is in use or inside a range Windows reserves for Hyper-V or WSL (`V1A-1121`). Run `netsh interface ipv4 show excludedportrange protocol=tcp` to list the reserved ranges, then pick another port with `--port` or `--pac-port`. * It would create a certificate but no one is at a terminal to accept the Windows warning (`V1A-1128`). For unattended installs, use an organization certificate. See [Roll out with MDM](/valarcode/mdm#windows). Each integration defines which local traffic reaches the proxy. It is not a system-wide VPN and does not cover remote sessions automatically. Check the harness documentation for coverage and connection settings. The proxy and PAC services listen only on this machine's loopback addresses, not on LAN interfaces. Setup configures IPv4 and IPv6 loopback listeners where available. On macOS, if a VPN, dock, or network change adds an enabled network service, rerun `valar proxy enable` to apply Valar-managed automatic-proxy settings to it. On Windows, the setting covers Wi-Fi and Ethernet, so a new network of either kind needs no action. A VPN built into Windows keeps its own proxy setting, which Valar does not change. With `--ingress external`, have IT update its managed proxy settings or PAC coverage instead. Verify routing on the new connection. These paths are relative to `~/.valar/` (`%USERPROFILE%\.valar\` on Windows) unless noted. The removal column describes a successful `valar proxy disable`; partial cleanup keeps the installation record for retry. | Files | Purpose | On removal | | - | - | - | | `config.json` | Coding key and shared CLI preferences | Kept; proxy setup records cleared | | `cli.log`, `cli-.log` | CLI support log and rotated history | Kept | | `proxy-pac-ownership.json` | Recognizes disabled Valar PAC settings during later setup | Kept | | `proxy/cert.pem`, `proxy/cert.key`
`proxy/leaf.pem`, `proxy/leaf.key` | Generated or copied device certificate and private key | Removed; original IT-supplied files untouched | | `proxy-state.json` | Installed listeners, certificate paths, ownership and generated-root fingerprint | Removed after cleanup succeeds | | `proxy-forwarding.json`, `proxy.version` | Forwarding status and running build | Removed | | `proxy.log`, `proxy.log.1`, `proxy.log.lock`
`pac.log`, `pac.log.1`
`proxy.boot.log`, `pac.boot.log` | Request, PAC and service-startup diagnostics | Removed | | `claude/desktop-quota-first-holds.json` | Saved temporary subscription holds | Removed | | `~/Library/LaunchAgents/ai.valarhq.valar.proxy.plist`
`~/Library/LaunchAgents/ai.valarhq.valar.pac.plist` | Per-user proxy and PAC services (macOS) | Removed | | Task Scheduler `\Valar\ai.valarhq.valar.proxy`
`\Valar\ai.valarhq.valar.pac` | Per-user proxy and PAC services (Windows) | Removed | | `%LOCALAPPDATA%\Valar\proxy\valar.exe` | The copy the Windows services run | Removed | Organization-managed certificate trust is preserved. Valar removes only the root certificate whose fingerprint this installation recorded as Valar-generated. Files you placed alongside the proxy's own files are not removed.
### Organization-managed certificates If your organization supplies the certificate, trust its root through your normal device-management process, then run: ```bash theme={"system"} valar proxy enable --cert /absolute/path/leaf-bundle.pem --key /absolute/path/leaf.key ``` Supply both flags. The bundle must contain a valid end-entity leaf first, followed by its intermediate certificates. The leaf must match the private key and chain to a root the operating system already trusts. It must cover every supported destination: `api.anthropic.com` and, for Cursor, `*.cursor.sh`, `*.api5.cursor.sh`, `*.us.api5.cursor.sh` and `*.global.api5.cursor.sh`. On macOS, root-run provisioning can read protected certificate/key sources without changing their ownership or permissions. Symlinked paths are supported; their targets must be regular PEM files. Without root, both sources must be readable by the user; otherwise rerun the same provisioning command with `sudo`. The installed copies remain private to the selected user. Valar copies the leaf and key into its protected local directory. It does not create, rotate, or remove your organization's root CA. Renew organization-managed certificates through your PKI and rerun setup with the replacement files. `valar proxy restart` keeps a healthy certificate. If a Valar-generated certificate needs replacement, it can request administrator access on macOS, or show the Windows certificate warning, to install a fresh one. A replacement certificate may require restarting affected apps. ### Create an organization-managed certificate Use your existing PKI when available. The example below creates a private root and one device leaf on a protected signing machine. Keep the root key there; distribute only the root's **public certificate** through MDM and the device's leaf bundle/key through a protected channel. Issue a separate leaf key for each device. Use a new directory for a new CA. Do not overwrite an existing CA when enrolling another device. ```bash theme={"system"} umask 077 mkdir organization-proxy-ca cd organization-proxy-ca cat > root.cnf <<'CONFIG' [req] prompt = no distinguished_name = root_name x509_extensions = root_ca [root_name] O = Example Organization CN = Example Organization Local Proxy Root [root_ca] basicConstraints = critical,CA:TRUE,pathlen:0 keyUsage = critical,keyCertSign,cRLSign subjectKeyIdentifier = hash CONFIG openssl genpkey -algorithm EC -pkeyopt ec_paramgen_curve:P-256 \ -pkeyopt ec_param_enc:named_curve -out root.key openssl req -new -x509 -key root.key -sha256 -days 3650 \ -config root.cnf -extensions root_ca -out root.pem ``` `root.key` can issue certificates trusted by enrolled devices. Store it offline or in your PKI's protected signing system. Never deploy it to user devices. The validity periods in this example are choices for your PKI to review, not Valar requirements. Run from the CA directory. Choose a new directory name for each device or renewal; this example uses `device-001`. ```bash theme={"system"} umask 077 mkdir device-001 cat > device-001/leaf.cnf <<'CONFIG' [server_leaf] basicConstraints = critical,CA:FALSE keyUsage = critical,digitalSignature extendedKeyUsage = serverAuth subjectKeyIdentifier = hash authorityKeyIdentifier = keyid,issuer subjectAltName = DNS:api.anthropic.com,DNS:*.cursor.sh,DNS:*.api5.cursor.sh,DNS:*.us.api5.cursor.sh,DNS:*.global.api5.cursor.sh CONFIG openssl genpkey -algorithm EC -pkeyopt ec_paramgen_curve:P-256 \ -pkeyopt ec_param_enc:named_curve -out device-001/leaf.key openssl req -new -key device-001/leaf.key \ -subj "/O=Example Organization/CN=Valar local proxy device-001" \ -out device-001/leaf.csr openssl x509 -req -in device-001/leaf.csr \ -CA root.pem -CAkey root.key -CAcreateserial -CAserial root.srl \ -sha256 -days 365 -extfile device-001/leaf.cnf \ -extensions server_leaf -out device-001/leaf.pem cat device-001/leaf.pem root.pem > device-001/leaf-bundle.pem ``` The supported destinations are `api.anthropic.com` and, for Cursor, `*.cursor.sh`, `*.api5.cursor.sh`, `*.us.api5.cursor.sh` and `*.global.api5.cursor.sh`. Update SANs when the supported harness destinations change. This example root signs device leaves directly; its `pathlen:0` constraint does not permit intermediate CAs. If you use an existing PKI with intermediates instead, include them after the leaf in the bundle. Root inclusion is optional; trust must still be installed separately. The recipe uses real extension files and an explicit serial-file path, so it works with macOS LibreSSL and OpenSSL. Serialize issuance when using this simple shared serial file; production PKI should manage issuance and serial numbers. ```bash theme={"system"} openssl verify -CAfile root.pem -purpose sslserver device-001/leaf.pem openssl x509 -in device-001/leaf.pem -noout -dates -subject -issuer openssl x509 -in device-001/leaf.pem -noout -text openssl x509 -in device-001/leaf.pem -pubkey -noout > device-001/cert-public.pem openssl pkey -in device-001/leaf.key -pubout > device-001/key-public.pem cmp device-001/cert-public.pem device-001/key-public.pem ``` Verification should report `OK`, and `cmp` should exit successfully. Confirm the leaf has `CA:FALSE`, server authentication usage, and the required SAN. These checks validate the issued files; they do not install macOS trust. `valar proxy enable` also checks hostname coverage, validity, matching keys, and the machine's trust store before accepting the certificate. ### Trust and distribute | File | Destination | | - | - | | `root.pem` | MDM trusted-root certificate profile on enrolled devices | | `leaf-bundle.pem` | The intended device; root-only staging is supported for root-run provisioning | | `leaf.key` | Protected staging on the same device; installed copy is private to the selected user | | `root.key` | Signing system only; never deploy | The leaf private key must be an unencrypted PEM file so the background service can start unattended. Protect it through filesystem permissions and secure delivery. PKCS#8, SEC1 EC, and PKCS#1 RSA encodings are supported. This file-based setup does not accept a keychain-only or hardware-key reference. Wait for the MDM trust profile to apply before provisioning. For a controlled manual pilot, your administrator can install the public root into the macOS System keychain: ```bash theme={"system"} sudo security add-trusted-cert -d -r trustRoot \ -k /Library/Keychains/System.keychain /absolute/path/root.pem ``` On Windows, deploy the root with an Intune trusted certificate profile or Group Policy into the computer's **Trusted Root Certification Authorities** store. For a manual pilot, from an administrator PowerShell: ```powershell theme={"system"} certutil -addstore Root C:\path\root.pem ``` This installs system trust; it is separate from copying the leaf files. Valar does not remove organization-managed trust during teardown. ### Renew certificates Issue a new device leaf before expiry, verify it, then rerun `proxy enable --cert ... --key ...` with the new files. Replacing source files or updating their symlinks alone does not replace Valar's installed copies. Reapply each routed harness's `on` command and follow any restart instructions; verify forwarding afterward. When rotating the root, distribute and verify the new trust profile before provisioning leaves signed by it. Keep the old root until devices have moved, then retire it through your MDM/PKI process. Valar does not manage organization certificate revocation or trust-profile removal. ## Deploy with MDM See [Roll out with MDM](/valarcode/mdm) for deploying the CLI, the proxy, and each harness from Jamf, Kandji, Intune, or other MDM tools. ## Remove the proxy 1. Run `valar proxy status` to see which harnesses use the proxy. 2. Run `valar off` for each harness still on the proxy lane. 3. Save active work. If IT manages proxy settings or PAC rules, have IT remove the local Valar pointers and reload affected apps. 4. Remove the shared component: ```bash theme={"system"} valar proxy disable ``` **OFF keeps the proxy installed.** An OFF harness can still send encrypted traffic through it. Removing the proxy may interrupt running sessions until they reload their connection settings. `proxy disable` refuses while any harness remains on the proxy lane, even with `--force`. Once all are OFF, it checks for running sessions that may still depend on the proxy and asks before proceeding. Read the restart instructions: automatic restart is available only where supported; other sessions need a manual restart. For unattended removal, `--force` approves this interruption. `--no-restart` suppresses automatic app restarts and leaves that work to your deployment process. Neither flag makes a running session independent of the proxy. Removal clears Valar-owned automatic-proxy settings, removes its services and proxy files, and removes only a root certificate recorded as generated by this installation. It preserves organization-managed trust and the CLI support log. With external steering, IT must remove its own proxy pointers. Check the exit code and run `valar proxy status`. If removal is incomplete, fix the reported cause and retry. Do not delete the installation records by hand while recovery is pending. For a Valar-generated root, record `root_fingerprint` from `~/.valar/proxy-state.json` **before** disabling the proxy. Successful cleanup removes that file. This check does not apply to organization-managed roots, which Valar leaves in place. After `valar proxy disable`, list matching certificates in the System keychain: ```bash theme={"system"} security find-certificate -a -Z -c "Valar Local Root CA" \ /Library/Keychains/System.keychain ``` Compare the reported SHA-256 hashes with the recorded fingerprint, ignoring colons and letter case. The recorded fingerprint should no longer appear. Another certificate with the same name may remain if it belongs to your organization or another installation; do not delete certificates by name alone. On Windows, list the matching certificates in your user store instead: ```powershell theme={"system"} Get-ChildItem Cert:\CurrentUser\Root | Where-Object Subject -like "*Valar Local Root CA*" | ForEach-Object { $_.GetCertHashString([System.Security.Cryptography.HashAlgorithmName]::SHA256) } ``` Compare each SHA-256 hash with the recorded fingerprint the same way. These commands only read the certificate store. If the recorded certificate remains, follow the removal error and retry `valar proxy disable`. A warning about a remaining trust setting after its certificate was removed is a separate cleanup issue; it does not mean the certificate is still installed. ## Deployment checks and diagnostics Use exit codes to detect failure, then read the full error and its next steps. A nonzero exit can describe partial completion; do not assume every file or service was rolled back. | Message/code | What to check | | - | - | | `V1A-1001` | Required administrator access was unavailable. Read which operation failed and whether any earlier work completed. | | `V1A-1004` | Root execution has no selected user. Run from that user's session or select the account in your root policy. | | `V1A-1008` | The selected user's home directory is absent. Enroll/log in the user before provisioning. | | `V1A-1120` | Provisioning could not confirm the listeners were ready. Check service startup logs and `valar proxy status`. | | `V1A-1121` | A listening port is in use, or reserved by Windows for Hyper-V or WSL. Choose another with `--port` or `--pac-port`. | | `V1A-1123` | Windows: the certificate warning was declined. Run `valar proxy enable` again and accept. | | `V1A-1124` | Windows: policy does not trust user-installed certificates. Deploy an organization certificate. | | `V1A-1125` | Windows: policy applies one proxy setting to every user. Provision with `--ingress external` and have IT add the Valar rule. | | `V1A-1128` | Windows: an unattended run would need someone to accept the certificate warning. Pass `--cert` and `--key`. | | `V1A-1800` / `V1A-1801` | Harness verification found unhealthy/unconfigured routing or no successful forwarded request. | | `V1A-1902` | Upgrade or migration did not finish. Resolve the reported cause and use the [upgrade recovery guidance](/valarcode/setup#if-an-upgrade-does-not-complete). | The local services run as the selected user under `~/Library/LaunchAgents/`. Their labels are `ai.valarhq.valar.proxy` and `ai.valarhq.valar.pac`. From that user's session, inspect them without changing state: ```bash theme={"system"} launchctl print "gui/$(id -u)/ai.valarhq.valar.proxy" launchctl print "gui/$(id -u)/ai.valarhq.valar.pac" scutil --proxy ``` On Windows, from that user's session: ```powershell theme={"system"} schtasks /Query /TN \Valar\ai.valarhq.valar.proxy schtasks /Query /TN \Valar\ai.valarhq.valar.pac reg query "HKCU\Software\Microsoft\Windows\CurrentVersion\Internet Settings" /v AutoConfigURL ``` After successful removal, both service lookups should report that the jobs are absent. With external ingress, the corporate PAC can remain configured; IT must remove the local Valar rules. With Valar-managed PAC, macOS may keep a disabled URL string. An enabled stale pointer still needs cleanup. ## Troubleshooting | Symptom | Next step | | - | - | | A proxy-supported harness uses the direct lane | Read its status explanation. Provision the proxy and run the harness's `on` command when ready. | | Certificate or connection settings changed | Re-run the harness's `on` command and follow any restart prompt. | | Proxy services are unavailable | Run `valar proxy status`, then `valar proxy restart`; inspect `valar proxy logs`. | | A corporate proxy is already configured | Keep the policy in place. Use an explicitly approved outbound proxy and/or IT-managed steering. | | `proxy enable` stops on a proxy set in network settings | Approve that exact proxy with `--upstream-proxy`. See [A proxy set in network settings](#a-proxy-set-in-network-settings). | | Status says other apps are bypassing the system proxy | Run the `--upstream-proxy` command status shows. | | Status warns about a published discovery script | Claude Desktop may not be routed on this network. See [Automatic proxy discovery (WPAD)](#automatic-proxy-discovery-wpad). | | Requests work but Valar verification fails | Check the selected lane and quota-first state; send a request that uses Valar, then retry. | | A session still uses its previous routing | Follow the harness's recovery warning and restart the affected session. | A session that still depends on an unavailable proxy can fail until the service is restored or its connection settings are reloaded. Some integrations can fall back to a direct connection; that does not prove Valar routing. Verify each harness after setup or repair. For command/setup failures, collect the [CLI support log](#support-logs). For service startup failures, inspect `~/.valar/proxy.boot.log` and `~/.valar/pac.boot.log` (under `%USERPROFILE%\.valar\` on Windows). ### Support logs The CLI writes command and setup diagnostics automatically to `~/.valar/cli.log`; no logging option is needed. Logging is best effort: a permissions problem or an early refusal can prevent a record, so keep the terminal error too. The CLI support log records command names, CLI/OS versions, exit codes and sanitized errors. Proxy setup failures can also include network-service names and proxy-setting classifications. It does not record full command arguments, configuration or environment contents, request bodies, or proxy traffic. Error text is sanitized to redact recognized credentials, argument values, URLs, account names and absolute paths. Review the file before sharing it. Rotation keeps the live log and one `~/.valar/cli-.log` backup, about 2 MiB combined. The live file has the latest records; include the backup if support needs older history. Both survive `valar proxy disable`. `valar proxy logs` and the service-startup logs are separate diagnostics; the CLI support log's redaction rules do not apply to them. Review those logs separately before sharing. # Model routing Source: https://docs.valarhq.ai/valarcode/routing How Valar decides which model serves each coding request Routing decides which model serves each request an engineer's harness makes. You set it once per coding key, and it applies to everyone using that key. ## How a request is routed Every request goes through three steps in the gateway: The harness asks for an Opus-, Sonnet-, Haiku-, or Fable-class model, and Valar sorts the request into one of those four **aliases**. Anything that matches none of them is treated as **Sonnet**, the working tier for most coding traffic. Using the client id on the request (most harnesses send it as an `X-Valar-Client-Id` header; Cursor, which cannot set headers, rides it as a token suffix), Valar assigns the engineer to one **cohort** in the key's split. The assignment is deterministic and sticky, so the same engineer always lands in the same cohort for a given key and split. Their experience stays consistent from one request to the next. The cohort pins each alias to a target model. Valar looks up the target for the request's alias and routes there. For example, an engineer whose cohort maps Sonnet to GLM-5.2 has their Sonnet-class calls served by GLM-5.2. If a key's routing cannot be read, ValarCode falls back to the frontier Claude tiers instead of failing the request, so a misconfiguration never blocks an engineer. ## Auto mode vs. Manual mode ValarCode has two ways to decide the split. Valar picks the target model for you. A new coding key starts on Auto, so routing works before you configure anything. Auto runs the objective you pick under **Optimize for**, set per key and per tool on the routing screen: * **Balanced** (default): the best answer for the lowest cost. * **Max savings**: the most cost-effective model that can do the work. A premium model serves the hardest requests, and steps in as a fallback when a request can't be served otherwise. * **Max quality**: the strongest models for each request's difficulty, prioritizing answer quality over cost, and asked to reason at a high effort level. It is available for every tool. Each tool has its own mix, because the tools reach different models; the dashboard shows the mix and the effort level for the tab you are on. Within each difficulty, a conversation is assigned one model from the mix and keeps it from turn to turn. For Claude Code: | Difficulty | Models | | - | - | | Hardest turns | Claude Opus 5.5 (80%), Claude Fable 5.1 (20%) | | Hard turns | GPT-6 Sol (80%), GLM-5.3 Fast (20%) | | Everyday turns | GPT-6 Luna (50%), GLM-5.3 Fast (50%) | | Simple turns | DeepSeek-V4.1-Flash | The effort level is a floor: a request that asks for more keeps its own setting, and a request that turns thinking off stays off. You choose which model serves each request class. In the routing editor you pick a target model for Opus, Sonnet, Haiku, and Fable from a dropdown of available models. Valar applies that selection to matching requests. A manual mapping is available for Claude Code today; the other harnesses route on Auto. A single selection is all you need to get started. To run more than one selection at once, split traffic across cohorts and give each cohort its own model choices (see [Cohorts](#cohorts)). Saving asks for a reason, which goes into an audited change log, and each save is versioned so you keep a full history of what changed and why. Manual suits you when you want direct control over which model handles each class: running a specific A/B, holding a known baseline, or rolling a model out on your own schedule. ## Editing routing Routing is set **per coding key**, in the dashboard under **ValarCode → Routing**: 1. Pick the coding key. 2. Choose the target model for each class (Opus, Sonnet, Haiku, Fable). To run more than one selection at once, split traffic across cohorts and set models per cohort. 3. Enter a reason and save. If you use multiple cohorts, their shares have to total 100%. Changes are versioned and take effect for new requests once you save. Because assignment is sticky to the current split, changing a cohort's share can move some engineers to a different cohort. That is expected when you rebalance an experiment. ## Cohorts A **cohort** is one arm of a split. Each cohort has: * a **share**, the percentage of engineers it covers, and * a **target model for each alias**: Opus, Sonnet, Haiku, and Fable. A split holds 1 to 5 cohorts, and their shares add up to 100%. Engineers are spread across cohorts by the sticky assignment above, so a cohort set to 30% gets about 30% of your engineers and keeps them there rather than re-rolling on every request. That is what makes a controlled comparison possible. Put 70% of engineers on a cohort that targets Claude and 30% on an open-weight cohort, then compare cost and outcomes between the two in [Analytics](/valarcode/analytics). ## Seeing which model answered Under Auto mode the model that answers is usually not the one the harness asked for, and the response keeps echoing the requested id so the harness can restore a session. To see the real answer, an organization admin turns on **Show serving model name in the response metadata headers** under **Settings → ValarCode settings**. Valar then adds both of these to every coding response, whichever tool sent it: * an `x-valar-served-model` response header, and * a `served_model` field in the response body, beside `model`. ```json theme={"system"} { "id": "msg_01ABC...", "model": "claude-sonnet-5", "served_model": "zai-org/GLM-5.3-fast", "role": "assistant", "content": [{ "type": "text", "text": "..." }] } ``` On a non-streaming response the field is top-level on every lane. On a stream it sits wherever that lane's dialect puts `model`: | Lane | Tools | Where the field sits on a stream | | - | - | - | | `/v1/messages` | Claude Code, Claude Desktop | on the `message_start` event's `message` object | | `/v1/chat/completions` | Cursor | on every chunk that carries `model` | | `/v1/responses` | Codex | under `response`, on every event that embeds the response object | Both surfaces name the same model, so use whichever your tooling can read. A body that reports an error carries no `served_model`, because no model answered it. ### Under each answer Claude Code, Cursor and Codex show neither response headers nor extra body fields. A second switch beneath the first, **Show serving model name in the response body**, adds a line at the end of each answer Auto routed: ```md theme={"system"} [served by GLM-5.3 Fast] ``` The line is added only to a finished answer. It is never added to a step where the agent calls a tool, or when you picked a specific model yourself. Valar also leaves it off the background requests it recognizes in Claude Code, Claude Desktop, OpenCode, Pi, Oh My Pi and GitHub Copilot, such as sub-agents, conversation titles and compaction summaries. One exception: in GitHub Copilot, a general sub-agent's answer can still carry the line. Valar removes the line from the conversation history before your next request reaches the model, so the model never sees it. Turning the first switch off turns this one off with it. Some harnesses, Claude Desktop among them, expose neither headers nor raw bodies. For those, run [`valar usage`](/valarcode/analytics#your-own-recent-requests), which lists your recent requests with the model that answered each one. ## Next steps Which models you can route to, and how each harness's traffic is served. Compare cohorts and see what each split saved. # Set up ValarCode Source: https://docs.valarhq.ai/valarcode/setup Install the Valar CLI, create a coding key, and connect Claude Code, Claude Desktop, Cursor, Codex, Pi, Oh My Pi, or VS Code Install the CLI, configure a **coding key**, and connect a harness. On macOS and Windows 11, Claude Code and Claude Desktop can share the optional [on-device proxy](/valarcode/proxy); on macOS, Cursor can use it too. ## Prerequisites * One of the supported harnesses installed locally: Claude Code, the Claude Desktop app, Cursor, Codex, Pi, Oh My Pi, or VS Code (via GitHub Copilot Chat). * Access to the [Valar Dashboard](https://app.valarhq.ai) to create a coding key. * macOS, Linux, or WSL for the CLI; on Windows 11, Claude Code and Claude Desktop. ## Install the CLI Install the `valar` CLI, a single binary with no runtime dependencies: ```bash theme={"system"} curl -fsSL https://raw.githubusercontent.com/valarhq/valar-code-cli/main/install.sh | sh ``` On Windows, from PowerShell: ```powershell theme={"system"} irm https://raw.githubusercontent.com/valarhq/valar-code-cli/main/install.ps1 | iex ``` Confirm it is on your path: ```bash theme={"system"} valar --version ``` ## Create a coding key In the [Valar Dashboard](https://app.valarhq.ai), open **ValarCode** and create a coding key. Coding keys are scoped to agentic harnesses like Claude Code, Cursor, Codex, and Pi, and they carry your routing split. They use the `vlrcode_` prefix, and the token is shown once, so copy it before you leave the page. A new key works right away. Until you set a split, it routes on Auto — Valar picks the target model — so it is safe to connect before you have tuned anything. See [Model routing](/valarcode/routing). Run `valar configure` once to store the key locally so you can drop `--api-key` from later commands: ```bash theme={"system"} valar configure --api-key ``` The key is resolved from, in order: `--api-key`, then `$VALAR_API_KEY`, then `~/.valar/config.json`, then an interactive prompt. ## Connect a harness Connect each harness from your own user account. Choose its guide for setup and restart requirements: `valar claude on` `valar claude-desktop on` `valar cursor on` `valar codex on` `valar opencode on` `valar pi on` `valar ohmypi on` `valar copilot on` Routing Cursor requires a paid Cursor plan (Pro, Business, or Enterprise) — `valar cursor on` fails on the free plan or when Cursor is not signed in. `valar cursor` routes the Cursor **editor** (the desktop app) on macOS and Linux; it is not available on Windows. On macOS it uses the [on-device proxy](/valarcode/proxy) when one is set up with `valar proxy enable` (Cursor 3.21.16 or newer); without it, Cursor is connected directly. The Cursor CLI (`cursor-agent`) is not supported. VS Code routing goes through GitHub Copilot Chat's custom-endpoint provider, so it needs a signed-in Copilot and a recent VS Code. `valar vscode` and `valar copilot` are the same command. ## On-device proxy Provision the shared proxy once on macOS or Windows 11, then enable each supported harness. Claude Code and Desktop retain their native login and can use subscription-first routing: ```bash theme={"system"} valar proxy enable valar claude on valar claude-desktop on ``` Enable only the apps you use. On macOS, proxy setup requests administrator access. On Windows it needs none; Windows asks you to confirm the local certificate instead. Setup does not turn either app on by itself. Each app has its own ON/OFF control. Without a usable proxy, Claude Code uses direct routing; Desktop uses a direct provider profile when no proxy is provisioned. You can also select direct routing with `--direct`. See [On-device proxy](/valarcode/proxy) for setup, quota-first, and corporate networks. To deploy with MDM, see [Roll out with MDM](/valarcode/mdm). The proxy runs on macOS and Windows 11; Claude Code on Linux uses direct routing. ## Per-engineer attribution To attribute usage to each engineer, the CLI attaches a **client id** to every request. Most harnesses (Claude Code, Claude Desktop, Codex, Pi, Oh My Pi, and VS Code/Copilot) send it as an `X-Valar-Client-Id` header; Cursor, which cannot set custom headers, rides it as a token suffix (`~`). By default the client id is your **OS username** (normalized to lowercase `a-z 0-9 . _ - @ +`, capped at 64 characters), so the analytics leaderboard shows real engineers rather than opaque strings. Change how it is derived with `--user-id`: | `--user-id` value | Client id | | - | - | | *(omitted)* or `USERNAME` | Your OS username (the default) | | `HOSTNAME` | The machine hostname | | `RANDOM` | An opaque `vc_…` hash of `username@hostname`, with no raw identity on the wire | | any other value | Used verbatim as the id (normalized the same way) | The chosen id is recorded on `on`, so you can print it any time: ```bash theme={"system"} valar id # e.g. jdoe (or vc_a713f7e67f3e8c06 under RANDOM) ``` A plain key with no client id still works, but those requests are not attributed to an engineer. The client id is supplied by the client and is used only for attribution and routing. It is not a security boundary; access is gated on the coding key itself. ## CLI reference The CLI follows a harness-first grammar. A bare harness command normally means `on`, so `valar claude` is the same as `valar claude on`. **Claude Desktop is the exception:** `valar claude-desktop` prints help; include `on`, `off`, or `status`. ```text theme={"system"} valar on|off|status [flags] valar claude reset valar claude-desktop reset valar proxy enable|disable|restart|status|logs valar status valar off [--force] valar configure [--api-key ] [--endpoint ] [--user-id ] valar configure --quota-first=true|false valar models list|refresh valar usage [--window mtd|7d|30d] [-n ] valar upgrade [--check] [--force] valar id valar version ``` Harness names are `claude`, `claude-desktop`, `cursor`, `codex`, `pi`, `ohmypi`, `opencode`, and `copilot`. `vscode` is an alias for `copilot`. `valar status` summarizes them; `valar off` attempts to turn off all configured harnesses. `valar update` is an alias for `valar upgrade`. **Per-harness commands** | Command | What it does | | - | - | | `valar on` | Configure Valar routing and the harness's model picker | | `valar off` | Stop Valar routing. Established proxy setups stay installed in passthrough; direct integrations restore their owned settings. | | `valar status` | Show the current routing state (endpoint, masked key, client id) | | `valar claude reset` / `valar claude-desktop reset` | Clear temporary subscription holds in the shared proxy; preserve counters and subscription allowance | **Flags** | Flag | Applies to | Meaning | | - | - | - | | `--api-key ` | `on`, `configure`, `models`, `proxy enable` | Your Valar coding key (see resolution order above) | | `--user-id ` | `on`, `configure` | How to derive the client id: `USERNAME` (default), `HOSTNAME`, `RANDOM`, or a literal | | `--direct` | `claude on`, `claude-desktop on`, `cursor on` | Use the Valar endpoint directly, without the on-device proxy (Claude: without subscription-first routing) | | `--valar-only` | `cursor on` | [Only Valar models in Cursor](/valarcode/cursor#two-modes); strict on macOS with the on-device proxy. Sticky: `--valar-only=false` turns it off; `off` clears it | | `--force` | Claude, Desktop, and Cursor `on`/`off`; global `off`; `upgrade`; `proxy disable` | Approve necessary session interruption without a prompt; does not bypass configuration checks | | `--no-restart` | `proxy disable` | Leave app restarts to you or IT; the removal confirmation still applies. | | `--quota-first=true\|false` | `configure`; Claude and Desktop `on` | Set the shared subscription-first preference; requires the proxy lane to take effect | | `--verify` | Claude and Desktop `status` | Require a healthy proxy lane and observed successful Valar forwarding | | `--single-session` | Claude, Desktop, and Codex `on` | Open an isolated routed session; Claude-family sessions use direct routing | | `--stable` / `--latest` | `claude on`, `configure` | [Hold Claude Code at Valar's stable release, or remove the hold and update](/valarcode/claude-code#stable-claude-code) | | `--check` | `upgrade` | Report whether a newer release exists without installing it | | `--no-requests` | `usage` | Print spend and savings only, skipping the recent-requests table | ## Keeping the CLI current ```bash theme={"system"} valar upgrade --check valar upgrade ``` `--check` reports available updates without installing them or migrating configuration. `upgrade` verifies the download's SHA256, replaces the CLI, and completes required configuration and service updates. Run it from your own terminal session, without adding `sudo`. If migration must refresh running Claude Code sessions, the upgrade asks before interrupting them. Use `valar upgrade --force` only when that interruption is acceptable. Saved conversations remain available; interrupted requests and tools are not replayed automatically. Follow any instructions to reopen `claude agents` or resume an interactive conversation. An older socket-based Claude setup moves to the proxy lane when the machine can support it, otherwise to direct routing. A saved quota-first preference remains saved, but direct requests use Valar without the Claude subscription. Installing the proxy is a separate step that may require administrator access on macOS. ### If an upgrade does not complete Read the full error to see whether installation failed or the new CLI could not finish updating configuration or services. Keep the recovery records and address the reported cause before retrying. `valar upgrade` checks for a newer repair first; repeating the same error without a fix will not help. Do not install an older CLI over partially migrated configuration. If you need to stop routing, use the harness's `off` command when the CLI says that recovery path is available, and restart affected sessions as instructed. Unknown or unreadable configurations can require repair before OFF is safe. ### Automatic update notices Interactive commands can offer an upgrade before continuing; non-interactive commands print a notice. The check is cached for up to a day. Set `VALAR_NO_UPDATE_CHECK=1` to disable these notices. ### Deprecated options `proxy enable --force`, `--no-restart`, and `--off` are deprecated no-ops with warnings. Use harness commands to change routing and approve session restarts. ## Managing keys * Revoke a coding key from the dashboard at any time. Requests using it stop resolving. * Issue a separate key per team or experiment to run different splits side by side. Each key can have its own routing policy. ## Next steps Set the split and choose target models. Watch usage and savings come in per engineer. # Visual Studio Code Source: https://docs.valarhq.ai/valarcode/vscode Route VS Code through Valar via GitHub Copilot Chat's custom-endpoint (BYOK) provider ValarCode routes VS Code through GitHub Copilot Chat's bring-your-own-key **Custom Endpoint** provider. Enabling registers a **Valar** provider group in Copilot Chat, and you pick a Valar model from its model picker. The group speaks the Anthropic Messages API; the gateway maps whatever model it sends to the target your routing resolves to. See [Set up ValarCode](/valarcode/setup) for the shared setup and the client-id model. The VS Code integration uses the `copilot` harness, so the command is `valar copilot`. `valar vscode` is an alias you can use interchangeably; attribution is recorded under the `copilot` harness either way. ## Prerequisites * VS Code (the stable build) with GitHub Copilot Chat signed in, and a version new enough to offer the Custom Endpoint provider. Insiders and forks keep their own config and are not routed. * A Valar coding key (`vlrcode_…`) and the `valar` CLI ([install](/valarcode/setup#install-the-cli)). ## Enable routing Run once per machine: ```bash theme={"system"} valar copilot on --api-key ``` Then **reload the VS Code window** and pick a Valar model in Copilot Chat's model picker. If the group is not listed yet, reload again. Check the state any time: ```bash theme={"system"} valar copilot status ``` ## What gets written The CLI registers a provider group in Copilot Chat's model config, `chatLanguageModels.json`, in the VS Code user directory (`~/Library/Application Support/Code/User` on macOS, `~/.config/Code/User` on Linux), plus one file per additional VS Code profile under `profiles/*/`. The file is a JSON array of provider groups; the CLI adds a group named **Valar**: ```json theme={"system"} [ { "vendor": "customendpoint", "name": "Valar", "apiType": "messages", "models": [ { "id": "", "name": "", "url": "https://api.valarhq.ai/v1/messages", "toolCalling": true, "vision": true, "thinking": true, "contextWindow": 200000, "maxOutputTokens": 64000, "requestHeaders": { "x-api-key": "vlrcode_…", "X-Valar-Harness": "copilot", "X-Valar-Client-Id": "jdoe", "X-Valar-Cli-Version": "1.4.2" } } ] } ] ``` In detail, enabling: * Registers a **Valar** group under the built-in `customendpoint` vendor, with `apiType: "messages"` so the group speaks the Anthropic Messages dialect. The group's `models` list **is** the picker; you choose a Valar model there yourself. * Writes the full per-model endpoint `https://api.valarhq.ai/v1/messages` on each model, since this vendor has no separate base-URL field. * Carries the coding key in the **`x-api-key` request header** rather than the vendor's `apiKey` field. That field is stored in the OS keychain and cannot be populated by an external writer, so the key rides as a request header instead; the gateway authenticates it like a bearer token. * Adds `X-Valar-Harness: copilot`, `X-Valar-Client-Id` and `X-Valar-Cli-Version` to each model's `requestHeaders` for attribution. On the first enable, the files it touches are snapshotted to `~/.valar/copilot/backup.json`. A VS Code profile you create later is routed by the next `on` and un-routed by `off` on the Valar marker, without a snapshot of its own. The client id defaults to your OS username, so the header reads `X-Valar-Client-Id: jdoe`. Use `--user-id RANDOM` for an opaque `vc_…` id instead. See [per-engineer attribution](/valarcode/setup#per-engineer-attribution). A `chatLanguageModels.json` that contains comments (JSONC) is skipped and reported rather than rewritten, so a hand-edited file is never clobbered. Remove the comments and re-run `on` to route that profile. ## Turn off routing ```bash theme={"system"} valar copilot off ``` This restores the pre-enable snapshot and removes the Valar group from any profile it was added to. **Reload the VS Code window** if the Valar models are still listed in Copilot Chat. ## Next steps How cohorts and the split decide which model serves each request. Per-engineer usage and savings, attributed by client id. # Webhooks Source: https://docs.valarhq.ai/webhooks Receive a callback when a background request finishes instead of polling for it When a request can take a while to finish, you don't have to poll `GET /v1/responses/{response_id}` in a loop. Attach a callback URL to the request and Valar delivers the finished result to your server as an HTTP POST. The feature works on `POST /v1/responses` and `POST /v1/chat/completions`, and is configured entirely through two `metadata` keys. ## Parameters Both keys live inside the `metadata` object on the create request. The destination URL for the callback. Must be an `http` or `https` URL. If you omit it or pass a value that can't be used, no callback is sent and the request itself is unaffected. An optional shared secret. Valar sends its value as a Bearer token in the callback's `Authorization` header so your endpoint can confirm the request is genuine. ## Set up an endpoint Your endpoint receives a JSON body and should respond as soon as it has accepted the payload. The body matches exactly what `GET /v1/responses/{response_id}` returns for the same request, whether the response completed or failed. ```python theme={"system"} from fastapi import FastAPI, Request app = FastAPI() @app.post("/hooks/inference-done") async def inference_done(request: Request): payload = await request.json() # payload is the same object GET /v1/responses/{id} returns enqueue_for_processing(payload["id"], payload) return {"ok": True} ``` If you set `webhook_token`, reject any incoming call whose header doesn't match. Compare against `Authorization: Bearer ` and return a 4xx for anything else. ```python theme={"system"} from fastapi import Header, HTTPException WEBHOOK_TOKEN = "whk_3f0a9c2e7b14" def assert_authorized(authorization: str = Header(default="")): if authorization != f"Bearer {WEBHOOK_TOKEN}": raise HTTPException(status_code=401) ``` Set `background=True` and add the callback URL (and token, if you use one) to `metadata`. ```python File 1 theme={"system"} from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key="YOUR_VALAR_API_KEY", # or read OPENAI_API_KEY from the environment ) response = client.responses.create( model="zai-org/GLM-5.2", input="Summarize this document.", background=True, metadata={ "completion_webhook": "https://app.example.com/hooks/inference-done", "webhook_token": "whk_3f0a9c2e7b14", }, ) ``` ## What Valar sends The callback is a single POST request. Always a POST to the URL in `completion_webhook`. `application/json`. Present only when `webhook_token` was set. Carries `Bearer `. The full response object, byte-for-byte identical to a `GET /v1/responses/{response_id}` call - including the `status` field and, for failed responses, the `error` envelope. The response `id` lives here and is your key for deduplication. ## When a callback fires Valar POSTs to your endpoint once a response reaches a terminal status: * **Completed.** The request finished normally. The body is the completed response object. * **Failed.** The body is the failed response object, including the `error.code` and `error.message` your `GET /v1/responses/{response_id}` would see. This includes responses Valar fails on your behalf when they get stranded short of terminal - for example, a background request that stays `in_progress` past its deadline arrives with `status: "failed"` and `error.code: "timeout"`. Use the `status` field to branch on success vs. failure in your handler. ## Delivery semantics The same callback can be delivered more than once. Treat the response `id` in the body as an idempotency key and ignore any `id` you have already processed. * **Retries.** A delivery that returns a non-2xx status or hits a network error is retried up to **3 times**. Return a **2xx** as soon as you accept the payload to stop the retries. * **Timeout.** Each attempt has a **30 second** ceiling. A timeout counts as a failed attempt and triggers the next retry. * **Best-effort.** Callbacks are best-effort. If every attempt fails the failure is logged but never touches the response record or the API, and the result stays available through `GET /v1/responses/{response_id}` regardless of whether delivery ever succeeded. ## Test it locally with ngrok You can point a real request at a listener on your own machine. ```bash theme={"system"} python -c " from http.server import HTTPServer, BaseHTTPRequestHandler; import json class H(BaseHTTPRequestHandler): def do_POST(self): print(json.dumps(json.loads(self.rfile.read(int(self.headers['Content-Length']))), indent=2)) self.send_response(200); self.end_headers() HTTPServer(('127.0.0.1', 8765), H).serve_forever() " ``` ```bash theme={"system"} ngrok http 8765 ``` Copy the `https://xxxx.ngrok-free.app` forwarding URL from the output. ```bash theme={"system"} curl -X POST https://api.valarhq.ai/v1/responses \ -H "Authorization: Bearer YOUR_VALAR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "zai-org/GLM-5.2", "input": "What is 2+2? Reply with just the number.", "background": true, "metadata": { "completion_webhook": "https://xxxx.ngrok-free.app" } }' ``` When the response finishes, Valar POSTs the full payload to your listener. ```bash theme={"system"} curl -X POST https://api.valarhq.ai/v1/responses \ -H "Authorization: Bearer YOUR_VALAR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "zai-org/GLM-5.2", "input": "What is 2+2? Reply with just the number.", "background": true, "metadata": { "completion_webhook": "https://xxxx.ngrok-free.app" } }' ``` When the response finishes, Valar POSTs the full payload to your listener. # Workload Tags Source: https://docs.valarhq.ai/workload-tags Name the workload behind each request so usage and spend split by job, not just by model ## Why tag requests Usage reporting groups by model. If you run invoice extraction and support-ticket triage against the same model, their spend lands in one bucket and you cannot tell which job costs what, or which one started failing. A workload tag fixes that. Send an optional name with each request and Valar records it on the usage ledger, so every job gets its own line in usage and spend even when several jobs share a model. ## How it works Send a `valar_workload` key inside the request's `metadata` object: ```python theme={"system"} from openai import OpenAI client = OpenAI( base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment api_key="YOUR_VALAR_API_KEY", # or read OPENAI_API_KEY from the environment ) response = client.chat.completions.create( model="google/gemini-3.1-flash-lite", messages=[{"role": "user", "content": "Extract the total from this invoice."}], metadata={"valar_workload": "invoice-extraction"}, ) ``` ```bash theme={"system"} curl https://api.valarhq.ai/v1/responses \ -H "Authorization: Bearer $VALAR_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "google/gemini-3.1-flash-lite", "input": "Extract the total from this invoice.", "metadata": {"valar_workload": "invoice-extraction"} }' ``` That's it. There is nothing to pre-register: the first tagged request creates the workload in your reporting. ## Rules Omit the key and nothing changes. Untagged requests are recorded exactly as before and appear as one "Untagged" bucket in the dashboard. Tags are lowercased, whitespace becomes `-`, and anything outside `a-z 0-9 . _ -` is dropped. So `"Invoice Extraction"` and `"invoice-extraction"` are the same workload. A caller who types a human name and one who sends a slug agree instead of splitting spend across two buckets. Reporting joins on the normalized name across time. Renaming a workload starts a new one; the old name keeps its history. The tag does not change routing, model selection, price, or output. The `metadata` echoed back on the response is byte-identical to what you sent, original casing included. ## Where you see it Open **Usage & Billing** in the dashboard and switch the Breakdown to **Workload**. Each workload shows its responses, tokens, success rate, latency (p50/p99), and spend over the last 30 days, for the selected workspace or across your whole organization. The view exports to CSV like the other breakdowns.