Skip to main content
Routing decides which model serves each request an engineer’s harness makes. You set it once per coding key, and it applies to everyone using that key.

How a request is routed

Every request goes through three steps in the gateway:
1

Classify the request

The harness asks for an Opus-, Sonnet-, Haiku-, or Fable-class model, and Valar sorts the request into one of those four aliases. Anything that matches none of them is treated as Sonnet, the working tier for most coding traffic.
2

Assign the engineer to a cohort

Using the client id on the request (most harnesses send it as an X-Valar-Client-Id header; Cursor, which cannot set headers, rides it as a token suffix), Valar assigns the engineer to one cohort in the key’s split. The assignment is deterministic and sticky, so the same engineer always lands in the same cohort for a given key and split. Their experience stays consistent from one request to the next.
3

Resolve the target model

The cohort pins each alias to a target model. Valar looks up the target for the request’s alias and routes there. For example, an engineer whose cohort maps Sonnet to GLM-5.2 has their Sonnet-class calls served by GLM-5.2.
If a key’s routing cannot be read, ValarCode falls back to the frontier Claude tiers instead of failing the request, so a misconfiguration never blocks an engineer.

Auto mode vs. Manual mode

ValarCode has two ways to decide the split.
Valar picks the target model for you. A new coding key starts on Auto, so routing works before you configure anything.Auto runs the objective you pick under Optimize for, set per key and per tool on the routing screen:
  • Balanced (default): the best answer for the lowest cost.
  • Max savings: the most cost-effective model that can do the work. A premium model serves the hardest requests, and steps in as a fallback when a request can’t be served otherwise.
  • Max quality: the strongest models for each request’s difficulty, prioritizing answer quality over cost, and asked to reason at a high effort level. It is available for every tool. Each tool has its own mix, because the tools reach different models; the dashboard shows the mix and the effort level for the tab you are on. Within each difficulty, a conversation is assigned one model from the mix and keeps it from turn to turn. For Claude Code: The effort level is a floor: a request that asks for more keeps its own setting, and a request that turns thinking off stays off.

Editing routing

Routing is set per coding key, in the dashboard under ValarCode → Routing:
  1. Pick the coding key.
  2. Choose the target model for each class (Opus, Sonnet, Haiku, Fable). To run more than one selection at once, split traffic across cohorts and set models per cohort.
  3. Enter a reason and save. If you use multiple cohorts, their shares have to total 100%.
Changes are versioned and take effect for new requests once you save. Because assignment is sticky to the current split, changing a cohort’s share can move some engineers to a different cohort. That is expected when you rebalance an experiment.

Cohorts

A cohort is one arm of a split. Each cohort has:
  • a share, the percentage of engineers it covers, and
  • a target model for each alias: Opus, Sonnet, Haiku, and Fable.
A split holds 1 to 5 cohorts, and their shares add up to 100%. Engineers are spread across cohorts by the sticky assignment above, so a cohort set to 30% gets about 30% of your engineers and keeps them there rather than re-rolling on every request. That is what makes a controlled comparison possible. Put 70% of engineers on a cohort that targets Claude and 30% on an open-weight cohort, then compare cost and outcomes between the two in Analytics.

Seeing which model answered

Under Auto mode the model that answers is usually not the one the harness asked for, and the response keeps echoing the requested id so the harness can restore a session. To see the real answer, an organization admin turns on Show serving model name in the response metadata headers under Settings → ValarCode settings. Valar then adds both of these to every coding response, whichever tool sent it:
  • an x-valar-served-model response header, and
  • a served_model field in the response body, beside model.
On a non-streaming response the field is top-level on every lane. On a stream it sits wherever that lane’s dialect puts model: Both surfaces name the same model, so use whichever your tooling can read. A body that reports an error carries no served_model, because no model answered it.

Under each answer

Claude Code, Cursor and Codex show neither response headers nor extra body fields. A second switch beneath the first, Show serving model name in the response body, adds a line at the end of each answer Auto routed:
The line is added only to a finished answer. It is never added to a step where the agent calls a tool, or when you picked a specific model yourself. Valar also leaves it off the background requests it recognizes in Claude Code, Claude Desktop, OpenCode, Pi, Oh My Pi and GitHub Copilot, such as sub-agents, conversation titles and compaction summaries. One exception: in GitHub Copilot, a general sub-agent’s answer can still carry the line. Valar removes the line from the conversation history before your next request reaches the model, so the model never sees it. Turning the first switch off turns this one off with it.
Some harnesses, Claude Desktop among them, expose neither headers nor raw bodies. For those, run valar usage, which lists your recent requests with the model that answered each one.

Next steps

Models

Which models you can route to, and how each harness’s traffic is served.

Analytics & savings

Compare cohorts and see what each split saved.