Key numbers
The setup view shows a 30-day, org-wide strip:- Engineers connected: distinct client ids seen. By default the client id is the engineer’s OS username, so the leaderboard reads as real names; teams that prefer opaque ids can connect with
--user-id RANDOM(see per-engineer attribution). - Requests routed: total coding requests.
- Tokens: input, cached, and output.
- Saved: dollars saved against a Claude-only baseline (never negative).
Savings and spend
A 30-day card breaks down the money:- Saved this period, with a ”% vs Claude-only” figure.
- Actual spend, what you paid across all served models.
- Claude-only cost, what the same traffic would have cost at Claude list prices.
- Spend vs. Claude-only, a per-day chart of actual against baseline. The gap between the two lines is your savings.
Operational metrics
With a window selector (1h, 6h, 24h, 7d):- Inference volume: requests by status class (2xx, 4xx, 5xx).
- Response time: end-to-end p50, p90, and p99 latency.
- Rate-limited requests: requests shed by plan limits.
- Token usage: throughput in tokens per second.
Where your requests went
A served-model mix table shows, for each model, its share of traffic, its request count, and its average savings against the Claude baseline. Rows served by a frontier Claude model show “n/a”, since they are the baseline. Models you have configured but that saw no traffic still show up, zero-filled, so nothing is missing from the picture.Your own recent requests
The dashboard aggregates. To attribute a single turn, runvalar usage on the machine that made
it. Under the spend summary it lists your last requests, newest first, with the routing tier each
request resolved into beside the model that actually answered.
-n for up to 100, or --no-requests to skip the table and see
the spend summary alone. The feed covers the last 7 days, is scoped to your own client id, and spans
every harness you use, so a Cursor or Codex request shows up beside a Claude Code one. That is what
lets it answer “which model produced that bad turn” for a tool that keeps no local log.
The class column is the routing tier: low, medium, high or max, corresponding to Haiku,
Sonnet, Opus and Fable. It is the tier the request resolved into, not the one it asked for. Under
Auto routing the router can re-class a request, so a Sonnet call sent to a low-tier model reads as
low. Read it as the routing decision rather than as what you typed.
The request id is Valar’s own, not the id your tool displays. Find the row by time, then quote that
id to us.
How savings are worked out
Savings compares two numbers for the same traffic, each priced on its own:- Actual spend: the real per-token rate of the model that served each request. Open-weight models use their catalog price, and the frontier Claude tiers use Claude list prices.
- Claude-only baseline: what the same traffic would have cost on Claude at list prices. Requests already served by Claude use that tier’s own list price. For an open-weight model, the tier it is priced against is the one that model stands in for — a top-tier open model is compared with a top-tier Claude, not with the same Claude for every model.
Claude-served traffic showing about $0 in savings is working as intended. Savings show up only when traffic routes to an open-weight model.
Next steps
Model routing
Adjust the split and watch savings respond.
ValarCode overview
Connect more engineers to grow the sample.