Skip to main content

GET /v1/usage

Summarizes spending and balance over the time range you request. By default, spend includes all workspaces in your organization, regardless of which workspace owns your API key.
Every monetary value is expressed in cents. The prior_period block mirrors the same-length window directly preceding your requested range, so you can compare the two.
When you supply workspace_id, the response echoes it as a top-level field. plan_type still describes your organization, not a separate workspace plan. Workspace grouping is available on /v1/usage/breakdown, not this endpoint.
The keys in sla_mix are completion window names: asap (the Now tier), standard, and flex.
balance, balance_unavailable, and days_remaining describe a prepaid-credit balance. Until prepaid credits are available, balance_unavailable is always true, balance is 0, and days_remaining is omitted - spend, burn rate, tokens, and the SLA mix are always reported.

GET /v1/usage/breakdown

Provides per-model rankings alongside time-series spend data - ideal for building charts.
For 7d, 30d, and period ranges granularity is "day"; for a 24h range or a day drill-down it is "hour".

Workspace filtering and breakdowns

To retrieve one workspace’s costs for a UTC day:
Use the same workspace_id parameter on /v1/usage for a filtered summary, including its prior-period comparison:
To separate costs by workspace in a single response, use group_by=workspace. This also gives you the workspace IDs to use in filtered requests:
For example, the command above can produce:
The API response keeps the existing top-level data and models fields and adds group_by: "workspace" and workspaces. Each workspace contains: Workspaces are sorted by total descending, then by workspace_id. Only workspaces with usage records in the selected range are included; a workspace with tokens but zero spend is still included. An empty result returns workspaces: []. BYOK usage remains excluded, as it is from the ungrouped spend response. You can combine workspace_id and group_by=workspace. The top-level totals and the workspace array then include only that workspace. The response echoes the filter in a top-level workspace_id field. Without either parameter, the response shape and organization-wide scope stay unchanged.
A workspace filter can only narrow your authenticated organization’s usage. An unknown workspace ID, a workspace outside your organization, or a workspace without usage returns an empty breakdown and zero summary usage metrics. An omitted or empty workspace_id includes all of your organization’s workspaces.
All monetary fields in the API response are integer cents. The dollars fields above are calculated by jq, not returned by the API. Values are rounded after aggregation, so summing separately rounded workspace or bucket totals can differ slightly from a combined total. Match the dashboard’s workspace scope and UTC date range when comparing costs; usage data can also lag slightly.

GET /v1/usage/activity

Surfaces operational metrics - request counts, token throughput, latency, and a list of recent requests.
A has_more value of true means more recent requests exist beyond the limit you asked for.
Like the other endpoints, activity data lags slightly and is not delivered as a real-time stream.

GET /v1/usage/activity/timeseries

Returns bucketed series for either requests or tokens - use it to draw throughput charts. Requests by model:
Token breakdown:

ValarCode analytics

The /v1/usage/coding/* endpoints report the traffic your ValarCode keys send through the coding lane: which key, which engineer, which model was served, and what it cost. They accept the same Authorization: Bearer API key as the endpoints above (a coding key or a standard key) and always cover the whole organization. Unlike the endpoints above, they also accept the organization automation credential as the bearer, so a provisioning service that manages keys can read their usage with the token it already holds. All three share the same window and filter parameters.

Window

Pick one of the two forms. Combining them returns 400.

Filters

Every filter narrows the rows before they are aggregated, so the filters compose with each other and with any window.

GET /v1/usage/coding/keys

One row per coding key, ranked by spend.

GET /v1/usage/coding/models

One row per served model, ranked by spend. Add key_id to get one key’s model mix.
served_model is the model that actually answered, which under Auto routing can differ from the model the coding tool asked for. share is the model’s fraction of the spend in the response (0–1), so with a key_id filter it is the share of that key’s spend.

GET /v1/usage/coding/requests

The last requests one engineer made, newest first. This is the per-request view: everything else on this page is an aggregate, so this is where you attribute a single turn to the model that answered it. engineer_id is required. limit defaults to 10 and accepts 1-100. The feed covers the last 7 days and spans every coding harness, so a Claude Code turn, a Cursor completion and a Codex run all appear in one list, each tagged by its harness.
class is the routing tier the request resolved into, one of low, medium, high or max, and served_model is what answered. Under Auto routing they differ by design. The tiers correspond to the Claude families: low to Haiku, medium to Sonnet, high to Opus and max to Fable. The feed reports the tier rather than the family because a request need not arrive as a Claude alias at all: an explicit pick of an open-weight model resolves to a tier too. The feed does not carry the exact model id the tool sent.
class is the tier a request resolved into, not the one it asked for. Under Auto routing the router can re-class a request, so a Sonnet call sent to a low-tier model is recorded as low. Read the field as the routing decision.
status is the request’s terminal state (completed, incomplete, failed, cancelled) and failure_code is present only when Valar itself refused the request, for example rate_limit for a quota shed. A request the upstream rejected shows failed with no failure_code. response_id is Valar’s own id for the request. It is not the id your client sees on the response, so use it the way this feed hands it to you: find the row by time, then quote its response_id to support.
valar usage prints this feed at the bottom of its output, so an engineer can read it without writing a request. Use -n to ask for more rows.

GET /v1/usage/coding/timeseries

A bucketed request series for charts. This endpoint takes a window (6h, 24h, 7d, 30d, default 24h) instead of range or start/end, because its buckets are aligned to the current clock and the final bucket is still filling. It accepts the same filters, so ?window=7d&key_id=… charts one key.

Putting it together

Per-key cost and token usage for a calendar month, then the model mix behind the top key:
Spend is integer cents, like the rest of the Usage API. The rollup behind these endpoints is per UTC day and lags live traffic by a few minutes; range=1h reads the live ledger instead and is the only sub-day view.

Errors

Errors share the format used throughout the inference API:

Response headers

Each response carries an X-Request-ID: <uuid>. Quote it in support requests to speed up troubleshooting.