> ## Documentation Index
> Fetch the complete documentation index at: https://docs.valarhq.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Prompt Caching

> How Valar reuses shared prompt prefixes, and what you can do to hit the cache more often

## How it works

Requests that share an opening are common: an agent that sends the same system prompt and tool definitions on every run, a chat that resends its history on every turn, a document a user asks several questions about. Paying full price to reprocess that shared text every time is waste, and it is usually the largest part of the bill.

Prompt caching removes it. Valar matches the shared prefix for you and bills those tokens at the **cached** rate, a fraction of the input rate on every model that offers one — so the repeated part of your prompt gets dramatically cheaper the more you send it. Responses also start sooner, since the model skips work it has already done.

Caching is on by default and there is nothing to configure. The rest of this page is about hitting it more often.

Two things decide whether a request hits:

1. **A prefix matches byte for byte, from the very start of the prompt.** Matching stops at the first byte that differs.
2. **A cache is local to the machine that served the request.** A later request hits only if it lands on that same machine.

You control the first completely. The second is what [session affinity](#routing-to-a-warm-cache) is for.

## Keep your prompt prefix stable

Because matching starts at the beginning of the prompt and stops at the first difference, the order of your prompt decides how much of it can be reused.

Put what stays the same at the front — system prompt, tool definitions, few-shot examples, a long document — and what varies at the end. Usually only the user's question changes, so it belongs last.

<Warning>
  Dynamic content at the *start* of a prompt kills your hit rate. A timestamp, a request id, or a
  session id in the opening line changes the first bytes on every request, so nothing behind it can
  be reused — including a system prompt that never changed at all.
</Warning>

```python ❌ Don't — the timestamp changes the first bytes, so nothing after it is reused theme={"system"}
response = client.responses.create(
    model="zai-org/GLM-5.3",
    instructions=f"Current time: {datetime.now()}\n\n{SYSTEM_PROMPT}",
    input=question,
)
```

```python ✅ Do — the prefix is identical every time; only the tail varies theme={"system"}
response = client.responses.create(
    model="zai-org/GLM-5.3",
    instructions=SYSTEM_PROMPT,
    input=f"{question}\n\nCurrent time: {datetime.now()}",
)
```

The same mistake wears other clothes:

* A request id or trace id interpolated into the system prompt.
* Tool definitions serialized in a different order between runs, or with object keys emitted in a different order.
* Retrieved documents or a conversation summary placed before the system prompt instead of after it.
* Few-shot examples sampled at random per request.

Caching pays in proportion to how long and how repeated your prefix is. A short prompt may not cache at all.

<Tip>
  If a prompt you believe is identical reports no cached tokens on its second call, diff two real
  request bodies before changing anything else. The difference is almost always in the first few
  lines.
</Tip>

## Routing to a warm cache

The second property is the one you cannot fix by editing your prompt. A cache belongs to the machine that built it, so two requests share a cache only if the same machine serves them.

Turns of one conversation usually land together. Separate runs of the same agent usually do not — which is why an agent with a long shared system prompt can cache well *within* a run and hardly at all *across* runs.

Send an `X-Session-Affinity` header to tell us which requests belong together. We use it to steer them toward a machine that may already have their prefix hot, instead of picking one blind. It is a hint, not a pin: capacity, load, and eviction all still get a vote, and a busy machine may hand your request to another one.

<CodeGroup>
  ```python OpenAI SDK theme={"system"}
  from openai import OpenAI

  client = OpenAI(
      base_url="https://api.valarhq.ai/v1", # or read OPENAI_BASE_URL from the environment
      api_key="YOUR_VALAR_API_KEY",  # or read OPENAI_API_KEY from the environment
  )

  response = client.responses.create(
      model="zai-org/GLM-5.3",
      instructions=SYSTEM_PROMPT,
      input="Review pull request #4127.",
      extra_headers={"X-Session-Affinity": "pr-review-agent"},
  )
  ```

  ```python Anthropic SDK theme={"system"}
  from anthropic import Anthropic

  client = Anthropic(
      auth_token="YOUR_VALAR_API_KEY",
      base_url="https://api.valarhq.ai",  # no /v1 — the SDK appends it
  )

  message = client.messages.create(
      model="zai-org/GLM-5.3",
      max_tokens=2048,
      system=SYSTEM_PROMPT,
      messages=[{"role": "user", "content": "Review pull request #4127."}],
      extra_headers={"X-Session-Affinity": "pr-review-agent"},
  )
  ```

  ```bash cURL theme={"system"}
  curl https://api.valarhq.ai/v1/responses \
    -H "Authorization: Bearer $VALAR_API_KEY" \
    -H "Content-Type: application/json" \
    -H "X-Session-Affinity: pr-review-agent" \
    -d '{
          "model": "zai-org/GLM-5.3",
          "instructions": "...your shared system prompt...",
          "input": "Review pull request #4127."
        }'
  ```
</CodeGroup>

Available on `/v1/responses`, `/v1/chat/completions`, and `/v1/messages`. There is nothing to pre-register and no value to look up: the string is yours to pick, up to 1024 bytes. Values are scoped to your organization, so yours never collide with another customer's.

The header is optional. Omit it and nothing changes: requests are routed as they are today, and automatic prefix caching still applies.

### Choosing a value

Use **one value per group of requests that share a prefix**.

For an agent, that is usually one value for the agent itself — every run of a PR reviewer that opens with the same system prompt and tools sends `pr-review-agent`. For a chat product it is usually one value per conversation or per user, because there the repeated prefix is the history rather than the system prompt. Group by whatever the requests actually share.

<Warning>
  Do not use a per-request value. A fresh UUID on every call puts every request in its own group,
  so nothing shares a cache and you end up worse off than sending no header at all. The same goes
  for routing traffic randomly.
</Warning>

### Rules

<ParamField path="A hint, not a guarantee" type="raises your hit rate; never promises one">
  Affinity only influences where a request goes. It cannot create a hit that would not otherwise
  exist — if the front of your prompt varies, it misses wherever it lands, so fix the prompt first.
  And even with a stable prefix, a hit is never certain: caches evict their oldest entries, machines
  restart, and a machine under load may not be the one that serves you.
</ParamField>

<ParamField path="Never affects results" type="routing only">
  Neither caching nor affinity changes model selection, output, or the price of a token. They change
  only how many of your input tokens qualify for the cached rate.
</ParamField>

## Measuring your hit rate

Every response reports its cached input tokens, in the shape its dialect uses:

| Endpoint | Field |
| - | - |
| `/v1/responses` | `usage.input_tokens_details.cached_tokens` |
| `/v1/chat/completions` | `usage.prompt_tokens_details.cached_tokens` |
| `/v1/messages` | `usage.cache_read_input_tokens` |

They also appear in the [usage endpoints](/usage-endpoints), which is the easier place to look at a hit rate across many requests.

Measure before and after a change rather than assuming. Cached tokens bill at the cached input rate on the [pricing page](/pricing).

## Batch requests

Batch requests take no affinity hint — a batch carries no per-request header. Prefix stability still applies to every item, so the ordering advice above is worth following there too.
