# MCP response caching

When agents call the same MCP tools repeatedly with similar arguments, every call travels to the upstream server and the identical result is sent back to the LLM context each time. Catalyst can cache those responses at the sidecar layer so repeated calls return instantly from memory.

Caching is **opt-in and per-server** — you enable it by adding a `cache` block to the `MCPServer` spec.

## What is cached

| MCP method | Default TTL |
|---|---|
| `tools/call` | 30 s |
| `resources/read` | 30 s |

Both are cached whether the upstream server answers with `application/json` or as a Streamable HTTP stream, which is what most MCP servers return.

Never cached:

- **Mutating and session operations** — `initialize`, `subscribe`, `unsubscribe`, and notifications.
- **Error responses** — a JSON-RPC error object, or a result with `isError: true`.
- **Calls that ask the client a question** — if a tool sends a sampling or elicitation request back to your agent while it runs, the result depended on the answer *that* agent gave, so it is never replayed to another caller.
- **Responses larger than `maxSize`** — see [Response size limit](#response-size-limit).

:::note
MCP List operations (`tools/list`, `prompts/list`, `resources/list`, `resources/templates/list`) are **not** cached. Each caller sees a list filtered to the tools its [access policy](https://docs.diagrid.io/develop/mcp/mcp-access-policies) grants, and that filtering happens in place of caching. The `metadataTTL` field is accepted in the spec but currently has no effect on caching.
:::

### Progress and log notifications

An MCP tool that reports progress or emits log notifications while it runs sends those as separate frames alongside its result. A cached reply carries **only the result** and no progress or log notifications.
If your agent depends on seeing those frames on every call, do not enable caching for that server.

## Enable caching

Add a `cache` block to your `MCPServer` spec:

```yaml
apiVersion: dapr.io/v1alpha1
kind: MCPServer
metadata:
  name: my-mcp
spec:
  endpoint:
    streamableHTTP:
      url: https://mcp.example.com/mcp
  cache:
    enabled: true
```

This enables caching with the default 30-second TTL. Apply the change:

```bash
diagrid apply -f my-mcp.yaml
```

### Custom TTL

Override the TTL to suit how quickly the upstream data changes:

```yaml
spec:
  cache:
    enabled: true
    ttl: 60s
```

A shorter `ttl` keeps data fresher at the cost of more upstream calls.

### Response size limit

`maxSize` caps the size of a single response that may be stored, as a quantity such as `256Ki` or `1Mi`:

```yaml
spec:
  cache:
    enabled: true
    ttl: 60s
    maxSize: 256Ki
```

The default is `64Ki` and note that any response above the limit **still reaches your agent in full**, it is just not stored, so every call for it goes upstream.

`maxSize` has a ceiling of `1Mi`. A larger value is clamped to that ceiling rather than refused, so an over-ambitious setting quietly degrades to 1 MiB instead of turning caching off. Size the value against the responses you want cached, and expect nothing above `1Mi` to take effect.

One cache serves every `MCPServer` the sidecar proxies, and it holds 1024 entries or 64 MiB in total, whichever it reaches first. Raising `maxSize` on one server buys larger entries for that server rather than more memory overall, so the servers share a fixed budget between them.

The limit applies to the whole response as it arrives, including any progress and log frames that precede the result, so a tool that logs heavily can exceed it even when its result is small.

:::warning
A large `maxSize` is a deliberate memory decision, and nothing caps it for you. A streamed response is held in memory for the life of the request, so the value bounds what one in-flight call costs as well as what a stored entry costs — concurrent calls to the same server multiply it. Raise it to fit the responses you actually want cached, not as a precaution.
:::

## Cache key and caller isolation

Each cache entry is keyed on:

- **MCP server name** — the `MCPServer` resource name
- **Method** — e.g. `tools/call`
- **Parameters** — the JSON-RPC `params` object (canonicalized), excluding `_meta.progressToken`
- **Caller** — the App ID of the calling agent
- **End user** — the verified end-user token the call carried, when there is one

Because the caller is part of the key, **caches are always isolated per agent**. Agent A's cached result is never served to Agent B, even for the same tool and arguments. The end-user token is in the key for the same reason: the proxy forwards that identity upstream, so an MCP server may answer differently per user, and two users behind a single App ID never share an entry.
`_meta` is the reserved object MCP clients use to attach out-of-band fields to a call. Only its `progressToken` member is dropped from the key: that is a fresh value on every call, so including it would give every request its own key and the cache would never return a hit. Every other member of `_meta` is kept, because an upstream server may read one and vary its result on it.

## Confirming a cache hit

A response served from the cache carries this header:

```
X-Diagrid-MCP-Cache: hit
```

A response fetched from the upstream server has no such header. The body of a hit is identical to the original result, with the JSON-RPC `id` rewritten to match your request, so the header is the only thing that distinguishes the two responses.

There is a second signal, and it is the one to know about when you are reading logs rather than responses: **a cache hit writes no entry to the App ID's API logs.** A hit never leaves the sidecar, so logging it would record an outbound call to the MCP server that did not happen. Do not read a missing log line as a dropped request — check the header, or the cache metrics, instead.

## Cache invalidation

Cache entries expire by TTL. Nothing else invalidates an entry, but an entry can still disappear early: when the cache reaches its entry or byte budget, storing a new response evicts the entries closest to expiring until the new one fits.

:::warning
Changing the `MCPServer` spec — the URL, its credentials, or the cache settings themselves — does **not** flush entries that are already cached. After such a change, allow up to one full `ttl` before every caller sees results from the new configuration. Set a short `ttl` on servers you expect to reconfigure often.
:::

## When not to use caching

Caching works best for tools that return the same result for the same input within a short window. Avoid enabling it for:

- **Side-effecting tools** that create, update, or delete resources (e.g. `create_issue`, `send_email`). While caching won't prevent the first call, it would serve a stale response on a retry.
- **Real-time data** where even a 30-second delay is unacceptable.
- **Tools whose progress or log output your agent consumes**, since a cache hit replays the result alone.

## See also

- [Manage MCP Servers](https://docs.diagrid.io/operate/project-operations/mcp-servers) — register, update, and delete `MCPServer` connections.
- [Connect an MCP client](https://docs.diagrid.io/develop/mcp/connect) — point an agent at the Catalyst MCP proxy.
