THE USEFUL DISTINCTION

Different caches reuse different things. Choose the layer you can control, isolate data correctly, and measure total cost and quality rather than cache-hit rate alone.

Three kinds of caching, three different jobs.

‘Add caching’ sounds like a single infrastructure decision. In a language-model system, it can mean reusing attention state inside an inference engine, reusing a shared prompt prefix through a provider, or reusing a completed answer for a similar request.

These techniques have different eligibility rules, memory costs, privacy implications, and effects on quality. A gateway can coordinate some of them, but it cannot create a cache capability that the underlying model service does not expose.

KV caching reuses computation inside inference.

During generation, a transformer can retain the attention keys and values computed for earlier tokens. This avoids recomputing that state at every generation step. The cache belongs to the inference engine and consumes memory; long contexts and many concurrent requests can make that memory significant.

Some serving engines also support reusing compatible prefix state across requests. Availability and isolation depend on the engine and configuration. In a self-hosted deployment, evaluate memory use, batching, concurrency, and time to first token together. An application proxy cannot inspect or control a hosted provider’s internal KV cache.

Provider prompt caching reuses eligible prefixes.

A model provider may offer caching for repeated prompt prefixes. A stable block of instructions, tool definitions, or reference material can be a candidate, provided it meets the provider’s requirements. The request can still produce a new answer; the cached object is not simply the previous response.

Eligibility, minimum prefix sizes, lifetimes, cache writes, and pricing vary. Keep stable content before per-request content where the provider’s format allows it. Avoid injecting a changing timestamp or request identifier into the reusable prefix unless it is necessary.

Measure the effective input cost using actual cache reads and writes. A workload that rarely repeats within the cache lifetime may see little benefit, even if its prompts are long.

Semantic caching reuses a response, so correctness matters.

A semantic cache looks for a prior request that is sufficiently similar and returns or adapts its stored response. It can avoid a new generation, but similarity is not the same as equivalence. Two questions can look alike and still require different answers because of user identity, date, permissions, or underlying data.

Tenant and permission boundaries, model and policy versions, freshness, and invalidation are part of the cache key and retrieval policy. Dynamic account information and consequential actions often need to bypass this layer. A previous tool result must not be mistaken for permission to repeat the action.

  • Choose workloads where response reuse is appropriate.
  • Partition by identity, tenant, permissions, and relevant context.
  • Set expiry and invalidate when source information changes.
  • Evaluate wrong-answer risk as well as saved requests.

Optimise for the completed task.

A high hit rate is only useful if the responses remain correct and the total economics improve. Include cache writes and reads, storage, embeddings, routing overhead, and invalidation work in the calculation. Compare end-to-end latency and quality with a baseline.

Begin with a representative workload and one cache strategy. Record cost per completed task, latency, answer quality, and where misses occur. This makes it easier to decide whether to tune a prefix, change a lifetime, narrow a semantic cache, or leave a workload uncached.

AIMS / PRACTICAL INTELLIGENCEMore field notes