Semantic Caching for LLM Apps: The Freshness, Safety, and Evaluation Playbook

Semantic Caching for LLM Apps: The Freshness, Safety, and Evaluation Playbook

Avatar of Do Quoc Viet

Written by

Do Quoc Viet

Published
Reading time

Listen to article

Ready to read

A semantic cache protects an LLM application with freshness, authorization, and evaluation gates

I once watched a support assistant answer a customer’s question in less than 100 milliseconds. The latency graph looked beautiful. The token bill had dropped. The cache hit rate was high enough to make the dashboard feel like a success story.

Then the policy team changed one sentence in the refund rules.

The assistant kept returning the old answer because the user’s new question was semantically close to a cached question from the previous week. Nothing crashed. No API returned a 500. The embedding search did exactly what it had been asked to do. The system was fast, inexpensive, and wrong in a way that was difficult to see from infrastructure metrics alone.

That is the production problem with semantic caching. It is not merely a clever key-value store with vectors. It is a decision system that decides when an old piece of model work is safe enough to reuse.

The thesis: A semantic cache should optimize a response only after it has established three things: the request is in the same policy scope, the cached evidence is fresh enough for this request, and the quality risk of reusing the result is below the application’s tolerance.

This article builds a practical playbook for LLM and RAG applications. It covers the basic mechanics, then spends most of its time on the parts that determine whether a cache is a performance feature or a silent correctness bug: freshness, invalidation, authorization, poisoning, intermediate context, evaluation, and rollout.

Semantic caching is similarity with a policy layer

Traditional caching uses a deterministic key. A request such as GET /products/4821?currency=VND maps to a known cache key, and the system either finds that exact representation or misses. The contract is relatively simple: the key describes the request, and the expiration policy describes how long the representation may be reused.

LLM requests are less repetitive at the string level. A customer may ask “Can I return this item?” or “What is the refund window for this order?” or “I changed my mind — how many days do I have to send it back?” The wording differs, but the intent may be similar. An embedding turns a text string into a vector, and a similarity search can locate previous requests that are close in meaning. OpenAI describes embeddings as vector representations used to measure relatedness between text strings, with cosine similarity as a common comparison function.

Redis describes the basic semantic-cache flow as embedding the incoming query, searching stored vectors, returning a cached response when the similarity is above a threshold, and calling the LLM on a miss. That is the useful starting point. It is not the full production contract.

A request moves through normalization, scope checks, semantic lookup, freshness validation, and either a safe cache hit or a new model run

A production cache has to answer questions that similarity alone cannot answer:

Question Why a vector score is not enough
Is this the same tenant, user, product, or permission scope? Two questions can be semantically identical but must not share an answer.
Is the source material still current? A high similarity score says nothing about document version or policy age.
Was the cached answer generated by the same prompt, model, and policy? Changes in instructions or tool semantics can make an old answer incompatible.
Is this a read-only explanation or a decision with side effects? Reusing a low-risk FAQ answer is not equivalent to replaying an approval decision.
Has the cached record been poisoned or contaminated? The cache becomes a durable storage layer for any mistake that passes the write path.

The important mental shift is to treat the vector score as one signal inside a cache admission policy, not as the policy itself.

Choose the right thing to cache

Teams often start by caching the final text because it is easy to store and easy to return. That is reasonable for a stable FAQ. It is a poor default for every RAG or agentic workflow.

There are at least four cache boundaries:

Boundary What is stored Good fit Main risk
Embedding lookup Query vector or normalized intent Avoiding repeated embedding work A vector is not an answer and still needs a safe retrieval policy.
Retrieved context Document chunks, summaries, or ranked evidence RAG systems where sources change independently of generation Stale or unauthorized evidence can be reused.
Intermediate computation Query rewrite, classification, extraction, or contextual summary Multi-step pipelines with repeated subproblems An error propagates into many downstream answers.
Final answer Model text plus evidence and metadata Stable, low-risk, read-only questions The answer may be stale, mis-scoped, or incompatible with a new policy.

A useful rule is to cache the lowest layer that is expensive and still safe to recompute into the current request. If a product catalog changes often, cache a normalized retrieval result with document versions rather than a final sentence that says “the product costs 599,000 VND.” If a classification step is stable and tenant-independent, cache the classification. If the output grants credit, changes account state, or exposes personal data, do not treat the final answer as a freely shareable object.

Research on semantic caching for contextual summaries makes a similar point: intermediate results can be reused across related requests and can better tolerate partial document updates and changing access patterns than only caching an end-to-end answer. The practical implication is that the cache boundary is an architecture decision, not a storage optimization.

Give every entry a cache envelope

A cache record should carry enough information for the read path to decide whether reuse is safe. Storing only {query, answer, embedding} is not enough for production.

type CacheEnvelope = {
  id: string;
  queryFingerprint: string;
  embedding: number[];
  answer?: string;
  context?: Array<{
    documentId: string;
    version: string;
    chunkId: string;
    contentHash: string;
  }>;
  tenantId: string;
  subjectScope: string;
  model: string;
  promptVersion: string;
  policyVersion: string;
  retrievalVersion: string;
  createdAt: string;
  expiresAt: string;
  riskClass: "low" | "medium" | "high";
  provenance: "human-reviewed" | "generated" | "imported";
  status: "active" | "stale" | "revoked";
};

The envelope is intentionally boring. Boring metadata prevents exciting incidents.

tenantId and subjectScope stop a response generated for one customer from becoming a neighbor’s answer. promptVersion, policyVersion, and retrievalVersion prevent a cache hit from silently bypassing a release boundary. Document versions and content hashes make invalidation explainable. riskClass allows the system to use a more conservative policy for financial, identity, medical, or state-changing answers.

The cache key should be derived from the envelope’s policy-relevant fields, not just the raw user question. One possible shape is:

semantic-cache:v3:
  tenant={tenantId}:
  scope={subjectScope}:
  model={model}:
  prompt={promptVersion}:
  policy={policyVersion}:
  retrieval={retrievalVersion}:
  boundary={cacheBoundary}

The semantic index can still search by embedding, but every candidate must pass the deterministic scope filter before its similarity score is considered. Similarity should never be a way to cross an authorization boundary.

Freshness is not the same as TTL

A time-to-live is useful, but it is only one expression of freshness. RFC 9111 makes the distinction clear for HTTP caches: a response is fresh when its age is within its freshness lifetime, and a stale response may require validation before reuse. The same mental model works for LLM caches, with one important addition: the origin is often a document store, policy service, database, or tool—not only a web server.

Imagine a policy answer generated at 09:00 with policyVersion=41. At 09:05, the policy service publishes version 42. The cached answer may have a one-hour TTL, but it is no longer fresh relative to the policy source. Waiting until 10:00 is not a freshness policy; it is delayed bug discovery.

Use several freshness signals together:

Signal Meaning Typical action
TTL A maximum age fallback Reject or revalidate after expiration.
Source version The exact version of a policy, document, price list, or schema Invalidate when the version changes.
Content hash Whether the material used by the answer changed Recompute affected entries.
Event timestamp When the source emitted an update Trigger targeted invalidation.
Risk class How costly a stale answer would be Use shorter TTL or no final-answer reuse for high risk.
Validation result Whether the candidate still matches current evidence Allow, downgrade to context-only, or miss.

A freshness matrix combines source version, TTL, risk class, and validation result before allowing reuse

A good invalidation rule is often more specific than “delete everything every hour.” If document refund-policy-v42 changes, invalidate entries whose provenance includes that document. If a user’s role is revoked, invalidate entries scoped to that subject. If the prompt changes from support-v7 to support-v8, either namespace the cache or deliberately run a migration job that regrades existing entries.

For RAG, you can model dependency edges explicitly:

cache_entry_9f2
  depends_on -> refund-policy:v42#chunk-7
  depends_on -> return-form:v12#chunk-2
  generated_by -> support-prompt:v7
  constrained_by -> policy-bundle:v19

The entry does not need to be deleted immediately for every update. A stale marker can route it to revalidation, a background refresh, or intermediate-context reuse. What matters is that the system knows why an entry is stale and can show that reason in a trace.

Thresholds are tuned with outcomes, not vibes

A similarity threshold is a useful control, but it is not a universal constant. A lower threshold tends to produce more hits and more false positives. A higher threshold tends to be safer but may miss useful reuse. The right value depends on language, domain, embedding model, query distribution, risk, and the amount of context preserved in the cached record. The semantic-caching research literature treats threshold selection as a trade-off between utility and hit rate rather than a one-number recipe.

Build an offline threshold set from real, sanitized traffic. Pair queries with labels such as:

Label Meaning
Safe reuse Same intent, same scope, same relevant evidence, and equivalent answer.
Context reuse only Similar question, but the final answer must be regenerated against current evidence.
Miss required Different intent, different authority, changed source, or high-risk ambiguity.
Adversarial The similarity is designed to lure the cache across a boundary or reuse poisoned content.

Then evaluate thresholds across more than hit rate. A useful scorecard includes cache precision, cache recall, answer correctness, stale-answer rate, unauthorized reuse rate, p50/p95 latency, token savings, and cost per successful task.

cache_precision = safe_reuses / all_cache_hits
cache_recall    = safe_reuses / all_requests_that_could_reuse
stale_rate      = stale_hits / all_cache_hits

quality_adjusted_savings =
  (baseline_cost - cache_cost) * correct_answer_rate

The last metric is deliberately conservative. A cheap wrong answer is not a successful optimization. If a cache reduces model spend by 60 percent while introducing a 2 percent unauthorized reuse rate in a sensitive workflow, the dashboard should not celebrate the raw savings.

Use separate thresholds for separate risk classes. A public product FAQ may tolerate a lower threshold than an identity-verification assistant. A final answer may require a higher threshold than a retrieved context candidate because the final answer has already committed to a conclusion.

Protect the cache from poisoning and replay

A semantic cache increases the half-life of model mistakes. Without a cache, a bad answer may disappear after the next request or after a prompt update. With a cache, the mistake can be retrieved repeatedly and look consistent. Consistency is not evidence of correctness.

There are several poisoning paths:

  1. A malicious user submits a prompt that causes the model to produce a dangerous answer, hoping a semantically similar future request receives it.
  2. A retrieved document contains instructions that are treated as trusted answer content and then stored with the result.
  3. An internal operator imports a cache snapshot from the wrong environment or tenant.
  4. A prompt or policy change makes old outputs invalid, but the old namespace remains searchable.

The mitigations are architectural. Separate cache write permissions from cache read permissions. Mark unreviewed entries as generated rather than trusted. Never cache a tool’s authorization decision as if it were a user-independent fact. Store provenance and keep a revocation path. For high-risk flows, use the cache only to retrieve evidence and force a fresh policy/model decision.

This also connects to prompt-injection defenses. A cached response is still model-produced data. If a malicious document can influence a cached answer, later users may encounter the payload without submitting the original malicious query. Treat cached text as untrusted input at the next prompt boundary, preserve source identifiers, and apply the same data/instruction separation used elsewhere in an agent system.

Scope is a security property, not a performance detail

A cache hit can be semantically perfect and still be a data breach. Consider two employees asking, “What is the status of my reimbursement?” The questions are close in vector space. Their answers must not be shared unless the cache record is explicitly scoped to a public, identical source.

At minimum, decide whether each cache boundary is:

Scope Example Safe default
Public A published return policy Shared, with source-version validation.
Tenant A company’s internal handbook Shared only inside the tenant.
User A person’s own order history Keyed to user and subject.
Session A temporary conversation assumption Keyed to session/thread and short-lived.
Action A decision or authorization Do not reuse as a general answer; recompute or revalidate.

The scope must be checked before vector similarity. This is analogous to double-keying in web caches to reduce privacy risk: the identity context becomes part of the lookup contract rather than an afterthought.

Do not let a model infer scope from the question. Scope should come from authenticated request context, server-side policy, and the current authorization snapshot. If the current user cannot access a source document, the cache must not reveal a summary of it, even if the summary itself appears harmless.

Observe the cache as a quality system

A cache dashboard with only hit rate and latency is an invitation to optimize the wrong thing. Add cache-specific telemetry to the same trace as the model and retrieval steps. The minimum useful events are:

Event Fields worth recording
cache.lookup boundary, tenant hash, candidate count, top score, threshold, scope result
cache.reject reason: scope, stale, policy version, low score, risk, missing provenance
cache.hit entry age, source age, model, prompt version, answer reuse or context reuse
cache.miss miss reason and downstream latency/cost
cache.write provenance, risk class, source dependencies, review status
cache.invalidate dependency, actor, reason, number of entries affected
cache.feedback user correction, grader outcome, incident link, regrade result

Do not put raw private prompts and answers into every trace by default. The folio’s existing observability principles apply here: record shape, versions, IDs, hashes, counts, and policy decisions first; keep content behind restricted access and an explicit break-glass path.

A cache-quality dashboard tracks hits, safe-reuse precision, stale responses, invalidations, and cost savings together

A cache hit should be considered successful only when the downstream quality signal agrees. Useful feedback loops include user corrections, citation checks, deterministic policy validators, sampled human review, and regression cases. When a cache hit fails, store a minimal failure fixture: query shape, scope, entry metadata, source versions, score, and observed outcome. The answer text may be restricted or redacted, but the failure should still become testable.

A rollout plan that does not begin with “turn it on”

A safe rollout can fit into a week if the cache boundary is small and the domain is low risk.

Day 1: define the eligible surface. Pick one read-only workflow with repeated questions. Write down what may be cached, which scope is allowed, what sources control freshness, and which answers must always miss.

Day 2: instrument a shadow lookup. Generate embeddings and search candidates, but never return a cached answer. Measure score distributions, candidate scopes, source ages, and likely reuse labels.

Day 3: build the envelope and invalidation path. Add versions, provenance, source dependencies, risk class, and a way to mark entries stale. If invalidation is not explainable, the cache is not ready.

Day 4: add offline and adversarial evaluation. Test paraphrases, ambiguous questions, changed documents, revoked access, prompt-version changes, poisoned records, and cross-tenant collisions.

Day 5: enable context-only reuse. Let the cache supply retrieved evidence or summaries while the current prompt and policy produce a fresh final answer. This usually gives a safer first win than returning old final text.

Day 6: enable low-risk final-answer reuse. Use a conservative threshold, a small traffic percentage, and an instant kill switch. Record every rejection reason, not only hits.

Day 7: review quality-adjusted savings. Promote only if the cache improves latency or cost without breaching correctness, freshness, privacy, or safety budgets.

Production checklist

Before enabling final-answer reuse, verify that the system can answer “why was this reused?” with evidence rather than a similarity score alone.

Area Production question
Boundary Are we caching the final answer, context, classification, or embedding?
Scope Can a candidate cross tenant, user, session, or authorization boundaries?
Freshness Which source version, event, hash, TTL, or validator controls reuse?
Compatibility Are model, prompt, policy, retrieval, and tool-schema versions recorded?
Risk Do high-impact decisions bypass final-answer cache reuse?
Poisoning Can entries be quarantined, revoked, regraded, and traced to provenance?
Evaluation Do we measure safe-reuse precision, stale hits, and unauthorized reuse?
Privacy Are raw prompts and answers protected, redacted, and retention-limited?
Operations Is there a kill switch, namespace rollback, and an owner for invalidation?
User experience Can the UI explain uncertainty or refresh when a cached result is not enough?

Conclusion: the cache is part of the answer’s trust boundary

Semantic caching is attractive because it attacks two painful properties of LLM systems at once: repeated work and unpredictable latency. Embeddings make it possible to recognize related requests even when the words change. Vector search makes the lookup practical. But neither creates a correctness guarantee.

The production-grade design is a policy pipeline. First, constrain the candidate by tenant and authority. Then validate model, prompt, policy, retrieval, and source versions. Apply a threshold calibrated against outcomes, not anecdotes. Prefer intermediate context reuse when final-answer reuse would freeze a conclusion. Track provenance and keep a revocation path. Finally, judge the cache on quality-adjusted savings rather than raw hit rate.

A cache hit is not the end of the decision. It is the moment when the system must prove that reusing old intelligence is safer than doing the work again.

References

Share this article

Found this post helpful? Feel free to share it with your network.

Comments (0)

Loading comments...