LLM caching stores a previous result so the application can reuse it instead of making the same expensive model call again.

Why is caching useful?

AI requests can be slower and more expensive than normal database queries.

If many users ask:

What is the company's leave policy?

and the answer is the same, generating it every time may waste time and money.

Basic cache flow

flowchart LR
    A["Request"] --> B{"Cached result?"}
    B -->|Yes| C["Return cached response"]
    B -->|No| D["Call LLM"]
    D --> E["Store response"]
    E --> C

Exact-match caching

Reuse a response only when the cache key is exactly the same.

A useful cache key may include:

model
+ prompt-template version
+ normalized user input
+ relevant settings
+ data or document version
+ user or tenant scope

If any important input changes, the cache should not return the old result.

Example

const cacheKey = createHash({
  model: "chosen-model",
  promptVersion: "faq-v2",
  question: normalizedQuestion,
  policyVersion: currentPolicyVersion
});

let answer = await cache.get(cacheKey);

if (!answer) {
  answer = await generateAnswer(question);
  await cache.set(cacheKey, answer, { ttl: 3600 });
}

return answer;

TTL and invalidation

TTL means time to live: how long a cached result remains valid.

Invalidation means removing the cache when its source data changes.

Example:

HR policy updated
    โ†“
Old RAG answer may be wrong
    โ†“
Delete cache or change policy version in cache key

What is semantic caching?

Semantic caching may reuse an answer for questions with similar meaning:

"What is the leave policy?"
"How does annual leave work?"

It can improve cache hits, but it is riskier because similar questions are not always identical.

Start with exact caching unless semantic reuse is clearly safe.

When should you avoid caching?

Be careful with:

Never let one user's private response appear for another user.

Common mistakes

Cache only when the same effective input should produce a reusable answer. Correct cache keys and invalidation matter more than simply adding Redis.