13. RAG Caching StrategiesΒΆ
Category: Production RAG Engineering
Module: Part VI β Production Deployment
Difficulty: Advanced
π OverviewΒΆ
Caching is one of the most important techniques for making production RAG systems:
A naive RAG request may execute:
User Query
β
Query Processing
β
Embedding
β
Vector Search
β
Keyword Search
β
Fusion
β
Reranking
β
Context Assembly
β
LLM
β
Validation
If the same or similar request is repeated, executing every stage again can waste:
A production RAG architecture should therefore consider caching at multiple levels:
RAG CACHE LAYERS
β
βββββββββββββββββββββΌββββββββββββββββββββ
βΌ βΌ βΌ
Embedding Cache Retrieval Cache Rerank Cache
β β β
βββββββββββββββββββββΌββββββββββββββββββββ
βΌ
Context Cache
β
βΌ
Semantic Cache
β
βΌ
Response Cache
However:
Caching in RAG is harder than caching a normal API response because knowledge, authorization, indexes, models, prompts, and retrieval strategies can all change.
The core challenge is therefore:
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand why caching is important in RAG
- Identify different RAG caching layers
- Design embedding caches
- Design retrieval caches
- Design reranking caches
- Design context caches
- Design semantic caches
- Design response caches
- Design tenant-aware caches
- Design authorization-aware caches
- Design version-aware cache keys
- Select appropriate TTL strategies
- Design cache invalidation
- Handle document updates
- Handle index updates
- Handle embedding model changes
- Handle prompt changes
- Prevent cache stampedes
- Handle cache penetration
- Handle cache avalanche
- Use cache warming
- Use distributed caching
- Understand cache consistency
- Measure cache effectiveness
- Optimize cache cost
- Design cache observability
- Integrate caching with CI/CD
- Design production-grade RAG caching architecture
π§ 1. Why Cache RAG?ΒΆ
Consider a request:
Suppose:
Total:
A cache hit could reduce the request to:
depending on the cache layer.
π§ 2. RAG Cost ModelΒΆ
A simplified request cost can be viewed as:
Caching can reduce repeated execution of some or all of these stages.
π§ 3. RAG Cache TaxonomyΒΆ
flowchart TD
A["User Query"] --> B["Query Cache"]
B -->|Miss| C["Embedding Cache"]
C -->|Miss| D["Embedding Model"]
D --> E["Retrieval Cache"]
E -->|Miss| F["Retrieval"]
F --> G["Rerank Cache"]
G -->|Miss| H["Reranker"]
H --> I["Context Cache"]
I -->|Miss| J["Context Assembly"]
J --> K["Semantic Cache"]
K -->|Miss| L["LLM"]
L --> M["Response Cache"] Different cache layers solve different problems.
π§ 4. Main RAG Cache LayersΒΆ
A production RAG platform may use:
1. Document Cache
2. Parsing Cache
3. Embedding Cache
4. Query Embedding Cache
5. Retrieval Cache
6. Reranking Cache
7. Context Cache
8. Semantic Cache
9. Response Cache
10. Model Output Cache
Not every system needs all of them.
π§ 5. Cache PlacementΒΆ
A useful architecture:
USER
β
βΌ
Response Cache
β
βββββββ΄ββββββ
β β
Hit Miss
β β
βΌ βΌ
RESPONSE Semantic Cache
β
βββββββ΄ββββββ
β β
Hit Miss
β β
βΌ βΌ
RESPONSE Retrieval
β
βββββββββββΌββββββββββ
βΌ βΌ βΌ
Embedding Retrieval Rerank
Cache Cache Cache
π§ 6. Cache Layer SelectionΒΆ
Do not cache everything.
Ask:
Is this operation expensive?
Is the result reusable?
How frequently does the input repeat?
How frequently does the result change?
Is the result security-sensitive?
Can stale data be tolerated?
What is the cost of storing it?
π§ 7. Embedding CacheΒΆ
Embedding generation is often deterministic for:
Therefore:
π§ 8. Embedding Cache ArchitectureΒΆ
flowchart LR
A["Text"] --> B["Content Hash"]
B --> C["Embedding Cache"]
C -->|Hit| D["Vector"]
C -->|Miss| E["Embedding Model"]
E --> F["Vector"]
F --> C π§ 9. Embedding Cache KeyΒΆ
A weak key:
A stronger key:
Example:
π§ 10. Why Model Version MattersΒΆ
Suppose:
creates:
Then:
creates:
The old cache must not accidentally return:
for a V2 request.
π§ 11. Embedding Cache ScopeΒΆ
Embedding caches are often good candidates for broader reuse because the vector represents content rather than a user's authorization.
However, sensitive systems should still consider:
especially when cached content itself is stored.
π§ 12. Query Embedding CacheΒΆ
User queries can also be cached:
This is useful when:
π§ 13. Retrieval CacheΒΆ
Retrieval caching stores search results:
Cache:
π§ 14. Retrieval Cache ArchitectureΒΆ
flowchart TD
A["Query"] --> B["Retrieval Cache"]
B -->|Hit| C["Cached Candidates"]
B -->|Miss| D["Dense / Sparse Retrieval"]
D --> E["Candidate Results"]
E --> B
E --> C π§ 15. Retrieval Cache KeyΒΆ
A production retrieval key should include the parameters that influence the result.
Example:
Conceptually:
π§ 16. Why Index Version MattersΒΆ
Suppose:
returns:
After deployment:
returns:
An old retrieval cache must not silently override the new index.
π§ 17. Reranking CacheΒΆ
Reranking can be expensive.
Example:
If the same candidate set is reranked repeatedly:
can avoid repeated computation.
π§ 18. Reranking Cache KeyΒΆ
Include:
Example:
π§ 19. Context CacheΒΆ
Context assembly may include:
The resulting evidence package can be cached.
π§ 20. Context Cache RiskΒΆ
Context can become stale when:
Document Changes
Index Changes
Authorization Changes
Retrieval Strategy Changes
Context Strategy Changes
Therefore context caches need stronger invalidation/versioning than simple application caches.
π§ 21. Semantic CacheΒΆ
A semantic cache attempts to reuse results for:
rather than exact queries.
Example:
These may be semantically equivalent.
π§ 22. Exact Cache vs Semantic CacheΒΆ
Exact CacheΒΆ
Semantic CacheΒΆ
π§ 23. Semantic Cache ArchitectureΒΆ
flowchart TD
A["New Query"] --> B["Query Embedding"]
B --> C["Semantic Cache Search"]
C --> D{"Similarity > Threshold?"}
D -->|Yes| E["Cached Result"]
D -->|No| F["Normal RAG Pipeline"]
F --> G["Store Result"] π§ 24. Semantic Cache ThresholdΒΆ
A semantic cache requires a similarity threshold.
Conceptually:
where:
A threshold that is too low may return incorrect answers.
A threshold that is too high reduces cache hits.
π§ 25. Semantic Cache Is Not Always SafeΒΆ
Consider:
and:
These may be semantically similar but require different answers.
Therefore semantic caching should consider:
π§ 26. Response CacheΒΆ
The simplest cache:
Example:
π§ 27. Response Cache ArchitectureΒΆ
flowchart LR
A["User"] --> B["Response Cache"]
B -->|Hit| C["Response"]
B -->|Miss| D["RAG Pipeline"]
D --> E["Response"]
E --> B π§ 28. Response Cache KeyΒΆ
A production response cache should consider:
Query
Tenant
Authorization Scope
Prompt Version
Model Version
Retriever Version
Index Version
Language
Application Version
Potentially:
if the response depends on them.
π§ 29. Response Cache SecurityΒΆ
This is one of the most important caching concerns.
Unsafe:
Example:
Then:
If access scopes differ, User B may receive unauthorized information.
π§ 30. Tenant-Aware CachingΒΆ
Use:
as part of the cache key.
Example:
π§ 31. Authorization-Aware CacheΒΆ
Tenant ID alone may not be enough.
Two users in the same tenant may have different permissions.
A stronger cache scope can include:
or a stable authorization-scope identifier.
π§ 32. Cache Isolation StrategiesΒΆ
Strategy 1 β Shared Cache + Strong KeyΒΆ
Strategy 2 β Namespace IsolationΒΆ
Strategy 3 β Dedicated CacheΒΆ
π§ 33. Shared vs Dedicated CacheΒΆ
| Strategy | Cost | Isolation | Complexity |
|---|---|---|---|
| Shared | Low | Medium | Low |
| Namespaced | Medium | High | Medium |
| Dedicated | High | Very High | High |
The appropriate choice depends on:
π§ 34. Cache TTLΒΆ
TTL means:
Example:
π§ 35. TTL Strategy by Cache TypeΒΆ
A possible starting point:
Embedding Cache
β Long TTL
Retrieval Cache
β Short / Medium TTL
Reranking Cache
β Short / Medium TTL
Semantic Cache
β Short / Medium TTL
Response Cache
β Depends heavily on freshness requirements
These are starting points, not universal values.
π§ 36. Freshness-Based TTLΒΆ
Instead of one global TTL:
Static Policy
β Long TTL
Frequently Changing Data
β Short TTL
Real-Time Data
β Very Short TTL / No Cache
π§ 37. Data VolatilityΒΆ
Classify knowledge:
LOW VOLATILITY
Policies
Documentation
MEDIUM VOLATILITY
Product Information
HIGH VOLATILITY
Inventory
Prices
Transactions
REAL-TIME
Account Balance
Market Data
Live Status
Caching strategy should reflect volatility.
π§ 38. Cacheability MatrixΒΆ
| Data | Cache? | Typical Strategy |
|---|---|---|
| Static Documentation | Yes | Long TTL |
| Enterprise Policies | Yes | Version-aware |
| Product Documentation | Yes | Version-aware |
| Frequently Updated Data | Carefully | Short TTL |
| Transaction Data | Carefully | Very short / bypass |
| User-Specific Data | Carefully | User-scoped |
| Highly Sensitive Data | Restricted | Strong isolation |
| Real-Time Data | Usually limited | Bypass / short TTL |
π§ 39. Cache InvalidationΒΆ
One of the hardest problems in production RAG is:
Potential invalidation triggers:
Document Update
Document Delete
Index Update
Embedding Model Update
Retriever Update
Reranker Update
Prompt Update
Model Update
Authorization Update
Tenant Policy Update
π§ 40. Invalidation StrategiesΒΆ
Common approaches:
TTL
Explicit Invalidation
Versioned Keys
Event-Driven Invalidation
Write-Through
Write-Behind
Cache Busting
π§ 41. Versioned Cache KeysΒΆ
One of the safest techniques:
When the index changes:
Old entries naturally become unused.
π§ 42. Versioned Cache ArchitectureΒΆ
No need to delete every old entry immediately.
π§ 43. Event-Driven InvalidationΒΆ
flowchart LR
A["Document Updated"] --> B["Change Event"]
B --> C["Cache Invalidation"]
C --> D["Affected Entries Removed"]
B --> E["Index Update"] π§ 44. Document-Level InvalidationΒΆ
Suppose:
changes.
Invalidate cache entries referencing:
This requires maintaining relationships:
π§ 45. Invalidation GranularityΒΆ
Possible levels:
Smaller invalidation scope generally reduces unnecessary cache misses but increases implementation complexity.
π§ 46. Cache Invalidation ArchitectureΒΆ
flowchart TD
A["Document Change"] --> B["Event Bus"]
B --> C["Identify Affected Index"]
C --> D["Identify Affected Cache Entries"]
D --> E["Invalidate"]
E --> F["Next Request"]
F --> G["Fresh Retrieval"] π§ 47. Cache StampedeΒΆ
A cache stampede occurs when many requests simultaneously miss the same cache entry.
1000 Requests
β
βΌ
Cache Miss
β
βββ Retrieval
βββ Retrieval
βββ Retrieval
βββ Retrieval
βββ ...
This can overload:
π§ 48. Preventing Cache StampedeΒΆ
Use:
Request Coalescing
Single Flight
Distributed Lock
Jittered TTL
Probabilistic Refresh
Background Refresh
π§ 49. Request CoalescingΒΆ
Request A ββ
Request B ββ€
Request C ββΌβββ One Computation
Request D ββ€
Request E ββ
Other requests wait for the same result.
π§ 50. Single-Flight PatternΒΆ
Conceptually:
if cache.exists(key):
return cache.get(key)
if computation_in_progress(key):
return await existing_computation(key)
create_computation(key)
result = compute()
cache.set(key, result)
return result
π§ 51. Distributed LockΒΆ
For distributed applications:
Other requests:
Use carefully to avoid deadlocks and excessive waiting.
π§ 52. TTL JitterΒΆ
If many entries expire simultaneously:
This can create a load spike.
Instead:
π§ 53. Cache AvalancheΒΆ
A cache avalanche occurs when many entries expire or become invalid at once.
Example:
Mitigation:
π§ 54. Cache PenetrationΒΆ
Cache penetration occurs when requests repeatedly ask for data that does not exist.
Repeated malicious or invalid queries can overload the backend.
π§ 55. Cache Penetration MitigationΒΆ
Use:
Example:
π§ 56. Negative CachingΒΆ
Example:
If retrieval repeatedly returns no evidence:
for a short TTL.
Do not use a long TTL because knowledge may later appear.
π§ 57. Semantic Cache False PositiveΒΆ
Suppose:
A naive semantic cache may consider them similar.
Result:
Therefore semantic caches should incorporate:
π§ 58. Semantic Cache GuardrailsΒΆ
A semantic cache hit should satisfy:
Semantic Similarity
+
Same Tenant
+
Compatible Authorization
+
Compatible Filters
+
Compatible Time Scope
+
Compatible Knowledge Version
π§ 59. Query NormalizationΒΆ
Before exact caching, normalize queries.
Examples:
Depending on application semantics, normalization may include:
Do not normalize away meaningful information.
π§ 60. Query FingerprintingΒΆ
Create a stable representation:
import hashlib
def fingerprint(query: str) -> str:
normalized = query.strip().lower()
return hashlib.sha256(
normalized.encode("utf-8")
).hexdigest()
Production implementations should normalize according to domain semantics.
π§ 61. Cache Key DesignΒΆ
A general cache key:
For example:
query
+
tenant
+
authorization_scope
+
retriever_version
+
index_version
+
prompt_version
+
model_version
π§ 62. Cache Key HierarchyΒΆ
This makes operational inspection easier.
π§ 63. Cache NamespacesΒΆ
Possible namespaces:
Example:
π§ 64. Distributed CacheΒΆ
A distributed cache such as Redis can provide:
Typical architecture:
π§ 65. Local vs Distributed CacheΒΆ
Local CacheΒΆ
Advantages:
Limitations:
Distributed CacheΒΆ
Advantages:
Trade-off:
π§ 66. Two-Level CacheΒΆ
A powerful architecture:
Request
β
L1 Local Cache
β
βββ Hit β Return
β
βββ Miss
β
L2 Distributed Cache
β
βββ Hit β Populate L1
β
βββ Miss β Compute
π§ 67. Two-Level CacheΒΆ
flowchart LR
A["Request"] --> B["L1 Local Cache"]
B -->|Hit| C["Response"]
B -->|Miss| D["L2 Distributed Cache"]
D -->|Hit| E["Populate L1"]
E --> C
D -->|Miss| F["RAG Pipeline"]
F --> G["Populate L2"]
G --> H["Populate L1"]
H --> C π§ 68. Cache SerializationΒΆ
Cache entries may contain:
Choose based on:
π§ 69. What Should Be Cached?ΒΆ
Good candidates:
Embeddings
Stable Retrieval Results
Reranking Results
Stable Evidence Packages
Repeated FAQ Responses
Poor candidates:
π§ 70. Cache CompressionΒΆ
Large cached evidence can consume significant memory.
Use compression when:
Trade-off:
π§ 71. Cache WarmingΒΆ
Pre-populate frequently requested entries.
Useful for:
π§ 72. Cache Warming PipelineΒΆ
flowchart LR
A["Golden Queries"] --> B["Warmup Job"]
B --> C["RAG Pipeline"]
C --> D["Cache"]
D --> E["Production"] π§ 73. Cache RefreshΒΆ
Instead of waiting for expiration:
Users continue receiving the previous valid value while the new result is computed.
π§ 74. Stale-While-RevalidateΒΆ
Conceptually:
Useful when:
Avoid for strict real-time or highly regulated data where stale information is unacceptable.
π§ 75. Cache Consistency ModelsΒΆ
Possible models:
RAG often uses:
for knowledge indexes and caches.
But some security-related state may require stronger guarantees.
π§ 76. Security State Should Not Be StaleΒΆ
Be particularly careful with:
A stale authorization cache can become a security vulnerability.
π§ 77. Authorization CacheΒΆ
If authorization decisions are cached:
should be considered in the key.
Also define:
for sensitive environments.
π§ 78. Cache and Document UpdatesΒΆ
Suppose:
Then:
is published.
Potential stale path:
Therefore:
π§ 79. Cache and Prompt UpdatesΒΆ
If:
produces:
then:
should not necessarily reuse the old response.
Use:
in the response cache key.
π§ 80. Cache and Model UpdatesΒΆ
Similarly:
and:
may generate different outputs.
Therefore include:
where response correctness depends on it.
π§ 81. Cache and Retriever UpdatesΒΆ
Changing:
can change:
Therefore retrieval caches should include:
π§ 82. Cache and Context StrategyΒΆ
Changing:
can change final evidence.
Therefore context cache keys should include:
π§ 83. Cache Dependency GraphΒΆ
flowchart TD
A["Document"] --> B["Index"]
B --> C["Retrieval"]
D["Retriever Version"] --> C
C --> E["Reranking"]
E --> F["Context"]
G["Prompt Version"] --> H["Response"]
F --> H
I["Model Version"] --> H A cache should be invalidated when one of its dependencies changes.
π§ 84. Dependency-Aware CacheΒΆ
Think of a cached response as:
Response
β
βββ Query
βββ Tenant
βββ Authorization
βββ Retrieval
βββ Index
βββ Context
βββ Prompt
βββ Model
Changing any critical dependency may invalidate the result.
π§ 85. Cache Dependency FingerprintΒΆ
A practical pattern:
dependency_fingerprint =
hash(
index_version
+
retriever_version
+
prompt_version
+
model_version
+
policy_version
)
Use the fingerprint as part of the cache key.
π§ 86. Cache Hit RateΒΆ
Basic metric:
Example:
π§ 87. Cache Miss RateΒΆ
Example:
π§ 88. Cache EffectivenessΒΆ
Hit rate alone is not enough.
Consider:
A cache with:
may still be poor if the cached operation is cheap.
π§ 89. Cache MetricsΒΆ
Monitor:
Hit Rate
Miss Rate
Eviction Rate
Entry Count
Memory Usage
Latency
Refresh Rate
Invalidation Rate
Error Rate
Stampede Events
π§ 90. RAG-Specific Cache MetricsΒΆ
Track:
Embedding Cache Hit Rate
Retrieval Cache Hit Rate
Reranking Cache Hit Rate
Semantic Cache Hit Rate
Response Cache Hit Rate
Also:
π§ 91. Cost SavingsΒΆ
Approximate:
Track actual savings rather than assuming every cache hit has the same value.
π§ 92. Cache LatencyΒΆ
Track:
The cache itself must not become a bottleneck.
π§ 93. Cache Capacity PlanningΒΆ
Estimate:
plus overhead.
π§ 94. ExampleΒΆ
Suppose:
Raw payload:
Actual memory requirement is higher due to:
π§ 95. Cache EvictionΒΆ
Common policies:
LRUΒΆ
Good for workloads where recent queries are more likely to repeat.
LFUΒΆ
Useful when popular queries should remain cached.
π§ 96. RAG Cache Eviction StrategyΒΆ
A combination can be useful:
For example:
π§ 97. Cache AdmissionΒΆ
Not every result deserves caching.
Example:
Potential admission signals:
π§ 98. Cost-Aware Cache AdmissionΒΆ
Cache expensive operations first.
Example:
Cheap Retrieval
β Low Priority
Expensive Reranking
β High Priority
Expensive LLM Response
β High Priority
π§ 99. Query FrequencyΒΆ
A simple strategy:
First Request
β
Compute
Second Request
β
Compute
Third Request
β
Cache
Repeated Requests
β
Cache
This avoids filling the cache with one-time queries.
π§ 100. Cache PollutionΒΆ
Cache pollution occurs when low-value entries consume memory.
Examples:
Mitigate with:
π§ 101. Bot TrafficΒΆ
Bots can generate:
which can cause:
Use:
π§ 102. Cache SecurityΒΆ
Protect cached data with:
π§ 103. Sensitive Cache DataΒΆ
Be careful caching:
Possible policies:
π§ 104. Cache EncryptionΒΆ
Consider:
π§ 105. Cache and ComplianceΒΆ
Compliance requirements may influence:
A cache is still a data store.
π§ 106. Cache DeletionΒΆ
When a user or document must be deleted:
Deletion workflows should account for derived cached data where required.
π§ 107. Cache Invalidation on DeletionΒΆ
flowchart TD
A["Document Deleted"] --> B["Deletion Event"]
B --> C["Delete From Index"]
B --> D["Invalidate Retrieval Cache"]
B --> E["Invalidate Context Cache"]
B --> F["Invalidate Response Cache"] π§ 108. Cache ObservabilityΒΆ
Every cache operation should ideally emit:
Avoid logging sensitive key contents.
π§ 109. Example Cache LogΒΆ
{
"cache": "retrieval",
"result": "hit",
"tenant": "tenant-a",
"latency_ms": 3,
"index_version": "v17"
}
π§ 110. Distributed Cache FailureΒΆ
What happens if Redis fails?
Do not assume:
Prefer:
when backend capacity allows.
π§ 111. Cache as an OptimizationΒΆ
A critical principle:
The cache should usually accelerate the system, not become the system's only source of truth.
Architecture:
π§ 112. Cache Failure StrategyΒΆ
flowchart TD
A["Request"] --> B["Cache"]
B -->|Available| C{"Hit?"}
C -->|Yes| D["Return"]
C -->|No| E["RAG Pipeline"]
B -->|Unavailable| E
E --> F["Response"] π§ 113. Circuit Breaker for CacheΒΆ
If cache infrastructure becomes unhealthy:
This prevents cache failure from increasing application latency.
π§ 114. Cache Warmup After RestartΒΆ
After a cache restart:
Mitigate with:
π§ 115. Cache Warmup PrioritiesΒΆ
Warm:
rather than everything.
π§ 116. Cache PrecomputationΒΆ
For known workloads:
Useful for:
π§ 117. Cache and StreamingΒΆ
Response caching can be more complicated when responses stream.
Possible approach:
Do not cache incomplete or failed responses.
π§ 118. Cache Only Validated ResponsesΒΆ
Prefer:
rather than:
Otherwise invalid output can be reused.
π§ 119. Cache PoisoningΒΆ
A cache poisoning scenario occurs when incorrect or malicious output becomes cached.
Potential causes:
Mitigation:
π§ 120. Semantic Cache PoisoningΒΆ
Semantic caches require extra caution.
A bad answer for:
could be incorrectly reused for:
Therefore semantic cache entries should carry:
π§ 121. Cache ProvenanceΒΆ
A response cache entry can store:
{
"response": "...",
"document_ids": [
"doc-123",
"doc-456"
],
"index_version": "v17",
"retriever_version": "v8",
"prompt_version": "v9",
"model_version": "v4",
"validated": true
}
This enables stronger invalidation and auditing.
π§ 122. Cache Dependency GraphΒΆ
The further downstream a cache is placed, the more dependencies it typically has.
π§ 123. Cache ComplexityΒΆ
Conceptually:
Embedding Cache
β
Few Dependencies
Retrieval Cache
β
More Dependencies
Context Cache
β
More Dependencies
Response Cache
β
Many Dependencies
Therefore:
Downstream caches generally require stronger invalidation and versioning strategies.
π§ 124. Cache Architecture RecommendationΒΆ
A mature production RAG system may use:
L1:
Local Cache
L2:
Distributed Cache
Pipeline:
Embedding Cache
Retrieval Cache
Reranking Cache
Application:
Semantic Cache
Optional:
Response Cache
Do not automatically enable every layer.
π§ 125. Recommended Cache SelectionΒΆ
Low TrafficΒΆ
Medium TrafficΒΆ
High TrafficΒΆ
FAQ WorkloadΒΆ
Highly Dynamic WorkloadΒΆ
π§ 126. Cache ArchitectureΒΆ
flowchart TD
A["User"] --> B["L1 Cache"]
B -->|Hit| C["Response"]
B -->|Miss| D["L2 Distributed Cache"]
D -->|Hit| E["Response"]
D -->|Miss| F["RAG Orchestrator"]
F --> G["Embedding Cache"]
G --> H["Retrieval Cache"]
H --> I["Reranking Cache"]
I --> J["Context Engine"]
J --> K["Semantic Cache"]
K --> L["LLM"]
L --> M["Validation"]
M --> N["Citation"]
N --> O["Response"]
O --> D
O --> B π§ 127. Cache Strategy by Pipeline StageΒΆ
| Stage | Cache Candidate | Main Concern |
|---|---|---|
| Document Parsing | Yes | Source version |
| Embedding | Yes | Model version |
| Retrieval | Yes | Index version |
| Reranking | Yes | Candidate/version changes |
| Context | Yes | Evidence freshness |
| Semantic | Yes | False positives |
| Response | Yes | Security/freshness |
π§ 128. Cache Decision TreeΒΆ
Is the operation expensive?
β
βββ No β Probably don't cache
β
βββ Yes
β
βΌ
Is the result reusable?
β
βββ No β Don't cache
β
βββ Yes
β
βΌ
Can stale results be tolerated?
β
ββββββ΄βββββ
βΌ βΌ
Yes No
β β
βΌ βΌ
Cache Short TTL /
Versioning /
Invalidation
π§ 129. Cache Strategy by RiskΒΆ
LOW RISK
β
Aggressive Caching
MEDIUM RISK
β
Version + TTL
HIGH RISK
β
Strict Invalidation
REAL-TIME / SECURITY CRITICAL
β
Bypass or Minimal Cache
π§ 130. Cache TestingΒΆ
Caching must be tested independently.
Test:
Hit
Miss
Expiration
Invalidation
Concurrent Requests
Cache Failure
Cache Restart
Version Change
Tenant Isolation
Authorization Change
Document Update
π§ͺ 131. Cache Unit TestsΒΆ
Test:
π§ͺ 132. Cache Integration TestsΒΆ
Verify:
Test:
π§ͺ 133. Cache Security TestsΒΆ
Test:
Tenant A β Tenant A Cache β
Tenant A β Tenant B Cache β
Authorized User β Response β
Unauthorized User β Response β
π§ͺ 134. Cache Stampede TestΒΆ
Simulate:
for the same missing key.
Expected:
rather than:
π§ͺ 135. Cache Invalidation TestΒΆ
Scenario:
Expected:
π§ͺ 136. Cache Failure TestΒΆ
Simulate:
Expected:
provided the backend can safely absorb the load.
π§ͺ 137. Cache Performance TestΒΆ
Measure:
π§ͺ 138. Cache Load TestΒΆ
Test:
π§ 139. Cache Monitoring DashboardΒΆ
A production dashboard should show:
Cache Hit Rate
Cache Miss Rate
Cache Latency
Eviction Rate
Memory Usage
Entry Count
Invalidation Rate
Stampede Events
Backend Load
LLM Calls Avoided
Cost Saved
π§ 140. Cache Cost ModelΒΆ
A distributed cache has its own cost:
Therefore:
should generally be the goal.
π§ 141. Cache ROIΒΆ
A simple conceptual model:
More sophisticated analysis should include:
π§ 142. Cache Anti-PatternsΒΆ
Anti-Pattern 1 β Global Response CacheΒΆ
without authorization-aware keys.
Anti-Pattern 2 β Cache Without VersioningΒΆ
Anti-Pattern 3 β Infinite TTLΒΆ
This creates stale knowledge.
Anti-Pattern 4 β Cache EverythingΒΆ
This causes:
Anti-Pattern 5 β No Stampede ProtectionΒΆ
Anti-Pattern 6 β Cache Before ValidationΒΆ
Invalid answers may become reusable.
Anti-Pattern 7 β Ignore DeletionΒΆ
Anti-Pattern 8 β Treat Cache as Source of TruthΒΆ
A cache should generally be a derived optimization.
π§ 143. Production Cache ChecklistΒΆ
β Cache layers identified
β Cache ownership defined
β Cache keys versioned
β Tenant isolation implemented
β Authorization scope considered
β TTL defined
β Invalidation strategy defined
β Document update invalidation handled
β Index version handled
β Embedding version handled
β Prompt version handled
β Model version handled
β Cache stampede protection
β Cache avalanche protection
β Cache penetration protection
β Negative caching considered
β Cache warming considered
β Cache failure fallback
β Cache encryption
β Cache observability
β Cache capacity planning
β Cache load testing
β Cache security testing
β Cache cost tracking
π§ 144. Recommended Production PatternΒΆ
A strong default architecture is:
REQUEST
β
βΌ
L1 Cache
β
ββββββ΄βββββ
βΌ βΌ
Hit Miss
β β
β βΌ
β L2 Distributed
β Cache
β β
β ββββββ΄βββββ
β βΌ βΌ
β Hit Miss
β β β
β β βΌ
β β RAG Pipeline
β β β
β β ββββββΌβββββ
β β βΌ βΌ βΌ
β β Embed Retrieve Rerank
β β Cache Cache Cache
β β β β β
β β ββββββΌβββββββ
β β βΌ
β β Context
β β β
β β βΌ
β β Semantic
β β Cache
β β β
β β ββββ΄βββ
β β βΌ βΌ
β β Hit Miss
β β β β
β β β LLM
β β β β
β β β Validate
β β β β
β β β Citation
β β β β
ββββββ΄βββββββ΄ββββββ
β
βΌ
RESPONSE
π§ 145. Final Mental ModelΒΆ
RAG caching should be thought of as:
RAG CACHING
β
βββββββββββββββββΌβββββββββββββββββ
βΌ βΌ βΌ
SPEED COST SCALABILITY
β β β
βββββββββββββββββΌβββββββββββββββββ
βΌ
CORRECTNESS
β
ββββββββββββββββΌβββββββββββββββ
βΌ βΌ βΌ
Freshness Security Versioning
β β β
ββββββββββββββββΌβββββββββββββββ
βΌ
INVALIDATION
β
βΌ
OBSERVABILITY
π§ 146. Cache Strategy FormulaΒΆ
A useful architectural mental model:
A cache that is fast but returns unauthorized or stale information is not a successful production cache.
π§ 147. Final Key TakeawaysΒΆ
- Caching can significantly reduce RAG latency and cost.
- RAG should generally use multiple cache layers selectively.
- Embedding caching avoids repeated embedding computation.
- Retrieval caching avoids repeated search operations.
- Reranking caching avoids repeated expensive ranking.
- Context caching can avoid repeated evidence assembly.
- Semantic caching enables reuse across similar queries.
- Response caching provides the largest potential savings but also carries the highest correctness and security risk.
- Exact caching is safer than semantic caching because the reuse condition is explicit.
- Semantic caching requires similarity thresholds and strong contextual guardrails.
- Cache keys must include every important dependency that can change the result.
- Tenant identity should generally be included in security-sensitive cache keys.
- Authorization scope must be considered when caching protected responses.
- Index version should be included in retrieval-related cache keys.
- Embedding version should be included in embedding-related cache keys.
- Retriever version should be included in retrieval cache keys.
- Prompt version and model version should be considered for response caches.
- TTL alone is rarely sufficient for enterprise RAG.
- Versioned cache namespaces provide a powerful invalidation mechanism.
- Event-driven invalidation is useful for knowledge-driven systems.
- Document-level invalidation can reduce unnecessary cache eviction.
- Cache stampedes can overload downstream RAG components.
- Single-flight, request coalescing, locks, jitter, and background refresh can mitigate stampedes.
- Cache avalanche can occur when many entries expire simultaneously.
- Cache penetration can occur when invalid or nonexistent queries repeatedly bypass the cache.
- Negative caching can reduce repeated no-result queries.
- Cache admission policies prevent cache pollution.
- Cache warming can reduce cold-start load.
- Stale-while-revalidate can improve latency when controlled staleness is acceptable.
- Cache failure should ideally degrade the system rather than bring down RAG.
- A cache should generally be an optimization layer, not the authoritative source of truth.
- Cached responses should preferably be validated before they become reusable.
- Cache entries can carry provenance and dependency metadata.
- Cache invalidation must account for document, index, model, prompt, retriever, and authorization changes.
- Two-level caches can combine local speed with distributed consistency.
- Cache eviction policies such as LRU and LFU help manage finite memory.
- Cache observability should include hits, misses, latency, evictions, invalidations, memory, and backend load.
- Measure LLM calls and tokens avoided to quantify RAG cache value.
- Cache cost must be compared against the cost saved.
- Sensitive information may require restricted or disabled caching.
- Cache deletion must be included in data deletion workflows.
- Cache testing should include concurrency, failure, invalidation, security, and stampede scenarios.
- The best cache architecture is not the one with the most cache layers.
- The best architecture is the one that maximizes safe reuse while preserving correctness, freshness, security, and operational simplicity.
π§ 148. Chapter NavigationΒΆ
Part VI β Production RAG Deployment & OperationsΒΆ
Previous:
12. RAG Deployment Patterns
Next:
14. Multi-Tenant RAG
Production RAG Engineering PathΒΆ
01 Prompt Assembly
β
02 Context Selection & Context Engineering
β
03 Response Validation
β
04 Citation & Source Attribution
β
05 Enterprise Response
β
06 RAG Evaluation & Benchmarking
β
07 RAG Observability
β
08 RAG Performance Optimization
β
09 RAG Cost Optimization
β
10 Production Retrieval Architecture
β
11 Building Production RAG Systems
β
12 RAG Deployment Patterns
β
13 RAG Caching Strategies
β
14 Multi-Tenant RAG
β
15 RAG Testing Frameworks
β
16 RAG Failure Patterns
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.