09. RAG Cost OptimizationΒΆ
Category: Production RAG Engineering
Module: Part V β Advanced Retrieval-Augmented Generation
Difficulty: Advanced
π OverviewΒΆ
A production RAG system must be optimized not only for accuracy and latency, but also for economic efficiency.
A system that produces excellent answers but costs:
may not be viable when serving:
Similarly, a system that is inexpensive but produces poor answers creates operational and business risk.
Production RAG cost optimization therefore focuses on the complete cost chain:
User Query
β
Query Processing
β
Embedding
β
Retrieval
β
Reranking
β
Context Processing
β
Prompt Assembly
β
LLM Generation
β
Validation
β
Citation
β
Observability
The objective is:
A useful production principle is:
Do not minimize cost blindly. Minimize the cost of achieving the required quality, latency, and reliability.
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand RAG cost architecture
- Identify major RAG cost drivers
- Calculate cost per request
- Calculate cost per user
- Calculate cost per tenant
- Calculate cost per workflow
- Understand LLM token economics
- Optimize input token usage
- Optimize output token usage
- Optimize retrieval costs
- Optimize embedding costs
- Optimize reranking costs
- Optimize validation costs
- Optimize agentic RAG costs
- Optimize Graph RAG costs
- Optimize SQL RAG costs
- Optimize multimodal RAG costs
- Implement caching strategies
- Implement model routing
- Implement model cascading
- Implement adaptive retrieval
- Reduce unnecessary LLM calls
- Optimize context size
- Optimize prompt size
- Optimize infrastructure costs
- Optimize vector database costs
- Optimize observability costs
- Implement cost budgets
- Implement cost guardrails
- Implement tenant-level cost controls
- Build cost dashboards
- Detect cost anomalies
- Perform cost attribution
- Perform cost forecasting
- Design cost-aware RAG architectures
π§ 1. What Is RAG Cost Optimization?ΒΆ
RAG cost optimization is the process of reducing the resources required to serve RAG requests while preserving acceptable:
A simplified objective is:
π§ 2. RAG Cost Is More Than LLM CostΒΆ
A common mistake is:
In reality:
Total RAG Cost
β
βββ LLM
βββ Embeddings
βββ Reranking
βββ Vector Database
βββ Search Infrastructure
βββ Compute
βββ Storage
βββ Network
βββ Observability
βββ Evaluation
βββ Background Processing
π§ 3. Cost ArchitectureΒΆ
flowchart TD
A["RAG Request"] --> B["Query Processing"]
B --> C["Embedding"]
C --> D["Retrieval"]
D --> E["Reranking"]
E --> F["Context Processing"]
F --> G["LLM"]
G --> H["Validation"]
H --> I["Citation"]
C --> J["Embedding Cost"]
D --> K["Vector/Search Cost"]
E --> L["Reranker Cost"]
G --> M["LLM Cost"]
H --> N["Validation Cost"]
I --> O["Processing Cost"]
P["Infrastructure"] --> Q["Compute"]
P --> R["Storage"]
P --> S["Network"]
P --> T["Observability"] π§ 4. Cost CategoriesΒΆ
A practical classification:
Variable Cost
β
Cost changes with requests/tokens
Fixed Cost
β
Infrastructure that exists regardless of request volume
Semi-Variable Cost
β
Resources that scale with workload
Examples:
VariableΒΆ
FixedΒΆ
Semi-VariableΒΆ
π§ 5. Cost Per RequestΒΆ
A simplified model:
C_request =
C_embedding
+ C_retrieval
+ C_reranking
+ C_context
+ C_generation
+ C_validation
+ C_observability
+ C_infrastructure
π§ 6. LLM CostΒΆ
LLM cost is commonly driven by:
Conceptually:
Actual pricing varies by provider, model, region, and pricing program.
π§ 7. RAG Token CompositionΒΆ
A request may contain:
Therefore:
π§ 8. Context Is Often the Largest Optimization OpportunityΒΆ
Example:
System Prompt 1,000 tokens
User Query 100 tokens
Conversation 900 tokens
Retrieved Context 8,000 tokens
βββββββββββββββββββββββββββββββ
Input 10,000 tokens
If the context is reduced:
the input token cost can fall significantly.
π§ 9. Context CostΒΆ
The goal is not:
but:
π§ 10. Context EfficiencyΒΆ
A useful conceptual metric:
Context Efficiency
=
Useful Evidence
ββββββββββββββββββ
Context Tokens
Higher is generally better.
This is not a universal standardized metric; use it as an engineering diagnostic.
π§ 11. Context WasteΒΆ
Example:
The system is paying for:
while deriving useful information from approximately:
π§ 12. Context OptimizationΒΆ
Use:
Top-K Tuning
+
Reranking
+
MMR
+
Context Compression
+
Deduplication
+
Metadata Filtering
+
Token Budgeting
π§ 13. Token BudgetΒΆ
Define:
Example:
Then:
π§ 14. Token Budget ArchitectureΒΆ
flowchart LR
A["Model Context Window"] --> B["System Instructions"]
A --> C["User Query"]
A --> D["Conversation"]
A --> E["Retrieved Context"]
A --> F["Output Budget"]
E --> G["Context Budget"]
G --> H["Relevant Evidence"] π§ 15. Dynamic Context BudgetΒΆ
Not every query needs the same amount of context.
Simple FAQ
β
2,000 tokens
Technical Query
β
4,000 tokens
Complex Multi-Hop Query
β
8,000 tokens
Use adaptive budgets where appropriate.
π§ 16. Cost of Top-KΒΆ
Increasing K can increase:
Example:
The quality improvement may not justify the additional cost.
π§ 17. Cost-Aware Top-KΒΆ
Instead of:
use:
π§ 18. Adaptive RetrievalΒΆ
Query
β
Retrieve Small Candidate Set
β
Confidence Check
β
βββ High Confidence β Generate
β
βββ Low Confidence β Expand Retrieval
This reduces expensive work for easy queries.
π§ 19. Early ExitΒΆ
Example:
Initial Retrieval
β
Top Result Score = 0.94
β
Evidence Sufficient?
β
βββ Yes β Generate
βββ No β Expand
Use calibrated signals rather than arbitrary score thresholds.
π§ 20. Query Rewriting CostΒΆ
Query rewriting may require an LLM.
If rewriting costs:
and is executed:
the rewriting layer alone contributes:
before the main generation cost.
π§ 21. Conditional Query RewritingΒΆ
Query
β
Complexity Detector
β
βββ Simple β Direct Retrieval
β
βββ Complex β Query Rewrite
This can significantly reduce unnecessary model calls.
π§ 22. Multi-Query CostΒΆ
Multi-query:
may multiply:
Use it when the quality improvement justifies the additional cost.
π§ 23. Multi-Query Cost OptimizationΒΆ
Query
β
Determine Need
β
βββ Low Ambiguity β Single Query
β
βββ High Ambiguity β Multi-Query
π§ 24. Embedding CostΒΆ
Embedding cost can come from:
Query embeddings happen frequently.
Document embeddings can become expensive during large ingestion operations.
π§ 25. Embedding Cost OptimizationΒΆ
Use:
where quality requirements permit.
π§ 26. Avoid Re-Embedding Unchanged DocumentsΒΆ
Use content hashes:
Pipeline:
Document
β
Hash
β
Compare Previous Hash
β
βββ Same β Skip
βββ Changed β Re-Embed
π§ 27. Incremental IndexingΒΆ
Bad:
Better:
π§ 28. Batch EmbeddingΒΆ
Document 1 β
Document 2 β
Document 3 βββ Batch β Embedding Model
Document 4 β
Document 5 β
Batching can reduce per-request overhead and improve accelerator utilization.
π§ 29. Embedding Model EconomicsΒΆ
Compare:
Model A
Quality = 91%
Cost = Low
Model B
Quality = 94%
Cost = Medium
Model C
Quality = 96%
Cost = High
Choose based on:
π§ 30. Local vs Hosted EmbeddingsΒΆ
HostedΒΆ
Costs may include:
LocalΒΆ
Costs shift toward:
Neither is universally cheaper.
π§ 31. Reranking CostΒΆ
Rerankers can be expensive because they may evaluate many query-document pairs.
Conceptually:
π§ 32. Candidate ReductionΒΆ
Instead of:
use:
π§ 33. Reranker SelectionΒΆ
Possible options:
Use the least expensive approach that meets the quality requirement.
π§ 34. Selective RerankingΒΆ
Query
β
Initial Retrieval
β
Confidence
β
βββ High β Skip Reranker
β
βββ Low β Rerank
This can reduce average cost.
π§ 35. LLM CostΒΆ
LLM cost often dominates production RAG economics.
RAG Request
β
βΌ
ββββββββββββββββββββββββββββ
β LLM COST β
ββββββββββββββββββββββββββββ€
β Input Tokens β
β Output Tokens β
β Number of Calls β
β Model Selection β
β Retries β
β Validation Calls β
β Agent Tool Calls β
ββββββββββββββββββββββββββββ
π§ 36. Reduce Number of LLM CallsΒΆ
One of the strongest cost optimizations:
Example:
Query
β
Classifier
β
βββ FAQ β Cached Answer
βββ Search β Retrieval + Small LLM
βββ Complex β Large LLM
βββ SQL β SQL Pipeline
π§ 37. Model RoutingΒΆ
flowchart TD
A["Query"] --> B["Model Router"]
B --> C["Small Model"]
B --> D["Medium Model"]
B --> E["Large Model"]
C --> F["Response"]
D --> F
E --> F The router should consider:
π§ 38. Model CascadeΒΆ
Average cost can decrease if most queries are successfully handled by the smaller model.
π§ 39. Model Routing ExampleΒΆ
π§ 40. Model Routing EconomicsΒΆ
Suppose:
instead of:
the average cost can be substantially lower, assuming quality remains acceptable.
π§ 41. Output Token OptimizationΒΆ
Output tokens directly affect:
Use:
π§ 42. Output BudgetΒΆ
Do not allocate the maximum output budget to every request.
π§ 43. Response ContractΒΆ
Example:
This reduces unnecessary output.
π§ 44. Prompt OptimizationΒΆ
Prompt cost includes:
Reduce:
π§ 45. Prompt VersioningΒΆ
Use:
Measure:
A shorter prompt is not automatically better if quality drops.
π§ 46. Prompt CachingΒΆ
If supported by the model/provider:
Potentially reduces:
π§ 47. Conversation CostΒΆ
Conversational RAG can become expensive because history grows.
Turn 1 β 500 tokens
Turn 2 β 1,200 tokens
Turn 3 β 2,500 tokens
Turn 4 β 4,500 tokens
Turn 5 β 7,000 tokens
π§ 48. Conversation CompressionΒΆ
Instead of passing all history:
π§ 49. Memory BudgetΒΆ
Set:
Use:
rather than blindly sending the complete conversation.
π§ 50. Semantic Conversation SelectionΒΆ
Retrieve only relevant historical messages:
This reduces token usage.
π§ 51. Retrieval CacheΒΆ
Cache retrieval results for repeated requests.
Key should account for:
π§ 52. Cache EconomicsΒΆ
If:
and:
then:
may avoid some downstream computation.
Actual savings depend on what the cache bypasses.
π§ 53. Semantic CacheΒΆ
Example:
Potentially reuse the result.
But semantic caching must validate:
π§ 54. Cache InvalidationΒΆ
Invalidate when:
Document Changes
Index Changes
Embedding Changes
Prompt Changes
Retriever Changes
Authorization Changes
π§ 55. Cache VersioningΒΆ
Example:
This prevents stale results from older pipeline versions.
π§ 56. Tenant Cost AttributionΒΆ
Enterprise RAG should answer:
How much does Tenant A cost?
How much does Tenant B cost?
Which tenant consumes the most tokens?
Which tenant generates the most expensive requests?
π§ 57. Cost Per TenantΒΆ
Example:
This enables:
π§ 58. Cost Per UserΒΆ
Track:
Avoid exposing user-level data broadly; apply appropriate privacy and access controls.
π§ 59. Cost by ApplicationΒΆ
Enterprise platforms may host:
Track cost separately.
π§ 60. Cost by WorkflowΒΆ
A single application may have:
Each can have a different cost profile.
π§ 61. Cost by ModelΒΆ
Track:
with:
π§ 62. Cost by ProviderΒΆ
For multi-cloud enterprise systems:
compare:
π§ 63. Reranking Cost by QueryΒΆ
Some queries may require:
others:
and complex queries:
Use query-aware policies.
π§ 64. Agentic RAG CostΒΆ
Agentic RAG can multiply cost.
Potentially:
π§ 65. Agentic RAG Cost GuardrailsΒΆ
Define:
Example:
π§ 66. Agentic Early TerminationΒΆ
Avoid unnecessary planning loops.
π§ 67. Agentic Loop DetectionΒΆ
Potential problem:
Use:
π§ 68. Graph RAG CostΒΆ
Graph RAG may require:
Cost can increase with traversal depth.
π§ 69. Graph Traversal BudgetΒΆ
Instead of:
use:
π§ 70. SQL RAG CostΒΆ
SQL RAG can generate expensive queries.
Potential risks:
Use:
π§ 71. SQL Result BudgetΒΆ
Instead of:
use:
Only return data required for reasoning.
π§ 72. Multimodal RAG CostΒΆ
Multimodal workloads may involve:
Optimize using:
π§ 73. Image Processing CostΒΆ
Avoid processing every image at maximum resolution.
π§ 74. Validation CostΒΆ
If every answer invokes:
the cost can multiply.
Possible alternatives:
π§ 75. Risk-Based ValidationΒΆ
flowchart TD
A["Generated Response"] --> B["Risk Classifier"]
B --> C["Low Risk"]
B --> D["Medium Risk"]
B --> E["High Risk"]
C --> F["Rule Validation"]
D --> G["Lightweight Validation"]
E --> H["Deep Validation"]
F --> I["Final Response"]
G --> I
H --> I π§ 76. Cost-Aware CitationΒΆ
Citations should ideally be produced using source metadata already carried through the pipeline.
Avoid unnecessary additional LLM calls just to reconstruct sources.
π§ 77. Infrastructure CostΒΆ
Infrastructure includes:
RAG Compute
Vector DB
Search Engine
Object Storage
Databases
GPU
Load Balancers
Network
Observability
π§ 78. Compute OptimizationΒΆ
Optimize:
π§ 79. CPU vs GPUΒΆ
Use GPU when:
CPU may be more appropriate for:
π§ 80. GPU UtilizationΒΆ
Poor:
while paying for a large GPU.
Potential solutions:
π§ 81. AutoscalingΒΆ
Scale based on:
π§ 82. Scale-to-ZeroΒΆ
For workloads that are:
scale-to-zero or scheduled compute can reduce infrastructure cost where supported.
π§ 83. Vector Database CostΒΆ
Vector DB cost depends on:
π§ 84. Vector Storage OptimizationΒΆ
Reduce:
where quality and operational requirements permit.
π§ 85. Data LifecycleΒΆ
Implement:
Not every document needs identical storage characteristics.
π§ 86. Document RetentionΒΆ
Enterprise knowledge bases may contain:
Remove expired data from the active retrieval path when policy allows.
π§ 87. Storage TieringΒΆ
Retrieve cold data only when required.
π§ 88. Network CostΒΆ
Network costs may come from:
Reduce unnecessary payloads and cross-region communication.
π§ 89. Region OptimizationΒΆ
Choose infrastructure locations based on:
Do not optimize cost by violating data residency or regulatory requirements.
π§ 90. Observability CostΒΆ
Observability can become expensive when capturing:
Use:
π§ 91. Evaluation CostΒΆ
RAG evaluation can itself consume LLM calls.
Large evaluation suites can become expensive.
π§ 92. Evaluation SamplingΒΆ
Instead of evaluating:
consider:
The right sampling policy depends on the application's risk.
π§ 93. Continuous Evaluation CostΒΆ
Use:
rather than evaluating every request with expensive models.
π§ 94. Cost-Aware EvaluationΒΆ
Prioritize:
π§ 95. Cost Anomaly DetectionΒΆ
Monitor:
π§ 96. Cost Spike ExampleΒΆ
Investigate:
π§ 97. Cost DashboardΒΆ
ββββββββββββββββββββββββββββββββββββββββββββββββ
β RAG COST DASHBOARD β
ββββββββββββββββββββββββββββββββββββββββββββββββ€
β Requests/day 120,000 β
β Avg Cost/request $0.018 β
β Daily Cost $2,160 β
β Monthly Forecast $64,800 β
β β
β LLM 69% β
β Embeddings 11% β
β Reranking 7% β
β Vector DB 6% β
β Infrastructure 4% β
β Observability 3% β
β β
β Cache Hit Rate 31% β
β Avg Context Tokens 3,900 β
β Avg Output Tokens 420 β
ββββββββββββββββββββββββββββββββββββββββββββββββ
Values are illustrative.
π§ 98. Cost BreakdownΒΆ
TOTAL COST
β
ββββββββββββββββββΌβββββββββββββββββ
βΌ βΌ βΌ
AI DATA PLATFORM
β β β
LLM Tokens Vector DB Compute
Embeddings Storage Network
Reranker Search Observability
Evaluation Object Storage
π§ 99. Cost per Request FormulaΒΆ
If:
then:
π§ 100. Monthly Cost ForecastΒΆ
A simple estimate:
Example:
This is a simple forecast and does not account for traffic growth or tiered pricing.
π§ 101. Cost ForecastingΒΆ
Forecast using:
Example:
But cost may not exactly double if:
π§ 102. Cost ElasticityΒΆ
Measure:
A highly elastic architecture may scale cost almost linearly.
An optimized architecture can sometimes achieve:
through:
π§ 103. Cost per Successful AnswerΒΆ
A more useful business metric than raw cost:
Cost per Successful Answer
=
Total Cost
ββββββββββββββββββββ
Successful Answers
This captures quality.
π§ 104. Cost per High-Quality AnswerΒΆ
Example:
Then:
π§ 105. Quality-Adjusted CostΒΆ
Conceptually:
This can help compare architectures.
However, quality scores must be defined consistently.
π§ 106. Cost vs QualityΒΆ
Quality
β²
β β
β β
β β
β β
β β
ββββββββββββββββββββββββββΊ Cost
There is often a point of diminishing returns.
π§ 107. Diminishing ReturnsΒΆ
Example:
The additional:
may not justify:
depending on the use case.
π§ 108. Cost Optimization FrontierΒΆ
Quality
β²
β
β β
β β
β β
β β
ββββββββββββββββββββΊ Cost
The goal is to operate near an efficient frontier rather than blindly choosing the cheapest or most accurate configuration.
π§ 109. Cost GuardrailsΒΆ
Production systems should have limits:
Max Cost / Request
Max Tokens / Request
Max Agent Iterations
Max Tool Calls
Max Retrieval Candidates
Max Context Tokens
π§ 110. Request BudgetΒΆ
Example:
Pipeline:
π§ 111. Budget-Aware RoutingΒΆ
If the request has already consumed:
a cheaper path may be selected.
π§ 112. Agent Cost BudgetΒΆ
Agent
β
βββ Step 1 β $0.01
βββ Step 2 β $0.02
βββ Step 3 β $0.02
βββ Step 4 β $0.03
β
βΌ
$0.08
β
βΌ
Budget = $0.10
Next expensive action may be blocked.
π§ 113. Tenant BudgetsΒΆ
Example:
Use:
according to business policy.
π§ 114. Soft vs Hard BudgetsΒΆ
Soft BudgetΒΆ
Hard BudgetΒΆ
π§ 115. Graceful DegradationΒΆ
When budget is constrained:
or:
or:
π§ 116. Cost-Aware ArchitectureΒΆ
flowchart TD
A["User Query"] --> B["Budget Manager"]
B --> C["Query Router"]
C --> D["Fast Path"]
C --> E["Standard RAG"]
C --> F["Advanced RAG"]
E --> G["Retrieval"]
F --> G
G --> H["Reranking"]
H --> I["Context Budget"]
I --> J["Model Router"]
J --> K["Small Model"]
J --> L["Large Model"]
K --> M["Validation"]
L --> M
M --> N["Response"]
B --> O["Cost Tracking"]
O --> P["Budget Enforcement"] π§ 117. Cost-Aware RetrievalΒΆ
A production retriever should consider:
not only:
π§ 118. Cost-Aware Query RoutingΒΆ
Use the expensive route only when required.
π§ 119. Cost-Aware ValidationΒΆ
π§ 120. Cost-Aware Agentic RAGΒΆ
Agent should understand:
Example:
π§ 121. Cost-Aware Graph RAGΒΆ
Use:
to prevent runaway graph exploration.
π§ 122. Cost-Aware SQL RAGΒΆ
Use:
where supported.
π§ 123. Cost-Aware Multimodal RAGΒΆ
Use:
π§ 124. Cost-Aware EvaluationΒΆ
Prioritize evaluation of:
π§ 125. Cost Optimization by LayerΒΆ
Layer Optimization
Query Routing / Rewrite selectively
Embedding Cache / Batch
Retrieval Top-K / ANN
Hybrid Parallel / Candidate reduction
Reranking Smaller candidate set
Context Compression / Deduplication
Prompt Shorter / Cache
LLM Model routing
Validation Risk-based
Citation Metadata propagation
Agent Step budgets
Graph Traversal limits
SQL Query limits
Multimodal Resolution / routing
Infrastructure Autoscaling
Observability Sampling
π§ 126. Cost Optimization PriorityΒΆ
A practical sequence:
1. Measure total cost
2. Identify largest cost component
3. Reduce unnecessary work
4. Reduce token volume
5. Reduce number of model calls
6. Add caching
7. Introduce model routing
8. Optimize retrieval
9. Optimize infrastructure
10. Add budget guardrails
11. Continuously benchmark
π§ 127. Cost Optimization ExampleΒΆ
Baseline:
Optimized:
Top-K = Adaptive
Reranker = Selective
Context = 4,000 tokens
LLM = Routed
Validation = Risk-Based
Cache = Enabled
π§ 128. Example Cost ComparisonΒΆ
| Configuration | Tokens | LLM Calls | Avg Cost | p95 Latency |
|---|---|---|---|---|
| Baseline | 8,500 | 3 | $0.052 | 2.8s |
| Context Optimized | 5,000 | 3 | $0.035 | 2.2s |
| Model Routing | 5,000 | 2 | $0.021 | 1.7s |
| Cached + Routed | 5,000 | 1.4 avg | $0.015 | 1.3s |
Values are illustrative.
π§ 129. Cost Optimization ExperimentΒΆ
Hypothesis:
Measure:
π§ 130. Cost Experiment MatrixΒΆ
| Experiment | Cost | Latency | Quality | Decision |
|---|---|---|---|---|
| Baseline | β | β | β | β |
| Context Reduction | β | β | β | β |
| Smaller Model | β | β | β | β |
| Caching | β | β | β | β |
| Selective Reranking | β | β | β | β |
| Model Routing | β | β | β | β |
| Adaptive Retrieval | β | β | β | β |
Populate using real benchmarks.
π§ͺ 131. Practical ProjectΒΆ
Build a:
Production RAG Cost Optimization Lab
Start with a baseline RAG system and progressively optimize:
π§ͺ 132. Baseline ProjectΒΆ
Query
β
Embedding
β
Vector Search
β
Top-10
β
Reranking
β
Large LLM
β
Validation LLM
β
Response
Measure:
π§ͺ 133. Optimization Stage 1 β ContextΒΆ
Change:
to:
Measure:
π§ͺ 134. Optimization Stage 2 β RerankingΒΆ
Change:
to:
Measure:
π§ͺ 135. Optimization Stage 3 β Model RoutingΒΆ
Measure:
π§ͺ 136. Optimization Stage 4 β CachingΒΆ
Add:
Measure:
π§ͺ 137. Optimization Stage 5 β Adaptive RetrievalΒΆ
Measure:
π§ͺ 138. Optimization Stage 6 β Budget GuardrailsΒΆ
Implement:
Test:
π§ͺ 139. Cost Test DatasetΒΆ
Include:
Simple Queries
Complex Queries
Long Queries
Multi-Hop Queries
No-Answer Queries
Repeated Queries
Ambiguous Queries
SQL Queries
Graph Queries
Multimodal Queries
Agentic Queries
π§ͺ 140. Cost Benchmark HarnessΒΆ
def benchmark_cost(rag, queries):
results = []
for query in queries:
result = rag.answer(query)
results.append({
"query": query,
"cost": result.cost,
"latency_ms": result.latency_ms,
"input_tokens": result.input_tokens,
"output_tokens": result.output_tokens,
"llm_calls": result.llm_calls
})
return results
π§ͺ 141. Cost MetricsΒΆ
Track:
Cost / Request
Cost / Successful Answer
Cost / High-Quality Answer
Cost / Tenant
Cost / User
Cost / Workflow
Cost / Model
Cost / Provider
π§ͺ 142. Token MetricsΒΆ
Track:
π§ͺ 143. Model MetricsΒΆ
Track:
π§ͺ 144. Cache MetricsΒΆ
Track:
π§ͺ 145. Budget MetricsΒΆ
Track:
π§ͺ 146. Cost DashboardΒΆ
A production dashboard should show:
ββββββββββββββββββββββββββββββββββββββββββββββ
β COST OVERVIEW β
ββββββββββββββββββββββββββββββββββββββββββββββ€
β Daily Cost $2,160 β
β Monthly Forecast $64,800 β
β Cost / Request $0.018 β
β Cost / Success $0.021 β
β β
β Input Tokens 4,100 β
β Output Tokens 390 β
β Cache Hit Rate 31% β
β β
β LLM 69% β
β Embedding 11% β
β Reranking 7% β
β Vector DB 6% β
β Infra 4% β
β Observability 3% β
ββββββββββββββββββββββββββββββββββββββββββββββ
Illustrative values.
π§ͺ 147. Tenant Cost DashboardΒΆ
This enables chargeback and optimization.
π§ͺ 148. Cost Anomaly DetectionΒΆ
Alert when:
or:
or:
π§ͺ 149. Example Cost AlertΒΆ
ALERT: RAG cost anomaly
Average cost/request:
$0.018
Current:
$0.029
Increase:
61%
Primary signal:
Context tokens +72%
Possible cause:
Top-K configuration changed.
π§ 150. Cost GovernanceΒΆ
Enterprise RAG should define:
π§ 151. Model GovernanceΒΆ
Define:
π§ 152. Cost Governance by RiskΒΆ
Different workloads can have different budgets:
π§ 153. FinOps for RAGΒΆ
RAG can adopt FinOps principles:
π§ 154. RAG FinOps ArchitectureΒΆ
flowchart TD
A["RAG Usage"] --> B["Telemetry"]
B --> C["Cost Attribution"]
C --> D["Tenant"]
C --> E["Application"]
C --> F["Model"]
C --> G["Workflow"]
D --> H["Budgets"]
E --> H
F --> H
G --> H
H --> I["Optimization"]
I --> J["Routing"]
I --> K["Caching"]
I --> L["Context Optimization"]
I --> M["Infrastructure Optimization"]
J --> N["Lower Cost"]
K --> N
L --> N
M --> N π§ 155. Production Cost Optimization ArchitectureΒΆ
USER
β
βΌ
βββββββββββββββ
β RAG API β
ββββββββ¬βββββββ
β
βΌ
βββββββββββββββββββ
β Budget Manager β
ββββββββββ¬βββββββββ
β
βΌ
Query Router
β
ββββββββββββββΌβββββββββββββ
βΌ βΌ βΌ
Fast Path Standard Advanced
β β β
ββββββββββββββΌβββββββββββββ
βΌ
Retrieval
β
βΌ
Reranking
β
βΌ
Context Budget
β
βΌ
Model Router
β
ββββββββββ΄βββββββββ
βΌ βΌ
Small LLM Large LLM
β β
ββββββββββ¬βββββββββ
βΌ
Validation
β
βΌ
Citation
β
βΌ
Response
β
βΌ
Cost Telemetry
β
ββββββββββββββΌβββββββββββββ
βΌ βΌ βΌ
Tenant Model Workflow
Cost Cost Cost
β β β
ββββββββββββββΌβββββββββββββ
βΌ
Cost Dashboard
β
βΌ
Optimization
π§ 156. Cost Optimization MaturityΒΆ
Level 1 β Basic VisibilityΒΆ
Level 2 β Token VisibilityΒΆ
Level 3 β Component CostΒΆ
Level 4 β Cost AttributionΒΆ
Level 5 β Cost ControlsΒΆ
Level 6 β Cost-Aware RAGΒΆ
Level 7 β Autonomous OptimizationΒΆ
π§ 157. Cost Optimization Mental ModelΒΆ
RAG COST
β
ββββββββββββββββββββββΌβββββββββββββββββββββ
βΌ βΌ βΌ
AI DATA PLATFORM
β β β
LLM Vector DB Compute
Embedding Storage Network
Reranker Search Observability
Evaluation
β β β
ββββββββββββββββββββββΌβββββββββββββββββββββ
βΌ
COST CONTROL
β
βββββββββββββββΌββββββββββββββ
βΌ βΌ βΌ
Reduce Reuse Route
Work Work Work
β β β
βββββββββββββββΌββββββββββββββ
βΌ
COST / QUALITY
π§ 158. Final Cost Optimization LoopΒΆ
Measure
β
Attribute
β
Find Largest Cost Driver
β
Remove Unnecessary Work
β
Reduce Tokens
β
Reduce Model Calls
β
Cache
β
Route
β
Optimize Infrastructure
β
Apply Budgets
β
Benchmark Quality
β
Deploy
β
Monitor
β
Repeat
π§ 159. Production PrinciplesΒΆ
Principle 1ΒΆ
The cheapest request is the request you do not need to execute.
Use:
Principle 2ΒΆ
The second-cheapest request is the one executed with the smallest suitable model.
Principle 3ΒΆ
Context is a cost center.
Do not treat retrieved tokens as free.
Principle 4ΒΆ
Every additional RAG stage has an economic cost.
Before adding:
measure the quality improvement.
Principle 5ΒΆ
Optimize cost per successful answer, not cost per request alone.
π§ 160. Production RAG Cost ChecklistΒΆ
β Calculate cost/request
β Calculate cost/successful answer
β Calculate cost/high-quality answer
β Track LLM input tokens
β Track LLM output tokens
β Track context tokens
β Track conversation tokens
β Track embedding cost
β Track reranking cost
β Track retrieval infrastructure
β Track vector DB cost
β Track compute cost
β Track network cost
β Track observability cost
β Track evaluation cost
β Optimize context size
β Optimize Top-K
β Optimize reranking candidates
β Optimize query rewriting
β Optimize multi-query
β Batch embeddings
β Cache embeddings
β Incrementally embed documents
β Use content hashes
β Optimize vector indexes
β Cache retrieval
β Implement semantic cache where appropriate
β Version caches
β Implement cache invalidation
β Preserve tenant isolation
β Implement model routing
β Implement model cascades
β Use smaller models where appropriate
β Limit output tokens
β Optimize prompts
β Use prompt caching where appropriate
β Implement selective validation
β Optimize citation processing
β Limit agent iterations
β Limit agent tool calls
β Limit agent token budgets
β Limit Graph RAG traversal
β Limit SQL result sets
β Optimize multimodal processing
β Optimize compute
β Optimize GPU utilization
β Optimize vector DB
β Optimize storage
β Optimize network
β Implement autoscaling
β Separate workloads
β Implement backpressure
β Implement tenant budgets
β Implement application budgets
β Implement workflow budgets
β Implement request budgets
β Implement soft limits
β Implement hard limits
β Implement graceful degradation
β Build cost dashboards
β Build tenant dashboards
β Build model dashboards
β Build workflow dashboards
β Detect anomalies
β Forecast costs
β Benchmark optimizations
β Monitor quality regressions
β Review cost continuously
π 161. Key TakeawaysΒΆ
- RAG cost is broader than LLM API cost.
- Cost should be measured across the entire architecture.
- LLM tokens are often a major cost driver.
- Context size is one of the most important optimization opportunities.
- Reduce unnecessary context before reducing model quality.
- Do not blindly increase Top-K.
- Use adaptive retrieval where appropriate.
- Query rewriting should be conditional when possible.
- Multi-query retrieval should justify its additional cost.
- Batch embedding operations.
- Cache repeated embeddings.
- Avoid re-embedding unchanged documents.
- Use incremental indexing.
- Use content hashes for change detection.
- Reranking cost grows with candidate volume.
- Reduce candidates before expensive reranking.
- Selective reranking can reduce cost.
- Model routing can significantly reduce average LLM cost.
- Model cascades can use expensive models only when required.
- Output token limits reduce both cost and latency.
- Prompt optimization reduces unnecessary input tokens.
- Conversation compression controls growing history cost.
- Retrieval caching can avoid repeated downstream computation.
- Semantic caching requires careful freshness and authorization controls.
- Cache versioning is essential.
- Tenant isolation must apply to caches.
- Agentic RAG requires explicit cost budgets.
- Agent loops must have limits.
- Graph traversal should have bounded depth and result size.
- SQL RAG requires execution and result-size controls.
- Multimodal RAG requires image and vision cost controls.
- Validation should be risk-based where appropriate.
- Citation processing should reuse source metadata.
- Infrastructure cost matters alongside model cost.
- Vector DB cost depends on storage, compute, replication, and workload.
- Autoscaling can reduce idle infrastructure costs.
- Storage tiering can reduce long-term knowledge-base costs.
- Observability can itself become a significant cost center.
- Evaluation should be sampled intelligently where appropriate.
- Cost should be attributable by tenant, application, model, and workflow.
- Cost budgets provide predictable governance.
- Cost guardrails protect against runaway agentic or high-token requests.
- Graceful degradation allows systems to remain useful under budget constraints.
- RAG FinOps combines visibility, allocation, optimization, and governance.
- Cost optimization must preserve quality and reliability.
- The correct target is not minimum cost.
- The correct target is minimum cost for the required production quality and service level.
π§ 162. Chapter NavigationΒΆ
Part V β Advanced Retrieval-Augmented GenerationΒΆ
Previous:
08. RAG Performance Optimization
Next:
10. Production Retrieval Architecture
Section:
06 β Production RAG Engineering
Production RAG Engineering PathΒΆ
01 Prompt Assembly
β
02 Context Selection & Context Engineering
β
03 Response Validation
β
04 Citation & Source Attribution
β
05 Enterprise Response
β
06 RAG Evaluation & Benchmarking
β
07 RAG Observability
β
08 RAG Performance Optimization
β
09 RAG Cost Optimization
β
10 Production Retrieval Architecture
β
11 Building Production RAG Systems
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.