Skip to content

09. RAG Cost OptimizationΒΆ

Category: Production RAG Engineering
Module: Part V β€” Advanced Retrieval-Augmented Generation
Difficulty: Advanced


πŸ“– OverviewΒΆ

A production RAG system must be optimized not only for accuracy and latency, but also for economic efficiency.

A system that produces excellent answers but costs:

$0.50 per request

may not be viable when serving:

10,000 requests/day

Similarly, a system that is inexpensive but produces poor answers creates operational and business risk.

Production RAG cost optimization therefore focuses on the complete cost chain:

User Query
    ↓
Query Processing
    ↓
Embedding
    ↓
Retrieval
    ↓
Reranking
    ↓
Context Processing
    ↓
Prompt Assembly
    ↓
LLM Generation
    ↓
Validation
    ↓
Citation
    ↓
Observability

The objective is:

Quality
   +
Performance
   +
Reliability
   +
Security
   +
Cost Efficiency

A useful production principle is:

Do not minimize cost blindly. Minimize the cost of achieving the required quality, latency, and reliability.


🎯 Learning Objectives¢

After completing this chapter, you will be able to:

  • Understand RAG cost architecture
  • Identify major RAG cost drivers
  • Calculate cost per request
  • Calculate cost per user
  • Calculate cost per tenant
  • Calculate cost per workflow
  • Understand LLM token economics
  • Optimize input token usage
  • Optimize output token usage
  • Optimize retrieval costs
  • Optimize embedding costs
  • Optimize reranking costs
  • Optimize validation costs
  • Optimize agentic RAG costs
  • Optimize Graph RAG costs
  • Optimize SQL RAG costs
  • Optimize multimodal RAG costs
  • Implement caching strategies
  • Implement model routing
  • Implement model cascading
  • Implement adaptive retrieval
  • Reduce unnecessary LLM calls
  • Optimize context size
  • Optimize prompt size
  • Optimize infrastructure costs
  • Optimize vector database costs
  • Optimize observability costs
  • Implement cost budgets
  • Implement cost guardrails
  • Implement tenant-level cost controls
  • Build cost dashboards
  • Detect cost anomalies
  • Perform cost attribution
  • Perform cost forecasting
  • Design cost-aware RAG architectures

🧠 1. What Is RAG Cost Optimization?¢

RAG cost optimization is the process of reducing the resources required to serve RAG requests while preserving acceptable:

Answer Quality
Latency
Reliability
Security

A simplified objective is:

Cost ↓
Quality ↔
Latency ↔
Reliability ↔

🧠 2. RAG Cost Is More Than LLM Cost¢

A common mistake is:

RAG Cost = LLM Cost

In reality:

Total RAG Cost
       β”‚
       β”œβ”€β”€ LLM
       β”œβ”€β”€ Embeddings
       β”œβ”€β”€ Reranking
       β”œβ”€β”€ Vector Database
       β”œβ”€β”€ Search Infrastructure
       β”œβ”€β”€ Compute
       β”œβ”€β”€ Storage
       β”œβ”€β”€ Network
       β”œβ”€β”€ Observability
       β”œβ”€β”€ Evaluation
       └── Background Processing

🧠 3. Cost Architecture¢

flowchart TD
    A["RAG Request"] --> B["Query Processing"]
    B --> C["Embedding"]
    C --> D["Retrieval"]
    D --> E["Reranking"]
    E --> F["Context Processing"]
    F --> G["LLM"]
    G --> H["Validation"]
    H --> I["Citation"]

    C --> J["Embedding Cost"]
    D --> K["Vector/Search Cost"]
    E --> L["Reranker Cost"]
    G --> M["LLM Cost"]
    H --> N["Validation Cost"]
    I --> O["Processing Cost"]

    P["Infrastructure"] --> Q["Compute"]
    P --> R["Storage"]
    P --> S["Network"]
    P --> T["Observability"]

🧠 4. Cost Categories¢

A practical classification:

Variable Cost
    ↓
Cost changes with requests/tokens

Fixed Cost
    ↓
Infrastructure that exists regardless of request volume

Semi-Variable Cost
    ↓
Resources that scale with workload

Examples:

VariableΒΆ

LLM Tokens
Embedding Requests
Reranker Requests
Evaluation Calls

FixedΒΆ

Base Infrastructure
Monitoring Platform
Reserved Capacity

Semi-VariableΒΆ

Vector DB
Compute
Storage
Network

🧠 5. Cost Per Request¢

A simplified model:

C_request =
C_embedding
+ C_retrieval
+ C_reranking
+ C_context
+ C_generation
+ C_validation
+ C_observability
+ C_infrastructure

🧠 6. LLM Cost¢

LLM cost is commonly driven by:

Input Tokens
+
Output Tokens

Conceptually:

LLM Cost
=
Input Tokens Γ— Input Price
+
Output Tokens Γ— Output Price

Actual pricing varies by provider, model, region, and pricing program.


🧠 7. RAG Token Composition¢

A request may contain:

System Prompt
+
User Query
+
Conversation History
+
Retrieved Context
+
Tool Results
+
Output

Therefore:

Input Tokens
=
Instructions
+
Query
+
History
+
Context
+
Tools

🧠 8. Context Is Often the Largest Optimization Opportunity¢

Example:

System Prompt       1,000 tokens
User Query            100 tokens
Conversation         900 tokens
Retrieved Context   8,000 tokens
───────────────────────────────
Input               10,000 tokens

If the context is reduced:

8,000 β†’ 3,000 tokens

the input token cost can fall significantly.


🧠 9. Context Cost¢

Retrieved Documents
        ↓
Filtering
        ↓
Reranking
        ↓
Compression
        ↓
Final Context
        ↓
LLM

The goal is not:

Retrieve Maximum Context

but:

Retrieve Sufficient Evidence

🧠 10. Context Efficiency¢

A useful conceptual metric:

Context Efficiency
=
Useful Evidence
──────────────────
Context Tokens

Higher is generally better.

This is not a universal standardized metric; use it as an engineering diagnostic.


🧠 11. Context Waste¢

Example:

Context:
10,000 tokens

Useful:
3,000 tokens

Potentially irrelevant:
7,000 tokens

The system is paying for:

10,000

while deriving useful information from approximately:

3,000

🧠 12. Context Optimization¢

Use:

Top-K Tuning
+
Reranking
+
MMR
+
Context Compression
+
Deduplication
+
Metadata Filtering
+
Token Budgeting

🧠 13. Token Budget¢

Define:

Maximum Context Tokens

Example:

Context Budget = 4,000

Then:

Retrieved Candidates
        ↓
Rank
        ↓
Select
        ↓
Fit within Budget

🧠 14. Token Budget Architecture¢

flowchart LR
    A["Model Context Window"] --> B["System Instructions"]
    A --> C["User Query"]
    A --> D["Conversation"]
    A --> E["Retrieved Context"]
    A --> F["Output Budget"]

    E --> G["Context Budget"]
    G --> H["Relevant Evidence"]

🧠 15. Dynamic Context Budget¢

Not every query needs the same amount of context.

Simple FAQ
    ↓
2,000 tokens

Technical Query
    ↓
4,000 tokens

Complex Multi-Hop Query
    ↓
8,000 tokens

Use adaptive budgets where appropriate.


🧠 16. Cost of Top-K¢

Increasing K can increase:

Retrieval Work
Reranking Work
Context Tokens
LLM Input Cost
LLM Latency

Example:

K = 5
    ↓
2,500 tokens

K = 20
    ↓
10,000 tokens

The quality improvement may not justify the additional cost.


🧠 17. Cost-Aware Top-K¢

Instead of:

Always K = 20

use:

Simple Query β†’ K = 5
Complex Query β†’ K = 10
High Uncertainty β†’ K = 20

🧠 18. Adaptive Retrieval¢

Query
 ↓
Retrieve Small Candidate Set
 ↓
Confidence Check
 β”‚
 β”œβ”€β”€ High Confidence β†’ Generate
 β”‚
 └── Low Confidence β†’ Expand Retrieval

This reduces expensive work for easy queries.


🧠 19. Early Exit¢

Example:

Initial Retrieval
      ↓
Top Result Score = 0.94
      ↓
Evidence Sufficient?
      β”‚
      β”œβ”€β”€ Yes β†’ Generate
      └── No  β†’ Expand

Use calibrated signals rather than arbitrary score thresholds.


🧠 20. Query Rewriting Cost¢

Query rewriting may require an LLM.

User Query
    ↓
Rewrite Model
    ↓
Retrieval

If rewriting costs:

$0.002/request

and is executed:

1,000,000 times

the rewriting layer alone contributes:

$2,000

before the main generation cost.


🧠 21. Conditional Query Rewriting¢

Query
 ↓
Complexity Detector
 β”‚
 β”œβ”€β”€ Simple β†’ Direct Retrieval
 β”‚
 └── Complex β†’ Query Rewrite

This can significantly reduce unnecessary model calls.


🧠 22. Multi-Query Cost¢

Multi-query:

Original
 β”œβ”€β”€ Query A
 β”œβ”€β”€ Query B
 β”œβ”€β”€ Query C
 └── Query D

may multiply:

Embedding Calls
Retrieval Calls
Reranking Candidates
Network Requests

Use it when the quality improvement justifies the additional cost.


🧠 23. Multi-Query Cost Optimization¢

Query
 ↓
Determine Need
 β”‚
 β”œβ”€β”€ Low Ambiguity β†’ Single Query
 β”‚
 └── High Ambiguity β†’ Multi-Query

🧠 24. Embedding Cost¢

Embedding cost can come from:

Document Ingestion
Query Embeddings
Re-indexing
Evaluation

Query embeddings happen frequently.

Document embeddings can become expensive during large ingestion operations.


🧠 25. Embedding Cost Optimization¢

Use:

Batching
Caching
Incremental Embedding
Change Detection
Smaller Models
Local Models

where quality requirements permit.


🧠 26. Avoid Re-Embedding Unchanged Documents¢

Use content hashes:

import hashlib


def content_hash(text):

    return hashlib.sha256(
        text.encode("utf-8")
    ).hexdigest()

Pipeline:

Document
   ↓
Hash
   ↓
Compare Previous Hash
   β”‚
   β”œβ”€β”€ Same β†’ Skip
   └── Changed β†’ Re-Embed

🧠 27. Incremental Indexing¢

Bad:

1,000,000 documents
        ↓
Document changes
        ↓
Re-embed 1,000,000

Better:

1,000,000 documents
        ↓
500 changed
        ↓
Re-embed 500

🧠 28. Batch Embedding¢

Document 1 ┐
Document 2 β”‚
Document 3 β”œβ”€β”€ Batch β†’ Embedding Model
Document 4 β”‚
Document 5 β”˜

Batching can reduce per-request overhead and improve accelerator utilization.


🧠 29. Embedding Model Economics¢

Compare:

Model A
Quality = 91%
Cost = Low

Model B
Quality = 94%
Cost = Medium

Model C
Quality = 96%
Cost = High

Choose based on:

Required Retrieval Quality
+
Cost Budget
+
Latency Target

🧠 30. Local vs Hosted Embeddings¢

HostedΒΆ

Application
    ↓
External Embedding API

Costs may include:

API Usage
Network

LocalΒΆ

Application
    ↓
Local Embedding Model

Costs shift toward:

Compute
GPU/CPU
Memory
Operations

Neither is universally cheaper.


🧠 31. Reranking Cost¢

Rerankers can be expensive because they may evaluate many query-document pairs.

Conceptually:

Cost ∝ Number of Candidates

🧠 32. Candidate Reduction¢

Instead of:

Retrieve 500
 ↓
Rerank 500

use:

Dense Retrieval β†’ 30
Sparse Retrieval β†’ 30
Merge β†’ 50
Rerank β†’ 50
Select β†’ 6

🧠 33. Reranker Selection¢

Possible options:

Lightweight Reranker
Cross-Encoder
LLM Reranker

Use the least expensive approach that meets the quality requirement.


🧠 34. Selective Reranking¢

Query
 ↓
Initial Retrieval
 ↓
Confidence
 β”‚
 β”œβ”€β”€ High β†’ Skip Reranker
 β”‚
 └── Low β†’ Rerank

This can reduce average cost.


🧠 35. LLM Cost¢

LLM cost often dominates production RAG economics.

RAG Request
     β”‚
     β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚        LLM COST          β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Input Tokens             β”‚
β”‚ Output Tokens            β”‚
β”‚ Number of Calls          β”‚
β”‚ Model Selection          β”‚
β”‚ Retries                  β”‚
β”‚ Validation Calls         β”‚
β”‚ Agent Tool Calls         β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

🧠 36. Reduce Number of LLM Calls¢

One of the strongest cost optimizations:

Do not call an LLM unless necessary.

Example:

Query
 ↓
Classifier
 β”‚
 β”œβ”€β”€ FAQ β†’ Cached Answer
 β”œβ”€β”€ Search β†’ Retrieval + Small LLM
 β”œβ”€β”€ Complex β†’ Large LLM
 └── SQL β†’ SQL Pipeline

🧠 37. Model Routing¢

flowchart TD
    A["Query"] --> B["Model Router"]

    B --> C["Small Model"]
    B --> D["Medium Model"]
    B --> E["Large Model"]

    C --> F["Response"]
    D --> F
    E --> F

The router should consider:

Complexity
Risk
Latency
Cost
Required Reasoning

🧠 38. Model Cascade¢

Small Model
     ↓
Confidence
 β”‚
 β”œβ”€β”€ High β†’ Final Answer
 β”‚
 └── Low β†’ Large Model

Average cost can decrease if most queries are successfully handled by the smaller model.


🧠 39. Model Routing Example¢

FAQ
   ↓
Small Model

Technical Design
   ↓
Medium Model

Complex Multi-Hop Reasoning
   ↓
Large Model

🧠 40. Model Routing Economics¢

Suppose:

80% queries β†’ Small Model
20% queries β†’ Large Model

instead of:

100% β†’ Large Model

the average cost can be substantially lower, assuming quality remains acceptable.


🧠 41. Output Token Optimization¢

Output tokens directly affect:

Cost
Latency

Use:

Response Length Policies
Max Output Tokens
Structured Output
Concise Instructions

🧠 42. Output Budget¢

Simple Answer
    ↓
200 tokens

Technical Explanation
    ↓
800 tokens

Detailed Analysis
    ↓
1500 tokens

Do not allocate the maximum output budget to every request.


🧠 43. Response Contract¢

Example:

Answer:
Maximum 5 bullet points.

Citations:
Only sources actually used.

Do not repeat the question.

This reduces unnecessary output.


🧠 44. Prompt Optimization¢

Prompt cost includes:

System Prompt
Examples
Instructions
Context
Conversation History
Tool Results

Reduce:

Repeated Instructions
Unused Examples
Duplicate Context
Excessive Formatting

🧠 45. Prompt Versioning¢

Use:

prompt-v1
prompt-v2
prompt-v3

Measure:

Quality
Tokens
Latency
Cost

A shorter prompt is not automatically better if quality drops.


🧠 46. Prompt Caching¢

If supported by the model/provider:

Static Prompt Prefix
        ↓
Cache
        ↓
Dynamic Context + Query

Potentially reduces:

Cost
Latency

🧠 47. Conversation Cost¢

Conversational RAG can become expensive because history grows.

Turn 1 β†’ 500 tokens
Turn 2 β†’ 1,200 tokens
Turn 3 β†’ 2,500 tokens
Turn 4 β†’ 4,500 tokens
Turn 5 β†’ 7,000 tokens

🧠 48. Conversation Compression¢

Instead of passing all history:

Full History
     ↓
Conversation Summary
     ↓
Relevant Recent Turns
     ↓
LLM

🧠 49. Memory Budget¢

Set:

Maximum Conversation Context

Use:

Summary
+
Relevant Turns
+
Current Query

rather than blindly sending the complete conversation.


🧠 50. Semantic Conversation Selection¢

Retrieve only relevant historical messages:

Conversation Memory
        ↓
Semantic Search
        ↓
Relevant Turns
        ↓
Prompt

This reduces token usage.


🧠 51. Retrieval Cache¢

Cache retrieval results for repeated requests.

Key should account for:

Normalized Query
Retriever Version
Index Version
Metadata Filter
Tenant
Authorization Context

🧠 52. Cache Economics¢

If:

100,000 requests

and:

30% cache hits

then:

30,000 requests

may avoid some downstream computation.

Actual savings depend on what the cache bypasses.


🧠 53. Semantic Cache¢

Example:

Query A:
"What database does payment use?"

Query B:
"Which DB is used by payments?"

Potentially reuse the result.

But semantic caching must validate:

Intent
Freshness
Authorization
Tenant
Knowledge Version

🧠 54. Cache Invalidation¢

Invalidate when:

Document Changes
Index Changes
Embedding Changes
Prompt Changes
Retriever Changes
Authorization Changes

🧠 55. Cache Versioning¢

Example:

tenant-a:
index:v12
retriever:v5
prompt:v8
embedding:v3

This prevents stale results from older pipeline versions.


🧠 56. Tenant Cost Attribution¢

Enterprise RAG should answer:

How much does Tenant A cost?

How much does Tenant B cost?

Which tenant consumes the most tokens?

Which tenant generates the most expensive requests?

🧠 57. Cost Per Tenant¢

Example:

Tenant A β†’ $120/month
Tenant B β†’ $480/month
Tenant C β†’ $75/month

This enables:

Chargeback
Budgeting
Quota Management
Optimization

🧠 58. Cost Per User¢

Track:

Requests/User
Tokens/User
Cost/User
Model Usage/User

Avoid exposing user-level data broadly; apply appropriate privacy and access controls.


🧠 59. Cost by Application¢

Enterprise platforms may host:

Support Copilot
Developer Assistant
Legal Assistant
Finance Assistant
Research Assistant

Track cost separately.

Application
    ↓
Requests
    ↓
Tokens
    ↓
Cost

🧠 60. Cost by Workflow¢

A single application may have:

Simple Search
Complex RAG
Agentic RAG
Graph RAG
SQL RAG
Document Analysis

Each can have a different cost profile.


🧠 61. Cost by Model¢

Track:

Model A
Model B
Model C

with:

Requests
Tokens
Latency
Quality
Cost

🧠 62. Cost by Provider¢

For multi-cloud enterprise systems:

AWS
Azure
GCP
External Providers
Self-Hosted Models

compare:

Cost
Latency
Quality
Availability

🧠 63. Reranking Cost by Query¢

Some queries may require:

No Reranking

others:

50 candidates

and complex queries:

100 candidates

Use query-aware policies.


🧠 64. Agentic RAG Cost¢

Agentic RAG can multiply cost.

User Query
   ↓
Planner
   ↓
Tool
   ↓
Observation
   ↓
Planner
   ↓
Tool
   ↓
Observation
   ↓
LLM

Potentially:

Multiple LLM Calls
+
Multiple Retrieval Calls
+
Multiple Tool Calls

🧠 65. Agentic RAG Cost Guardrails¢

Define:

Max Iterations
Max Tool Calls
Max Tokens
Max Execution Time
Max Cost

Example:

Max Steps = 5
Max Tools = 8
Max Cost = $0.10

🧠 66. Agentic Early Termination¢

Agent
 ↓
Evidence Sufficient?
 β”‚
 β”œβ”€β”€ Yes β†’ Final Answer
 └── No β†’ Continue

Avoid unnecessary planning loops.


🧠 67. Agentic Loop Detection¢

Potential problem:

Plan
 ↓
Search
 ↓
Plan
 ↓
Search
 ↓
Plan
 ↓
Search

Use:

Iteration Limit
Repeated Action Detection
Budget Limit

🧠 68. Graph RAG Cost¢

Graph RAG may require:

Graph Query
Entity Retrieval
Relationship Traversal
Subgraph Construction
LLM Reasoning

Cost can increase with traversal depth.


🧠 69. Graph Traversal Budget¢

Instead of:

Unlimited Traversal

use:

Maximum Hops
Maximum Nodes
Maximum Edges
Maximum Graph Query Time

🧠 70. SQL RAG Cost¢

SQL RAG can generate expensive queries.

Potential risks:

Large Table Scan
Complex Join
Repeated Query
Unbounded Result Set

Use:

Query Limits
Timeouts
Read-Only Access
Pagination
Query Validation

🧠 71. SQL Result Budget¢

Instead of:

100,000 rows

use:

Top-N
Aggregations
Pagination
Server-Side Filtering

Only return data required for reasoning.


🧠 72. Multimodal RAG Cost¢

Multimodal workloads may involve:

OCR
Vision Models
Image Embeddings
Text Embeddings
Image Storage
Large Context

Optimize using:

Selective OCR
Image Resizing
Compression
Cached Embeddings
Model Routing

🧠 73. Image Processing Cost¢

Avoid processing every image at maximum resolution.

Image
 ↓
Determine Required Resolution
 ↓
Resize
 ↓
Vision Model

🧠 74. Validation Cost¢

If every answer invokes:

Primary LLM
+
Validation LLM
+
Citation LLM

the cost can multiply.

Possible alternatives:

Rules
+
Cheap Model
+
Selective Deep Validation

🧠 75. Risk-Based Validation¢

flowchart TD
    A["Generated Response"] --> B["Risk Classifier"]

    B --> C["Low Risk"]
    B --> D["Medium Risk"]
    B --> E["High Risk"]

    C --> F["Rule Validation"]
    D --> G["Lightweight Validation"]
    E --> H["Deep Validation"]

    F --> I["Final Response"]
    G --> I
    H --> I

🧠 76. Cost-Aware Citation¢

Citations should ideally be produced using source metadata already carried through the pipeline.

Retriever
 ↓
Chunk Metadata
 ↓
Context
 ↓
Response
 ↓
Citation

Avoid unnecessary additional LLM calls just to reconstruct sources.


🧠 77. Infrastructure Cost¢

Infrastructure includes:

RAG Compute
Vector DB
Search Engine
Object Storage
Databases
GPU
Load Balancers
Network
Observability

🧠 78. Compute Optimization¢

Optimize:

CPU Utilization
Memory
GPU Utilization
Autoscaling
Instance Size
Workload Scheduling

🧠 79. CPU vs GPU¢

Use GPU when:

High Model Throughput
Large Embedding Workloads
Local Reranking
Local LLM Inference

CPU may be more appropriate for:

Lightweight Retrieval
Metadata Filtering
Small Embedding Models
Low-Volume Workloads

🧠 80. GPU Utilization¢

Poor:

GPU Utilization = 15%

while paying for a large GPU.

Potential solutions:

Batching
Smaller GPU
Higher Concurrency
Model Quantization
Autoscaling

🧠 81. Autoscaling¢

Scale based on:

Requests/sec
Queue Depth
CPU
Memory
GPU
Token Throughput
Latency

🧠 82. Scale-to-Zero¢

For workloads that are:

Low Traffic
Batch
Development
Evaluation

scale-to-zero or scheduled compute can reduce infrastructure cost where supported.


🧠 83. Vector Database Cost¢

Vector DB cost depends on:

Data Size
Vector Dimensions
Replication
Queries
Storage
Compute
Index Type
High Availability

🧠 84. Vector Storage Optimization¢

Reduce:

Vector Dimensions
Duplicate Vectors
Unused Metadata
Old Versions
Redundant Indexes

where quality and operational requirements permit.


🧠 85. Data Lifecycle¢

Implement:

Hot
 ↓
Warm
 ↓
Cold
 ↓
Archive
 ↓
Delete

Not every document needs identical storage characteristics.


🧠 86. Document Retention¢

Enterprise knowledge bases may contain:

Active Documents
Historical Documents
Expired Documents
Archived Documents

Remove expired data from the active retrieval path when policy allows.


🧠 87. Storage Tiering¢

Hot Knowledge
    ↓
Fast Vector Search

Cold Knowledge
    ↓
Lower-Cost Storage

Retrieve cold data only when required.


🧠 88. Network Cost¢

Network costs may come from:

LLM API Calls
Vector DB
Object Storage
Cross-Region Traffic
Microservice Calls
Telemetry

Reduce unnecessary payloads and cross-region communication.


🧠 89. Region Optimization¢

Choose infrastructure locations based on:

Latency
Data Residency
Compliance
Availability
Provider Pricing

Do not optimize cost by violating data residency or regulatory requirements.


🧠 90. Observability Cost¢

Observability can become expensive when capturing:

Full Prompts
Full Context
Full Documents
Every Trace
Every Token

Use:

Sampling
Redaction
Aggregation
Retention
Selective Payload Capture

🧠 91. Evaluation Cost¢

RAG evaluation can itself consume LLM calls.

Production Dataset
      ↓
Evaluation Model
      ↓
Scores

Large evaluation suites can become expensive.


🧠 92. Evaluation Sampling¢

Instead of evaluating:

100% of requests

consider:

100% Critical Requests
100% Failures
100% Low-Quality Signals
Sample Normal Requests

The right sampling policy depends on the application's risk.


🧠 93. Continuous Evaluation Cost¢

Use:

Offline Evaluation
+
Production Sampling
+
Human Review

rather than evaluating every request with expensive models.


🧠 94. Cost-Aware Evaluation¢

Prioritize:

New Model
New Retriever
New Prompt
New Knowledge Base
High-Risk Domain
Production Regression

🧠 95. Cost Anomaly Detection¢

Monitor:

Average Cost / Request
Tokens / Request
Requests / Tenant
Model Distribution
Cache Hit Rate

🧠 96. Cost Spike Example¢

Normal:

$0.018/request

Suddenly:

$0.031/request

Investigate:

Context Tokens ↑
Model Routing Changed
Reranking Enabled
Validation Added
Cache Hit Rate ↓

🧠 97. Cost Dashboard¢

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              RAG COST DASHBOARD             β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Requests/day                    120,000      β”‚
β”‚ Avg Cost/request                $0.018       β”‚
β”‚ Daily Cost                      $2,160       β”‚
β”‚ Monthly Forecast                $64,800      β”‚
β”‚                                              β”‚
β”‚ LLM                              69%         β”‚
β”‚ Embeddings                       11%         β”‚
β”‚ Reranking                         7%         β”‚
β”‚ Vector DB                         6%         β”‚
β”‚ Infrastructure                    4%         β”‚
β”‚ Observability                     3%         β”‚
β”‚                                              β”‚
β”‚ Cache Hit Rate                   31%         β”‚
β”‚ Avg Context Tokens              3,900        β”‚
β”‚ Avg Output Tokens                 420        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Values are illustrative.


🧠 98. Cost Breakdown¢

                    TOTAL COST
                         β”‚
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β–Ό                β–Ό                β–Ό
       AI              DATA             PLATFORM
        β”‚                β”‚                β”‚
    LLM Tokens       Vector DB         Compute
    Embeddings       Storage           Network
    Reranker         Search            Observability
    Evaluation       Object Storage

🧠 99. Cost per Request Formula¢

If:

Daily Cost = $2,160
Daily Requests = 120,000

then:

Cost/request
=
2160 / 120000
=
$0.018

🧠 100. Monthly Cost Forecast¢

A simple estimate:

Monthly Cost
=
Average Daily Cost Γ— Number of Days

Example:

$2,160 Γ— 30
=
$64,800

This is a simple forecast and does not account for traffic growth or tiered pricing.


🧠 101. Cost Forecasting¢

Forecast using:

Requests
+
Tokens
+
Model Mix
+
Infrastructure

Example:

Current:
100K requests/day

Projected:
200K requests/day

But cost may not exactly double if:

Caching improves
Model mix changes
Reserved capacity applies

🧠 102. Cost Elasticity¢

Measure:

Traffic +10%
Cost +?

A highly elastic architecture may scale cost almost linearly.

An optimized architecture can sometimes achieve:

Traffic ↑
Cost ↑ slower than traffic

through:

Caching
Batching
Autoscaling
Efficient Models

🧠 103. Cost per Successful Answer¢

A more useful business metric than raw cost:

Cost per Successful Answer
=
Total Cost
────────────────────
Successful Answers

This captures quality.


🧠 104. Cost per High-Quality Answer¢

Example:

Total Cost:
$10,000

High-Quality Answers:
800,000

Then:

$10,000 / 800,000
=
$0.0125

🧠 105. Quality-Adjusted Cost¢

Conceptually:

Quality-Adjusted Cost
=
Cost
────────────
Quality Score

This can help compare architectures.

However, quality scores must be defined consistently.


🧠 106. Cost vs Quality¢

Quality
  β–²
  β”‚                 ●
  β”‚             ●
  β”‚         ●
  β”‚      ●
  β”‚   ●
  └────────────────────────► Cost

There is often a point of diminishing returns.


🧠 107. Diminishing Returns¢

Example:

Cost       Quality

$0.01       85%
$0.015      91%
$0.02       94%
$0.04       95%
$0.08       95.5%

The additional:

$0.04

may not justify:

+0.5%

depending on the use case.


🧠 108. Cost Optimization Frontier¢

                 Quality
                    β–²
                    β”‚
                    β”‚       ●
                    β”‚     ●
                    β”‚   ●
                    β”‚ ●
                    └──────────────────► Cost

The goal is to operate near an efficient frontier rather than blindly choosing the cheapest or most accurate configuration.


🧠 109. Cost Guardrails¢

Production systems should have limits:

Max Cost / Request
Max Tokens / Request
Max Agent Iterations
Max Tool Calls
Max Retrieval Candidates
Max Context Tokens

🧠 110. Request Budget¢

Example:

Request Budget = $0.10

Pipeline:

Query
 ↓
Budget Check
 ↓
Retrieval
 ↓
Reranking
 ↓
Generation
 ↓
Remaining Budget?

🧠 111. Budget-Aware Routing¢

Budget = $0.10

Small Model:
$0.01

Reranker:
$0.02

Large Model:
$0.08

If the request has already consumed:

$0.07

a cheaper path may be selected.


🧠 112. Agent Cost Budget¢

Agent
 β”‚
 β”œβ”€β”€ Step 1 β†’ $0.01
 β”œβ”€β”€ Step 2 β†’ $0.02
 β”œβ”€β”€ Step 3 β†’ $0.02
 └── Step 4 β†’ $0.03
              β”‚
              β–Ό
           $0.08
              β”‚
              β–Ό
        Budget = $0.10

Next expensive action may be blocked.


🧠 113. Tenant Budgets¢

Example:

Tenant A
Monthly Budget:
$5,000

Tenant B
Monthly Budget:
$2,000

Use:

Quota
Warning Threshold
Hard Limit

according to business policy.


🧠 114. Soft vs Hard Budgets¢

Soft BudgetΒΆ

80%
 ↓
Warning

Hard BudgetΒΆ

100%
 ↓
Block / Degrade

🧠 115. Graceful Degradation¢

When budget is constrained:

Large Model
    ↓
Small Model

or:

Advanced RAG
    ↓
Fast Retrieval

or:

Deep Validation
    ↓
Light Validation

🧠 116. Cost-Aware Architecture¢

flowchart TD
    A["User Query"] --> B["Budget Manager"]

    B --> C["Query Router"]

    C --> D["Fast Path"]
    C --> E["Standard RAG"]
    C --> F["Advanced RAG"]

    E --> G["Retrieval"]
    F --> G

    G --> H["Reranking"]

    H --> I["Context Budget"]

    I --> J["Model Router"]

    J --> K["Small Model"]
    J --> L["Large Model"]

    K --> M["Validation"]
    L --> M

    M --> N["Response"]

    B --> O["Cost Tracking"]
    O --> P["Budget Enforcement"]

🧠 117. Cost-Aware Retrieval¢

A production retriever should consider:

Quality
Latency
Cost

not only:

Similarity Score

🧠 118. Cost-Aware Query Routing¢

Query
 β”‚
 β”œβ”€β”€ Cheap Route
 β”‚
 β”œβ”€β”€ Standard Route
 β”‚
 └── Expensive Route

Use the expensive route only when required.


🧠 119. Cost-Aware Validation¢

Low Risk
  ↓
No expensive validator

Medium Risk
  ↓
Cheap validator

High Risk
  ↓
Deep validator

🧠 120. Cost-Aware Agentic RAG¢

Agent should understand:

Remaining Budget
Remaining Steps
Remaining Tokens

Example:

{
  "max_cost": 0.10,
  "spent": 0.063,
  "remaining": 0.037,
  "max_steps": 5,
  "steps_used": 3
}

🧠 121. Cost-Aware Graph RAG¢

Use:

Maximum Hops
Maximum Nodes
Maximum Edges
Maximum Query Time

to prevent runaway graph exploration.


🧠 122. Cost-Aware SQL RAG¢

Use:

Query Timeout
Row Limit
Read-Only Mode
Cost Estimation
Execution Plan

where supported.


🧠 123. Cost-Aware Multimodal RAG¢

Use:

Image Resolution Policy
OCR Selection
Vision Model Routing
Image Cache
Embedding Cache

🧠 124. Cost-Aware Evaluation¢

Prioritize evaluation of:

New Models
New Prompts
New Retrieval
High-Risk Queries
Production Failures
User Negative Feedback

🧠 125. Cost Optimization by Layer¢

Layer                Optimization

Query                Routing / Rewrite selectively
Embedding            Cache / Batch
Retrieval            Top-K / ANN
Hybrid               Parallel / Candidate reduction
Reranking            Smaller candidate set
Context              Compression / Deduplication
Prompt               Shorter / Cache
LLM                  Model routing
Validation           Risk-based
Citation             Metadata propagation
Agent                Step budgets
Graph                Traversal limits
SQL                  Query limits
Multimodal           Resolution / routing
Infrastructure       Autoscaling
Observability         Sampling

🧠 126. Cost Optimization Priority¢

A practical sequence:

1. Measure total cost
2. Identify largest cost component
3. Reduce unnecessary work
4. Reduce token volume
5. Reduce number of model calls
6. Add caching
7. Introduce model routing
8. Optimize retrieval
9. Optimize infrastructure
10. Add budget guardrails
11. Continuously benchmark

🧠 127. Cost Optimization Example¢

Baseline:

Top-K              = 20
Reranker           = 20
Context            = 8,000 tokens
LLM                = Large
Validation         = Large LLM
Cache              = None

Optimized:

Top-K              = Adaptive
Reranker           = Selective
Context            = 4,000 tokens
LLM                = Routed
Validation         = Risk-Based
Cache              = Enabled

🧠 128. Example Cost Comparison¢

Configuration Tokens LLM Calls Avg Cost p95 Latency
Baseline 8,500 3 $0.052 2.8s
Context Optimized 5,000 3 $0.035 2.2s
Model Routing 5,000 2 $0.021 1.7s
Cached + Routed 5,000 1.4 avg $0.015 1.3s

Values are illustrative.


🧠 129. Cost Optimization Experiment¢

Hypothesis:

Reducing context from 8K
to 4K tokens will reduce cost
without materially reducing quality.

Measure:

Cost
Latency
Faithfulness
Answer Relevance
Citation Accuracy

🧠 130. Cost Experiment Matrix¢

Experiment Cost Latency Quality Decision
Baseline β€” β€” β€” β€”
Context Reduction β€” β€” β€” β€”
Smaller Model β€” β€” β€” β€”
Caching β€” β€” β€” β€”
Selective Reranking β€” β€” β€” β€”
Model Routing β€” β€” β€” β€”
Adaptive Retrieval β€” β€” β€” β€”

Populate using real benchmarks.


πŸ§ͺ 131. Practical ProjectΒΆ

Build a:

Production RAG Cost Optimization Lab

Start with a baseline RAG system and progressively optimize:

Token Usage
LLM Calls
Retrieval
Reranking
Context
Caching
Model Selection
Infrastructure

πŸ§ͺ 132. Baseline ProjectΒΆ

Query
 ↓
Embedding
 ↓
Vector Search
 ↓
Top-10
 ↓
Reranking
 ↓
Large LLM
 ↓
Validation LLM
 ↓
Response

Measure:

Cost
Latency
Tokens
Quality

πŸ§ͺ 133. Optimization Stage 1 β€” ContextΒΆ

Change:

10 chunks

to:

Adaptive context selection

Measure:

Tokens
Cost
Quality

πŸ§ͺ 134. Optimization Stage 2 β€” RerankingΒΆ

Change:

Rerank 100

to:

Retrieve 50
Rerank 30
Select 6

Measure:

Latency
Cost
Recall
Answer Quality

πŸ§ͺ 135. Optimization Stage 3 β€” Model RoutingΒΆ

Simple
 ↓
Small Model

Complex
 ↓
Large Model

Measure:

Model Distribution
Cost
Quality
Latency

πŸ§ͺ 136. Optimization Stage 4 β€” CachingΒΆ

Add:

Embedding Cache
Retrieval Cache
Semantic Cache

Measure:

Hit Rate
Cost Reduction
Latency Reduction

πŸ§ͺ 137. Optimization Stage 5 β€” Adaptive RetrievalΒΆ

Initial Retrieval
      ↓
Confidence
 β”‚
 β”œβ”€β”€ High β†’ Generate
 └── Low β†’ Expand

Measure:

Average Retrieval Work
Cost
Recall
Quality

πŸ§ͺ 138. Optimization Stage 6 β€” Budget GuardrailsΒΆ

Implement:

Max Cost
Max Tokens
Max LLM Calls
Max Retrieval Candidates
Max Agent Steps

Test:

Normal Query
Complex Query
Adversarial Query
Runaway Agent

πŸ§ͺ 139. Cost Test DatasetΒΆ

Include:

Simple Queries
Complex Queries
Long Queries
Multi-Hop Queries
No-Answer Queries
Repeated Queries
Ambiguous Queries
SQL Queries
Graph Queries
Multimodal Queries
Agentic Queries

πŸ§ͺ 140. Cost Benchmark HarnessΒΆ

def benchmark_cost(rag, queries):

    results = []

    for query in queries:

        result = rag.answer(query)

        results.append({
            "query": query,
            "cost": result.cost,
            "latency_ms": result.latency_ms,
            "input_tokens": result.input_tokens,
            "output_tokens": result.output_tokens,
            "llm_calls": result.llm_calls
        })

    return results

πŸ§ͺ 141. Cost MetricsΒΆ

Track:

Cost / Request
Cost / Successful Answer
Cost / High-Quality Answer
Cost / Tenant
Cost / User
Cost / Workflow
Cost / Model
Cost / Provider

πŸ§ͺ 142. Token MetricsΒΆ

Track:

Input Tokens
Output Tokens
Context Tokens
Conversation Tokens
Tool Tokens
Total Tokens

πŸ§ͺ 143. Model MetricsΒΆ

Track:

Model
Requests
Tokens
Latency
Quality
Cost
Fallback Rate

πŸ§ͺ 144. Cache MetricsΒΆ

Track:

Cache Hit Rate
Cache Miss Rate
Saved Requests
Saved Tokens
Saved Cost
Stale Results
Invalidations

πŸ§ͺ 145. Budget MetricsΒΆ

Track:

Budget Utilization
Requests Over Budget
Soft Limit Warnings
Hard Limit Blocks
Graceful Degradations

πŸ§ͺ 146. Cost DashboardΒΆ

A production dashboard should show:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚             COST OVERVIEW                 β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Daily Cost                    $2,160      β”‚
β”‚ Monthly Forecast             $64,800      β”‚
β”‚ Cost / Request                 $0.018      β”‚
β”‚ Cost / Success                 $0.021      β”‚
β”‚                                            β”‚
β”‚ Input Tokens                    4,100      β”‚
β”‚ Output Tokens                     390      β”‚
β”‚ Cache Hit Rate                    31%      β”‚
β”‚                                            β”‚
β”‚ LLM                             69%        β”‚
β”‚ Embedding                       11%        β”‚
β”‚ Reranking                        7%        β”‚
β”‚ Vector DB                        6%        β”‚
β”‚ Infra                            4%        β”‚
β”‚ Observability                    3%        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Illustrative values.


πŸ§ͺ 147. Tenant Cost DashboardΒΆ

Tenant      Requests     Tokens       Cost

Tenant A      40K        120M       $1,200
Tenant B      20K         90M         $900
Tenant C      10K         30M         $300

This enables chargeback and optimization.


πŸ§ͺ 148. Cost Anomaly DetectionΒΆ

Alert when:

Cost/request > baseline Γ— 1.30

or:

Daily cost > expected budget

or:

Token usage increases unexpectedly

πŸ§ͺ 149. Example Cost AlertΒΆ

ALERT: RAG cost anomaly

Average cost/request:
$0.018

Current:
$0.029

Increase:
61%

Primary signal:
Context tokens +72%

Possible cause:
Top-K configuration changed.

🧠 150. Cost Governance¢

Enterprise RAG should define:

Budget
Quota
Model Policy
Token Policy
Retention Policy
Cache Policy
Routing Policy

🧠 151. Model Governance¢

Define:

Approved Models
Maximum Model Tier
Allowed Providers
Allowed Regions
Fallback Models

🧠 152. Cost Governance by Risk¢

Different workloads can have different budgets:

Low Risk
   ↓
Low-Cost Model

Medium Risk
   ↓
Standard Model

High Risk
   ↓
Premium Model + Validation

🧠 153. FinOps for RAG¢

RAG can adopt FinOps principles:

Visibility
      ↓
Allocation
      ↓
Optimization
      ↓
Governance
      ↓
Continuous Improvement

🧠 154. RAG FinOps Architecture¢

flowchart TD
    A["RAG Usage"] --> B["Telemetry"]

    B --> C["Cost Attribution"]

    C --> D["Tenant"]
    C --> E["Application"]
    C --> F["Model"]
    C --> G["Workflow"]

    D --> H["Budgets"]
    E --> H
    F --> H
    G --> H

    H --> I["Optimization"]

    I --> J["Routing"]
    I --> K["Caching"]
    I --> L["Context Optimization"]
    I --> M["Infrastructure Optimization"]

    J --> N["Lower Cost"]
    K --> N
    L --> N
    M --> N

🧠 155. Production Cost Optimization Architecture¢

                         USER
                           β”‚
                           β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚ RAG API     β”‚
                    β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”˜
                           β”‚
                           β–Ό
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β”‚ Budget Manager  β”‚
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β”‚
                           β–Ό
                    Query Router
                           β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β–Ό            β–Ό            β–Ό
           Fast Path    Standard     Advanced
              β”‚            β”‚            β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β–Ό
                      Retrieval
                           β”‚
                           β–Ό
                       Reranking
                           β”‚
                           β–Ό
                    Context Budget
                           β”‚
                           β–Ό
                      Model Router
                           β”‚
                  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
                  β–Ό                 β–Ό
              Small LLM         Large LLM
                  β”‚                 β”‚
                  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β–Ό
                       Validation
                           β”‚
                           β–Ό
                        Citation
                           β”‚
                           β–Ό
                       Response

                           β”‚
                           β–Ό
                    Cost Telemetry
                           β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β–Ό            β–Ό            β–Ό
           Tenant       Model        Workflow
            Cost         Cost           Cost
              β”‚            β”‚            β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                           β–Ό
                    Cost Dashboard
                           β”‚
                           β–Ό
                     Optimization

🧠 156. Cost Optimization Maturity¢

Level 1 β€” Basic VisibilityΒΆ

Track LLM Cost

Level 2 β€” Token VisibilityΒΆ

Input
Output
Context

Level 3 β€” Component CostΒΆ

Embedding
Retrieval
Reranking
LLM

Level 4 β€” Cost AttributionΒΆ

Tenant
User
Application
Workflow

Level 5 β€” Cost ControlsΒΆ

Budgets
Quotas
Guardrails

Level 6 β€” Cost-Aware RAGΒΆ

Adaptive Retrieval
Model Routing
Caching
Context Optimization

Level 7 β€” Autonomous OptimizationΒΆ

Measure
   ↓
Detect
   ↓
Optimize
   ↓
Benchmark
   ↓
Deploy

🧠 157. Cost Optimization Mental Model¢

                         RAG COST
                            β”‚
       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
       β–Ό                    β–Ό                    β–Ό
      AI                  DATA                PLATFORM
       β”‚                    β”‚                    β”‚
      LLM               Vector DB            Compute
      Embedding         Storage              Network
      Reranker          Search               Observability
      Evaluation
       β”‚                    β”‚                    β”‚
       β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                            β–Ό
                       COST CONTROL
                            β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β–Ό             β–Ό             β–Ό
           Reduce         Reuse         Route
           Work           Work          Work
              β”‚             β”‚             β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                            β–Ό
                       COST / QUALITY

🧠 158. Final Cost Optimization Loop¢

Measure
   ↓
Attribute
   ↓
Find Largest Cost Driver
   ↓
Remove Unnecessary Work
   ↓
Reduce Tokens
   ↓
Reduce Model Calls
   ↓
Cache
   ↓
Route
   ↓
Optimize Infrastructure
   ↓
Apply Budgets
   ↓
Benchmark Quality
   ↓
Deploy
   ↓
Monitor
   ↓
Repeat

🧠 159. Production Principles¢

Principle 1ΒΆ

The cheapest request is the request you do not need to execute.

Use:

Cache
Early Exit
Routing
Deduplication

Principle 2ΒΆ

The second-cheapest request is the one executed with the smallest suitable model.


Principle 3ΒΆ

Context is a cost center.

Do not treat retrieved tokens as free.


Principle 4ΒΆ

Every additional RAG stage has an economic cost.

Before adding:

Reranking
Query Rewriting
Validation
Agent Planning

measure the quality improvement.


Principle 5ΒΆ

Optimize cost per successful answer, not cost per request alone.


🧠 160. Production RAG Cost Checklist¢

☐ Calculate cost/request
☐ Calculate cost/successful answer
☐ Calculate cost/high-quality answer
☐ Track LLM input tokens
☐ Track LLM output tokens
☐ Track context tokens
☐ Track conversation tokens
☐ Track embedding cost
☐ Track reranking cost
☐ Track retrieval infrastructure
☐ Track vector DB cost
☐ Track compute cost
☐ Track network cost
☐ Track observability cost
☐ Track evaluation cost

☐ Optimize context size
☐ Optimize Top-K
☐ Optimize reranking candidates
☐ Optimize query rewriting
☐ Optimize multi-query
☐ Batch embeddings
☐ Cache embeddings
☐ Incrementally embed documents
☐ Use content hashes
☐ Optimize vector indexes

☐ Cache retrieval
☐ Implement semantic cache where appropriate
☐ Version caches
☐ Implement cache invalidation
☐ Preserve tenant isolation

☐ Implement model routing
☐ Implement model cascades
☐ Use smaller models where appropriate
☐ Limit output tokens
☐ Optimize prompts
☐ Use prompt caching where appropriate

☐ Implement selective validation
☐ Optimize citation processing
☐ Limit agent iterations
☐ Limit agent tool calls
☐ Limit agent token budgets
☐ Limit Graph RAG traversal
☐ Limit SQL result sets
☐ Optimize multimodal processing

☐ Optimize compute
☐ Optimize GPU utilization
☐ Optimize vector DB
☐ Optimize storage
☐ Optimize network
☐ Implement autoscaling
☐ Separate workloads
☐ Implement backpressure

☐ Implement tenant budgets
☐ Implement application budgets
☐ Implement workflow budgets
☐ Implement request budgets
☐ Implement soft limits
☐ Implement hard limits
☐ Implement graceful degradation

☐ Build cost dashboards
☐ Build tenant dashboards
☐ Build model dashboards
☐ Build workflow dashboards
☐ Detect anomalies
☐ Forecast costs
☐ Benchmark optimizations
☐ Monitor quality regressions
☐ Review cost continuously

πŸ“š 161. Key TakeawaysΒΆ

  • RAG cost is broader than LLM API cost.
  • Cost should be measured across the entire architecture.
  • LLM tokens are often a major cost driver.
  • Context size is one of the most important optimization opportunities.
  • Reduce unnecessary context before reducing model quality.
  • Do not blindly increase Top-K.
  • Use adaptive retrieval where appropriate.
  • Query rewriting should be conditional when possible.
  • Multi-query retrieval should justify its additional cost.
  • Batch embedding operations.
  • Cache repeated embeddings.
  • Avoid re-embedding unchanged documents.
  • Use incremental indexing.
  • Use content hashes for change detection.
  • Reranking cost grows with candidate volume.
  • Reduce candidates before expensive reranking.
  • Selective reranking can reduce cost.
  • Model routing can significantly reduce average LLM cost.
  • Model cascades can use expensive models only when required.
  • Output token limits reduce both cost and latency.
  • Prompt optimization reduces unnecessary input tokens.
  • Conversation compression controls growing history cost.
  • Retrieval caching can avoid repeated downstream computation.
  • Semantic caching requires careful freshness and authorization controls.
  • Cache versioning is essential.
  • Tenant isolation must apply to caches.
  • Agentic RAG requires explicit cost budgets.
  • Agent loops must have limits.
  • Graph traversal should have bounded depth and result size.
  • SQL RAG requires execution and result-size controls.
  • Multimodal RAG requires image and vision cost controls.
  • Validation should be risk-based where appropriate.
  • Citation processing should reuse source metadata.
  • Infrastructure cost matters alongside model cost.
  • Vector DB cost depends on storage, compute, replication, and workload.
  • Autoscaling can reduce idle infrastructure costs.
  • Storage tiering can reduce long-term knowledge-base costs.
  • Observability can itself become a significant cost center.
  • Evaluation should be sampled intelligently where appropriate.
  • Cost should be attributable by tenant, application, model, and workflow.
  • Cost budgets provide predictable governance.
  • Cost guardrails protect against runaway agentic or high-token requests.
  • Graceful degradation allows systems to remain useful under budget constraints.
  • RAG FinOps combines visibility, allocation, optimization, and governance.
  • Cost optimization must preserve quality and reliability.
  • The correct target is not minimum cost.
  • The correct target is minimum cost for the required production quality and service level.

🧭 162. Chapter Navigation¢

Part V β€” Advanced Retrieval-Augmented GenerationΒΆ

Previous:
08. RAG Performance Optimization

Next:
10. Production Retrieval Architecture

Section:
06 β€” Production RAG Engineering

Production RAG Engineering PathΒΆ

01 Prompt Assembly
        ↓
02 Context Selection & Context Engineering
        ↓
03 Response Validation
        ↓
04 Citation & Source Attribution
        ↓
05 Enterprise Response
        ↓
06 RAG Evaluation & Benchmarking
        ↓
07 RAG Observability
        ↓
08 RAG Performance Optimization
        ↓
09 RAG Cost Optimization
        ↓
10 Production Retrieval Architecture
        ↓
11 Building Production RAG Systems

Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β€” One Chapter at a Time.