06. RAG Evaluation and BenchmarkingΒΆ
Category: Production RAG Engineering
Module: Part V β Advanced Retrieval-Augmented Generation
Difficulty: Advanced
π OverviewΒΆ
Building a RAG system that works is relatively easy.
Building a RAG system that can be measured, evaluated, compared, monitored, and continuously improved is much harder.
A production RAG system contains multiple components:
Query
β
Query Transformation
β
Retriever
β
Reranker
β
Context Selection
β
Prompt Assembly
β
LLM
β
Response Validation
β
Citation
β
Enterprise Response
A poor answer may therefore originate from:
Bad Query
Bad Retrieval
Bad Ranking
Bad Context
Bad Prompt
Bad Model
Bad Grounding
Bad Citation
Bad Response Processing
Therefore, evaluating only the final answer is insufficient.
A mature RAG evaluation framework must evaluate the system at multiple layers:
Retrieval Quality
+
Context Quality
+
Generation Quality
+
Groundedness
+
Citation Quality
+
Response Quality
+
Performance
+
Cost
+
Reliability
The central principle is:
Production RAG systems should be evaluated as end-to-end systems as well as individual components.
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand why RAG evaluation is different from traditional ML evaluation
- Understand RAG evaluation dimensions
- Design retrieval evaluation datasets
- Evaluate retriever quality
- Evaluate ranking quality
- Evaluate context quality
- Evaluate generation quality
- Evaluate groundedness
- Evaluate answer relevance
- Evaluate faithfulness
- Evaluate citation quality
- Evaluate response completeness
- Design golden datasets
- Design benchmark datasets
- Understand Recall@K
- Understand Precision@K
- Understand Hit Rate
- Understand MRR
- Understand MAP
- Understand NDCG
- Understand Context Recall
- Understand Context Precision
- Understand Answer Relevance
- Understand Faithfulness
- Understand Groundedness
- Understand Citation Accuracy
- Understand Citation Coverage
- Understand end-to-end evaluation
- Design LLM-as-a-Judge evaluation
- Understand judge calibration
- Reduce evaluator bias
- Perform human evaluation
- Perform automated evaluation
- Design regression testing
- Design RAG benchmarks
- Evaluate latency
- Evaluate throughput
- Evaluate token usage
- Evaluate cost
- Evaluate reliability
- Design production RAG evaluation pipelines
- Build continuous RAG evaluation systems
- Compare RAG architectures quantitatively
- Build enterprise-grade RAG evaluation dashboards
π§ 1. Why RAG Evaluation Is DifferentΒΆ
Traditional machine learning often evaluates:
RAG introduces additional stages:
The final answer depends on both:
Therefore:
A strong LLM cannot compensate indefinitely for poor retrieval.
π§© 2. RAG Evaluation StackΒΆ
flowchart TD
A["User Query"] --> B["Query Processing"]
B --> C["Retriever"]
C --> D["Reranker"]
D --> E["Context Selection"]
E --> F["Prompt Assembly"]
F --> G["LLM"]
G --> H["Response Validation"]
H --> I["Citation"]
I --> J["Final Response"]
C --> K["Retrieval Evaluation"]
D --> L["Ranking Evaluation"]
E --> M["Context Evaluation"]
G --> N["Generation Evaluation"]
H --> O["Validation Evaluation"]
I --> P["Citation Evaluation"]
J --> Q["End-to-End Evaluation"] π§ 3. Evaluation DimensionsΒΆ
A production RAG system can be evaluated across:
1. Retrieval
2. Ranking
3. Context
4. Generation
5. Groundedness
6. Answer Relevance
7. Citation
8. Completeness
9. Safety
10. Latency
11. Throughput
12. Cost
13. Reliability
π§ 4. Component-Level vs End-to-End EvaluationΒΆ
Component-LevelΒΆ
Evaluate:
End-to-EndΒΆ
Evaluate:
Both are required.
π§ 5. Evaluation PyramidΒΆ
βββββββββββββββββββ
β End-to-End β
β Evaluation β
ββββββββββ¬βββββββββ
β
ββββββββββ΄βββββββββ
β Response / β
β Generation β
ββββββββββ¬βββββββββ
β
ββββββββββ΄βββββββββ
β Context β
β Evaluation β
ββββββββββ¬βββββββββ
β
ββββββββββ΄βββββββββ
β Retrieval / β
β Ranking β
ββββββββββ¬βββββββββ
β
ββββββββββ΄βββββββββ
β Data / β
β Ground Truth β
βββββββββββββββββββ
π§ 6. What Are We Actually Measuring?ΒΆ
A RAG evaluation should answer:
Did we retrieve the right evidence?
Did we rank the evidence correctly?
Did we provide enough context?
Did the model use the context correctly?
Did the model answer the question?
Did the answer remain grounded?
Were citations correct?
Was the response complete?
Was the response safe?
Was the system fast enough?
Was the system cost-efficient?
π§ 7. RAG Evaluation DatasetΒΆ
A basic evaluation record can contain:
{
"question": "What database does the payment service use?",
"ground_truth_answer": "The payment service uses PostgreSQL.",
"relevant_documents": [
"DOC-1042"
],
"relevant_chunks": [
"DOC-1042-C17"
]
}
π§ 8. Golden DatasetΒΆ
A golden dataset is a curated set of questions and expected evidence or answers used for repeatable evaluation.
Example:
Golden datasets are essential for:
π§ 9. Golden Dataset StructureΒΆ
{
"id": "QA-001",
"question": "What database does the payment service use?",
"expected_answer":
"The payment service uses PostgreSQL.",
"relevant_sources": [
"DOC-1042"
],
"relevant_chunks": [
"DOC-1042-C17"
],
"expected_claims": [
"Payment Service uses PostgreSQL"
]
}
π§ 10. Types of Evaluation DataΒΆ
A production evaluation suite should include:
Simple Questions
Complex Questions
Multi-Hop Questions
Ambiguous Questions
No-Answer Questions
Conflicting Evidence
Historical Questions
Numerical Questions
Long-Context Questions
Multi-Document Questions
Metadata Queries
SQL Questions
Graph Questions
Multimodal Questions
Adversarial Questions
π§ 11. Evaluation Dataset SplitΒΆ
Use separate datasets for:
Example:
Development
β
Prompt / Retriever Tuning
Validation
β
Architecture Selection
Regression
β
Release Testing
Production
β
Continuous Monitoring
π§ 12. Avoid Evaluation LeakageΒΆ
Do not continuously tune the system against exactly the same benchmark used for final reporting.
Otherwise:
A separate holdout dataset should be maintained.
π§ 13. Retrieval EvaluationΒΆ
The first major evaluation layer is retrieval.
Question:
Did the retriever return the evidence needed to answer the query?
We then compare:
π§ 14. Retrieval Ground TruthΒΆ
Suppose:
Retriever returns:
Then:
The evaluation system can calculate retrieval metrics.
π§ 15. Precision@KΒΆ
Precision@K measures how many retrieved results are relevant.
Conceptually:
Relevant Results in Top K
βββββββββββββββββββββββββ
K
Example:
π§ 16. Recall@KΒΆ
Recall@K measures how many of the relevant documents were retrieved.
Conceptually:
Relevant Documents Retrieved in Top K
ββββββββββββββββββββββββββββββββββββββ
Total Relevant Documents
Example:
π§ 17. Precision vs RecallΒΆ
Precision
β
"How much of what I retrieved is useful?"
Recall
β
"How much of what I needed did I retrieve?"
For RAG:
is often important because missing the key evidence can make the final answer impossible.
π§ 18. Hit RateΒΆ
Hit Rate asks:
Did at least one relevant result appear in the retrieved set?
Example:
Result:
If no relevant document appears:
π§ 19. Hit Rate@KΒΆ
Across many queries:
Queries With At Least One Relevant Result
βββββββββββββββββββββββββββββββββββββββββ
Total Queries
Example:
π§ 20. Mean Reciprocal RankΒΆ
MRR focuses on the position of the first relevant result.
For a query:
Reciprocal rank:
Across queries:
MRR is useful when finding the first useful result quickly matters.
π§ 21. MRR ExampleΒΆ
Query 1 β Relevant at rank 1 β 1.00
Query 2 β Relevant at rank 2 β 0.50
Query 3 β Relevant at rank 4 β 0.25
Average:
π§ 22. MAPΒΆ
Mean Average Precision considers the positions of multiple relevant results.
Useful when:
are expected for a query.
MAP is especially useful for information-retrieval benchmarking.
π§ 23. NDCGΒΆ
Normalized Discounted Cumulative Gain considers both:
Higher-ranked relevant documents receive more importance.
This makes NDCG useful for:
π§ 24. Graded RelevanceΒΆ
Not all results are simply:
A better model may use:
This is useful for NDCG and ranking evaluation.
π§ 25. Retrieval Evaluation ExampleΒΆ
Query:
"What database does Payment Service use?"
Results:
Rank 1 β Database Architecture β 3
Rank 2 β Kafka Architecture β 1
Rank 3 β Deployment Guide β 0
Rank 4 β Payment Overview β 2
Rank 5 β Security Policy β 0
The evaluation system can determine:
π§ 26. Reranker EvaluationΒΆ
For reranking:
Evaluate:
π§ 27. Reranking BenchmarkΒΆ
Example:
This indicates that reranking improved ordering even if retrieval recall remained unchanged.
π§ 28. Context EvaluationΒΆ
Retrieval is not the same as context quality.
The system may retrieve:
but include only:
in the final prompt.
Therefore:
must be evaluated separately.
π§ 29. Context RecallΒΆ
Context recall asks:
Did the retrieved context contain the information required to answer the question?
Example:
Context contains the required evidence.
π§ 30. Context PrecisionΒΆ
Context precision asks:
How much of the provided context is actually relevant?
Example:
The context has low precision.
π§ 31. Context Recall vs Context PrecisionΒΆ
Context Recall
β
Did we include the evidence?
Context Precision
β
Did we avoid unnecessary evidence?
An effective RAG pipeline needs both.
π§ 32. Context QualityΒΆ
A useful mental model:
π§ 33. Context OrderingΒΆ
Even when the right evidence is present, ordering can matter.
Example:
versus:
Context ordering should therefore be evaluated when relevant to the model and prompt architecture.
π§ 34. Context Compression EvaluationΒΆ
For contextual compression:
Evaluate:
A compression system should not remove evidence required for answering the question.
π§ 35. Parent-Child Retrieval EvaluationΒΆ
Parent-child retrieval should be evaluated on:
π§ 36. Multi-Query Retrieval EvaluationΒΆ
Multi-query retrieval can increase recall.
Compare:
Metrics:
Improvement in recall is not enough if cost and latency become unacceptable.
π§ 37. Hybrid Search EvaluationΒΆ
Compare:
Measure:
π§ 38. Metadata Filtering EvaluationΒΆ
Metadata-aware retrieval should be evaluated for:
A filter that improves precision but accidentally removes valid evidence can reduce recall.
π 39. Security EvaluationΒΆ
For enterprise RAG:
must never become:
Evaluation should therefore include:
Authorization Tests
Tenant Isolation Tests
Permission Tests
Metadata Leakage Tests
Citation Leakage Tests
π§ 40. Generation EvaluationΒΆ
After retrieval and context evaluation, evaluate generation.
Questions:
Did the answer address the question?
Was it grounded?
Was it complete?
Was it accurate?
Was it concise?
Was it consistent?
π§ 41. Answer RelevanceΒΆ
Answer relevance measures whether the generated answer actually addresses the user's question.
Example:
High relevance.
But:
Low relevance.
π§ 42. FaithfulnessΒΆ
Faithfulness asks:
Does the answer remain supported by the provided context?
Context:
Answer:
High faithfulness.
But:
If throughput was not present in the context, the second claim is unsupported.
π§ 43. GroundednessΒΆ
Groundedness measures whether generated claims are supported by retrieved evidence.
A grounded answer should not introduce unsupported facts.
π§ 44. Faithfulness vs GroundednessΒΆ
These concepts overlap but can be operationalized differently.
Faithfulness
β
Does the answer faithfully use the supplied context?
Groundedness
β
Are the claims supported by the evidence?
Organizations should define their exact metric semantics consistently.
π§ 45. CompletenessΒΆ
An answer can be grounded but incomplete.
Question:
Answer:
The answer may be grounded but incomplete.
π§ 46. Completeness EvaluationΒΆ
Evaluate:
Example:
Required:
Database β
Messaging β
Deployment β
Generated:
Database β
Messaging β
Deployment β
π§ 47. Citation EvaluationΒΆ
Citation evaluation should measure:
Citation Validity
Citation Accuracy
Citation Coverage
Citation Completeness
Source Authority
Source Freshness
π§ 48. Citation ValidityΒΆ
A citation is invalid when:
π§ 49. Citation AccuracyΒΆ
Correctly Supporting Citations
βββββββββββββββββββββββββββββββ
Total Citations
A citation should actually support the associated claim.
π§ 50. Citation CoverageΒΆ
Cited Factual Claims
ββββββββββββββββββββ
Total Factual Claims
Example:
π§ 51. Citation CompletenessΒΆ
Citation completeness evaluates whether important claims received appropriate evidence attribution.
Example:
The response is incomplete from an attribution perspective.
π§ 52. Source QualityΒΆ
Source quality can consider:
Example:
π§ 53. Source FreshnessΒΆ
For changing knowledge:
source freshness becomes especially important.
Evaluation can include:
π§ 54. Response QualityΒΆ
Final response evaluation can include:
π§ 55. LLM-as-a-JudgeΒΆ
A powerful approach is to use another LLM to evaluate the response.
The judge may score:
π§© 56. LLM-as-a-Judge ArchitectureΒΆ
flowchart TD
A["Evaluation Dataset"] --> B["RAG System"]
B --> C["Generated Answer"]
A --> D["Question"]
A --> E["Reference Answer"]
C --> F["Judge Model"]
D --> F
E --> F
F --> G["Evaluation Score"]
G --> H["Evaluation Store"] π§ 57. Judge PromptΒΆ
A judge prompt can specify:
Evaluate whether the answer is supported
by the provided context.
Score:
0 = Unsupported
1 = Partially Supported
2 = Fully Supported
Return structured JSON only.
π§© 58. Structured Judge OutputΒΆ
{
"faithfulness": 2,
"answer_relevance": 2,
"completeness": 1,
"reason": "The answer is supported by the retrieved context but omits one required detail."
}
π§ 59. Why Structured Evaluation MattersΒΆ
Free-form judge output:
is difficult to aggregate.
Structured output:
is easier to:
π§ 60. LLM Judge RisksΒΆ
LLM-as-a-Judge can introduce:
Bias
Position Bias
Verbosity Bias
Model Bias
Prompt Sensitivity
Self-Preference
Inconsistent Scoring
Therefore:
LLM evaluation itself must be evaluated.
π§ 61. Judge CalibrationΒΆ
Compare judge scores with human evaluations.
Measure agreement.
If agreement is poor:
π§ 62. Evaluation RubricΒΆ
Example:
Faithfulness
0 β Completely unsupported
1 β Mostly unsupported
2 β Partially supported
3 β Mostly supported
4 β Fully supported
A detailed rubric makes judging more consistent.
π§ 63. Human EvaluationΒΆ
Human evaluation remains important for:
Ambiguous Questions
Complex Reasoning
Enterprise Workflows
High-Risk Domains
User Experience
Judge Calibration
π§ 64. Human Evaluation FormΒΆ
Example:
Question:
_____________________
Answer:
_____________________
Was the answer correct?
[ ] Yes
[ ] Partially
[ ] No
Was it grounded?
[ ] Yes
[ ] Partially
[ ] No
Was it complete?
[ ] Yes
[ ] Partially
[ ] No
Were citations correct?
[ ] Yes
[ ] No
π§ 65. Human + Automated EvaluationΒΆ
A mature evaluation architecture combines:
π§© 66. Evaluation StrategyΒΆ
flowchart TD
A["Evaluation Dataset"] --> B["Automated Metrics"]
A --> C["LLM Judge"]
A --> D["Human Evaluation"]
B --> E["Evaluation Aggregator"]
C --> E
D --> E
E --> F["Quality Report"] π§ 67. Deterministic MetricsΒΆ
Use deterministic metrics when possible.
Examples:
These are generally easier to reproduce.
π§ 68. Semantic MetricsΒΆ
Semantic evaluation can assess:
These often require:
π§ 69. Exact MatchΒΆ
For structured answers:
Exact match:
But:
would fail exact matching despite being semantically correct.
π§ 70. F1 for Extractive AnswersΒΆ
For structured text or entity extraction, token-level precision, recall, and F1 can be useful.
Example:
The shared token:
can contribute to precision and recall.
π§ 71. Semantic SimilarityΒΆ
Embedding-based evaluation can compare:
However:
Semantic similarity does not guarantee factual correctness.
Two statements can be semantically similar while containing a subtle incorrect number or date.
π§ 72. Numerical EvaluationΒΆ
Numbers should often be evaluated separately.
Expected:
Generated:
Semantic similarity may be high.
Factual correctness:
Therefore evaluation should preserve exact facts.
π§ 73. Date EvaluationΒΆ
Expected:
Generated:
A date-aware evaluator should detect the mismatch.
π§ 74. Identifier EvaluationΒΆ
Important identifiers:
These should often use exact matching.
π§ 75. Multi-Hop EvaluationΒΆ
Multi-hop questions require multiple pieces of evidence.
Example:
Evaluation should verify:
Missing one critical hop can invalidate the final conclusion.
π§ 76. Multi-Hop RecallΒΆ
For multi-hop RAG:
The answer may still fail because:
is missing.
Therefore evaluate evidence coverage per hop.
π§ 77. No-Answer EvaluationΒΆ
A strong RAG system should know when evidence is unavailable.
Question:
Evidence:
Expected behavior:
not:
π§ 78. Abstention AccuracyΒΆ
Evaluate:
Correct Abstentions
ββββββββββββββββββββ
Total No-Answer Cases
Also measure:
where the system abstains even though sufficient evidence exists.
π§ 79. Selective PredictionΒΆ
A production RAG system can choose between:
based on evidence quality.
π§ 80. CalibrationΒΆ
If a system reports:
then approximately 90% of similarly scored responses should ideally be correct under the chosen definition.
Confidence should therefore be evaluated for calibration rather than treated as truth.
π§ 81. RAG Evaluation MatrixΒΆ
| Dimension | Example Metrics |
|---|---|
| Retrieval | Recall@K, Precision@K |
| Ranking | MRR, NDCG, MAP |
| Context | Context Recall, Context Precision |
| Generation | Relevance, Completeness |
| Grounding | Faithfulness, Groundedness |
| Citation | Accuracy, Coverage, Validity |
| Safety | Policy Violations, Leakage |
| Reliability | Failure Rate, Availability |
| Performance | Latency, Throughput |
| Cost | Token Cost, Request Cost |
π§ 82. Retrieval BenchmarkΒΆ
Example:
Baseline Hybrid
Recall@5 0.72 0.86
Recall@10 0.81 0.92
MRR 0.64 0.77
NDCG@5 0.58 0.74
Latency 80ms 130ms
This makes architecture trade-offs measurable.
π§ 83. End-to-End BenchmarkΒΆ
System A System B
Answer Relevance 0.82 0.89
Faithfulness 0.84 0.93
Citation Accuracy 0.88 0.96
Completeness 0.76 0.87
p95 Latency 1.4s 1.9s
Cost / Request $0.012 $0.021
The best system is not necessarily the one with the highest quality score alone.
π§ 84. Quality vs CostΒΆ
A system may improve:
while increasing:
This may or may not be worthwhile.
Enterprise evaluation must therefore consider:
together.
π§ 85. Quality-Cost FrontierΒΆ
Quality
β
β β
β β
β β
β β
β β
ββββββββββββββββββββββ Cost
The goal is often to identify the best trade-off rather than maximize one metric blindly.
π§ 86. Latency EvaluationΒΆ
Measure:
not just average latency.
Example:
The average hides tail latency.
π§ 87. RAG Latency BreakdownΒΆ
Total Latency
β
βββ Query Processing
βββ Embedding
βββ Vector Search
βββ Keyword Search
βββ Reranking
βββ Context Compression
βββ Prompt Assembly
βββ LLM Generation
βββ Validation
βββ Citation
βββ Response Rendering
π§ 88. ThroughputΒΆ
Measure:
or:
under realistic concurrency.
Example:
π§ 89. Token EvaluationΒΆ
Track:
Example:
Large contexts can become expensive quickly.
π§ 90. Cost EvaluationΒΆ
Request cost can be decomposed:
Embedding Cost
+
Retrieval Infrastructure
+
Reranking
+
LLM Input Tokens
+
LLM Output Tokens
+
Evaluation
+
Observability
π§ 91. Evaluation CostΒΆ
Evaluation itself costs money.
For example:
may become expensive.
Therefore use a layered strategy:
π§ 92. Sampling StrategyΒΆ
Not every production request needs full expensive evaluation.
Possible approach:
The actual sampling percentages should be based on system risk and operational requirements.
π§ 93. Continuous EvaluationΒΆ
A production RAG system should be evaluated continuously.
Production Requests
β
Sampling
β
Evaluation
β
Metrics
β
Dashboard
β
Alerts
β
Improvement
π§© 94. Continuous Evaluation ArchitectureΒΆ
flowchart LR
A["Production RAG"] --> B["Evaluation Sampler"]
B --> C["Evaluation Pipeline"]
C --> D["Deterministic Metrics"]
C --> E["LLM Judge"]
C --> F["Human Review"]
D --> G["Evaluation Store"]
E --> G
F --> G
G --> H["Dashboard"]
H --> I["Alerts"]
H --> J["Model / Retrieval Improvements"] π§ 95. Evaluation StoreΒΆ
Store:
Question
Retrieved Documents
Selected Context
Answer
Citations
Metrics
Judge Scores
Latency
Cost
Model Version
Prompt Version
Retriever Version
This enables historical comparison.
π§ 96. Experiment TrackingΒΆ
Every experiment should record:
Experiment ID
Dataset Version
Model
Embedding Model
Retriever
Reranker
Prompt Version
Chunking Strategy
Top-K
Metrics
Cost
Latency
π§© 97. Experiment ConfigurationΒΆ
experiment:
name: hybrid-reranker-v2
dataset:
version: "2026-08"
retrieval:
strategy: hybrid
top_k: 20
reranking:
enabled: true
top_n: 5
generation:
model: enterprise-model
temperature: 0.1
π§ 98. Experiment ComparisonΒΆ
Experiment A
β
Dense Retrieval
β
No Reranker
Experiment B
β
Hybrid Retrieval
β
Reranker
Experiment C
β
Hybrid Retrieval
β
Reranker
β
Context Compression
Compare them using the same evaluation dataset.
π§ 99. ReproducibilityΒΆ
A benchmark should be reproducible.
Record:
Dataset Version
Model Version
Embedding Version
Prompt Version
Retriever Version
Reranker Version
Configuration
Evaluation Version
Without this, historical scores become difficult to interpret.
π§ 100. Dataset VersioningΒΆ
Example:
Dataset changes can otherwise make:
from one period incomparable with:
from another period.
π§ 101. Evaluation Configuration VersioningΒΆ
All should be recorded.
π§ 102. Regression TestingΒΆ
Every important system change should trigger evaluation.
Examples:
Retriever Changed
β
Run RAG Benchmark
Chunking Changed
β
Run RAG Benchmark
Prompt Changed
β
Run RAG Benchmark
Model Changed
β
Run RAG Benchmark
π§ 103. Regression GateΒΆ
A deployment can require:
If a metric falls below its threshold:
π§© 104. Evaluation CI/CDΒΆ
flowchart LR
A["Code Change"] --> B["Build"]
B --> C["Unit Tests"]
C --> D["RAG Evaluation"]
D --> E{"Thresholds Met?"}
E -->|Yes| F["Deploy"]
E -->|No| G["Block Deployment"] π§ 105. RAG Quality GateΒΆ
Example:
quality_gate:
retrieval:
recall_at_10: ">=0.90"
ranking:
ndcg_at_5: ">=0.75"
generation:
faithfulness: ">=0.90"
citation:
accuracy: ">=0.95"
performance:
p95_latency_ms: "<=2000"
Thresholds are illustrative and should be calibrated against business requirements.
π§ 106. Benchmark CategoriesΒΆ
A robust benchmark should contain:
Category A β Simple Retrieval
Category B β Semantic Retrieval
Category C β Metadata Filtering
Category D β Multi-Hop
Category E β Long Context
Category F β Conflicting Evidence
Category G β No Answer
Category H β Historical
Category I β Numerical
Category J β Security
π§ 107. Benchmark by DifficultyΒΆ
Level 1
Simple factual retrieval
Level 2
Multi-document retrieval
Level 3
Multi-hop reasoning
Level 4
Conflicting evidence
Level 5
Complex enterprise reasoning
π§ 108. Benchmark by DomainΒΆ
Enterprise evaluation can be separated by domain:
This can reveal domain-specific weaknesses.
π§ 109. Slice-Based EvaluationΒΆ
Overall score can hide important failures.
Example:
But:
Therefore:
Always evaluate important slices separately.
π§ 110. Slice DimensionsΒΆ
Possible slices:
Question Type
Domain
Language
User Role
Document Type
Query Length
Difficulty
Retrieval Strategy
Source Type
Tenant
π§ 111. Query Complexity EvaluationΒΆ
Measure performance across:
Short Query
Long Query
Keyword Query
Natural Language Query
Multi-Intent Query
Multi-Hop Query
Ambiguous Query
π§ 112. Language EvaluationΒΆ
For multilingual systems:
Evaluate each language separately.
A system with:
should not report only:
without exposing the language difference.
π§ 113. Document-Type EvaluationΒΆ
Evaluate:
Different document types can produce different retrieval behavior.
π§ 114. Long-Context EvaluationΒΆ
Long documents can create:
Evaluate:
π§ 115. Needle-in-a-Haystack EvaluationΒΆ
A useful benchmark places a small relevant piece of information inside a large context.
Large Document
β
βββ Noise
βββ Noise
βββ Noise
βββ Critical Evidence
βββ Noise
βββ Noise
Evaluate whether the system can retrieve and use the critical evidence.
π§ 116. Context Position EvaluationΒΆ
Place evidence at:
Then compare answer quality.
This can expose positional weaknesses in long-context generation.
π§ 117. Adversarial EvaluationΒΆ
Test:
Prompt Injection
Malicious Documents
Conflicting Sources
Fake Citations
Sensitive Documents
Irrelevant Documents
Instruction Hijacking
π 118. RAG Prompt Injection EvaluationΒΆ
Example malicious document:
Expected behavior:
π§ 119. Retrieval Poisoning EvaluationΒΆ
A malicious document may contain:
and rank highly.
Evaluation should test whether:
can reduce the impact.
π§ 120. Conflicting Evidence BenchmarkΒΆ
Example:
Evaluate whether the system:
π§ 121. Noisy Context BenchmarkΒΆ
Measure:
π§ 122. Duplicate Evidence BenchmarkΒΆ
If the same information appears repeatedly:
the system should not unnecessarily amplify it.
Evaluate:
π§ 123. Contradiction BenchmarkΒΆ
Test whether the answer introduces contradictions:
This should fail:
π§ 124. Evaluation of Agentic RAGΒΆ
Agentic RAG introduces:
Evaluate:
π§ 125. Agentic RAG MetricsΒΆ
Example:
Task Success Rate
Tool Selection Accuracy
Average Retrieval Steps
Average Tool Calls
Failed Tool Calls
Token Usage
Cost
Latency
A more capable agent is not necessarily better if it performs unnecessary retrieval loops.
π§ 126. Graph RAG EvaluationΒΆ
Graph RAG should evaluate:
Entity Retrieval
Relationship Retrieval
Path Accuracy
Subgraph Relevance
Graph Coverage
Final Answer Groundedness
π§ 127. SQL RAG EvaluationΒΆ
SQL RAG should evaluate:
SQL Generation Accuracy
SQL Safety
Query Execution Success
Result Correctness
Schema Selection
Column Selection
Final Answer Accuracy
π§ 128. Multimodal RAG EvaluationΒΆ
Evaluate:
Image Retrieval
OCR Accuracy
Table Retrieval
Visual Evidence
Cross-Modal Grounding
Citation Accuracy
Answer Accuracy
π§ 129. Production BenchmarkingΒΆ
Benchmark production-like workloads:
Do not benchmark only:
and call the system production-ready.
π§ 130. Load TestingΒΆ
Example:
Measure:
π§ 131. Stress TestingΒΆ
Push the system beyond normal expected load.
Goal:
Measure:
π§ 132. Failure TestingΒΆ
Simulate:
Vector DB unavailable
LLM unavailable
Embedding service unavailable
Reranker timeout
Network timeout
Source unavailable
Citation service failure
Evaluate:
π§ 133. Reliability MetricsΒΆ
Track:
Availability
Error Rate
Timeout Rate
Retry Rate
Fallback Rate
Abstention Rate
Validation Failure Rate
π§ 134. Evaluation DashboardΒΆ
ββββββββββββββββββββββββββββββββββββββββββββββββ
β RAG QUALITY DASHBOARD β
ββββββββββββββββββββββββββββββββββββββββββββββββ€
β Retrieval Recall@10 92.4% β
β NDCG@5 81.7% β
β Context Precision 89.3% β
β Context Recall 94.1% β
β Answer Relevance 93.8% β
β Faithfulness 96.1% β
β Citation Accuracy 97.4% β
β Citation Coverage 95.9% β
β β
β p95 Latency 1.82 sec β
β Avg Cost / Request $0.018 β
β Error Rate 0.4% β
ββββββββββββββββββββββββββββββββββββββββββββββββ
Values are illustrative only.
π§ 135. Evaluation TrendΒΆ
Track metrics over time:
Faithfulness
100% β€
95% β€ βββββββ
90% β€ ββββ
85% β€ β
80% β€
βββββββββββββββββββ
v1 v2 v3 v4 v5
A sudden decline can indicate a regression.
π§ 136. Regression DetectionΒΆ
Example:
Potential causes:
The evaluation system should make this visible.
π§ 137. Evaluation CorrelationΒΆ
Compare:
Example:
Over many examples, evaluate whether the judge tracks human judgment reliably.
π§ 138. Judge AgreementΒΆ
Possible measurements:
The exact statistical method should depend on the evaluation design and scoring scale.
π§ 139. Evaluation Rubric ExampleΒΆ
Question:
What database does Payment Service use?
Answer:
The payment service uses PostgreSQL.
Evidence:
The payment service uses PostgreSQL.
Evaluation:
Relevance:
4/4
Faithfulness:
4/4
Completeness:
4/4
Citation:
4/4
π§ 140. Evaluation RecordΒΆ
A production evaluation record can contain:
{
"evaluation_id": "EVAL-1042",
"query": "What database does Payment Service use?",
"retrieval": {
"recall_at_5": 1.0,
"precision_at_5": 0.4,
"mrr": 1.0
},
"generation": {
"faithfulness": 1.0,
"answer_relevance": 1.0,
"completeness": 1.0
},
"citation": {
"accuracy": 1.0,
"coverage": 1.0
}
}
π§ 141. Evaluation PipelineΒΆ
class RAGEvaluator:
def evaluate(
self,
query,
retrieved_documents,
context,
answer,
citations,
ground_truth
):
retrieval = self.evaluate_retrieval(
retrieved_documents,
ground_truth
)
context_score = self.evaluate_context(
context,
ground_truth
)
generation = self.evaluate_generation(
query,
context,
answer,
ground_truth
)
citation = self.evaluate_citations(
answer,
citations,
ground_truth
)
return {
"retrieval": retrieval,
"context": context_score,
"generation": generation,
"citation": citation
}
π§ 142. Evaluation InterfacesΒΆ
class GenerationEvaluator:
def evaluate(
self,
question,
context,
answer,
reference
):
raise NotImplementedError
class CitationEvaluator:
def evaluate(
self,
claims,
citations,
sources
):
raise NotImplementedError
π§ 143. Evaluation Strategy PatternΒΆ
RAGEvaluator
β
βββ RetrievalEvaluator
βββ RankingEvaluator
βββ ContextEvaluator
βββ GenerationEvaluator
βββ GroundingEvaluator
βββ CitationEvaluator
βββ SafetyEvaluator
βββ PerformanceEvaluator
Each evaluator can evolve independently.
π§ 144. Aggregating ScoresΒΆ
Do not blindly average every metric.
Example:
A simple average:
may hide a critical weakness.
For example:
may be unacceptable regardless of other scores.
π§ 145. Weighted EvaluationΒΆ
A business may define:
But safety may also be treated as a hard gate instead of a weighted metric.
π§ 146. Hard Quality GatesΒΆ
Example:
If security violations are non-zero:
This is often more appropriate than averaging safety into an overall score.
π§ 147. Evaluation ScorecardΒΆ
βββββββββββββββββββββββββββββββββββββββ
β RAG SCORECARD β
βββββββββββββββββββββββββββββββββββββββ€
β Retrieval Recall@10 92% β
β NDCG@5 82% β
β Context Recall 94% β
β Context Precision 89% β
β Faithfulness 96% β
β Answer Relevance 94% β
β Citation Accuracy 97% β
β Citation Coverage 96% β
β Safety Violations 0 β
β p95 Latency 1.8 sec β
β Cost / Request $0.018 β
βββββββββββββββββββββββββββββββββββββββ
π§ 148. Evaluation Trade-OffsΒΆ
Improving one metric can hurt another.
Example:
Increase Top-K
β
Recall β
β
Context Size β
β
Cost β
β
Latency β
β
Potential Context Noise β
Therefore:
RAG optimization is a multi-objective problem.
π§ 149. Top-K EvaluationΒΆ
Test:
Measure:
Choose K based on measured trade-offs.
π§ 150. Chunk Size EvaluationΒΆ
Benchmark:
Measure:
π§ 151. Overlap EvaluationΒΆ
Test:
chunk overlap.
Measure:
π§ 152. Embedding Model BenchmarkΒΆ
Compare:
using the same:
π§ 153. Vector Store BenchmarkΒΆ
Compare:
based on:
π§ 154. Retriever BenchmarkΒΆ
Compare:
using a common evaluation dataset.
π§ 155. Reranker BenchmarkΒΆ
Compare:
Measure:
π§ 156. Prompt BenchmarkΒΆ
Compare:
while holding constant:
This isolates the prompt effect.
π§ 157. Model BenchmarkΒΆ
Compare:
with identical:
Measure:
π§ 158. Full-System BenchmarkΒΆ
The most realistic benchmark evaluates the complete pipeline:
This measures actual production behavior.
π§ 159. Benchmark Experiment MatrixΒΆ
Retriever
β
βββββββββββΌββββββββββ
βΌ βΌ βΌ
Dense Hybrid BM25
β β β
βββββββββββΌββββββββββ
βΌ
Reranker
β
βΌ
Context
β
βΌ
LLM
β
βΌ
Evaluation
π§ 160. Evaluation and ObservabilityΒΆ
Evaluation asks:
Observability asks:
Both are required.
provide a complete production quality system.
π§ 161. Evaluation and RAG ObservabilityΒΆ
Evaluation metrics:
Observability:
Together:
π§ 162. Production Evaluation ArchitectureΒΆ
RAG SYSTEM
β
βΌ
Production Query
β
ββββββββββββββββ΄βββββββββββββββ
βΌ βΌ
Response Trace
β β
βΌ βΌ
Evaluation Observability
β β
ββββββββΌβββββββ βββββββΌββββββ
βΌ βΌ βΌ βΌ βΌ βΌ
Metrics Judge Human Latency Tokens Errors
β β β β β β
ββββββββΌβββββββ βββββββΌββββββ
βΌ βΌ
Evaluation Store Observability Store
β β
ββββββββββββββββ¬βββββββββββββββ
βΌ
Dashboard
π§ 163. Evaluation AlertsΒΆ
Alert when:
Recall drops
Faithfulness drops
Citation accuracy drops
Latency increases
Cost increases
Error rate increases
Abstention rate increases
Example:
π§ 164. Evaluation DriftΒΆ
Performance may degrade because:
Knowledge Base Changes
Model Changes
Embedding Changes
User Query Distribution Changes
Document Distribution Changes
Prompt Changes
Therefore evaluation should run continuously.
π§ 165. Data DriftΒΆ
Production queries may evolve.
Example:
The benchmark should evolve with real user behavior.
π§ 166. Query Distribution MonitoringΒΆ
Track:
Use production data to expand the benchmark.
π§ 167. Evaluation Feedback LoopΒΆ
flowchart LR
A["Production Queries"] --> B["Sampling"]
B --> C["Evaluation"]
C --> D["Quality Metrics"]
D --> E["Failure Analysis"]
E --> F["System Improvement"]
F --> G["New Benchmark Cases"]
G --> C This creates continuous improvement.
π§ 168. Failure AnalysisΒΆ
Do not stop at:
Investigate:
Which queries failed?
Why?
Retrieval failure?
Ranking failure?
Context failure?
Generation failure?
Citation failure?
π§ 169. Failure TaxonomyΒΆ
RETRIEVAL_FAILURE
RANKING_FAILURE
CONTEXT_FAILURE
GENERATION_FAILURE
GROUNDING_FAILURE
CITATION_FAILURE
SECURITY_FAILURE
POLICY_FAILURE
PERFORMANCE_FAILURE
COST_FAILURE
π§ 170. Failure Analysis ExampleΒΆ
Question:
"What database does Payment Service use?"
Retrieved:
Security Policy
Kafka Architecture
Deployment Guide
Expected:
Database Architecture
Classification:
RETRIEVAL_FAILURE
π§ 171. Failure AttributionΒΆ
A useful production framework:
Final Failure
β
βββ Retrieval?
βββ Ranking?
βββ Context?
βββ Prompt?
βββ Model?
βββ Validation?
βββ Citation?
This helps engineering teams fix the right layer.
π§ 172. RAG Evaluation LifecycleΒΆ
Define Metrics
β
Build Dataset
β
Run Baseline
β
Identify Failures
β
Improve System
β
Run Benchmark
β
Regression Test
β
Deploy
β
Monitor Production
β
Sample & Evaluate
β
Update Dataset
β
Repeat
π§ 173. Production Evaluation MaturityΒΆ
Level 1 β Manual TestingΒΆ
Level 2 β Golden DatasetΒΆ
Level 3 β Automated MetricsΒΆ
Level 4 β Continuous EvaluationΒΆ
Level 5 β Enterprise Evaluation PlatformΒΆ
π§ 174. Evaluation Maturity ModelΒΆ
Enterprise AI
β²
β
Continuous Evaluation
β
Automated Metrics
β
Golden Dataset
β
Manual QA
β
ββββββββββββββββΊ
π§ͺ 175. Practical ProjectΒΆ
Build a Production RAG Evaluation Framework.
InputΒΆ
ExecuteΒΆ
EvaluateΒΆ
OutputΒΆ
π§ͺ 176. Suggested Project StructureΒΆ
rag-evaluation/
β
βββ datasets/
β βββ golden/
β βββ regression/
β βββ benchmark/
β βββ production-samples/
β
βββ evaluators/
β βββ retrieval/
β βββ ranking/
β βββ context/
β βββ generation/
β βββ grounding/
β βββ citation/
β βββ safety/
β βββ performance/
β
βββ judges/
β βββ prompts/
β βββ schemas/
β
βββ experiments/
β
βββ reports/
β
βββ dashboards/
β
βββ pipelines/
π§ͺ 177. Evaluation ConfigurationΒΆ
dataset:
name: enterprise-rag-golden
version: "v3"
retrieval:
top_k: 10
reranking:
enabled: true
top_n: 5
generation:
temperature: 0.1
evaluation:
retrieval: true
context: true
generation: true
grounding: true
citation: true
performance: true
cost: true
π§ͺ 178. Evaluation RunnerΒΆ
class EvaluationRunner:
def run(
self,
dataset,
rag_system
):
results = []
for item in dataset:
result = rag_system.answer(
item.question
)
evaluation = self.evaluate(
item,
result
)
results.append(
evaluation
)
return results
π§ͺ 179. Evaluation ReportΒΆ
=========================================
RAG EVALUATION REPORT
=========================================
Dataset:
enterprise-rag-golden-v3
Queries:
1,000
Retrieval
-----------------------------------------
Recall@5 88.2%
Recall@10 94.1%
MRR 81.4%
NDCG@5 79.2%
Generation
-----------------------------------------
Faithfulness 95.3%
Answer Relevance 93.7%
Completeness 90.1%
Citation
-----------------------------------------
Citation Accuracy 97.2%
Citation Coverage 95.8%
Performance
-----------------------------------------
p50 Latency 0.91 sec
p95 Latency 1.82 sec
p99 Latency 3.74 sec
Cost
-----------------------------------------
Average / Request $0.018
=========================================
Values are illustrative.
π§ͺ 180. Regression ReportΒΆ
=========================================
REGRESSION TEST
=========================================
Previous Version:
v2.4
Candidate Version:
v2.5
Recall@10:
92.1% β 93.4% PASS
Faithfulness:
95.2% β 96.0% PASS
Citation Accuracy:
97.1% β 97.4% PASS
p95 Latency:
1.7s β 1.9s PASS
Cost:
$0.017 β $0.019 WARNING
=========================================
Decision:
PASS
π§ 181. Advanced Evaluation ExerciseΒΆ
Extend the evaluation platform with:
β Golden datasets
β Dataset versioning
β Retrieval metrics
β Ranking metrics
β Context metrics
β Generation metrics
β Grounding metrics
β Citation metrics
β Safety evaluation
β LLM-as-a-Judge
β Human evaluation
β Experiment tracking
β Regression testing
β Quality gates
β Production sampling
β Failure taxonomy
β Evaluation dashboards
β Evaluation alerts
β Cost analysis
β Latency analysis
β Slice-based evaluation
β Multi-language evaluation
β Multi-tenant evaluation
π§ 182. Evaluation Best PracticesΒΆ
1. Evaluate Retrieval SeparatelyΒΆ
Do not assume:
2. Evaluate Grounding SeparatelyΒΆ
A correct answer may still be unsupported by the retrieved context.
3. Evaluate Citations SeparatelyΒΆ
A cited answer can still contain incorrect citations.
4. Use a Golden DatasetΒΆ
Create stable test cases.
5. Version EverythingΒΆ
Track:
6. Use Multiple Evaluation MethodsΒΆ
7. Evaluate SlicesΒΆ
Do not rely only on one aggregate score.
8. Track Cost and LatencyΒΆ
Quality without operational feasibility is not enough.
9. Include Negative CasesΒΆ
Test:
10. Run Regression Tests Before DeploymentΒΆ
Every important RAG change should be evaluated.
π§ 183. Production Evaluation ChecklistΒΆ
β Define evaluation objectives
β Define quality metrics
β Define performance metrics
β Define cost metrics
β Build golden dataset
β Build regression dataset
β Build benchmark dataset
β Build negative examples
β Build adversarial examples
β Define retrieval ground truth
β Define expected claims
β Define expected citations
β Implement Recall@K
β Implement Precision@K
β Implement Hit Rate
β Implement MRR
β Implement MAP
β Implement NDCG
β Implement Context Recall
β Implement Context Precision
β Implement Answer Relevance
β Implement Faithfulness
β Implement Groundedness
β Implement Completeness
β Implement Citation Validity
β Implement Citation Accuracy
β Implement Citation Coverage
β Implement LLM Judge
β Calibrate LLM Judge
β Implement Human Evaluation
β Implement Latency Metrics
β Implement Throughput Metrics
β Implement Token Metrics
β Implement Cost Metrics
β Implement Experiment Tracking
β Version Evaluation Datasets
β Version Evaluation Configurations
β Implement Regression Testing
β Implement Quality Gates
β Implement CI/CD Integration
β Implement Production Sampling
β Implement Continuous Evaluation
β Implement Evaluation Dashboard
β Implement Evaluation Alerts
β Implement Failure Taxonomy
β Implement Failure Analysis
β Implement Slice-Based Evaluation
β Evaluate Security
β Evaluate Tenant Isolation
β Evaluate Prompt Injection
β Evaluate Data Leakage
β Evaluate Multilingual Queries
β Evaluate Multi-Hop Queries
β Evaluate No-Answer Queries
β Evaluate Conflicting Evidence
β Evaluate Long Context
β Track Quality vs Cost
β Track Quality vs Latency
π§ 184. Final Production ArchitectureΒΆ
PRODUCTION RAG
β
βΌ
βββββββββββββββββ
β User Request β
βββββββββ¬ββββββββ
β
βΌ
RAG PIPELINE
β
βββββββββββββββββΌββββββββββββββββ
βΌ βΌ βΌ
Retrieval Context Generation
β β β
βββββββββββββββββΌββββββββββββββββ
βΌ
Final Response
β
ββββββββββββββββββΌβββββββββββββββββ
βΌ βΌ βΌ
Retrieval Generation Citation
Evaluation Evaluation Evaluation
β β β
ββββββββββββββββββΌβββββββββββββββββ
βΌ
Quality Aggregator
β
ββββββββββββββΌβββββββββββββ
βΌ βΌ βΌ
Metrics Judge Human
β β β
ββββββββββββββΌβββββββββββββ
βΌ
Evaluation Store
β
βΌ
Dashboard
β
ββββββββββββ΄βββββββββββ
βΌ βΌ
Alerts Analysis
β β
ββββββββββββ¬βββββββββββ
βΌ
System Improvement
β
βΌ
New Benchmark
β
ββββββββββββΊ
π§ 185. Final Mental ModelΒΆ
RAG QUALITY
β
βββββββββββββββββββββββΌββββββββββββββββββββββ
βΌ βΌ βΌ
RETRIEVAL GENERATION OPERATIONS
β β β
ββ Recall ββ Relevance ββ Latency
ββ Precision ββ Faithfulness ββ Throughput
ββ MRR ββ Groundedness ββ Cost
ββ NDCG ββ Completeness ββ Reliability
ββ Hit Rate ββ Safety ββ Scalability
β β β
βββββββββββββββββββββββΌββββββββββββββββββββββ
βΌ
CITATIONS
β
ββββββββββΌβββββββββ
βΌ βΌ βΌ
Validity Accuracy Coverage
β
βΌ
END-TO-END QUALITY
β
βΌ
PRODUCTION DECISION
β
βββββββββββββΌββββββββββββ
βΌ βΌ βΌ
Deploy Improve Reject
The fundamental production loop is:
Build
β
Measure
β
Analyze
β
Improve
β
Benchmark
β
Validate
β
Deploy
β
Monitor
β
Evaluate
β
Repeat
A RAG system should therefore never be considered "good" simply because it produces convincing answers.
A production-grade RAG system is one whose retrieval quality, evidence grounding, response quality, citation correctness, security, latency, cost, and reliability can all be measured and continuously improved.
π 186. Key TakeawaysΒΆ
- RAG evaluation must cover the complete retrieval-to-response pipeline.
- Retrieval quality and generation quality should be evaluated separately.
- Recall@K measures how much relevant evidence was retrieved.
- Precision@K measures how much retrieved evidence was relevant.
- Hit Rate measures whether at least one relevant result was retrieved.
- MRR measures the rank of the first relevant result.
- MAP evaluates ranking quality across multiple relevant results.
- NDCG evaluates graded relevance while considering ranking position.
- Context Recall measures whether required evidence is present.
- Context Precision measures whether unnecessary context is minimized.
- Answer Relevance measures whether the response addresses the question.
- Faithfulness measures whether the response appropriately uses supplied context.
- Groundedness measures whether claims are supported by evidence.
- Completeness measures whether required information was actually addressed.
- Citation validity, accuracy, and coverage are separate dimensions.
- Source authority and freshness can materially affect enterprise answer quality.
- LLM-as-a-Judge can scale semantic evaluation but must itself be calibrated.
- Human evaluation remains valuable for complex and high-risk cases.
- Deterministic metrics should be preferred where exact evaluation is possible.
- Numeric values, dates, identifiers, and versions require specialized validation.
- No-answer and abstention cases are important benchmark scenarios.
- Multi-hop, Graph RAG, SQL RAG, Agentic RAG, and Multimodal RAG require specialized evaluation.
- Security and tenant isolation should be part of RAG evaluation.
- Evaluation datasets must be versioned.
- Evaluation configurations must be versioned.
- Experiments should be reproducible.
- Regression testing should be integrated into CI/CD.
- Quality gates can prevent degraded RAG systems from reaching production.
- Production sampling enables continuous evaluation.
- Slice-based evaluation exposes weaknesses hidden by aggregate metrics.
- Quality should be evaluated alongside latency and cost.
- Evaluation itself has an operational cost and should therefore use layered strategies.
- Failure analysis should identify the specific RAG component responsible for poor results.
- Continuous evaluation creates a feedback loop between production behavior and system improvement.
- The goal is not a single "RAG score."
- The goal is a measurable, explainable, reproducible, continuously improving RAG system.
π§ Production RAG Evaluation Mental ModelΒΆ
βββββββββββββββββββββββββββ
β USER QUERY β
ββββββββββββββ¬βββββββββββββ
β
βΌ
ββββββββββββββββββββ
β RETRIEVAL β
ββββββββββ¬ββββββββββ
β
Evaluate Recall
Precision / MRR
NDCG / Hit Rate
β
βΌ
ββββββββββββββββββββ
β CONTEXT β
ββββββββββ¬ββββββββββ
β
Evaluate Context
Recall / Precision
β
βΌ
ββββββββββββββββββββ
β GENERATION β
ββββββββββ¬ββββββββββ
β
Evaluate Relevance
Faithfulness
Groundedness
Completeness
β
βΌ
ββββββββββββββββββββ
β CITATION β
ββββββββββ¬ββββββββββ
β
Evaluate Validity
Accuracy / Coverage
β
βΌ
ββββββββββββββββββββ
β SECURITY β
ββββββββββ¬ββββββββββ
β
Evaluate Authorization
Leakage / Injection
β
βΌ
ββββββββββββββββββββ
β OPERATIONS β
ββββββββββ¬ββββββββββ
β
Evaluate Latency
Throughput / Cost
β
βΌ
ββββββββββββββββββββ
β END-TO-END β
β QUALITY β
ββββββββββ¬ββββββββββ
β
βΌ
ββββββββββββββββββββ
β CONTINUOUS LOOP β
ββββββββββ¬ββββββββββ
β
βΌ
Improve System
π§ Chapter NavigationΒΆ
Part V β Advanced Retrieval-Augmented GenerationΒΆ
Previous:
05. Enterprise Response
Next:
07. RAG Observability
Section:
06 β Production RAG Engineering
Production RAG Engineering PathΒΆ
01 Prompt Assembly
β
02 Context Selection & Context Engineering
β
03 Response Validation
β
04 Citation & Source Attribution
β
05 Enterprise Response
β
06 RAG Evaluation & Benchmarking
β
07 RAG Observability
β
08 RAG Performance Optimization
β
09 RAG Cost Optimization
β
10 Production Retrieval Architecture
β
11 Building Production RAG Systems
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.