19 β RAG Evaluation FundamentalsΒΆ
Learn how to systematically evaluate Retrieval-Augmented Generation (RAG) systems across retrieval quality, context quality, answer quality, grounding, citations, latency, cost, and overall user experience.
π OverviewΒΆ
Building a RAG pipeline is only the beginning.
A system can successfully execute:
and still produce poor results.
For example:
Therefore, production RAG systems require systematic evaluation.
The central question is not:
"Did the pipeline run successfully?"
It is:
"Did the system retrieve the right evidence and use that evidence to produce a correct, relevant, grounded, and useful answer?"
1. Why RAG Evaluation MattersΒΆ
A RAG system contains multiple stages:
Query
β
Query Processing
β
Embedding
β
Retrieval
β
Filtering
β
Ranking
β
Context Assembly
β
Prompt
β
LLM
β
Answer
A failure at any stage can affect the final response.
For example:
Therefore, evaluating only the final answer is insufficient.
2. RAG Evaluation DimensionsΒΆ
A practical evaluation model is:
RAG Evaluation
βββββββββββββββββββββββββββ
β Retrieval Quality β
ββββββββββββββ¬βββββββββββββ
β
βββββββββββββββββββββββββββ
β Context Quality β
ββββββββββββββ¬βββββββββββββ
β
βββββββββββββββββββββββββββ
β Generation Quality β
ββββββββββββββ¬βββββββββββββ
β
βββββββββββββββββββββββββββ
β System Quality β
βββββββββββββββββββββββββββ
The major dimensions are:
3. Retrieval vs Generation EvaluationΒΆ
One of the most important distinctions is:
versus:
Example:
Question:
"What is the annual leave entitlement?"
Retrieved:
"Employees receive 25 days."
Generated:
"Employees receive 30 days."
Retrieval succeeded.
Generation failed.
4. Retrieval Failure ExampleΒΆ
Question:
"What is the annual leave entitlement?"
Retrieved:
"Employees receive 10 sick days."
Generated:
"Employees receive 10 sick days."
The LLM may have correctly summarized the supplied context.
But the RAG system still failed.
The problem is:
5. RAG Evaluation PipelineΒΆ
flowchart TD
A["Evaluation Dataset"] --> B["RAG Pipeline"]
B --> C["Retrieved Context"]
B --> D["Generated Answer"]
A --> E["Expected Evidence"]
A --> F["Reference Answer"]
C --> G["Retrieval Evaluation"]
C --> H["Context Evaluation"]
D --> I["Generation Evaluation"]
D --> J["Grounding Evaluation"]
D --> K["Citation Evaluation"]
G --> L["Evaluation Report"]
H --> L
I --> L
J --> L
K --> L
B --> M["Latency / Cost Metrics"]
M --> L 6. Evaluation DatasetΒΆ
A RAG evaluation requires representative questions.
Example:
[
{
"question": "What is the annual leave entitlement?",
"expected_source": "employee-handbook",
"expected_section": "Annual Leave",
"reference_answer": "Employees receive 25 days of annual leave."
},
{
"question": "How should annual leave be requested?",
"expected_source": "employee-handbook",
"expected_section": "Annual Leave",
"reference_answer": "Employees should submit requests through the employee portal."
}
]
The dataset becomes the foundation for repeatable evaluation.
7. What an Evaluation Dataset Should ContainΒΆ
A useful evaluation record may include:
Question
Expected Answer
Expected Source
Expected Section
Expected Evidence
Metadata Filters
Difficulty
Question Type
Example:
{
"id": "rag-001",
"question": "What is the annual leave entitlement?",
"expected_answer": "25 days",
"expected_source": "employee-handbook",
"expected_section": "Annual Leave",
"category": "policy"
}
8. Evaluation Dataset TypesΒΆ
A production dataset should contain different question types.
Fact Questions
Policy Questions
Multi-Sentence Questions
Comparison Questions
Procedural Questions
Ambiguous Questions
Unanswerable Questions
For example:
"What is the annual leave entitlement?"
"How do I request annual leave?"
"Can unused leave be carried forward?"
"What is the policy for contractors?"
"Does the handbook mention remote work?"
9. Answerable vs Unanswerable QuestionsΒΆ
A strong evaluation dataset should include questions for which the knowledge base contains no answer.
Example:
Expected behavior:
The system should not fabricate an answer.
10. Retrieval EvaluationΒΆ
Retrieval evaluation asks:
Did the retriever return the relevant evidence?
Important metrics include:
These metrics measure different aspects of retrieval quality.
11. Precision@KΒΆ
Precision@K measures how many of the top K retrieved results are relevant.
Conceptually:
For example:
A higher value means the retrieved result set contains more relevant documents.
12. Precision@K FormulaΒΆ
Example:
Precision is especially useful when irrelevant context is expensive.
13. Recall@KΒΆ
Recall@K measures whether the relevant evidence was retrieved within the top K results.
Conceptually:
Recall@K =
Relevant Results Retrieved in Top-K
-----------------------------------
Total Relevant Results
Example:
14. Recall@K FormulaΒΆ
Recall is important because:
If the correct evidence is never retrieved, the generation stage cannot use it.
15. Precision vs RecallΒΆ
Precision
β
"How much of what I retrieved was relevant?"
Recall
β
"How much of the relevant information did I retrieve?"
Example:
Retriever A:
5 results
4 relevant
High Precision
Retriever B:
10 results
4 relevant
Lower Precision
Higher opportunity for Recall
The correct balance depends on the application.
16. Hit RateΒΆ
Hit Rate measures whether at least one relevant result appears within the top K.
Example:
The query is a hit.
If no relevant result appears:
17. Mean Reciprocal RankΒΆ
MRR considers the position of the first relevant result.
If the first relevant document appears at:
the reciprocal rank is:
If it appears at:
then:
18. MRR FormulaΒΆ
MRR is useful when the position of the first relevant result matters.
19. NDCGΒΆ
NDCG stands for:
It considers both:
and:
A highly relevant result near the top contributes more than the same result appearing much later.
NDCG is useful when relevance can have multiple levels.
For example:
20. Retrieval Metrics SummaryΒΆ
| Metric | Main Question |
|---|---|
| Precision@K | How many retrieved results are relevant? |
| Recall@K | Did we retrieve the relevant evidence? |
| Hit Rate | Did at least one relevant result appear? |
| MRR | How early is the first relevant result? |
| NDCG | Are highly relevant results ranked near the top? |
No single metric completely describes retrieval quality.
21. Context EvaluationΒΆ
Retrieval can succeed while context assembly fails.
Example:
Therefore context should also be evaluated.
Important dimensions include:
22. Context RelevanceΒΆ
Context relevance asks:
Does the retrieved context actually help answer the question?
Example:
Question:
"What is the annual leave entitlement?"
Context:
"Employees receive 25 days of annual leave."
High relevance.
But:
23. Context CompletenessΒΆ
Context completeness asks:
Does the retrieved context contain enough information to answer the question?
Example:
Question:
"How many leave days can be carried forward?"
Context:
"Employees may carry forward unused leave."
If the number of days is missing, the context may be incomplete.
24. Context NoiseΒΆ
Too much irrelevant context can reduce generation quality.
The goal is not:
but:
25. Context Quality PipelineΒΆ
flowchart LR
A["Retrieved Chunks"] --> B["Relevance"]
B --> C["Completeness"]
C --> D["Deduplication"]
D --> E["Token Budget"]
E --> F["Final Context"] 26. Generation EvaluationΒΆ
Generation evaluation asks:
Did the LLM produce a useful answer from the supplied evidence?
Important dimensions include:
27. Answer CorrectnessΒΆ
Answer correctness asks:
Does the generated answer correctly answer the question?
Example:
Question:
"What is the annual leave entitlement?"
Expected:
25 days
Generated:
Employees receive 25 days of annual leave.
Result:
Correct
28. Answer RelevanceΒΆ
An answer can be factually correct but still poorly focused.
Question:
Answer:
Employees receive 25 days of annual leave.
The company also has policies for sick leave,
remote work, travel expenses, and office security...
The answer begins correctly but contains unnecessary information.
Answer relevance measures whether the response directly addresses the user's question.
29. FaithfulnessΒΆ
Faithfulness asks:
Is the generated answer supported by the retrieved context?
Example:
Context:
Employees receive 25 days of annual leave.
Answer:
Employees receive 25 days of annual leave.
Faithful
But:
Context:
Employees receive 25 days of annual leave.
Answer:
Employees receive 30 days of annual leave.
Not faithful
30. GroundednessΒΆ
Groundedness is closely related to faithfulness.
A grounded answer should be supported by the evidence provided to the LLM.
If the answer contains unsupported claims:
31. Hallucination in RAGΒΆ
RAG reduces the need for the model to rely entirely on pretrained knowledge, but it does not eliminate hallucinations.
Example:
Context:
Employees receive 25 days of annual leave.
Answer:
Employees receive 25 days and can
automatically carry forward 15 days.
If the context does not mention the carry-forward limit, that claim is unsupported.
32. Groundedness EvaluationΒΆ
A practical evaluation process can be:
Generated Answer
β
Split into Claims
β
Check Each Claim
β
Supported by Context?
β
Groundedness Score
Example:
33. Claim-Level EvaluationΒΆ
Suppose the answer is:
Employees receive 25 days of annual leave.
Requests must be submitted through the portal.
Unused leave may be carried forward.
Break it into:
Then evaluate each claim independently.
This gives a more detailed view than simply marking the entire answer correct or incorrect.
34. Citation EvaluationΒΆ
Enterprise RAG systems often return citations.
Example:
Citation evaluation should verify:
Citation Exists
Citation Points to Correct Source
Citation Supports the Claim
Citation Is Not Fabricated
35. Citation AccuracyΒΆ
Bad example:
If page 90 contains unrelated content:
The answer may still be correct, but the source attribution is wrong.
36. Citation CoverageΒΆ
If an answer contains several factual claims:
citation coverage is incomplete.
A production system should determine which claims require source attribution.
37. Answer CompletenessΒΆ
An answer can be correct but incomplete.
Question:
Context:
Answer:
Correct, but incomplete.
A better answer:
38. Answer Quality ModelΒΆ
Answer Quality
ββββββββββββββββββ
β Correctness β
βββββββββ¬βββββββββ
β
ββββββββββββββββββ
β Relevance β
βββββββββ¬βββββββββ
β
ββββββββββββββββββ
β Groundedness β
βββββββββ¬βββββββββ
β
ββββββββββββββββββ
β Completeness β
βββββββββ¬βββββββββ
β
ββββββββββββββββββ
β Citation β
ββββββββββββββββββ
A strong answer should perform well across these dimensions.
39. Human EvaluationΒΆ
Human evaluation remains important.
A reviewer can score an answer on:
For example:
Score: 1 β Poor
Score: 2 β Weak
Score: 3 β Acceptable
Score: 4 β Good
Score: 5 β Excellent
40. Human Evaluation RubricΒΆ
| Dimension | Question |
|---|---|
| Correctness | Is the answer factually correct? |
| Relevance | Does it answer the question directly? |
| Groundedness | Is it supported by retrieved evidence? |
| Completeness | Does it include important information? |
| Clarity | Is it easy to understand? |
| Citation Quality | Are sources accurate and useful? |
41. LLM-as-a-JudgeΒΆ
An LLM can also evaluate generated responses.
Conceptually:
Example evaluation prompt:
Evaluate whether the answer is supported
by the provided context.
Question:
{question}
Context:
{context}
Answer:
{answer}
Return:
{
"grounded": true,
"score": 0.95,
"reason": "The answer is directly supported by the context."
}
LLM-based evaluation should itself be validated.
42. Human Evaluation vs Automated EvaluationΒΆ
Human EvaluationΒΆ
Advantages:
Limitations:
Automated EvaluationΒΆ
Advantages:
Limitations:
A production evaluation strategy often combines both.
43. Reference-Based EvaluationΒΆ
If a reference answer exists:
This can evaluate:
However, exact string matching is usually insufficient for natural-language answers.
44. Semantic Answer EvaluationΒΆ
Two answers may express the same meaning differently.
Reference:
Generated:
String comparison would show a difference.
Semantic evaluation can recognize that the answers convey the same information.
45. Exact MatchΒΆ
Exact Match is useful when the expected output has a deterministic form.
Example:
Generated:
Exact match succeeds.
It is less suitable for open-ended answers.
46. Token-Level MetricsΒΆ
Traditional NLP metrics can compare generated text against references.
Examples include:
However, they should not be treated as complete measures of RAG quality.
A response can use different words while still being correct.
47. RAG-Specific EvaluationΒΆ
RAG evaluation should focus on:
rather than relying solely on text-overlap metrics.
48. Evaluation FrameworksΒΆ
RAG evaluation can be implemented manually or using specialized tools.
Examples include:
Frameworks can help automate metrics and evaluation workflows.
The important principle is:
Understand the metric before relying on a framework to calculate it.
49. Evaluation Without a FrameworkΒΆ
A simple custom evaluator can be implemented.
def evaluate_answer(
expected_answer,
generated_answer
):
return {
"contains_expected": (
expected_answer.lower()
in generated_answer.lower()
)
}
This is intentionally simple.
Production evaluation should use richer semantic and grounding checks.
50. Retrieval Evaluation ExampleΒΆ
def evaluate_retrieval(
retrieved_ids,
expected_ids,
k=5
):
retrieved = set(
retrieved_ids[:k]
)
expected = set(
expected_ids
)
relevant = (
retrieved.intersection(expected)
)
precision = (
len(relevant) / k
if k
else 0
)
recall = (
len(relevant) / len(expected)
if expected
else 0
)
return {
"precision_at_k": precision,
"recall_at_k": recall
}
This demonstrates the basic idea behind retrieval evaluation.
51. Evaluation RecordΒΆ
A useful evaluation result might look like:
{
"question_id": "rag-001",
"retrieval": {
"precision_at_5": 0.8,
"recall_at_5": 1.0,
"hit": true
},
"generation": {
"correct": true,
"grounded": true,
"relevant": true
},
"citation": {
"accurate": true
},
"latency_ms": 910
}
This makes results machine-readable.
52. Batch EvaluationΒΆ
Instead of evaluating one question:
run the complete evaluation dataset.
results = []
for item in evaluation_dataset:
result = rag_pipeline.answer(
item["question"]
)
evaluation = evaluate(
item,
result
)
results.append(
evaluation
)
This allows regression testing.
53. Evaluation DashboardΒΆ
A production team can track:
Retrieval Recall@K
Retrieval Precision@K
Hit Rate
Answer Correctness
Groundedness
Citation Accuracy
Latency
Cost
Failure Rate
Example:
Retrieval Recall@5 0.91
Retrieval Precision@5 0.78
Groundedness 0.94
Answer Correctness 0.89
Citation Accuracy 0.97
P95 Latency 1.8s
These numbers are illustrative.
54. Evaluation TrendΒΆ
Evaluation should be continuous.
flowchart LR
A["Baseline"] --> B["Change"]
B --> C["Run Evaluation"]
C --> D["Compare Metrics"]
D --> E["Accept / Reject"]
E --> F["Production"]
F --> G["Monitor"]
G --> B This turns RAG evaluation into an engineering feedback loop.
55. Regression TestingΒΆ
Suppose the current system has:
After changing the chunking strategy:
The change should be investigated before deployment.
Similarly:
could indicate that the new prompt or context strategy is causing problems.
56. Evaluation Before and After ChangesΒΆ
Common changes requiring evaluation include:
Embedding Model
Chunk Size
Chunk Overlap
Metadata Filters
Top-K
Similarity Threshold
Prompt
LLM Model
Context Formatting
Vector Database
Retrieval Strategy
Every change can affect downstream quality.
57. Evaluation MatrixΒΆ
| Change | Retrieval Impact | Generation Impact |
|---|---|---|
| Embedding Model | High | Indirect |
| Chunk Size | High | High |
| Chunk Overlap | Medium | Medium |
| Top-K | High | High |
| Similarity Threshold | High | High |
| Prompt | Low | High |
| LLM Model | None | High |
| Context Formatting | Low | High |
| Vector Database | Potentially High | Indirect |
58. Evaluation by Question TypeΒΆ
Different questions may produce different results.
Example:
Fact Questions
β
High Accuracy
Multi-hop Questions
β
Lower Accuracy
Ambiguous Questions
β
Variable Accuracy
Unanswerable Questions
β
Hallucination Risk
Therefore aggregate scores alone can hide important weaknesses.
59. Segment EvaluationΒΆ
Evaluate by categories.
Policy Questions
Accuracy = 95%
Procedural Questions
Accuracy = 89%
Unanswerable Questions
Safe Response = 96%
Multi-document Questions
Accuracy = 78%
This reveals where improvements are needed.
60. Difficulty LevelsΒΆ
Evaluation questions can also be classified:
Example:
EasyΒΆ
MediumΒΆ
HardΒΆ
Difficulty-based evaluation helps identify system limitations.
61. Negative TestingΒΆ
RAG systems should be tested with intentionally difficult or unsupported queries.
Examples:
Unknown Topic
Wrong Department
Outdated Policy
Unauthorized Information
Ambiguous Query
Contradictory Documents
Expected behavior should be defined beforehand.
62. Security EvaluationΒΆ
Security should be evaluated separately.
Example:
If the user does not have access:
The test should verify that restricted content never reaches the generation layer.
63. Evaluation of Empty RetrievalΒΆ
Test:
Expected:
Not:
This is an important production safety test.
64. Evaluation of Citation IntegrityΒΆ
Test:
If not:
Citation integrity is particularly important for:
applications.
65. Latency EvaluationΒΆ
RAG latency should be measured end-to-end.
Measure:
rather than relying only on average latency.
66. Why P95 MattersΒΆ
Suppose:
Most users may receive fast responses, but a significant minority experience much slower requests.
Production systems should therefore monitor latency distributions.
67. Cost EvaluationΒΆ
A RAG request can incur:
Track cost per:
where useful.
68. Cost vs QualityΒΆ
Reducing cost aggressively can reduce answer quality.
Example:
versus:
The goal is not minimum cost.
It is:
69. Evaluation ScorecardΒΆ
A practical scorecard can be:
Retrieval
Recall@5
Precision@5
MRR
Context
Relevance
Completeness
Generation
Correctness
Groundedness
Relevance
Citations
Accuracy
Coverage
System
P95 Latency
Cost / Query
Error Rate
70. Example ScorecardΒΆ
| Category | Metric | Example |
|---|---|---|
| Retrieval | Recall@5 | 0.91 |
| Retrieval | Precision@5 | 0.82 |
| Retrieval | MRR | 0.88 |
| Context | Relevance | 0.93 |
| Generation | Correctness | 0.90 |
| Generation | Groundedness | 0.95 |
| Citations | Accuracy | 0.97 |
| System | P95 Latency | 1.8s |
The values are illustrative.
71. Weighted EvaluationΒΆ
Some applications may combine metrics.
For example:
Overall Score =
0.30 Γ Retrieval
+
0.30 Γ Groundedness
+
0.20 Γ Correctness
+
0.10 Γ Citation
+
0.10 Γ Relevance
The weights should reflect business priorities.
For high-risk applications, correctness and grounding may deserve greater importance.
72. Why a Single Score Is DangerousΒΆ
A single score can hide failures.
Example:
but:
The system may still be unsuitable for an enterprise application requiring reliable citations.
Always inspect individual metrics.
73. Evaluation ArchitectureΒΆ
flowchart TD
A["Evaluation Dataset"] --> B["RAG System"]
B --> C["Retrieval Results"]
B --> D["Generated Answers"]
B --> E["Telemetry"]
C --> F["Retrieval Metrics"]
C --> G["Context Metrics"]
D --> H["Correctness"]
D --> I["Groundedness"]
D --> J["Relevance"]
D --> K["Citation Accuracy"]
E --> L["Latency"]
E --> M["Cost"]
E --> N["Reliability"]
F --> O["Evaluation Report"]
G --> O
H --> O
I --> O
J --> O
K --> O
L --> O
M --> O
N --> O 74. Production Evaluation WorkflowΒΆ
1. Define evaluation goals.
2. Build representative questions.
3. Identify expected evidence.
4. Define reference answers where appropriate.
5. Run the RAG system.
6. Capture retrieved documents.
7. Measure retrieval quality.
8. Evaluate context quality.
9. Evaluate generated answers.
10. Measure groundedness.
11. Validate citations.
12. Measure latency.
13. Measure cost.
14. Segment results by question type.
15. Compare against baseline.
16. Investigate regressions.
17. Approve or reject changes.
18. Continuously monitor production behavior.
75. Baseline EvaluationΒΆ
Before optimizing the system, establish a baseline.
Example:
Baseline:
Future changes should be compared against this baseline.
76. Experiment TrackingΒΆ
A RAG experiment should record:
Embedding Model
Chunk Size
Chunk Overlap
Top-K
Similarity Threshold
Vector Store
Prompt Version
LLM Model
Evaluation Dataset Version
Evaluation Metrics
Example:
{
"experiment": "rag-v12",
"embedding_model": "embedding-v3",
"chunk_size": 500,
"chunk_overlap": 50,
"top_k": 5,
"llm": "enterprise-model-x",
"prompt_version": "prompt-v4"
}
This enables reproducible experimentation.
77. Evaluation Dataset VersioningΒΆ
The dataset itself should be versioned.
If questions change, metric comparisons may no longer be directly comparable.
Therefore:
should be recorded together.
78. Production Monitoring vs Offline EvaluationΒΆ
These are complementary.
Offline EvaluationΒΆ
Production MonitoringΒΆ
Offline evaluation provides controlled comparison.
Production monitoring reveals actual system behavior.
79. Continuous EvaluationΒΆ
flowchart LR
A["Code Change"] --> B["Offline Evaluation"]
B --> C{"Pass?"}
C -->|No| D["Reject Change"]
C -->|Yes| E["Deploy"]
E --> F["Production Monitoring"]
F --> G["New Evaluation Data"]
G --> B This creates a continuous quality loop.
80. Evaluation AlertsΒΆ
Production systems can alert when:
Groundedness drops
Retrieval recall drops
Empty retrieval increases
Latency increases
Cost increases
Citation failures increase
Error rate increases
Example:
This should trigger investigation.
81. Root-Cause AnalysisΒΆ
When quality decreases, inspect the pipeline stage.
Do not immediately change the model.
The failure may be caused by:
82. Example Root-Cause InvestigationΒΆ
Suppose:
Check:
This suggests:
rather than an LLM generation problem.
83. Another Root-Cause ExampleΒΆ
Suppose:
Retrieval remains strong.
The problem may be:
This illustrates why multiple metrics are required.
84. Evaluation Anti-PatternsΒΆ
84.1 Only Measuring Final AnswersΒΆ
This hides retrieval failures.
84.2 Using Only One QuestionΒΆ
One question cannot represent an enterprise workload.
84.3 Using Only Exact String MatchingΒΆ
Equivalent answers may use different wording.
84.4 Ignoring Unanswerable QuestionsΒΆ
This hides hallucination behavior.
84.5 Ignoring CitationsΒΆ
Correct answers can still have incorrect sources.
84.6 Ignoring LatencyΒΆ
A high-quality system can still be unusable if responses are too slow.
84.7 Ignoring CostΒΆ
An accurate system may still be economically impractical.
84.8 Changing Multiple Variables at OnceΒΆ
If you change:
at the same time, it becomes difficult to identify what caused the improvement or regression.
85. Controlled ExperimentsΒΆ
Prefer:
Example:
Keep other major variables stable.
86. Evaluation of Chunk SizeΒΆ
Example experiment:
The best chunk size is not necessarily the largest or smallest.
It should be selected based on evaluation results and workload characteristics.
87. Evaluation of Top-KΒΆ
Example:
Increasing K improves retrieval coverage but may introduce context noise.
This demonstrates why retrieval and generation metrics should be evaluated together.
88. Quality Trade-OffΒΆ
The optimal value is application-specific.
89. Evaluation as an Engineering DisciplineΒΆ
RAG evaluation should be treated like software testing.
RAG quality should not depend on manual experimentation alone.
90. Test Pyramid for RAGΒΆ
βββββββββββββββββ
β Human Review β
βββββββββββββββββ
βββββββββββββββββββββββ
β End-to-End Evaluationβ
βββββββββββββββββββββββ
βββββββββββββββββββββββββββββ
β Retrieval / Generation β
βββββββββββββββββββββββββββββ
βββββββββββββββββββββββββββββββββββ
β Unit / Integration Tests β
βββββββββββββββββββββββββββββββββββ
The lower levels should provide fast feedback.
The higher levels provide broader quality validation.
91. Evaluation Strategy by StageΒΆ
| Stage | Example Evaluation |
|---|---|
| Document Processing | Extraction accuracy |
| Chunking | Context completeness |
| Embeddings | Retrieval benchmark |
| Retrieval | Recall@K / Precision@K |
| Ranking | MRR / NDCG |
| Context | Relevance / completeness |
| Generation | Correctness / relevance |
| Grounding | Faithfulness |
| Citations | Source accuracy |
| System | Latency / cost / reliability |
92. RAG Quality LifecycleΒΆ
Build
β
Evaluate
β
Identify Weakness
β
Change One Component
β
Evaluate Again
β
Compare
β
Deploy
β
Monitor
β
Collect New Cases
β
Expand Dataset
β
Evaluate Again
This turns evaluation into a continuous improvement process.
93. Key TakeawaysΒΆ
- RAG evaluation must measure more than whether the pipeline executes successfully.
- Retrieval and generation should be evaluated separately.
- Retrieval evaluation determines whether the right evidence was found.
- Precision@K measures how many retrieved results are relevant.
- Recall@K measures whether relevant evidence was retrieved.
- Hit Rate measures whether at least one relevant result was retrieved.
- MRR evaluates the position of the first relevant result.
- NDCG evaluates both relevance and ranking position.
- Context evaluation measures relevance, completeness, and noise.
- Generation evaluation measures correctness, relevance, completeness, and clarity.
- Faithfulness and groundedness measure whether answers are supported by retrieved evidence.
- Citation evaluation verifies that sources are accurate and actually support generated claims.
- Evaluation datasets should contain representative, difficult, and unanswerable questions.
- Human evaluation provides nuanced judgment but is expensive.
- Automated evaluation provides scale and repeatability but requires validation.
- LLM-as-a-judge can help evaluate semantic qualities but should not be blindly trusted.
- Exact-match metrics are useful for deterministic outputs but are insufficient for open-ended RAG answers.
- Offline evaluation and production monitoring serve different purposes and should complement each other.
- Evaluation datasets should be versioned.
- System configurations should be tracked with evaluation results.
- Baselines are essential for measuring improvements and regressions.
- Changes to chunking, embeddings, retrieval, prompts, or LLMs should be evaluated before production rollout.
- Aggregate scores can hide important weaknesses, so metrics should also be segmented by question type.
- RAG evaluation should include latency, cost, reliability, and security behavior.
- Empty retrieval and unanswerable questions are important tests for hallucination resistance.
- Evaluation should become part of the RAG engineering lifecycle rather than a final one-time activity.
The central principle is:
A production RAG system should be evaluated as a complete pipeline: retrieve the right evidence, construct useful context, generate a correct and grounded answer, provide trustworthy citations, and meet the required latency, reliability, security, and cost targets.
94. Chapter NavigationΒΆ
Part IV β Prompt Engineering & RAG FundamentalsΒΆ
Previous Chapter: 18. Building Your First RAG Pipeline
Current Chapter: 19 β RAG Evaluation Fundamentals
Next Chapter: 20. Enterprise Generative AI Application Architecture
Part IV ChaptersΒΆ
- 01. Introduction to Prompt Engineering
- 02. Prompt Engineering Fundamentals
- 03. Advanced Prompt Engineering
- 04. Prompt Design Patterns
- 05. Zero-shot, One-shot & Few-shot Prompting
- 06. Chain-of-Thought Prompting
- 07. ReAct Prompting
- 08. Structured Outputs & Output Parsing
- 09. Function Calling & Tool Calling
- 10. Embeddings in Practice
- 11. Document Processing & Vectorization
- 12. Document Chunking Strategies
- 13. Vector Database Fundamentals
- 14. Similarity Search Techniques
- 15. RAG Pipeline Components
- 16. Retrieval and Generation Pipeline
- 17. Vector Databases in RAG
- 18. Building Your First RAG Pipeline
- 19. RAG Evaluation Fundamentals
- 20. Enterprise Generative AI Application Architecture
- 21. Deploying AI Applications with Gradio
ReferencesΒΆ
- Retrieval-Augmented Generation evaluation research
- Information retrieval evaluation documentation
- Precision and Recall documentation
- MRR and NDCG evaluation documentation
- RAG evaluation methodology documentation
- Ragas documentation
- DeepEval documentation
- TruLens documentation
- LangChain evaluation documentation
- LlamaIndex evaluation documentation
- LLM evaluation documentation
- Enterprise AI evaluation and observability documentation
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.