Retriever ComparisonΒΆ
π OverviewΒΆ
Retrieval is one of the most important components of a Retrieval-Augmented Generation (RAG) system.
A basic RAG pipeline may begin with a simple VectorStore Retriever, but production systems often require different retrieval strategies depending on the nature of the query, document structure, metadata, and business requirements.
In the previous chapters, we explored:
01 β VectorStore Retriever
02 β Multi-Query Retriever
03 β Self-Query Retriever
04 β Parent-Document Retriever
Each technique solves a different retrieval problem.
The goal of this chapter is to understand when to use each retriever, what problem it solves, its trade-offs, and how these techniques can be composed into production retrieval architectures.
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand the role of different retriever strategies
- Compare VectorStore, Multi-Query, Self-Query, and Parent-Document retrieval
- Understand the strengths and limitations of each approach
- Select an appropriate retriever for a given problem
- Understand retrieval precision vs recall trade-offs
- Understand how retrievers can be composed
- Design a retrieval decision strategy
- Understand when to move from basic to advanced retrieval
- Evaluate retrieval strategies using production metrics
- Design a framework-agnostic retrieval architecture
1. Why Do We Need Multiple Retrievers?ΒΆ
A single retrieval strategy cannot optimally solve every retrieval problem.
Consider these queries:
Query A:
"What is our remote work policy?"
Query B:
"What is our cloud strategy?"
Query C:
"Show HR policies for Germany from 2025."
Query D:
"Who is responsible for approving exceptions?"
Each query may require a different retrieval behavior.
flowchart TD
A["User Query"] --> B{"Retrieval Requirement"}
B -->|"Simple semantic search"| C["VectorStore Retriever"]
B -->|"Multiple interpretations"| D["Multi-Query Retriever"]
B -->|"Metadata constraints"| E["Self-Query Retriever"]
B -->|"Context surrounding match"| F["Parent-Document Retriever"]
C --> G["Retrieved Context"]
D --> G
E --> G
F --> G The important architectural principle is:
Retriever selection should be driven by the retrieval problem, not by the popularity of a framework or technique.
2. The Four Core RetrieversΒΆ
We have now covered four foundational retrieval patterns.
| Retriever | Primary Purpose |
|---|---|
| VectorStore Retriever | Semantic similarity search |
| Multi-Query Retriever | Multiple semantic perspectives |
| Self-Query Retriever | Semantic search + metadata filtering |
| Parent-Document Retriever | Fine-grained retrieval + broader context |
A simplified view:
VectorStore
β
βββ Basic Semantic Retrieval
β
βββ Multi-Query
β βββ Multiple Query Perspectives
β
βββ Self-Query
β βββ Semantic + Metadata Constraints
β
βββ Parent-Document
βββ Child Retrieval + Parent Context
3. VectorStore RetrieverΒΆ
The VectorStore Retriever is the simplest baseline.
Example:
retriever = vector_store.as_retriever(
search_kwargs={
"k": 5
}
)
documents = retriever.invoke(
"What is the remote work policy?"
)
Best suited forΒΆ
- Simple semantic questions
- Well-structured knowledge bases
- Short, self-contained chunks
- Initial RAG implementations
- Baseline retrieval evaluation
Main limitationΒΆ
It may miss relevant information when the query is ambiguous or broad.
4. Multi-Query RetrieverΒΆ
Multi-Query Retrieval generates multiple alternative queries from the original user question.
User Query
β
LLM Query Generator
β
ββββββββββ¬βββββββββ¬βββββββββ
Query A Query B Query C
β β β
Search Search Search
ββββββββββΌβββββββββ
β
Candidate Pool
Example:
Generated queries:
Technical risks of cloud migration
Security risks of cloud migration
Financial risks of cloud migration
Operational risks of cloud migration
Best suited forΒΆ
- Ambiguous questions
- Broad questions
- Multiple semantic perspectives
- Queries requiring evidence from different areas
Main limitationΒΆ
It can also introduce query drift.
5. Self-Query RetrieverΒΆ
Self-Query Retrieval converts natural language into:
Example:
becomes conceptually:
Architecture:
flowchart LR
A["Natural Language Query"] --> B["LLM Query Analyzer"]
B --> C["Semantic Query"]
B --> D["Metadata Filters"]
C --> E["Vector Search"]
D --> F["Metadata Filtering"]
E --> G["Relevant Documents"]
F --> G Best suited forΒΆ
- Rich metadata
- Enterprise document repositories
- Natural-language filtering
- Multi-tenant knowledge bases
- Date, region, department, and document-type filtering
Main limitationΒΆ
The LLM can generate incorrect filters.
Therefore:
should be used in production.
6. Parent-Document RetrieverΒΆ
Parent-Document Retrieval separates:
from:
The system retrieves a small child chunk:
and then resolves the parent:
Architecture:
flowchart TD
A["Document"] --> B["Parent"]
B --> C["Child Chunks"]
C --> D["Embedding"]
D --> E["Vector Store"]
F["Query"] --> E
E --> G["Matching Child"]
G --> H["Parent ID"]
H --> I["Parent Store"]
I --> J["Parent Context"]
J --> K["LLM"] Best suited forΒΆ
- Long documents
- Technical manuals
- Policies
- Legal documents
- Research documents
- Documents where surrounding context matters
Main limitationΒΆ
Large parents can create:
7. Side-by-Side ComparisonΒΆ
| Capability | VectorStore | Multi-Query | Self-Query | Parent-Document |
|---|---|---|---|---|
| Semantic Search | β | β | β | β |
| Multiple Query Perspectives | β | β | β | β |
| Metadata Filtering | Basic | Possible | β | Possible |
| Small Retrieval Units | β | β | β | β |
| Larger Generation Context | β | β | β | β |
| Query Generation LLM | β | β | β | β |
| Retrieval Coverage | Baseline | High | High with filters | Baseline |
| Context Preservation | Basic | Basic | Basic | High |
| Additional LLM Cost | β | β | β | β |
| Additional Storage | Usually no | No | No | Usually yes |
| Implementation Complexity | Low | Medium | Medium | Medium |
8. The Core Trade-OffΒΆ
Retriever selection can be understood using two major dimensions:
Retrieval Coverage
β
β
Multi-Query β
β
Self-Query β
β
β
VectorStore βββββββββββββββββΌβββββββββββββ
β
β
β
β
β
But retrieval quality is not one-dimensional.
A production system needs to balance:
Therefore, the "best" retriever is workload-specific.
9. Precision vs RecallΒΆ
Retrieval systems often have a fundamental trade-off.
PrecisionΒΆ
How much of the retrieved information is relevant?
High precision means:
RecallΒΆ
How much of the relevant information was successfully retrieved?
High recall means:
A simplified view:
quadrantChart
title Retrieval Strategy Trade-offs
x-axis Lower Precision --> Higher Precision
y-axis Lower Recall --> Higher Recall
quadrant-1 High Recall / High Precision
quadrant-2 High Recall / Lower Precision
quadrant-3 Lower Recall / Lower Precision
quadrant-4 Higher Precision / Lower Recall
VectorStore: [0.65, 0.60]
MultiQuery: [0.55, 0.82]
SelfQuery: [0.76, 0.78]
ParentDocument: [0.62, 0.70] The positions in this conceptual chart are illustrative rather than measured benchmark results.
In production, these values should be determined through evaluation.
10. Which Retriever Should You Start With?ΒΆ
For a new RAG application, a sensible progression is:
Start
β
VectorStore Retriever
β
Evaluate
β
Identify Retrieval Problem
β
Add Specialized Retriever
For example:
Baseline
β
VectorStore
β
Problem: Query Ambiguity
β
Multi-Query
Problem: Metadata Constraints
β
Self-Query
Problem: Context Loss
β
Parent-Document
This is preferable to implementing every advanced technique immediately.
11. Decision TreeΒΆ
A practical decision tree:
flowchart TD
A["Start with User Query"] --> B{"Is semantic search sufficient?"}
B -->|Yes| C["VectorStore Retriever"]
B -->|No| D{"Does query contain metadata constraints?"}
D -->|Yes| E["Self-Query Retriever"]
D -->|No| F{"Does query have multiple interpretations?"}
F -->|Yes| G["Multi-Query Retriever"]
F -->|No| H{"Does surrounding context matter?"}
H -->|Yes| I["Parent-Document Retriever"]
H -->|No| J["Improve Baseline Retrieval"] In practice, more than one answer can be true.
For example:
may require a composed retrieval pipeline.
12. Retriever CompositionΒΆ
The most important lesson from these chapters is that retrievers do not have to operate independently.
They can be composed.
For example:
Architecture:
flowchart TD
A["User Query"] --> B["Multi-Query"]
B --> C["Query A"]
B --> D["Query B"]
B --> E["Query C"]
C --> F["Self-Query"]
D --> F
E --> F
F --> G["Vector Retrieval"]
G --> H["Child Candidates"]
H --> I["Parent Resolution"]
I --> J["Candidate Parents"]
J --> K["Re-ranking"]
K --> L["Context Selection"]
L --> M["LLM"] This is closer to how enterprise retrieval architectures evolve.
13. Retrieval as a PipelineΒΆ
Instead of thinking:
think:
For example:
Query Understanding
β
Query Expansion
β
Metadata Filtering
β
Semantic Retrieval
β
Parent Resolution
β
Re-ranking
β
Context Compression
β
Context Selection
Each stage addresses a different problem.
14. Capability-Based Retrieval ArchitectureΒΆ
A framework-agnostic architecture might expose:
Then specialized capabilities can implement the interface:
class VectorStoreRetriever(Retriever):
...
class MultiQueryRetriever(Retriever):
...
class SelfQueryRetriever(Retriever):
...
class ParentDocumentRetriever(Retriever):
...
A higher-level pipeline can compose them:
retrieval_pipeline = RetrievalPipeline(
query_processor=query_processor,
retriever=retriever,
reranker=reranker,
context_selector=context_selector
)
This keeps application logic independent of the retrieval implementation.
15. Retrieval Strategy MatrixΒΆ
A useful engineering matrix is:
| Requirement | Recommended Starting Point |
|---|---|
| Simple semantic question | VectorStore |
| Broad question | Multi-Query |
| Ambiguous question | Multi-Query |
| Metadata constraints | Self-Query |
| Long documents | Parent-Document |
| Need surrounding context | Parent-Document |
| Multiple semantic perspectives | Multi-Query |
| Region / department / year filters | Self-Query |
| Small, self-contained documents | VectorStore |
| Large structured documents | Parent-Document |
| Multiple requirements | Composition |
This matrix should be treated as a starting point, not a fixed rule.
16. Example: HR Knowledge AssistantΒΆ
Suppose an HR assistant contains:
with metadata:
User asks:
This query has multiple requirements:
German
β
Metadata Constraint
Latest
β
Version / Date Constraint
Parental Leave
β
Semantic Retrieval
Exceptions
β
Context Requirement
A composed architecture could be:
17. Example: Architecture Knowledge AssistantΒΆ
User:
This query has multiple dimensions:
A suitable strategy could be:
Here Multi-Query improves retrieval coverage.
18. Example: Technical DocumentationΒΆ
User:
If the documentation is organized into:
and the relevant information is distributed across a section, Parent-Document Retrieval may be useful:
19. Combining All FourΒΆ
A sophisticated pipeline can use all four techniques.
flowchart TD
A["User Query"] --> B["Query Understanding"]
B --> C["Multi-Query Generation"]
C --> D["Query A"]
C --> E["Query B"]
C --> F["Query C"]
D --> G["Self-Query"]
E --> G
F --> G
G --> H["Vector Retrieval"]
H --> I["Child Chunks"]
I --> J["Parent Resolution"]
J --> K["Candidate Parents"]
K --> L["Re-ranking"]
L --> M["Context Selection"]
M --> N["Prompt Assembly"]
N --> O["LLM"]
O --> P["Response Validation"] This is powerful but also significantly more complex.
Therefore:
Complexity should be introduced only when evaluation shows that the simpler architecture is insufficient.
20. Complexity vs BenefitΒΆ
Every retrieval enhancement introduces additional complexity.
VectorStore
β
Low Complexity
Multi-Query
β
+ LLM Query Generation
+ Multiple Retrieval Calls
Self-Query
β
+ Query Parsing
+ Filter Validation
Parent-Document
β
+ Parent Store
+ Parent Resolution
Advanced Composition
β
+ Multiple Dependencies
+ More Latency
+ More Observability Requirements
The engineering goal is not maximum complexity.
The goal is:
21. Retrieval Decision ExampleΒΆ
Consider three queries.
Query AΒΆ
Recommended:
Query BΒΆ
Recommended starting point:
Potential generated perspectives:
Query CΒΆ
Recommended:
Potential filters:
Query DΒΆ
Recommended:
because surrounding policy context may be important.
22. Retrieval Quality MetricsΒΆ
Retriever comparison should be based on measurable outcomes.
Important retrieval metrics include:
Recall@KΒΆ
Measures whether relevant documents appear in the top K results.
Precision@KΒΆ
Measures how many of the top K results are relevant.
MRRΒΆ
Mean Reciprocal Rank measures how high the first relevant result appears.
NDCGΒΆ
Normalized Discounted Cumulative Gain evaluates ranking quality while giving higher weight to results near the top.
These metrics help determine whether a retrieval strategy actually improves retrieval.
23. End-to-End RAG MetricsΒΆ
Retriever metrics alone are not enough.
The final system should also measure:
Useful metrics include:
The exact evaluation framework can vary by system.
24. Latency ComparisonΒΆ
Conceptually:
Multi-Query:
Self-Query:
Parent-Document:
Therefore, every advanced strategy can introduce additional latency.
Production systems should measure:
rather than relying only on average latency.
25. Cost ComparisonΒΆ
A simplified view:
| Retriever | LLM Cost | Retrieval Cost | Storage Complexity |
|---|---|---|---|
| VectorStore | Low | Low | Low |
| Multi-Query | Higher | Higher | Low |
| Self-Query | Higher | Medium | Medium |
| Parent-Document | Low | Medium | Higher |
These are conceptual comparisons.
Actual costs depend on:
- Model choice
- Number of queries
- Vector database
- Document size
- Storage architecture
- Query volume
- Caching
26. ObservabilityΒΆ
Advanced retrieval requires strong observability.
At minimum, capture:
Original Query
β
Generated Queries
β
Filters
β
Retriever Type
β
Retrieved Documents
β
Scores
β
Parent IDs
β
Re-ranking
β
Final Context
Example trace:
trace_id: 8f3a21
query:
"Find German HR policies from 2025."
retriever:
SelfQueryRetriever
semantic_query:
"HR policies"
filters:
country = Germany
department = HR
year = 2025
candidate_count:
37
final_count:
5
latency:
428 ms
This information is extremely useful when diagnosing production retrieval failures.
27. Security ComparisonΒΆ
Retriever type does not determine authorization.
Every retriever should operate inside the same security boundary.
This applies to:
For Multi-Query:
For Parent-Document:
must also be validated.
A user should never gain access to a parent document merely because an authorized child matched.
28. Retriever Selection PrinciplesΒΆ
Principle 1ΒΆ
Start with the simplest strategy that satisfies the requirements.
Principle 2ΒΆ
Measure retrieval quality before adding complexity.
Principle 3ΒΆ
Solve a specific retrieval problem with each technique.
Principle 4ΒΆ
Keep retrieval capabilities composable.
Principle 5ΒΆ
Separate retrieval from authorization.
Principle 6ΒΆ
Treat generated queries and filters as untrusted intermediate data.
Principle 7ΒΆ
Measure latency and cost alongside retrieval quality.
π― Advanced Retriever Decision MatrixΒΆ
| Requirement | LlamaIndex | LangChain / Generic Pattern |
|---|---|---|
| Semantic search | Vector Index Retriever | VectorStore Retriever |
| Exact terminology | BM25 Retriever | BM25 / keyword retriever |
| Multiple query perspectives | Query Fusion Retriever | MultiQueryRetriever |
| Metadata-aware retrieval | Metadata filters | SelfQueryRetriever |
| Small chunks + large context | Auto-Merging Retriever | ParentDocumentRetriever |
| Follow references between nodes | Recursive Retriever | Custom graph / reference traversal |
| Retrieve whole relevant documents | Document Summary Retriever | Custom document-level retrieval |
| Relevance + diversity | MMR / postprocessing | MMR search |
| Dense + sparse retrieval | Hybrid retrieval | Ensemble / hybrid retrieval |
| Improve candidate ranking | Re-ranker | Re-ranker |
| Complex retrieval selection | Router / Agentic retrieval | Router / Agentic retrieval |
Selection PrincipleΒΆ
Do not ask:
"What is the best retriever?"
Ask:
"What retrieval failure am I trying to solve?"
---
# 29. Recommended Evolution Path
A practical enterprise retrieval journey is:
```mermaid
flowchart LR
A["VectorStore Baseline"] --> B["Evaluate"]
B --> C{"Problem Identified?"}
C -->|"Query Ambiguity"| D["Multi-Query"]
C -->|"Metadata Constraints"| E["Self-Query"]
C -->|"Context Loss"| F["Parent-Document"]
D --> G["Evaluate"]
E --> G
F --> G
G --> H["Hybrid / Re-ranking"]
H --> I["Production RAG"]
This avoids premature optimization.
30. Beyond the Four RetrieversΒΆ
These four retrievers are only the beginning of advanced retrieval engineering.
Later techniques can address additional problems:
Hybrid Search
β
Semantic + Keyword Retrieval
Re-ranking
β
Better Candidate Ordering
Contextual Compression
β
Reduce Context Noise
Graph RAG
β
Relationship-Based Retrieval
SQL RAG
β
Structured Data Retrieval
Knowledge Graphs
β
Entity + Relationship Retrieval
Agentic RAG
β
Iterative Retrieval Decisions
The important point is that these techniques build upon the retrieval foundations established here.
31. Production Retrieval ArchitectureΒΆ
A mature enterprise architecture may eventually evolve into:
flowchart TD
A["User"] --> B["RAG API"]
B --> C["Authentication"]
C --> D["Authorization"]
D --> E["Query Processing"]
E --> F["Retriever Router"]
F --> G["VectorStore"]
F --> H["MultiQuery"]
F --> I["SelfQuery"]
F --> J["ParentDocument"]
F --> K["Hybrid"]
F --> L["Graph"]
F --> M["SQL"]
G --> N["Candidate Pool"]
H --> N
I --> N
J --> N
K --> N
L --> N
M --> N
N --> O["Deduplication"]
O --> P["Re-ranking"]
P --> Q["Context Selection"]
Q --> R["Prompt Assembly"]
R --> S["LLM"]
S --> T["Response Validation"]
T --> U["Citations"]
U --> V["Enterprise Response"]
B --> W["Observability"]
F --> W
N --> W
P --> W
S --> W This is the direction toward production retrieval engineering.
32. Framework-Agnostic Retrieval InterfaceΒΆ
An enterprise platform should ideally expose a stable abstraction.
from abc import ABC, abstractmethod
class Retriever(ABC):
@abstractmethod
def retrieve(self, query: str):
pass
Implementations can include:
class VectorStoreRetriever(Retriever):
...
class MultiQueryRetriever(Retriever):
...
class SelfQueryRetriever(Retriever):
...
class ParentDocumentRetriever(Retriever):
...
class HybridRetriever(Retriever):
...
class GraphRetriever(Retriever):
...
class SQLRetriever(Retriever):
...
The application can therefore remain independent of the underlying framework.
Enterprise RAG Application
β
Retriever Interface
β
βββββββββββΌββββββββββ
β β β
Vector Hybrid Graph
This is particularly useful when building reusable enterprise AI platforms.
33. Retrieval Strategy Cheat SheetΒΆ
βββββββββββββββββββββββββββββββββββββββββββββββββ
β RETRIEVER CHEAT SHEET β
βββββββββββββββββββββββββββββββββββββββββββββββββ€
β VectorStore β
β β Simple semantic retrieval β
β β
β Multi-Query β
β β Multiple semantic perspectives β
β β
β Self-Query β
β β Natural language + metadata filters β
β β
β Parent-Document β
β β Small retrieval + large context β
β β
β Hybrid β
β β Semantic + lexical search β
β β
β Re-ranking β
β β Improve candidate ordering β
β β
β Graph RAG β
β β Relationship-based retrieval β
β β
β SQL RAG β
β β Structured data retrieval β
βββββββββββββββββββββββββββββββββββββββββββββββββ
34. Key TakeawaysΒΆ
- There is no universally best retriever.
- VectorStore Retrieval is the natural baseline.
- Multi-Query Retrieval improves retrieval coverage through multiple query perspectives.
- Self-Query Retrieval combines semantic search with structured metadata filtering.
- Parent-Document Retrieval separates retrieval granularity from generation context.
- Retriever selection should be driven by the problem being solved.
- Advanced retrievers can be composed into larger retrieval pipelines.
- More retrieval complexity does not automatically mean better results.
- Every advanced technique introduces additional latency, cost, or operational complexity.
- Retrieval quality should be evaluated using metrics such as Recall@K, Precision@K, MRR, and NDCG.
- End-to-end RAG quality must also be measured through context and answer-quality metrics.
- Authorization must remain independent of retrieval strategy.
- LLM-generated queries and filters should be treated as untrusted intermediate data.
- Parent documents must be protected by the same access-control model as child chunks.
- Production retrieval requires strong observability.
- Framework-independent retrieval interfaces help reduce application coupling.
- The four retrievers covered so far provide the foundation for more advanced techniques such as Hybrid Search, Re-ranking, Graph RAG, SQL RAG, and Agentic RAG.
The central engineering principle is:
Don't ask:
"Which retriever is the best?"
Ask:
"What retrieval problem are we solving?"
β
"What capability addresses it?"
β
"Does evaluation prove the improvement?"
β
"Is the additional complexity justified?"
π§ Chapter NavigationΒΆ
Part V β Advanced Retrieval-Augmented GenerationΒΆ
Previous:
04. Parent-Document Retriever
Next:
Enterprise-retrieval-engineering/01 Contextual Compression Retriever
Section:
01 β Core Retrieval Engineering
Retrieval Engineering PathΒΆ
01 VectorStore Retriever
β
02 Multi-Query Retriever
β
03 Self-Query Retriever
β
04 Parent-Document Retriever
β
05 Retriever Comparison
β
06 Hybrid Search
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.