Skip to content

Retriever ComparisonΒΆ

πŸ“– OverviewΒΆ

Retrieval is one of the most important components of a Retrieval-Augmented Generation (RAG) system.

A basic RAG pipeline may begin with a simple VectorStore Retriever, but production systems often require different retrieval strategies depending on the nature of the query, document structure, metadata, and business requirements.

In the previous chapters, we explored:

01 β€” VectorStore Retriever
02 β€” Multi-Query Retriever
03 β€” Self-Query Retriever
04 β€” Parent-Document Retriever

Each technique solves a different retrieval problem.

The goal of this chapter is to understand when to use each retriever, what problem it solves, its trade-offs, and how these techniques can be composed into production retrieval architectures.


🎯 Learning Objectives¢

After completing this chapter, you will be able to:

  • Understand the role of different retriever strategies
  • Compare VectorStore, Multi-Query, Self-Query, and Parent-Document retrieval
  • Understand the strengths and limitations of each approach
  • Select an appropriate retriever for a given problem
  • Understand retrieval precision vs recall trade-offs
  • Understand how retrievers can be composed
  • Design a retrieval decision strategy
  • Understand when to move from basic to advanced retrieval
  • Evaluate retrieval strategies using production metrics
  • Design a framework-agnostic retrieval architecture

1. Why Do We Need Multiple Retrievers?ΒΆ

A single retrieval strategy cannot optimally solve every retrieval problem.

Consider these queries:

Query A:
"What is our remote work policy?"

Query B:
"What is our cloud strategy?"

Query C:
"Show HR policies for Germany from 2025."

Query D:
"Who is responsible for approving exceptions?"

Each query may require a different retrieval behavior.

flowchart TD
    A["User Query"] --> B{"Retrieval Requirement"}

    B -->|"Simple semantic search"| C["VectorStore Retriever"]
    B -->|"Multiple interpretations"| D["Multi-Query Retriever"]
    B -->|"Metadata constraints"| E["Self-Query Retriever"]
    B -->|"Context surrounding match"| F["Parent-Document Retriever"]

    C --> G["Retrieved Context"]
    D --> G
    E --> G
    F --> G

The important architectural principle is:

Retriever selection should be driven by the retrieval problem, not by the popularity of a framework or technique.


2. The Four Core RetrieversΒΆ

We have now covered four foundational retrieval patterns.

Retriever Primary Purpose
VectorStore Retriever Semantic similarity search
Multi-Query Retriever Multiple semantic perspectives
Self-Query Retriever Semantic search + metadata filtering
Parent-Document Retriever Fine-grained retrieval + broader context

A simplified view:

VectorStore
     β”‚
     β”œβ”€β”€ Basic Semantic Retrieval
     β”‚
     β”œβ”€β”€ Multi-Query
     β”‚      └── Multiple Query Perspectives
     β”‚
     β”œβ”€β”€ Self-Query
     β”‚      └── Semantic + Metadata Constraints
     β”‚
     └── Parent-Document
            └── Child Retrieval + Parent Context

3. VectorStore RetrieverΒΆ

The VectorStore Retriever is the simplest baseline.

User Query
    ↓
Embedding
    ↓
Vector Search
    ↓
Top-K Chunks

Example:

retriever = vector_store.as_retriever(
    search_kwargs={
        "k": 5
    }
)

documents = retriever.invoke(
    "What is the remote work policy?"
)

Best suited forΒΆ

  • Simple semantic questions
  • Well-structured knowledge bases
  • Short, self-contained chunks
  • Initial RAG implementations
  • Baseline retrieval evaluation

Main limitationΒΆ

One Query
    ↓
One Semantic Representation

It may miss relevant information when the query is ambiguous or broad.


4. Multi-Query RetrieverΒΆ

Multi-Query Retrieval generates multiple alternative queries from the original user question.

User Query
     ↓
LLM Query Generator
     ↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”
Query A  Query B  Query C
   ↓        ↓        ↓
 Search   Search   Search
   β””β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”˜
            ↓
      Candidate Pool

Example:

Original:

"What are the risks of cloud migration?"

Generated queries:

Technical risks of cloud migration
Security risks of cloud migration
Financial risks of cloud migration
Operational risks of cloud migration

Best suited forΒΆ

  • Ambiguous questions
  • Broad questions
  • Multiple semantic perspectives
  • Queries requiring evidence from different areas

Main limitationΒΆ

More Queries
     ↓
More Retrieval Calls
     ↓
More Latency + Cost

It can also introduce query drift.


5. Self-Query RetrieverΒΆ

Self-Query Retrieval converts natural language into:

Semantic Query
+
Metadata Filters

Example:

"Find HR policies for Germany from 2025."

becomes conceptually:

Query:
HR policies

Filters:
department = HR
country = Germany
year = 2025

Architecture:

flowchart LR
    A["Natural Language Query"] --> B["LLM Query Analyzer"]

    B --> C["Semantic Query"]
    B --> D["Metadata Filters"]

    C --> E["Vector Search"]
    D --> F["Metadata Filtering"]

    E --> G["Relevant Documents"]
    F --> G

Best suited forΒΆ

  • Rich metadata
  • Enterprise document repositories
  • Natural-language filtering
  • Multi-tenant knowledge bases
  • Date, region, department, and document-type filtering

Main limitationΒΆ

The LLM can generate incorrect filters.

Therefore:

LLM Output
    ↓
Validation
    ↓
Authorization
    ↓
Retrieval

should be used in production.


6. Parent-Document RetrieverΒΆ

Parent-Document Retrieval separates:

Retrieval Granularity

from:

Generation Context

The system retrieves a small child chunk:

Query
 ↓
Child Chunk

and then resolves the parent:

Child Chunk
     ↓
Parent ID
     ↓
Parent Document

Architecture:

flowchart TD
    A["Document"] --> B["Parent"]
    B --> C["Child Chunks"]

    C --> D["Embedding"]
    D --> E["Vector Store"]

    F["Query"] --> E
    E --> G["Matching Child"]

    G --> H["Parent ID"]
    H --> I["Parent Store"]
    I --> J["Parent Context"]
    J --> K["LLM"]

Best suited forΒΆ

  • Long documents
  • Technical manuals
  • Policies
  • Legal documents
  • Research documents
  • Documents where surrounding context matters

Main limitationΒΆ

Large parents can create:

Context Explosion
      ↓
More Tokens
      ↓
Higher Cost
      ↓
Potential Context Noise

7. Side-by-Side ComparisonΒΆ

Capability VectorStore Multi-Query Self-Query Parent-Document
Semantic Search βœ… βœ… βœ… βœ…
Multiple Query Perspectives ❌ βœ… ❌ ❌
Metadata Filtering Basic Possible βœ… Possible
Small Retrieval Units βœ… βœ… βœ… βœ…
Larger Generation Context ❌ ❌ ❌ βœ…
Query Generation LLM ❌ βœ… βœ… ❌
Retrieval Coverage Baseline High High with filters Baseline
Context Preservation Basic Basic Basic High
Additional LLM Cost ❌ βœ… βœ… ❌
Additional Storage Usually no No No Usually yes
Implementation Complexity Low Medium Medium Medium

8. The Core Trade-OffΒΆ

Retriever selection can be understood using two major dimensions:

                     Retrieval Coverage
                            ↑
                            β”‚
                Multi-Query β”‚
                            β”‚
              Self-Query    β”‚
                            β”‚
                            β”‚
VectorStore ────────────────┼────────────→
                            β”‚
                            β”‚
                            β”‚
                            β”‚
                            ↓

But retrieval quality is not one-dimensional.

A production system needs to balance:

Recall
Precision
Latency
Cost
Context Quality
Security
Scalability

Therefore, the "best" retriever is workload-specific.


9. Precision vs RecallΒΆ

Retrieval systems often have a fundamental trade-off.

PrecisionΒΆ

How much of the retrieved information is relevant?

Retrieved Documents
        ↓
Relevant Documents

High precision means:

Less Noise

RecallΒΆ

How much of the relevant information was successfully retrieved?

Relevant Documents
        ↓
Retrieved Relevant Documents

High recall means:

Less Missing Evidence

A simplified view:

quadrantChart
    title Retrieval Strategy Trade-offs
    x-axis Lower Precision --> Higher Precision
    y-axis Lower Recall --> Higher Recall
    quadrant-1 High Recall / High Precision
    quadrant-2 High Recall / Lower Precision
    quadrant-3 Lower Recall / Lower Precision
    quadrant-4 Higher Precision / Lower Recall
    VectorStore: [0.65, 0.60]
    MultiQuery: [0.55, 0.82]
    SelfQuery: [0.76, 0.78]
    ParentDocument: [0.62, 0.70]

The positions in this conceptual chart are illustrative rather than measured benchmark results.

In production, these values should be determined through evaluation.


10. Which Retriever Should You Start With?ΒΆ

For a new RAG application, a sensible progression is:

Start
  ↓
VectorStore Retriever
  ↓
Evaluate
  ↓
Identify Retrieval Problem
  ↓
Add Specialized Retriever

For example:

Baseline
   ↓
VectorStore
   ↓
Problem: Query Ambiguity
   ↓
Multi-Query

Problem: Metadata Constraints
   ↓
Self-Query

Problem: Context Loss
   ↓
Parent-Document

This is preferable to implementing every advanced technique immediately.


11. Decision TreeΒΆ

A practical decision tree:

flowchart TD
    A["Start with User Query"] --> B{"Is semantic search sufficient?"}

    B -->|Yes| C["VectorStore Retriever"]

    B -->|No| D{"Does query contain metadata constraints?"}

    D -->|Yes| E["Self-Query Retriever"]

    D -->|No| F{"Does query have multiple interpretations?"}

    F -->|Yes| G["Multi-Query Retriever"]

    F -->|No| H{"Does surrounding context matter?"}

    H -->|Yes| I["Parent-Document Retriever"]

    H -->|No| J["Improve Baseline Retrieval"]

In practice, more than one answer can be true.

For example:

Ambiguous query
+
Metadata constraints
+
Long documents

may require a composed retrieval pipeline.


12. Retriever CompositionΒΆ

The most important lesson from these chapters is that retrievers do not have to operate independently.

They can be composed.

For example:

Multi-Query
     ↓
Self-Query
     ↓
Vector Retrieval
     ↓
Parent Resolution
     ↓
Re-ranking

Architecture:

flowchart TD
    A["User Query"] --> B["Multi-Query"]

    B --> C["Query A"]
    B --> D["Query B"]
    B --> E["Query C"]

    C --> F["Self-Query"]
    D --> F
    E --> F

    F --> G["Vector Retrieval"]

    G --> H["Child Candidates"]
    H --> I["Parent Resolution"]
    I --> J["Candidate Parents"]
    J --> K["Re-ranking"]
    K --> L["Context Selection"]
    L --> M["LLM"]

This is closer to how enterprise retrieval architectures evolve.


13. Retrieval as a PipelineΒΆ

Instead of thinking:

Which retriever should I use?

think:

Which retrieval capabilities do I need?

For example:

Query Understanding
        ↓
Query Expansion
        ↓
Metadata Filtering
        ↓
Semantic Retrieval
        ↓
Parent Resolution
        ↓
Re-ranking
        ↓
Context Compression
        ↓
Context Selection

Each stage addresses a different problem.


14. Capability-Based Retrieval ArchitectureΒΆ

A framework-agnostic architecture might expose:

class Retriever:

    def retrieve(self, query: str):
        raise NotImplementedError

Then specialized capabilities can implement the interface:

class VectorStoreRetriever(Retriever):
    ...


class MultiQueryRetriever(Retriever):
    ...


class SelfQueryRetriever(Retriever):
    ...


class ParentDocumentRetriever(Retriever):
    ...

A higher-level pipeline can compose them:

retrieval_pipeline = RetrievalPipeline(
    query_processor=query_processor,
    retriever=retriever,
    reranker=reranker,
    context_selector=context_selector
)

This keeps application logic independent of the retrieval implementation.


15. Retrieval Strategy MatrixΒΆ

A useful engineering matrix is:

Requirement Recommended Starting Point
Simple semantic question VectorStore
Broad question Multi-Query
Ambiguous question Multi-Query
Metadata constraints Self-Query
Long documents Parent-Document
Need surrounding context Parent-Document
Multiple semantic perspectives Multi-Query
Region / department / year filters Self-Query
Small, self-contained documents VectorStore
Large structured documents Parent-Document
Multiple requirements Composition

This matrix should be treated as a starting point, not a fixed rule.


16. Example: HR Knowledge AssistantΒΆ

Suppose an HR assistant contains:

50,000 documents

with metadata:

Country
Department
Document Type
Year
Version

User asks:

"Find the latest German parental leave
policy and explain the exceptions."

This query has multiple requirements:

German
   ↓
Metadata Constraint

Latest
   ↓
Version / Date Constraint

Parental Leave
   ↓
Semantic Retrieval

Exceptions
   ↓
Context Requirement

A composed architecture could be:

Self-Query
     ↓
Metadata Filtering
     ↓
Vector Retrieval
     ↓
Parent Document
     ↓
Context Selection

17. Example: Architecture Knowledge AssistantΒΆ

User:

"What are the main risks of migrating
our payment platform to the cloud?"

This query has multiple dimensions:

Technical
Security
Financial
Operational
Compliance

A suitable strategy could be:

Multi-Query
     ↓
Multiple Perspectives
     ↓
Vector Retrieval
     ↓
Candidate Pool
     ↓
Re-ranking

Here Multi-Query improves retrieval coverage.


18. Example: Technical DocumentationΒΆ

User:

"How does the API gateway authenticate
requests?"

If the documentation is organized into:

Architecture
Authentication
Authorization
API Gateway
Security
Troubleshooting

and the relevant information is distributed across a section, Parent-Document Retrieval may be useful:

Child Match
    ↓
Authentication Section
    ↓
Parent Context
    ↓
LLM

19. Combining All FourΒΆ

A sophisticated pipeline can use all four techniques.

flowchart TD
    A["User Query"] --> B["Query Understanding"]

    B --> C["Multi-Query Generation"]

    C --> D["Query A"]
    C --> E["Query B"]
    C --> F["Query C"]

    D --> G["Self-Query"]
    E --> G
    F --> G

    G --> H["Vector Retrieval"]

    H --> I["Child Chunks"]

    I --> J["Parent Resolution"]

    J --> K["Candidate Parents"]

    K --> L["Re-ranking"]

    L --> M["Context Selection"]

    M --> N["Prompt Assembly"]

    N --> O["LLM"]

    O --> P["Response Validation"]

This is powerful but also significantly more complex.

Therefore:

Complexity should be introduced only when evaluation shows that the simpler architecture is insufficient.


20. Complexity vs BenefitΒΆ

Every retrieval enhancement introduces additional complexity.

VectorStore
    ↓
Low Complexity

Multi-Query
    ↓
+ LLM Query Generation
+ Multiple Retrieval Calls

Self-Query
    ↓
+ Query Parsing
+ Filter Validation

Parent-Document
    ↓
+ Parent Store
+ Parent Resolution

Advanced Composition
    ↓
+ Multiple Dependencies
+ More Latency
+ More Observability Requirements

The engineering goal is not maximum complexity.

The goal is:

Required Quality
      +
Acceptable Cost
      +
Acceptable Latency
      +
Operational Simplicity

21. Retrieval Decision ExampleΒΆ

Consider three queries.

Query AΒΆ

"What is the vacation policy?"

Recommended:

VectorStore Retriever

Query BΒΆ

"What are the security risks
of cloud migration?"

Recommended starting point:

Multi-Query Retriever

Potential generated perspectives:

Technical risks
Security risks
Operational risks
Compliance risks

Query CΒΆ

"Find Germany HR policies from 2025."

Recommended:

Self-Query Retriever

Potential filters:

country = Germany
department = HR
year = 2025

Query DΒΆ

"Explain the exception process
described in the remote-work policy."

Recommended:

Parent-Document Retriever

because surrounding policy context may be important.


22. Retrieval Quality MetricsΒΆ

Retriever comparison should be based on measurable outcomes.

Important retrieval metrics include:

Recall@KΒΆ

Measures whether relevant documents appear in the top K results.

Relevant Retrieved
────────────────────
Total Relevant

Precision@KΒΆ

Measures how many of the top K results are relevant.

Relevant Retrieved
────────────────────
Retrieved Documents

MRRΒΆ

Mean Reciprocal Rank measures how high the first relevant result appears.

NDCGΒΆ

Normalized Discounted Cumulative Gain evaluates ranking quality while giving higher weight to results near the top.

These metrics help determine whether a retrieval strategy actually improves retrieval.


23. End-to-End RAG MetricsΒΆ

Retriever metrics alone are not enough.

The final system should also measure:

Retrieval Quality
       ↓
Context Quality
       ↓
Answer Quality

Useful metrics include:

Faithfulness
Answer Relevance
Context Relevance
Context Recall
Citation Accuracy
Groundedness

The exact evaluation framework can vary by system.


24. Latency ComparisonΒΆ

Conceptually:

VectorStore

Query
 ↓
Embedding
 ↓
Search
 ↓
Response

Multi-Query:

Query
 ↓
LLM Query Generation
 ↓
Search Γ— N
 ↓
Aggregation
 ↓
Response

Self-Query:

Query
 ↓
LLM Query Construction
 ↓
Validation
 ↓
Filtered Search
 ↓
Response

Parent-Document:

Query
 ↓
Search
 ↓
Parent Lookup
 ↓
Response

Therefore, every advanced strategy can introduce additional latency.

Production systems should measure:

P50
P95
P99

rather than relying only on average latency.


25. Cost ComparisonΒΆ

A simplified view:

Retriever LLM Cost Retrieval Cost Storage Complexity
VectorStore Low Low Low
Multi-Query Higher Higher Low
Self-Query Higher Medium Medium
Parent-Document Low Medium Higher

These are conceptual comparisons.

Actual costs depend on:

  • Model choice
  • Number of queries
  • Vector database
  • Document size
  • Storage architecture
  • Query volume
  • Caching

26. ObservabilityΒΆ

Advanced retrieval requires strong observability.

At minimum, capture:

Original Query
        ↓
Generated Queries
        ↓
Filters
        ↓
Retriever Type
        ↓
Retrieved Documents
        ↓
Scores
        ↓
Parent IDs
        ↓
Re-ranking
        ↓
Final Context

Example trace:

trace_id: 8f3a21

query:
"Find German HR policies from 2025."

retriever:
SelfQueryRetriever

semantic_query:
"HR policies"

filters:
country = Germany
department = HR
year = 2025

candidate_count:
37

final_count:
5

latency:
428 ms

This information is extremely useful when diagnosing production retrieval failures.


27. Security ComparisonΒΆ

Retriever type does not determine authorization.

Every retriever should operate inside the same security boundary.

Authentication
      ↓
Authorization
      ↓
Retrieval

This applies to:

VectorStore
Multi-Query
Self-Query
Parent-Document

For Multi-Query:

Every generated query
        ↓
Same authorization boundary

For Parent-Document:

Child access
        ↓
Parent access

must also be validated.

A user should never gain access to a parent document merely because an authorized child matched.


28. Retriever Selection PrinciplesΒΆ

Principle 1ΒΆ

Start with the simplest strategy that satisfies the requirements.

Principle 2ΒΆ

Measure retrieval quality before adding complexity.

Principle 3ΒΆ

Solve a specific retrieval problem with each technique.

Principle 4ΒΆ

Keep retrieval capabilities composable.

Principle 5ΒΆ

Separate retrieval from authorization.

Principle 6ΒΆ

Treat generated queries and filters as untrusted intermediate data.

Principle 7ΒΆ

Measure latency and cost alongside retrieval quality.

🎯 Advanced Retriever Decision Matrix¢

Requirement LlamaIndex LangChain / Generic Pattern
Semantic search Vector Index Retriever VectorStore Retriever
Exact terminology BM25 Retriever BM25 / keyword retriever
Multiple query perspectives Query Fusion Retriever MultiQueryRetriever
Metadata-aware retrieval Metadata filters SelfQueryRetriever
Small chunks + large context Auto-Merging Retriever ParentDocumentRetriever
Follow references between nodes Recursive Retriever Custom graph / reference traversal
Retrieve whole relevant documents Document Summary Retriever Custom document-level retrieval
Relevance + diversity MMR / postprocessing MMR search
Dense + sparse retrieval Hybrid retrieval Ensemble / hybrid retrieval
Improve candidate ranking Re-ranker Re-ranker
Complex retrieval selection Router / Agentic retrieval Router / Agentic retrieval

Selection PrincipleΒΆ

Do not ask:

"What is the best retriever?"

Ask:

"What retrieval failure am I trying to solve?"

---

# 29. Recommended Evolution Path

A practical enterprise retrieval journey is:

```mermaid
flowchart LR
    A["VectorStore Baseline"] --> B["Evaluate"]

    B --> C{"Problem Identified?"}

    C -->|"Query Ambiguity"| D["Multi-Query"]
    C -->|"Metadata Constraints"| E["Self-Query"]
    C -->|"Context Loss"| F["Parent-Document"]

    D --> G["Evaluate"]
    E --> G
    F --> G

    G --> H["Hybrid / Re-ranking"]
    H --> I["Production RAG"]

This avoids premature optimization.


30. Beyond the Four RetrieversΒΆ

These four retrievers are only the beginning of advanced retrieval engineering.

Later techniques can address additional problems:

Hybrid Search
        ↓
Semantic + Keyword Retrieval

Re-ranking
        ↓
Better Candidate Ordering

Contextual Compression
        ↓
Reduce Context Noise

Graph RAG
        ↓
Relationship-Based Retrieval

SQL RAG
        ↓
Structured Data Retrieval

Knowledge Graphs
        ↓
Entity + Relationship Retrieval

Agentic RAG
        ↓
Iterative Retrieval Decisions

The important point is that these techniques build upon the retrieval foundations established here.


31. Production Retrieval ArchitectureΒΆ

A mature enterprise architecture may eventually evolve into:

flowchart TD
    A["User"] --> B["RAG API"]

    B --> C["Authentication"]
    C --> D["Authorization"]

    D --> E["Query Processing"]

    E --> F["Retriever Router"]

    F --> G["VectorStore"]
    F --> H["MultiQuery"]
    F --> I["SelfQuery"]
    F --> J["ParentDocument"]
    F --> K["Hybrid"]
    F --> L["Graph"]
    F --> M["SQL"]

    G --> N["Candidate Pool"]
    H --> N
    I --> N
    J --> N
    K --> N
    L --> N
    M --> N

    N --> O["Deduplication"]
    O --> P["Re-ranking"]
    P --> Q["Context Selection"]

    Q --> R["Prompt Assembly"]
    R --> S["LLM"]

    S --> T["Response Validation"]
    T --> U["Citations"]
    U --> V["Enterprise Response"]

    B --> W["Observability"]
    F --> W
    N --> W
    P --> W
    S --> W

This is the direction toward production retrieval engineering.


32. Framework-Agnostic Retrieval InterfaceΒΆ

An enterprise platform should ideally expose a stable abstraction.

from abc import ABC, abstractmethod


class Retriever(ABC):

    @abstractmethod
    def retrieve(self, query: str):
        pass

Implementations can include:

class VectorStoreRetriever(Retriever):
    ...


class MultiQueryRetriever(Retriever):
    ...


class SelfQueryRetriever(Retriever):
    ...


class ParentDocumentRetriever(Retriever):
    ...


class HybridRetriever(Retriever):
    ...


class GraphRetriever(Retriever):
    ...


class SQLRetriever(Retriever):
    ...

The application can therefore remain independent of the underlying framework.

Enterprise RAG Application
           ↓
     Retriever Interface
           ↓
 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 ↓         ↓         ↓
Vector   Hybrid    Graph

This is particularly useful when building reusable enterprise AI platforms.


33. Retrieval Strategy Cheat SheetΒΆ

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ RETRIEVER CHEAT SHEET                         β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ VectorStore                                   β”‚
β”‚ β†’ Simple semantic retrieval                   β”‚
β”‚                                               β”‚
β”‚ Multi-Query                                   β”‚
β”‚ β†’ Multiple semantic perspectives              β”‚
β”‚                                               β”‚
β”‚ Self-Query                                    β”‚
β”‚ β†’ Natural language + metadata filters         β”‚
β”‚                                               β”‚
β”‚ Parent-Document                               β”‚
β”‚ β†’ Small retrieval + large context             β”‚
β”‚                                               β”‚
β”‚ Hybrid                                        β”‚
β”‚ β†’ Semantic + lexical search                   β”‚
β”‚                                               β”‚
β”‚ Re-ranking                                    β”‚
β”‚ β†’ Improve candidate ordering                  β”‚
β”‚                                               β”‚
β”‚ Graph RAG                                     β”‚
β”‚ β†’ Relationship-based retrieval                β”‚
β”‚                                               β”‚
β”‚ SQL RAG                                       β”‚
β”‚ β†’ Structured data retrieval                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

34. Key TakeawaysΒΆ

  • There is no universally best retriever.
  • VectorStore Retrieval is the natural baseline.
  • Multi-Query Retrieval improves retrieval coverage through multiple query perspectives.
  • Self-Query Retrieval combines semantic search with structured metadata filtering.
  • Parent-Document Retrieval separates retrieval granularity from generation context.
  • Retriever selection should be driven by the problem being solved.
  • Advanced retrievers can be composed into larger retrieval pipelines.
  • More retrieval complexity does not automatically mean better results.
  • Every advanced technique introduces additional latency, cost, or operational complexity.
  • Retrieval quality should be evaluated using metrics such as Recall@K, Precision@K, MRR, and NDCG.
  • End-to-end RAG quality must also be measured through context and answer-quality metrics.
  • Authorization must remain independent of retrieval strategy.
  • LLM-generated queries and filters should be treated as untrusted intermediate data.
  • Parent documents must be protected by the same access-control model as child chunks.
  • Production retrieval requires strong observability.
  • Framework-independent retrieval interfaces help reduce application coupling.
  • The four retrievers covered so far provide the foundation for more advanced techniques such as Hybrid Search, Re-ranking, Graph RAG, SQL RAG, and Agentic RAG.

The central engineering principle is:

Don't ask:

"Which retriever is the best?"

Ask:

"What retrieval problem are we solving?"
                ↓
"What capability addresses it?"
                ↓
"Does evaluation prove the improvement?"
                ↓
"Is the additional complexity justified?"

🧭 Chapter Navigation¢

Part V β€” Advanced Retrieval-Augmented GenerationΒΆ

Previous:
04. Parent-Document Retriever

Next:
Enterprise-retrieval-engineering/01 Contextual Compression Retriever

Section:
01 β€” Core Retrieval Engineering

Retrieval Engineering PathΒΆ

01 VectorStore Retriever
          ↓
02 Multi-Query Retriever
          ↓
03 Self-Query Retriever
          ↓
04 Parent-Document Retriever
          ↓
05 Retriever Comparison
          ↓
06 Hybrid Search

Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β€” One Chapter at a Time.