Multi-Query RetrieverΒΆ
π OverviewΒΆ
A Multi-Query Retriever improves retrieval by generating multiple alternative queries from a single user question and using those queries to search the knowledge base.
A standard VectorStore Retriever performs:
A Multi-Query Retriever changes this to:
βββ Query 1 βββ Retrieval βββ
β β
User Query βββ LLM βββΌββ Query 2 βββ Retrieval βββΌβββ Combine
β β
βββ Query 3 βββ Retrieval βββ
β
Unique Relevant Documents
The key idea is:
One user question can have multiple semantic interpretations. Generate multiple perspectives, retrieve for each perspective, and combine the results.
This is particularly useful when a query is ambiguous, underspecified, conversational, or likely to miss relevant documents when represented by only one embedding.
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand the limitations of single-query retrieval
- Understand how Multi-Query Retrieval works
- Understand query expansion and query diversification
- Design a Multi-Query Retrieval pipeline
- Implement a Multi-Query Retriever
- Understand result aggregation and deduplication
- Understand the role of the LLM in query generation
- Configure the number of generated queries
- Evaluate Multi-Query Retrieval
- Understand its advantages and limitations
- Identify when Multi-Query Retrieval should be used
- Understand how Multi-Query Retrieval fits into production RAG
1. The Problem with Single-Query RetrievalΒΆ
Consider the user query:
This query could refer to several different concepts:
Cloud Strategy
βββ Cloud Migration
βββ Cloud Architecture
βββ Cloud Security
βββ Cloud Governance
βββ Cloud Cost Optimization
βββ Multi-Cloud Strategy
A traditional VectorStore Retriever creates one embedding:
The problem is that the embedding may not sufficiently represent every possible interpretation.
2. The Multi-Query IdeaΒΆ
Instead of searching with one query, the system asks an LLM to generate several alternative formulations.
For example:
Original queryΒΆ
Generated queriesΒΆ
1. What is the company's cloud migration strategy?
2. What are the organization's cloud architecture principles?
3. What is the company's multi-cloud strategy?
4. What are the organization's cloud governance policies?
5. What is the company's approach to cloud security and cost management?
Each query is then used independently for retrieval.
flowchart TD
A["Original User Query"] --> B["Query Generation LLM"]
B --> C["Query 1"]
B --> D["Query 2"]
B --> E["Query 3"]
B --> F["Query 4"]
C --> G["Vector Search"]
D --> H["Vector Search"]
E --> I["Vector Search"]
F --> J["Vector Search"]
G --> K["Results"]
H --> K
I --> K
J --> K
K --> L["Deduplication"]
L --> M["Combined Retrieval Set"] 3. Single Query vs Multi-QueryΒΆ
Single-query retrievalΒΆ
Multi-query retrievalΒΆ
βββ Query A βββ Search βββ
β β
Original Query ββ LLM ββ Query B βββ Search βββΌβ Combine
β β
βββ Query C βββ Search βββ€
β β
βββ Query D βββ Search βββ
The difference is:
4. Multi-Query Retrieval PipelineΒΆ
A complete pipeline looks like:
flowchart LR
A["User Query"] --> B["LLM Query Generator"]
B --> C["Alternative Query 1"]
B --> D["Alternative Query 2"]
B --> E["Alternative Query 3"]
B --> F["Alternative Query N"]
C --> G["Retriever"]
D --> G
E --> G
F --> G
G --> H["Candidate Documents"]
H --> I["Deduplication"]
I --> J["Combined Document Set"]
J --> K["Context Selection"]
K --> L["LLM"] There are therefore two different LLM/retrieval responsibilities:
5. Why Multiple Queries Improve RetrievalΒΆ
Consider an enterprise knowledge base:
Documents
βββββββββ
Cloud Migration Strategy
Cloud Architecture Standards
Cloud Security Policy
Cloud Governance Framework
Cloud Cost Optimization
Multi-Cloud Strategy
User:
One query might retrieve:
A Multi-Query Retriever could additionally retrieve:
The result is greater retrieval coverage.
Single Query
β
Narrow Retrieval Space
Multi-Query
β
Multiple Semantic Paths
β
Broader Retrieval Coverage
6. Query GenerationΒΆ
The LLM receives the original query and a query-generation instruction.
Conceptually:
System:
Generate several alternative search queries
that represent different interpretations of
the user's question.
User:
How does our organization approach cloud?
The LLM might produce:
How does the company approach cloud migration?
What are the company's cloud architecture principles?
What is the organization's multi-cloud strategy?
What are the company's cloud governance policies?
The generated queries should ideally be:
- Relevant
- Diverse
- Concise
- Search-oriented
- Grounded in the original intent
7. Query DiversityΒΆ
Generating the same query multiple times provides little benefit.
Bad example:
Query 1:
What is the cloud strategy?
Query 2:
What is the company's cloud strategy?
Query 3:
What is the organization's cloud strategy?
Query 4:
Tell me about the cloud strategy.
These queries are linguistically different but semantically almost identical.
A better approach is to generate different retrieval perspectives:
Query 1 β Cloud migration
Query 2 β Cloud architecture
Query 3 β Cloud governance
Query 4 β Multi-cloud
Query 5 β Cloud security
Therefore:
Query diversity matters as much as query count.
8. Query Expansion vs Multi-Query RetrievalΒΆ
These concepts are related but not identical.
| Technique | Purpose |
|---|---|
| Query Expansion | Add terms or concepts to a query |
| Query Rewriting | Transform the query into a better search query |
| Multi-Query Retrieval | Generate multiple alternative queries |
| Query Decomposition | Break one complex question into smaller questions |
Example:
Query expansionΒΆ
Query rewritingΒΆ
Multi-queryΒΆ
1. What are the technical risks of migrating
the payment platform to AWS?
2. What are the expected costs of migrating
the payment platform to AWS?
3. What are the operational risks of AWS migration?
4. What are the security risks of the migration?
5. What is the expected infrastructure cost?
Query decompositionΒΆ
Multi-Query Retrieval therefore focuses specifically on retrieving information through multiple alternative search formulations.
9. Basic LangChain ImplementationΒΆ
LangChain provides a Multi-Query Retriever abstraction.
A simplified example:
from langchain.retrievers.multi_query import MultiQueryRetriever
retriever = MultiQueryRetriever.from_llm(
retriever=vector_store.as_retriever(
search_kwargs={"k": 5}
),
llm=llm
)
documents = retriever.invoke(
"How does our organization approach cloud?"
)
for document in documents:
print(document.page_content)
Conceptually:
The important point is that the Multi-Query Retriever generally sits above another retriever.
It is therefore a retrieval orchestration pattern rather than a replacement for the vector database.
10. Framework-Agnostic ImplementationΒΆ
A simple framework-independent implementation can make the architecture easier to understand.
class MultiQueryRetriever:
def __init__(
self,
query_generator,
base_retriever
):
self.query_generator = query_generator
self.base_retriever = base_retriever
def retrieve(self, query: str):
queries = self.query_generator.generate(query)
documents = []
for generated_query in queries:
results = self.base_retriever.retrieve(
generated_query
)
documents.extend(results)
return self._deduplicate(documents)
def _deduplicate(self, documents):
seen = set()
unique_documents = []
for document in documents:
document_id = document.get("id")
if document_id not in seen:
seen.add(document_id)
unique_documents.append(document)
return unique_documents
Architecture:
flowchart LR
A["Application"] --> B["MultiQueryRetriever"]
B --> C["Query Generator"]
B --> D["Base Retriever"]
C --> E["Query 1"]
C --> F["Query 2"]
C --> G["Query N"]
E --> D
F --> D
G --> D
D --> H["Results"]
H --> I["Deduplication"]
I --> J["Final Documents"] This separation is valuable for enterprise AI architecture because the query-generation model and the retrieval implementation can evolve independently.
11. Result AggregationΒΆ
Each generated query produces its own result set.
For example:
Query 1
βββ Document A
βββ Document B
βββ Document C
Query 2
βββ Document B
βββ Document D
βββ Document E
Query 3
βββ Document A
βββ Document E
βββ Document F
A naΓ―ve combination gives:
The system therefore needs deduplication:
Conceptually:
flowchart TD
A["Query 1 Results"] --> D["Result Aggregator"]
B["Query 2 Results"] --> D
C["Query 3 Results"] --> D
D --> E["Deduplicate"]
E --> F["Combined Candidate Set"] 12. DeduplicationΒΆ
Documents may be returned by multiple generated queries.
A simple document identifier can be used:
def deduplicate(documents):
seen = set()
unique = []
for document in documents:
doc_id = document.metadata.get(
"document_id"
)
if doc_id not in seen:
seen.add(doc_id)
unique.append(document)
return unique
In a production system, the identity might be based on:
The correct identity strategy depends on the ingestion architecture.
13. The Importance of Result RankingΒΆ
Deduplication alone does not determine which documents should appear first.
Suppose:
A document appearing across multiple query result sets may be a strong candidate.
A β appears 2 times
B β appears 2 times
C β appears 1 time
D β appears 1 time
E β appears 2 times
F β appears 1 time
This provides one possible signal:
However, frequency alone is not a sufficient ranking strategy.
Production systems may combine:
This becomes especially important when Multi-Query Retrieval is combined with Re-ranking, covered later in Part V.
14. Multi-Query with Re-rankingΒΆ
A powerful production architecture is:
flowchart TD
A["User Query"] --> B["Query Generator"]
B --> C["Query 1"]
B --> D["Query 2"]
B --> E["Query 3"]
C --> F["Retriever"]
D --> F
E --> F
F --> G["Candidate Pool"]
G --> H["Deduplication"]
H --> I["Re-ranker"]
I --> J["Top Relevant Documents"]
J --> K["Context Builder"]
K --> L["LLM"] This creates a two-stage retrieval strategy:
Stage 1
βββββββ
Multi-Query Retrieval
β
High Recall
Stage 2
βββββββ
Re-ranking
β
High Precision
This is often more useful than simply increasing k.
15. Multi-Query with Metadata FilteringΒΆ
Multi-query retrieval can also be combined with metadata constraints.
Example:
Generated queries:
1. What is the parental leave policy?
2. What are employee parental leave benefits?
3. What is the maternity and paternity leave policy?
Metadata constraint:
Pipeline:
Original Query
β
Query Generation
β
Multiple Queries
β
Semantic Retrieval
β
Metadata Filtering
β
Candidate Documents
This is particularly useful for enterprise multi-tenant systems.
16. Multi-Query and Access ControlΒΆ
A critical production consideration is security.
Suppose the user asks:
The system generates multiple queries.
Every generated query must still operate within the user's authorized data boundary.
flowchart TD
A["User"] --> B["Authorization Context"]
B --> C["Original Query"]
C --> D["Multi-Query Generator"]
D --> E["Query 1"]
D --> F["Query 2"]
D --> G["Query 3"]
E --> H["Authorized Retrieval"]
F --> H
G --> H
H --> I["Documents"] The important principle is:
Query expansion must never bypass authorization boundaries.
Security filtering should be enforced independently of the generated query text.
17. Controlling the Number of QueriesΒΆ
More queries do not automatically mean better retrieval.
For example:
1 query
β
1 retrieval operation
5 queries
β
5 retrieval operations
10 queries
β
10 retrieval operations
This can increase:
- Latency
- LLM query-generation cost
- Vector database load
- Result-processing overhead
- Context-management complexity
A production system therefore needs a balance:
A typical starting configuration might be:
but this should be validated using the application's evaluation dataset rather than treated as a universal rule.
18. Failure ModesΒΆ
Multi-Query Retrieval introduces new failure modes.
Failure 1 β Query DriftΒΆ
The generated query may move away from the user's actual intent.
Original:
"What is the company's remote work policy?"
Generated:
"What are the future trends in remote work?"
The second query is related but not equivalent.
Failure 2 β Redundant QueriesΒΆ
Query 1:
Cloud migration strategy
Query 2:
Company cloud migration strategy
Query 3:
Cloud migration approach
These may retrieve almost identical results.
Failure 3 β Hallucinated Search ConceptsΒΆ
The LLM may introduce concepts that were never present in the original question.
If Kubernetes was never part of the user's intent, this can introduce retrieval drift.
Failure 4 β Retrieval ExplosionΒΆ
If:
then the initial candidate pool could contain up to:
before deduplication.
This increases downstream processing.
Failure 5 β Duplicate ContextΒΆ
Multiple queries may retrieve the same chunks.
Without deduplication:
19. When Multi-Query Retrieval Works WellΒΆ
Multi-Query Retrieval is especially useful for:
Ambiguous QueriesΒΆ
Broad QuestionsΒΆ
Conceptual QuestionsΒΆ
Queries with Multiple PerspectivesΒΆ
Possible dimensions:
20. When Multi-Query May Not Be NecessaryΒΆ
A simple vector retriever may be sufficient for highly specific questions.
Example:
If the knowledge base contains a clearly indexed policy document, multiple generated queries may add unnecessary latency.
The principle is:
Do not introduce Multi-Query Retrieval simply because it is available.
21. Multi-Query vs Other Retrieval StrategiesΒΆ
| Technique | Main Problem Solved |
|---|---|
| VectorStore Retriever | Basic semantic retrieval |
| Multi-Query | Multiple semantic interpretations |
| Self-Query | Metadata-aware queries |
| Parent-Document | Context preservation |
| Contextual Compression | Remove irrelevant content |
| Ensemble | Combine retrieval strategies |
| Hybrid Search | Semantic + lexical retrieval |
| Re-ranking | Improve candidate ordering |
| Router Retriever | Select retrieval source |
| Agentic Retrieval | Iterative retrieval decisions |
Multi-Query is therefore one component in the broader retrieval toolbox.
22. Multi-Query as a Recall OptimizationΒΆ
The central goal can be represented as:
Retrieval Quality
β
βββββββββββββββ΄ββββββββββββββ
β β
Recall Precision
β β
Multi-Query Re-ranking
Query Expansion Filtering
Query Rewriting Scoring
Multi-Query Retrieval primarily helps increase the probability that relevant information appears in the candidate set.
It should therefore generally be thought of as a recall-oriented retrieval technique.
23. Production ArchitectureΒΆ
A production Multi-Query RAG pipeline can look like:
flowchart TD
A["User Request"] --> B["Authentication"]
B --> C["Query Processing"]
C --> D["Multi-Query Generator"]
D --> E["Query A"]
D --> F["Query B"]
D --> G["Query C"]
E --> H["Retriever"]
F --> H
G --> H
H --> I["Candidate Pool"]
I --> J["Deduplication"]
J --> K["Metadata / ACL Filtering"]
K --> L["Re-ranking"]
L --> M["Context Selection"]
M --> N["Prompt Assembly"]
N --> O["LLM"]
O --> P["Response Validation"]
P --> Q["Enterprise Response"]
C --> R["Tracing"]
H --> R
L --> R
O --> R Observability should capture:
Original Query
Generated Queries
Number of Queries
Retrieval Latency per Query
Candidate Count
Deduplicated Count
Re-ranking Scores
Final Context
Token Usage
LLM Latency
Final Response
This makes debugging possible.
24. EvaluationΒΆ
A Multi-Query Retriever should be evaluated against a baseline VectorStore Retriever.
BaselineΒΆ
Multi-QueryΒΆ
Compare:
| Metric | Baseline | Multi-Query |
|---|---|---|
| Recall@K | Measure | Measure |
| Precision@K | Measure | Measure |
| MRR | Measure | Measure |
| NDCG | Measure | Measure |
| Answer Correctness | Measure | Measure |
| Latency | Measure | Measure |
| Token Cost | Measure | Measure |
| Retrieval Operations | Measure | Measure |
The correct question is not:
"Does Multi-Query retrieve more documents?"
It is:
"Does Multi-Query improve useful retrieval enough to justify its latency and cost?"
25. Simple Evaluation ExampleΒΆ
Suppose the evaluation dataset contains:
Baseline:
Multi-Query:
Then:
But suppose latency changes from:
and query-generation cost also increases.
The production decision must consider the complete trade-off:
This is why evaluation and observability become essential in production RAG.
26. Practical ExampleΒΆ
Consider an enterprise knowledge assistant.
UserΒΆ
Query GeneratorΒΆ
1. What are the technical risks of migrating
the payment platform to the cloud?
2. What are the security risks associated
with payment platform cloud migration?
3. What are the financial risks of migrating
the payment platform?
4. What are the operational risks of cloud
migration for payment systems?
5. What compliance risks exist when moving
payment systems to the cloud?
RetrievalΒΆ
Query 1 β Architecture Documents
Query 2 β Security Policies
Query 3 β Financial Analysis
Query 4 β Operations Documents
Query 5 β Compliance Policies
AggregationΒΆ
Re-rankingΒΆ
GenerationΒΆ
This is where Multi-Query Retrieval becomes especially valuable: one user question can require evidence from several knowledge perspectives.
27. Multi-Query Retriever in an Enterprise Retrieval AbstractionΒΆ
A capability-oriented enterprise architecture might define:
from abc import ABC, abstractmethod
from typing import List
class Retriever(ABC):
@abstractmethod
def retrieve(self, query: str) -> List[dict]:
pass
The Multi-Query Retriever can then compose another retriever:
class MultiQueryRetriever(Retriever):
def __init__(
self,
query_generator,
base_retriever
):
self.query_generator = query_generator
self.base_retriever = base_retriever
def retrieve(self, query: str):
generated_queries = (
self.query_generator.generate(query)
)
all_documents = []
for generated_query in generated_queries:
all_documents.extend(
self.base_retriever.retrieve(
generated_query
)
)
return self._deduplicate(all_documents)
def _deduplicate(self, documents):
seen = set()
results = []
for document in documents:
key = document["id"]
if key not in seen:
seen.add(key)
results.append(document)
return results
This creates a composable architecture:
flowchart LR
A["RAG Application"] --> B["Retriever Interface"]
B --> C["MultiQueryRetriever"]
C --> D["Query Generator"]
C --> E["Base Retriever"]
E --> F["VectorStore Retriever"] Later, the same composition model can support:
without changing the application-facing interface.
28. Design PrincipleΒΆ
A useful enterprise design principle is:
Retrieval strategies should be composable capabilities rather than hard-coded application logic.
For example:
Retriever
β
βββ VectorStoreRetriever
βββ MultiQueryRetriever
βββ SelfQueryRetriever
βββ ParentDocumentRetriever
βββ HybridRetriever
βββ EnsembleRetriever
βββ AgenticRetriever
This enables different applications to select different retrieval pipelines.
29. Key TakeawaysΒΆ
- Multi-Query Retrieval generates multiple search queries from one user query.
- It is designed primarily to improve retrieval coverage and recall.
- An LLM can generate alternative semantic perspectives.
- Each generated query is passed through a base retriever.
- Results are combined and usually deduplicated.
- Query diversity is more important than simply generating many queries.
- Multi-Query Retrieval is different from query expansion, rewriting, and decomposition.
- Multi-Query works particularly well for ambiguous and broad questions.
- It can be combined with metadata filtering and access-control constraints.
- Re-ranking can refine the larger candidate pool generated by Multi-Query Retrieval.
- Generating more queries increases latency and retrieval cost.
- Generated queries can suffer from query drift or hallucinated concepts.
- Authorization boundaries must apply to every generated query.
- Production systems should evaluate recall, precision, latency, cost, and answer quality.
- Multi-Query Retrieval should be introduced when its retrieval-quality improvement justifies its operational overhead.
- It can be implemented as a composable layer above a base Retriever.
The central idea is:
One Query
β
Multiple Retrieval Perspectives
β
Broader Candidate Coverage
β
Deduplication
β
Re-ranking / Context Selection
β
Better Evidence
β
Better RAG Response
π§ Chapter NavigationΒΆ
Part V β Advanced Retrieval-Augmented GenerationΒΆ
Previous:
01. VectorStore Retriever
Next:
03. Self-Query Retriever
Section:
01 β Core Retrieval Engineering
Retrieval Engineering PathΒΆ
01 VectorStore Retriever
β
02 Multi-Query Retriever
β
03 Self-Query Retriever
β
04 Parent-Document Retriever
β
05 Retriever Comparison
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.