Ensemble RetrieverΒΆ
π OverviewΒΆ
An Ensemble Retriever combines multiple retrieval strategies and merges their results to improve retrieval quality.
A single retriever often performs well for a particular type of query, but different retrieval techniques have different strengths.
For example:
- Vector search is strong at semantic similarity.
- BM25 is strong at exact keyword matching.
- Metadata filtering is strong when structured constraints are available.
- Specialized retrievers may perform better for particular document structures.
Instead of relying on one retrieval strategy, an ensemble retriever combines multiple retrievers:
User Query
β
ββββββββββββΌβββββββββββ
β β β
Vector Search BM25 Metadata Search
β β β
β β β
Results A Results B Results C
β β β
ββββββββββββΌβββββββββββ
β
Result Fusion
β
Final Ranking
β
Retrieved Context
The central idea is:
Use multiple retrieval signals instead of depending on a single retrieval algorithm.
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand ensemble retrieval
- Understand why multiple retrievers can outperform a single retriever
- Compare semantic and lexical retrieval
- Understand rank fusion
- Understand Reciprocal Rank Fusion (RRF)
- Configure weighted retrievers
- Implement an ensemble retrieval pipeline
- Combine vector and BM25 retrieval
- Understand score normalization challenges
- Design framework-independent ensemble retrieval
- Evaluate ensemble retrieval quality
- Identify common failure modes
- Design production-ready ensemble retrieval architectures
1. Why Use Multiple Retrievers?ΒΆ
Consider the query:
A semantic vector retriever may perform very well.
Now consider:
A lexical retriever such as BM25 may be better because the exact token:
is important.
Now consider:
The query contains both:
Different retrieval mechanisms may contribute different signals.
Therefore:
versus:
2. Semantic vs Lexical RetrievalΒΆ
Two of the most common retrieval approaches are:
Semantic RetrievalΒΆ
Uses embeddings to find conceptually similar content.
Example:
Even though the words differ, the meaning is similar.
Lexical RetrievalΒΆ
Uses keyword or term matching.
Example:
Exact terminology can be extremely important.
3. Why Combine Them?ΒΆ
Semantic retrieval can miss exact identifiers.
Lexical retrieval can miss semantic relationships.
Consider:
Semantic retrieval may find:
while lexical retrieval may focus on exact terms such as:
Combining the two gives:
4. Ensemble Retrieval ArchitectureΒΆ
A typical ensemble architecture looks like:
flowchart TD
A["User Query"] --> B["Retriever A"]
A --> C["Retriever B"]
A --> D["Retriever C"]
B --> E["Results A"]
C --> F["Results B"]
D --> G["Results C"]
E --> H["Result Fusion"]
F --> H
G --> H
H --> I["Final Ranking"]
I --> J["Top-K Results"] Each retriever operates independently.
The ensemble layer then combines the results.
5. Common Ensemble CombinationsΒΆ
Some common combinations include:
A production architecture may therefore look like:
Query
β
βββββββββββββββββ
β β
β β
Vector BM25
Search Search
β β
βββββββββ¬ββββββββ
β
Result Fusion
β
Re-ranking
β
Top-K
6. Basic LangChain ExampleΒΆ
LangChain provides an ensemble retriever abstraction.
A simplified implementation:
from langchain.retrievers import EnsembleRetriever
ensemble_retriever = EnsembleRetriever(
retrievers=[
vector_retriever,
bm25_retriever
],
weights=[
0.7,
0.3
]
)
documents = ensemble_retriever.invoke(
"How do I configure OAuth authentication?"
)
for document in documents:
print(document.page_content)
The important concept is:
Vector Retriever
β
βββ Weight 0.7
β
β
Ensemble
β
β
βββ Weight 0.3
β
BM25 Retriever
The weights determine how strongly each retrieval strategy contributes to the final result.
7. Vector + BM25 EnsembleΒΆ
One of the most useful combinations in RAG systems is:
For example:
ensemble_retriever = EnsembleRetriever(
retrievers=[
vector_retriever,
bm25_retriever
],
weights=[0.7, 0.3]
)
Architecture:
flowchart LR
A["Query"] --> B["Vector Retriever"]
A --> C["BM25 Retriever"]
B --> D["Semantic Results"]
C --> E["Keyword Results"]
D --> F["Fusion"]
E --> F
F --> G["Final Ranking"]
G --> H["Top-K Documents"] This is commonly referred to as hybrid retrieval.
However, ensemble retrieval is the broader concept because it can combine more than just dense and sparse retrieval.
8. Understanding Retrieval WeightsΒΆ
Suppose we have:
Conceptually:
Another configuration could be:
or:
The correct values depend on the query distribution and evaluation results.
There is no universally optimal weighting.
9. Weighting StrategyΒΆ
A useful starting point might be:
For identifier-heavy workloads:
For balanced workloads:
These should be treated as starting points rather than production defaults.
The final configuration should be determined through evaluation.
10. The Score Compatibility ProblemΒΆ
Different retrievers often produce different score ranges.
For example:
BM25 may produce:
Directly adding these values would be problematic.
does not have a meaningful interpretation.
Therefore, ensemble systems need a way to combine retrieval signals.
11. Rank-Based FusionΒΆ
One solution is to use the rank position instead of raw retrieval scores.
For example:
BM25 Results:
Instead of comparing:
we compare:
This makes different retrieval strategies easier to combine.
12. Reciprocal Rank FusionΒΆ
A common rank fusion method is Reciprocal Rank Fusion (RRF).
The basic formula is:
where:
dis a documentrank(d)is the document's rank in a retrieverkis a constant used to reduce the impact of very high rankings
Conceptually:
Retriever A
β
Rank List A
β
βββββββββββββββ
β β
Retriever B β
β β
Rank List B β
β β
ββββββββ¬βββββββ
β
RRF Fusion
β
Combined Ranking
13. RRF ExampleΒΆ
Suppose:
and:
Document A appears:
Document C appears:
Therefore both documents receive strong combined rankings.
This is one of the reasons RRF is useful:
A document does not need to be ranked first by every retriever to become highly ranked in the combined result.
14. Rank Fusion VisualizationΒΆ
Vector BM25
β β
β β
Rank List A Rank List B
β β
βββββββ¬ββββββ
β
RRF
β
Combined Scores
β
Final Ranking
The fusion stage effectively asks:
"Which documents consistently appear near the top across multiple retrieval strategies?"
15. Weighted EnsembleΒΆ
Not all retrievers need to have equal importance.
Suppose:
The architecture becomes:
Query
β
βββββββββββ΄ββββββββββ
β β
Vector BM25
Retriever Retriever
β β
Weight Weight
0.7 0.3
β β
βββββββββββ¬ββββββββββ
β
Fusion
β
Final Ranking
Weighted ensembles allow the system to reflect the characteristics of the workload.
16. Three-Retriever EnsembleΒΆ
Ensemble retrieval is not limited to two retrievers.
For example:
Architecture:
flowchart TD
A["User Query"] --> B["Vector Retriever"]
A --> C["BM25 Retriever"]
A --> D["Metadata Retriever"]
B --> E["Semantic Results"]
C --> F["Lexical Results"]
D --> G["Filtered Results"]
E --> H["Weighted Fusion"]
F --> H
G --> H
H --> I["Final Ranking"]
I --> J["Top-K"] Example weights:
The weights should again be validated through evaluation.
17. Ensemble with Multiple Vector StoresΒΆ
An enterprise system may also combine different vector stores.
For example:
Query
βββ Product Knowledge Index
βββ Technical Documentation Index
βββ Support Knowledge Index
Each index may have a different retrieval strategy.
Query
β
ββββββββββββΌβββββββββββ
β β β
Product Technical Support
Retriever Retriever Retriever
β β β
ββββββββββββΌβββββββββββ
β
Fusion
β
Final Results
This is particularly useful when enterprise knowledge is partitioned into different domains.
18. Domain-Specific Ensemble RetrievalΒΆ
Consider an enterprise platform containing:
A router or classifier may determine the likely domain.
Alternatively, an ensemble can search multiple domain indexes:
Query
β
ββββββββββ¬βββββββββββ¬βββββββββββ¬βββββββββ
β β β β
HR Finance Engineering Legal
β β β β
ββββββββββ΄βββββββββββ΄βββββββββββ΄βββββββββ
β
Fusion
β
Final Results
This can increase recall when the query spans multiple domains.
19. Ensemble Retrieval vs Router RetrievalΒΆ
These concepts should not be confused.
Ensemble RetrieverΒΆ
Multiple retrievers are queried:
Router RetrieverΒΆ
The system chooses one or more retrievers based on the query:
Therefore:
A router can itself be used together with an ensemble.
20. Ensemble + RerankingΒΆ
A strong retrieval pipeline may combine ensemble retrieval with reranking.
flowchart LR
A["User Query"] --> B["Vector Retriever"]
A --> C["BM25 Retriever"]
B --> D["Semantic Results"]
C --> E["Lexical Results"]
D --> F["Result Fusion"]
E --> F
F --> G["Candidate Pool"]
G --> H["Reranker"]
H --> I["Top-K Results"] The responsibilities are:
Vector/BM25
β Generate candidates
Fusion
β Combine retrieval signals
Reranker
β Perform deeper relevance evaluation
This is often stronger than simply increasing k on a single retriever.
21. Ensemble + Contextual CompressionΒΆ
Ensemble retrieval can also feed into contextual compression.
Query
β
Vector Retriever
β
BM25 Retriever
β
Result Fusion
β
Top Candidates
β
Contextual Compression
β
Relevant Context
β
LLM
Architecture:
flowchart TD
A["Query"] --> B["Vector Retriever"]
A --> C["BM25 Retriever"]
B --> D["Vector Results"]
C --> E["BM25 Results"]
D --> F["Result Fusion"]
E --> F
F --> G["Candidate Documents"]
G --> H["Contextual Compression"]
H --> I["Relevant Context"]
I --> J["LLM"] This is useful when the ensemble improves recall but produces a large candidate context.
22. Ensemble + Reranking + CompressionΒΆ
A more advanced pipeline is:
Query
β
Multiple Retrievers
β
Result Fusion
β
Candidate Pool
β
Reranking
β
Contextual Compression
β
Context Selection
β
Prompt Assembly
β
LLM
Architecture:
flowchart TD
A["User Query"] --> B["Vector Retriever"]
A --> C["BM25 Retriever"]
A --> D["Additional Retriever"]
B --> E["Vector Results"]
C --> F["BM25 Results"]
D --> G["Additional Results"]
E --> H["Result Fusion"]
F --> H
G --> H
H --> I["Candidate Pool"]
I --> J["Reranker"]
J --> K["Contextual Compressor"]
K --> L["Context Selector"]
L --> M["Prompt Assembly"]
M --> N["LLM"] This creates a layered retrieval architecture:
23. BM25 ExampleΒΆ
A simple BM25 retriever can be created from documents.
from langchain_community.retrievers import BM25Retriever
bm25_retriever = BM25Retriever.from_documents(
documents
)
bm25_retriever.k = 5
results = bm25_retriever.invoke(
"OAuth authentication"
)
This gives the application a lexical retrieval path.
24. Combining BM25 and Vector RetrievalΒΆ
A simplified example:
from langchain.retrievers import EnsembleRetriever
vector_retriever = vector_store.as_retriever(
search_kwargs={"k": 5}
)
bm25_retriever = BM25Retriever.from_documents(
documents
)
bm25_retriever.k = 5
ensemble_retriever = EnsembleRetriever(
retrievers=[
vector_retriever,
bm25_retriever
],
weights=[
0.7,
0.3
]
)
results = ensemble_retriever.invoke(
"OAuth authentication requirements"
)
The application now has a single retriever interface:
even though multiple retrieval systems are operating underneath it.
25. Framework-Agnostic Ensemble InterfaceΒΆ
For an enterprise AI platform, it is useful to abstract the concept.
from abc import ABC, abstractmethod
class Retriever(ABC):
@abstractmethod
def retrieve(
self,
query: str,
top_k: int
) -> list:
pass
An ensemble can then operate against this interface:
class EnsembleRetriever(Retriever):
def __init__(
self,
retrievers,
weights
):
self.retrievers = retrievers
self.weights = weights
def retrieve(
self,
query: str,
top_k: int
) -> list:
results = []
for retriever, weight in zip(
self.retrievers,
self.weights
):
results.extend(
retriever.retrieve(
query,
top_k
)
)
return self.fuse(results)
The key architectural principle is:
26. DeduplicationΒΆ
Different retrievers may return the same document.
For example:
Without deduplication:
A fusion layer should identify duplicates.
Possible deduplication keys include:
- Document ID
- Chunk ID
- Source ID + page
- Content hash
Example:
def deduplicate(documents):
seen = set()
unique = []
for document in documents:
document_id = document.metadata["chunk_id"]
if document_id not in seen:
seen.add(document_id)
unique.append(document)
return unique
27. Duplicate Content ProblemΒΆ
Duplicate results can distort the ensemble.
For example:
The same document may appear multiple times.
If the system treats each occurrence independently, the document may receive excessive influence.
Therefore:
is often safer.
28. Metadata PreservationΒΆ
Ensemble retrieval should preserve metadata.
Example:
{
"content": "...",
"metadata": {
"source": "security-policy.pdf",
"page": 14,
"chunk_id": "security-14-03",
"retriever": "bm25"
}
}
The metadata can help with:
- Debugging
- Evaluation
- Citation
- Observability
- Retriever analysis
For production systems, consider adding:
29. Observability for Ensemble RetrievalΒΆ
An ensemble system should make individual retriever contributions visible.
Example:
Query
β
Vector Retriever
βββ 10 results
βββ latency: 42 ms
βββ top score: 0.91
BM25 Retriever
βββ 10 results
βββ latency: 18 ms
βββ top score: 14.8
β
Fusion
βββ candidates: 20
βββ duplicates: 6
βββ final candidates: 14
β
Reranker
βββ top 5
This information is extremely useful when debugging retrieval quality.
30. Ensemble Retrieval MetricsΒΆ
Important evaluation metrics include:
Recall@KΒΆ
Measures whether relevant documents appear in the top K results.
is especially important when evaluating candidate generation.
Precision@KΒΆ
Measures how many of the retrieved documents are relevant.
helps evaluate retrieval relevance.
MRRΒΆ
Mean Reciprocal Rank measures how highly the first relevant result appears.
is useful when the first relevant document matters significantly.
NDCGΒΆ
Normalized Discounted Cumulative Gain considers both relevance and ranking position.
This is useful for comparing ranking quality across retrieval strategies.
31. Comparing Single vs Ensemble RetrievalΒΆ
A useful evaluation experiment is:
versus:
Compare:
Example:
| Retrieval Strategy | Recall@10 | MRR | Latency |
|---|---|---|---|
| Vector | 0.78 | 0.69 | 40 ms |
| BM25 | 0.71 | 0.63 | 20 ms |
| Ensemble | 0.86 | 0.77 | 58 ms |
The numbers above are illustrative.
The important point is to measure whether the additional retrieval strategy actually improves the system.
32. Query-Type AnalysisΒΆ
A useful production evaluation approach is to classify queries.
For example:
Semantic Queries
Exact Identifier Queries
Short Queries
Long Queries
Technical Queries
Natural Language Questions
Metadata-Heavy Queries
Then compare retrieval performance.
Example:
| Query Type | Vector | BM25 | Ensemble |
|---|---|---|---|
| Semantic | High | Medium | High |
| Exact Identifier | Medium | High | High |
| Technical | High | High | Very High |
| Metadata-heavy | Medium | Medium | High |
This helps determine whether ensemble retrieval provides meaningful value across different query classes.
33. Latency ConsiderationsΒΆ
Ensemble retrieval usually increases retrieval work.
For example:
Approximate pipeline:
If the retrievers run sequentially.
However, they can often be executed concurrently:
Query
β
βββββββββ΄ββββββββ
β β
Vector BM25
40 ms 20 ms
β β
βββββββββ¬ββββββββ
β
Fusion
The overall retrieval latency can then approach:
rather than:
This is an important production optimization.
34. Parallel RetrievalΒΆ
Conceptually:
from concurrent.futures import ThreadPoolExecutor
def retrieve_parallel(query):
with ThreadPoolExecutor(
max_workers=2
) as executor:
vector_future = executor.submit(
vector_retriever.invoke,
query
)
bm25_future = executor.submit(
bm25_retriever.invoke,
query
)
vector_results = vector_future.result()
bm25_results = bm25_future.result()
return vector_results, bm25_results
Production implementations should use the concurrency model appropriate to the runtime and infrastructure.
35. When Ensemble Retrieval Helps MostΒΆ
Ensemble retrieval is especially useful when:
- Queries contain exact identifiers
- Documents contain technical terminology
- Semantic and lexical signals are both important
- Different retrievers have complementary strengths
- Retrieval recall is insufficient
- The knowledge base contains heterogeneous documents
- Queries vary significantly in style
- A single retrieval strategy has known blind spots
36. When Ensemble Retrieval May Not HelpΒΆ
An ensemble may not be worthwhile when:
or:
or:
or:
Therefore:
Do not add retrievers simply because more retrieval algorithms exist.
Measure the incremental value.
37. Common Failure ModesΒΆ
37.1 Poor Weight SelectionΒΆ
Incorrect weights may cause one retriever to dominate.
may effectively behave like a vector-only system.
37.2 Duplicate ResultsΒΆ
Multiple retrievers may return the same chunks.
Without deduplication:
37.3 Score IncompatibilityΒΆ
Raw scores from different retrievers may not be directly comparable.
Rank-based fusion can help address this problem.
37.4 Increased LatencyΒΆ
More retrievers mean more retrieval operations.
37.5 Increased CostΒΆ
Cloud-hosted retrieval services, rerankers, or additional model calls can increase operational cost.
37.6 No Measurable Quality ImprovementΒΆ
The most important failure mode:
Always compare against a baseline.
38. Production Design PatternΒΆ
A practical enterprise architecture can look like:
flowchart TD
A["User Query"] --> B["Query Processing"]
B --> C["Vector Retriever"]
B --> D["BM25 Retriever"]
B --> E["Domain Retriever"]
C --> F["Result Normalization"]
D --> F
E --> F
F --> G["Deduplication"]
G --> H["Weighted Rank Fusion"]
H --> I["Candidate Pool"]
I --> J["Reranker"]
J --> K["Contextual Compression"]
K --> L["Context Selection"]
L --> M["Prompt Assembly"]
M --> N["LLM"] This architecture separates:
Candidate Generation
β
Result Fusion
β
Precision Optimization
β
Context Optimization
β
Generation
39. Enterprise Retrieval InterfaceΒΆ
A production platform may expose a single interface:
Internally:
EnterpriseRetriever
β
βββ VectorRetriever
βββ BM25Retriever
βββ MetadataRetriever
βββ DomainRetriever
β
Fusion Layer
β
Reranker
β
Context Selection
The application does not need to know which individual retrievers are being used.
This is useful for evolving retrieval architecture without changing downstream application code.
40. Configuration-Driven EnsembleΒΆ
Retriever configuration can be externalized.
Example:
retrieval:
ensemble:
enabled: true
retrievers:
- name: vector
weight: 0.7
top_k: 10
- name: bm25
weight: 0.3
top_k: 10
fusion:
strategy: reciprocal_rank_fusion
deduplication:
enabled: true
This allows retrieval strategies and weights to be changed without modifying application logic.
41. Testing Ensemble RetrievalΒΆ
A production test suite should include:
Example:
def test_ensemble_returns_relevant_document():
results = ensemble_retriever.invoke(
"OAuth authentication"
)
assert len(results) > 0
assert any(
"OAuth" in doc.page_content
for doc in results
)
For production evaluation, use a curated query-answer-document dataset rather than relying only on assertions like the example above.
42. Ensemble Retrieval Evaluation PipelineΒΆ
flowchart LR
A["Evaluation Dataset"] --> B["Baseline Retriever"]
A --> C["Ensemble Retriever"]
B --> D["Baseline Metrics"]
C --> E["Ensemble Metrics"]
D --> F["Comparison"]
E --> F
F --> G["Decision"] Compare:
The ensemble should be adopted only when the improvement justifies the additional complexity.
43. Recommended Starting ConfigurationΒΆ
For a general enterprise knowledge assistant, a reasonable starting architecture is:
Query
β
Vector Retriever
β
BM25 Retriever
β
RRF / Weighted Fusion
β
Deduplication
β
Reranker
β
Contextual Compression
β
Prompt Assembly
β
LLM
This should be treated as an architecture to evaluate rather than a universal production recipe.
44. Decision FrameworkΒΆ
Does the current retriever miss relevant documents?
β
βββββββ΄ββββββ
No Yes
β β
β Is the missing
β information lexical?
β β
β ββββββ΄βββββ
β Yes No
β β β
β Add BM25 Try another
β retrieval signal
β β β
β ββββββ¬ββββββ
β β
β Evaluate
β β
βββββββββββββββ€
β
Does quality improve?
β
βββββββ΄ββββββ
Yes No
β β
Keep Ensemble Reconsider
The core engineering principle is:
Ensemble retrieval should be evidence-driven and evaluation-driven.
45. Production ChecklistΒΆ
Before deploying an ensemble retriever:
β Baseline retriever has been evaluated
β Additional retriever provides complementary recall
β Retriever weights are configurable
β Fusion strategy is defined
β Score incompatibility is handled
β Duplicate documents are removed
β Document IDs are normalized
β Source metadata is preserved
β Individual retriever latency is measured
β Fusion latency is measured
β Individual retriever errors are handled
β Parallel retrieval is considered
β Retrieval metrics are tracked
β Regression tests exist
β Cost impact is measured
β End-to-end RAG quality is evaluated
46. Key TakeawaysΒΆ
- Ensemble retrieval combines multiple retrieval strategies.
- Different retrievers provide different retrieval signals.
- Vector retrieval is strong for semantic similarity.
- BM25 is strong for lexical and exact-term matching.
- Combining dense and sparse retrieval can improve recall.
- Ensemble retrieval can combine more than two retrievers.
- Retrieval weights control the contribution of individual retrievers.
- Raw retrieval scores may not be directly comparable.
- Rank-based fusion provides a useful alternative.
- Reciprocal Rank Fusion is a common rank-fusion approach.
- Duplicate results should be removed before or during fusion.
- Metadata should be preserved throughout the retrieval pipeline.
- Ensemble retrieval can be combined with reranking.
- Ensemble retrieval can be combined with contextual compression.
- Parallel execution can reduce additional retrieval latency.
- More retrievers do not automatically mean better retrieval.
- Ensemble architectures should always be evaluated against a baseline.
- Production decisions should consider quality, latency, cost, and complexity together.
The central pattern is:
Multiple Retrieval Signals
β
Fusion
β
Better Candidates
β
Re-ranking
β
Context Optimization
β
Generation
Or simply:
π§ Chapter NavigationΒΆ
Part V β Advanced Retrieval-Augmented GenerationΒΆ
Previous:
01. Contextual Compression Retriever
Next:
03. Multi-Vector Retriever
Section:
02 β Enterprise Retrieval Engineering
Enterprise Retrieval Engineering PathΒΆ
01 Contextual Compression Retriever
β
02 Ensemble Retriever
β
03 Multi-Vector Retriever
β
04 Time-Weighted Retriever
β
05 Hybrid Search Retriever
β
06 HyDE Retriever
β
07 Router Retriever
β
08 Multi-Stage Retrieval
β
09 Agentic Retrieval
β
10 Re-ranking Techniques
β
11 MMR & Diversity-Aware Retrieval
β
12 Metadata-Aware Retrieval
β
13 Advanced Query Rewriting
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.