Skip to content

Contextual Compression RetrieverΒΆ

πŸ“– OverviewΒΆ

A Contextual Compression Retriever improves RAG retrieval by reducing irrelevant information from documents returned by a retriever.

A traditional retriever may successfully identify relevant documents, but those documents can still contain a large amount of information that is unrelated to the user's question.

For example:

User Query

"What authentication mechanism is required
for production APIs?"

The retriever may return:

Production API Security Guide
β”œβ”€β”€ Authentication
β”œβ”€β”€ Authorization
β”œβ”€β”€ Network Security
β”œβ”€β”€ Rate Limiting
β”œβ”€β”€ Logging
β”œβ”€β”€ Monitoring
β”œβ”€β”€ Deployment
└── Incident Response

The document is relevant, but the LLM may only need:

Authentication
Authorization

Contextual compression introduces an additional processing stage:

Query
  ↓
Retriever
  ↓
Candidate Documents
  ↓
Contextual Compression
  ↓
Relevant Content
  ↓
LLM

The goal is not simply to retrieve fewer documents.

The goal is to provide the LLM with more relevant context with less unnecessary information.


🎯 Learning Objectives¢

After completing this chapter, you will be able to:

  • Understand contextual compression in RAG
  • Understand why retrieved documents may contain irrelevant information
  • Differentiate retrieval from context compression
  • Understand the role of a base retriever
  • Understand document compressors
  • Implement a contextual compression pipeline
  • Understand embedding-based compression
  • Understand LLM-based compression
  • Understand reranking-based compression
  • Combine compression with advanced retrievers
  • Evaluate compression quality
  • Identify common compression failure modes
  • Design production-ready contextual compression architectures

1. The Context Problem in RAGΒΆ

A basic RAG pipeline typically looks like:

User Query
    ↓
Retriever
    ↓
Retrieved Documents
    ↓
Prompt
    ↓
LLM
    ↓
Response

Suppose the retriever returns:

Document A β†’ 1,800 tokens
Document B β†’ 1,500 tokens
Document C β†’ 1,700 tokens
Document D β†’ 1,200 tokens

Total:

6,200 tokens

But only:

900 tokens

may actually be useful for answering the question.

The LLM still receives:

6,200 tokens

unless the application introduces another mechanism to reduce the context.

This can lead to:

  • Higher token usage
  • Higher cost
  • Higher latency
  • More irrelevant context
  • More context competition
  • Lower answer quality

2. What Is Contextual Compression?ΒΆ

Contextual compression takes retrieved documents and keeps only the information relevant to the current query.

The basic architecture is:

User Query
     ↓
Base Retriever
     ↓
Candidate Documents
     ↓
Contextual Compressor
     ↓
Compressed Context
     ↓
Prompt Assembly
     ↓
LLM
     ↓
Response

The important distinction is:

Retriever
    ↓
Find potentially relevant information

Compressor
    ↓
Remove information that is not useful for this query

3. Retrieval vs CompressionΒΆ

Retrieval asks:

"Which documents might contain the answer?"

Compression asks:

"Which parts of those documents are useful for answering this specific query?"

Therefore:

Retrieval
    ↓
Candidate Evidence

Compression
    ↓
Focused Evidence

This creates a two-stage retrieval architecture:

Broad Retrieval
      ↓
Context Refinement

4. Why Do We Need Contextual Compression?ΒΆ

Consider an enterprise policy document:

Remote Work Policy

1. Introduction
2. Eligibility
3. Working Arrangements
4. Approval Process
5. Exceptions
6. Security Requirements
7. Manager Responsibilities
8. Compliance
9. Reporting

User query:

"Who approves exceptions to remote work?"

A retriever may return the entire policy.

But the answer may primarily exist in:

Section 4 β€” Approval Process
Section 5 β€” Exceptions
Section 7 β€” Manager Responsibilities

Contextual compression can reduce:

Entire Policy

to:

Relevant Sections

This allows the generation layer to work with a smaller and more focused context.


5. Core ArchitectureΒΆ

flowchart LR
    A["User Query"] --> B["Base Retriever"]
    B --> C["Candidate Documents"]
    C --> D["Contextual Compressor"]
    D --> E["Relevant Passages"]
    E --> F["Context Selector"]
    F --> G["Prompt Assembly"]
    G --> H["LLM"]
    H --> I["Response"]

The architecture separates:

Candidate Discovery

from:

Evidence Refinement

This separation is important in production RAG systems because retrieval and context optimization have different responsibilities.


6. Contextual Compression ComponentsΒΆ

A typical implementation contains two primary components:

Contextual Compression Retriever
        β”‚
        β”œβ”€β”€ Base Retriever
        β”‚
        └── Base Compressor

Base RetrieverΒΆ

Responsible for:

Query
 ↓
Candidate Documents

Base CompressorΒΆ

Responsible for:

Candidate Documents
 ↓
Relevant Content

This allows compression to be added without completely redesigning the existing retrieval layer.


7. Basic LangChain ExampleΒΆ

LangChain provides a ContextualCompressionRetriever abstraction.

A simplified example:

from langchain.retrievers import ContextualCompressionRetriever

compression_retriever = ContextualCompressionRetriever(
    base_compressor=compressor,
    base_retriever=retriever
)

documents = compression_retriever.invoke(
    "What authentication mechanism is required?"
)

for document in documents:
    print(document.page_content)

The important architecture is:

ContextualCompressionRetriever
            β”‚
            β”œβ”€β”€ Base Retriever
            β”‚
            └── Compressor

The application can continue using the retriever interface while the compression stage remains encapsulated.


8. Base RetrieverΒΆ

For example:

base_retriever = vector_store.as_retriever(
    search_kwargs={
        "k": 10
    }
)

The pipeline becomes:

Query
 ↓
Retrieve Top 10
 ↓
Compression
 ↓
Relevant Context

The advantage is that retrieval can remain relatively broad.

The compressor then controls what is actually passed forward.


9. Why Retrieve More Before Compressing?ΒΆ

Suppose:

Top-K = 3

The correct evidence may not appear in those three documents.

Increasing retrieval to:

Top-K = 10

can improve recall.

However, sending all 10 documents to the LLM may increase context size significantly.

Contextual compression provides a middle ground:

Retrieve 10
     ↓
Compress
     ↓
Keep Relevant Evidence
     ↓
LLM

This gives the system an opportunity to optimize:

Recall
+
Context Relevance

10. Extractive CompressionΒΆ

The simplest form of compression is extractive compression.

The system keeps relevant sentences or passages from the original document.

Example:

Original Document

The platform supports OAuth 2.0 authentication.
Applications must register with the identity service.
API requests require valid access tokens.
Logs are retained for 90 days.
Monitoring is enabled across production environments.

Query:

"How are API requests authenticated?"

Compressed context:

The platform supports OAuth 2.0 authentication.

API requests require valid access tokens.

No new information is generated.

The compressor simply selects relevant evidence.


11. Why Extractive Compression Is ValuableΒΆ

Extractive compression has an important advantage:

Original Evidence
       ↓
Selected Evidence

rather than:

Original Evidence
       ↓
Generated Summary

This reduces the risk of introducing new information.

For enterprise systems, this can be especially useful for:

Legal
Compliance
Finance
Security
Healthcare

where preserving the exact source meaning is important.


12. Embedding-Based CompressionΒΆ

Another approach is to calculate semantic similarity between the query and smaller pieces of retrieved documents.

The pipeline becomes:

flowchart TD
    A["Retrieved Document"] --> B["Passage Splitter"]
    B --> C["Individual Passages"]
    C --> D["Passage Embeddings"]

    E["Query Embedding"] --> F["Similarity Calculation"]
    D --> F

    F --> G["Similarity Scores"]
    G --> H["Relevant Passages"]

Conceptually:

for passage in passages:

    score = similarity(
        query_embedding,
        passage_embedding
    )

    if score >= threshold:
        selected.append(passage)

13. Embedding Compression ExampleΒΆ

Suppose the query is:

"What is the API rate limit?"

Retrieved document contains:

P1 β†’ Authentication
P2 β†’ Rate Limits
P3 β†’ Logging
P4 β†’ Deployment
P5 β†’ Monitoring

Similarity scores:

P1 β†’ 0.42
P2 β†’ 0.91
P3 β†’ 0.31
P4 β†’ 0.22
P5 β†’ 0.37

If:

threshold = 0.70

then:

P2

is retained.

The compression flow becomes:

Large Document
      ↓
Passages
      ↓
Embedding Similarity
      ↓
High-Scoring Passages

14. Similarity ThresholdsΒΆ

Embedding-based compression requires threshold tuning.

For example:

threshold = 0.80

may remove too much information.

While:

threshold = 0.40

may keep too much information.

There is no universal threshold.

It depends on:

  • Embedding model
  • Similarity metric
  • Passage size
  • Document type
  • Query distribution
  • Evaluation dataset

Therefore, threshold values should be established through evaluation rather than copied blindly between systems.


15. LLM-Based CompressionΒΆ

An LLM can also extract query-relevant information.

Conceptually:

Query
+
Retrieved Document
        ↓
Compression LLM
        ↓
Relevant Evidence

A simple prompt might be:

compression_prompt = """
Extract only the information relevant to the question.

Question:
{query}

Document:
{document}

Rules:
- Do not introduce new information.
- Do not infer missing facts.
- Preserve important conditions and exceptions.
- Preserve numbers and dates.
"""

The compressor is not supposed to answer the question.

It is supposed to extract evidence.


16. Compression LLM vs Answer LLMΒΆ

It is important to separate these responsibilities.

Compression LLMΒΆ

Document
   ↓
Relevant Evidence

Answer LLMΒΆ

Relevant Evidence
   ↓
Final Answer

Architecture:

flowchart LR
    A["Retrieved Documents"] --> B["Compression LLM"]
    B --> C["Relevant Evidence"]
    C --> D["Prompt Assembly"]
    D --> E["Answer LLM"]
    E --> F["Final Response"]

This separation can make the pipeline easier to reason about and evaluate.


17. Reranking-Based CompressionΒΆ

A reranker can score passages based on their relevance to the query.

For example:

50 candidate passages
        ↓
Reranker
        ↓
Top 10 passages

Pipeline:

flowchart LR
    A["Query"] --> B["Retriever"]
    B --> C["Candidate Documents"]
    C --> D["Passage Extraction"]
    D --> E["Reranker"]
    E --> F["Top Relevant Passages"]
    F --> G["LLM"]

Reranking is often used before compression when the candidate set is large.


18. Reranking vs CompressionΒΆ

These concepts are related but different.

RerankingΒΆ

Changes the ordering:

A
B
C
D
E

↓

C
A
E
B
D

CompressionΒΆ

Reduces the content:

A
B
C
D
E

↓

C
A

Therefore:

Reranking
    ↓
Which candidates are most relevant?

Compression
    ↓
Which content should remain?

They can be used together.


19. Combined Reranking and CompressionΒΆ

A production pipeline could be:

Query
 ↓
Retriever
 ↓
Candidate Documents
 ↓
Reranker
 ↓
Top Candidates
 ↓
Contextual Compressor
 ↓
Relevant Evidence
 ↓
LLM

Architecture:

flowchart TD
    A["User Query"] --> B["Retriever"]
    B --> C["Candidate Documents"]
    C --> D["Reranker"]
    D --> E["Top Candidates"]
    E --> F["Contextual Compressor"]
    F --> G["Relevant Evidence"]
    G --> H["Prompt Assembly"]
    H --> I["LLM"]

This combination becomes increasingly useful as retrieval systems scale.


20. Contextual Compression with Multi-QueryΒΆ

Contextual compression can also be combined with Multi-Query Retrieval.

User Query
    ↓
Multi-Query Generation
    ↓
Query A / B / C
    ↓
Multiple Retrieval Paths
    ↓
Candidate Pool
    ↓
Compression
    ↓
Final Context

Architecture:

flowchart TD
    A["User Query"] --> B["Multi-Query Generator"]

    B --> C["Query A"]
    B --> D["Query B"]
    B --> E["Query C"]

    C --> F["Retriever"]
    D --> F
    E --> F

    F --> G["Candidate Pool"]
    G --> H["Deduplication"]
    H --> I["Contextual Compression"]
    I --> J["Context"]
    J --> K["LLM"]

This is useful when retrieval recall is important but the resulting candidate context is too large.


21. Contextual Compression with Self-QueryΒΆ

Self-Query Retrieval can be used before compression.

For example:

"Find German HR policies from 2025
about parental leave."

Self-Query may produce:

Semantic Query:
parental leave

Filters:
country = Germany
department = HR
year = 2025

Then:

Self-Query
    ↓
Filtered Retrieval
    ↓
Contextual Compression
    ↓
LLM

Architecture:

flowchart LR
    A["Natural Language Query"] --> B["Self-Query"]

    B --> C["Semantic Query"]
    B --> D["Metadata Filters"]

    C --> E["Vector Search"]
    D --> E

    E --> F["Retrieved Documents"]
    F --> G["Compression"]
    G --> H["LLM"]

This can reduce the amount of irrelevant content entering the compression stage.


22. Contextual Compression with Parent-Document RetrievalΒΆ

Parent-Document Retrieval can also be combined with compression.

Query
 ↓
Child Retrieval
 ↓
Parent Resolution
 ↓
Parent Documents
 ↓
Compression
 ↓
Relevant Parent Sections
 ↓
LLM

Architecture:

flowchart TD
    A["Query"] --> B["Child Retriever"]
    B --> C["Child Matches"]
    C --> D["Parent IDs"]
    D --> E["Parent Store"]
    E --> F["Parent Documents"]
    F --> G["Contextual Compressor"]
    G --> H["Relevant Sections"]
    H --> I["LLM"]

This can provide:

Precise Retrieval
+
Broader Context
+
Controlled Context Size

23. Compression OrderingΒΆ

There are multiple possible pipeline designs.

Option AΒΆ

Retrieve
 ↓
Rerank
 ↓
Compress

Option BΒΆ

Retrieve
 ↓
Compress
 ↓
Rerank

Option CΒΆ

Retrieve
 ↓
Compress
 ↓
Context Selection

There is no universal ordering.

The correct architecture depends on:

  • Candidate Count
  • Compression Cost
  • Reranker Cost
  • Latency Requirements
  • Context Size

For example, if a retriever returns thousands of candidates, an inexpensive filtering stage may be necessary before applying an expensive LLM compressor.


24. Contextual Compression and ChunkingΒΆ

Chunking and compression happen at different stages.

ChunkingΒΆ

Usually occurs during ingestion:

Document
 ↓
Chunks
 ↓
Embeddings
 ↓
Vector Store

CompressionΒΆ

Occurs during retrieval:

Query
 ↓
Retrieved Documents
 ↓
Relevant Content

Therefore:

Chunking
    ↓
Controls retrieval units

Compression
    ↓
Controls generation context

25. Contextual Compression vs SummarizationΒΆ

These are also different.

SummarizationΒΆ

Document
 ↓
Summary

The objective is to represent the overall document.

Contextual CompressionΒΆ

Query + Document
 ↓
Query-Relevant Evidence

The objective is to preserve information relevant to the current query.

Therefore:

Contextual compression is query-aware, while ordinary summarization does not necessarily depend on the user's question.


26. Compression and Information LossΒΆ

Compression introduces an important risk:

Too Much Compression
        ↓
Evidence Loss
        ↓
Incorrect Answer

Consider:

Employees may work remotely up to three days
per week, subject to manager approval.

An aggressive compressor might return:

Employees may work remotely up to three days.

The condition:

subject to manager approval

has been removed.

The meaning has changed.

Therefore, compression must preserve:

  • Conditions
  • Exceptions
  • Restrictions
  • Numbers
  • Dates
  • Thresholds
  • Qualifications
  • Definitions

27. Compression and FaithfulnessΒΆ

The compressor should not invent information.

Bad:

Source:

Employees may work remotely up to three days.

Compressed:

Employees are entitled to three remote-work days.

The word:

entitled

introduces a stronger interpretation.

Better:

Employees may work remotely up to three days.

The compression stage should therefore behave as an evidence transformation, not an answer-generation stage.


28. High-Stakes Enterprise ApplicationsΒΆ

Contextual compression requires additional care for:

Legal
Compliance
Financial
Healthcare
Security

A safer architecture may prefer:

Retrieve
 ↓
Rerank
 ↓
Extractive Compression
 ↓
Citation
 ↓
LLM
 ↓
Response Validation

rather than unrestricted generative summarization.

The objective should be:

Reduce Noise
without
Changing Meaning

29. Preserving CitationsΒΆ

Compression must preserve source metadata.

For example:

{
  "content": "Employees may work remotely up to three days per week.",
  "source": "remote-work-policy.pdf",
  "page": 14,
  "section": "Working Arrangements"
}

The compressed context can retain:

Source:
remote-work-policy.pdf

Page:
14

Section:
Working Arrangements

Content:
Employees may work remotely up to three days per week.

This allows downstream components to generate reliable citations.


30. Compression with Citation MetadataΒΆ

Architecture:

flowchart TD
    A["Retrieved Document"] --> B["Compressor"]
    B --> C["Relevant Evidence"]

    A --> D["Source Metadata"]

    C --> E["Evidence + Metadata"]
    D --> E

    E --> F["Prompt Assembly"]
    F --> G["LLM"]
    G --> H["Citation Validation"]
    H --> I["Enterprise Response"]

The key rule is:

Compression should never break source traceability.


31. Structured Compression OutputΒΆ

For production systems, structured output can be useful.

Example:

{
  "relevant": true,
  "content": "Employees may work remotely up to three days per week, subject to manager approval.",
  "source": "remote-work-policy.pdf",
  "page": 14,
  "section": "Working Arrangements"
}

This gives downstream systems access to:

Content
+
Source
+
Page
+
Section

rather than only an unstructured string.


32. Framework-Agnostic Compressor InterfaceΒΆ

An enterprise AI platform can define a generic compressor interface.

from abc import ABC, abstractmethod


class DocumentCompressor(ABC):

    @abstractmethod
    def compress(
        self,
        query: str,
        documents: list
    ) -> list:
        pass

Possible implementations:

class EmbeddingCompressor(DocumentCompressor):
    ...


class RerankingCompressor(DocumentCompressor):
    ...


class LLMCompressor(DocumentCompressor):
    ...


class ExtractiveCompressor(DocumentCompressor):
    ...

This keeps the application independent from a particular AI framework.


33. Composable Retrieval ArchitectureΒΆ

The retrieval architecture can now be expressed as capabilities:

Retriever
    ↓
Candidate Documents
    ↓
Reranker
    ↓
Compressor
    ↓
Context Selector
    ↓
Prompt Builder
    ↓
LLM

Architecture:

flowchart LR
    A["Query"] --> B["Retriever"]
    B --> C["Reranker"]
    C --> D["Compressor"]
    D --> E["Context Selector"]
    E --> F["Prompt Builder"]
    F --> G["LLM"]

Each component can be independently replaced.

This is a useful design for enterprise AI platforms.


34. Context Compression MetricsΒΆ

Compression should not be evaluated only by how many tokens it removes.

Important metrics include:

Compression Ratio
Context Relevance
Information Retention
Answer Accuracy
Faithfulness
Latency
Cost

Compression RatioΒΆ

Conceptually:

Compressed Tokens
──────────────────
Original Tokens

For example:

Original Context = 5,000 tokens

Compressed Context = 1,000 tokens

Compression Ratio = 20%

But:

Lower ratio
β‰ 
Better system

If important evidence is removed, answer quality can decrease.


35. Quality vs CompressionΒΆ

The relationship can be visualized conceptually:

Answer Quality
     ↑
     β”‚
     β”‚              ●
     β”‚           ●
     β”‚        ●
     β”‚      ●
     β”‚    ●
     β”‚  ●
     β”‚ ●
     └────────────────────────→
       Compression Aggressiveness

Initially:

Compression
    ↓
Less Noise
    ↓
Better Context

Eventually:

Compression
    ↓
Evidence Loss
    ↓
Lower Answer Quality

Therefore, the goal is:

Maximum useful context reduction without losing answer-critical evidence.


36. Common Failure ModesΒΆ

36.1 Over-CompressionΒΆ

Too Much Information Removed
        ↓
Missing Evidence
        ↓
Incorrect Answer

36.2 Under-CompressionΒΆ

Almost Entire Document Retained
        ↓
Minimal Token Savings
        ↓
Limited Benefit

36.3 Qualification LossΒΆ

Original:

Employees may work remotely up to three days,
subject to manager approval.

Compressed:

Employees may work remotely up to three days.

The qualification was lost.


36.4 Citation LossΒΆ

The compressor removes:

Document
Page
Section
Chunk ID

making citation generation difficult.


36.5 Hallucinated CompressionΒΆ

The LLM compressor adds facts that were not present in the source.

This is especially dangerous for high-stakes applications.


37. Guardrails for LLM CompressionΒΆ

A production compression prompt should contain explicit constraints.

Example:

You are a retrieval context compressor.

Given a user query and retrieved document:

1. Extract only information relevant to the query.
2. Do not introduce new facts.
3. Do not infer missing information.
4. Preserve numbers, dates, conditions,
   exceptions, and qualifications.
5. Preserve source metadata.
6. Do not answer the user's question.
7. If no relevant information exists,
   return no evidence.

The compressor should therefore be treated as:

Evidence Extraction

rather than:

Answer Generation

38. Production RAG ArchitectureΒΆ

A mature enterprise architecture may look like:

flowchart TD
    A["Client"] --> B["RAG API"]

    B --> C["Authentication"]
    C --> D["Authorization"]
    D --> E["Query Processing"]

    E --> F["Retriever"]
    F --> G["Candidate Documents"]

    G --> H["Deduplication"]
    H --> I["Reranker"]
    I --> J["Contextual Compressor"]
    J --> K["Context Selector"]

    K --> L["Prompt Assembly"]
    L --> M["LLM"]

    M --> N["Response Validation"]
    N --> O["Citation"]
    O --> P["Enterprise Response"]

    F --> Q["Observability"]
    I --> Q
    J --> Q
    M --> Q

Observability should span the complete pipeline:

Retrieval
   ↓
Ranking
   ↓
Compression
   ↓
Generation
   ↓
Validation

39. Decision FlowΒΆ

flowchart TD
    A["Retrieved Documents"] --> B{"Too Much Irrelevant Context?"}

    B -->|No| C["Use Retrieved Context"]

    B -->|Yes| D{"Need Better Candidate Ordering?"}

    D -->|Yes| E["Add Re-ranking"]

    D -->|No| F{"Need Passage-Level Filtering?"}

    F -->|Yes| G["Add Contextual Compression"]

    F -->|No| H["Improve Retrieval"]

    E --> I{"Still Too Much Context?"}

    I -->|Yes| G
    I -->|No| J["Context Selection"]

This highlights an important engineering principle:

Compression should solve a measured context problem rather than being added simply because it is an advanced technique.


40. Production ChecklistΒΆ

Before deploying contextual compression:

☐ Compression preserves source metadata
☐ Compression does not invent facts
☐ Conditions and exceptions are preserved
☐ Numeric values are preserved
☐ Dates are preserved
☐ Citation mapping is maintained
☐ Compression latency is measured
☐ Compression cost is measured
☐ Information loss is evaluated
☐ Empty compression results are handled
☐ High-stakes use cases have stronger validation
☐ Similarity thresholds are evaluated
☐ Retrieval baseline is available
☐ Answer quality is compared before and after compression

41. When to Use Contextual CompressionΒΆ

Contextual compression is especially useful when:

  • Retrieved documents are large
  • Retrieved chunks contain significant irrelevant information
  • The application has strict context limits
  • Token costs are important
  • Retrieval recall needs to remain high
  • Enterprise documents contain multiple unrelated sections
  • The LLM performs poorly with noisy context

Typical use cases include:

Enterprise Knowledge Assistants
Technical Documentation
Legal Document Search
Compliance Systems
Research Assistants
Financial Knowledge Systems
Security Knowledge Bases
Enterprise Copilots

42. When It May Not Be NecessaryΒΆ

Contextual compression may provide little benefit when:

Documents are already very small

or:

Retrieved chunks are highly focused

or:

The context window is sufficiently large

or:

Compression latency costs more than it saves

For example, a simple FAQ:

Question
+
Answer

may already provide highly focused retrieval context.


A strong production starting architecture is:

flowchart LR
    A["Query"] --> B["Retriever"]
    B --> C["Top-K Candidates"]
    C --> D["Re-ranker"]
    D --> E["Contextual Compressor"]
    E --> F["Context Selector"]
    F --> G["Prompt Assembly"]
    G --> H["LLM"]
    H --> I["Response Validation"]
    I --> J["Enterprise Response"]

The important point is that each layer has a clear responsibility:

Retriever
β†’ Find candidates

Re-ranker
β†’ Order candidates

Compressor
β†’ Remove irrelevant content

Context Selector
β†’ Control final context

LLM
β†’ Generate answer

44. Key TakeawaysΒΆ

  • Contextual Compression reduces irrelevant information from retrieved documents.
  • Retrieval and compression solve different problems.
  • A retriever finds potentially relevant documents.
  • A compressor extracts or filters useful content from those documents.
  • Extractive compression selects existing evidence without generating new facts.
  • Embedding-based compression can filter passages using semantic similarity.
  • LLM-based compression provides strong semantic understanding but adds cost and latency.
  • Reranking and compression are complementary.
  • Compression can be combined with Multi-Query Retrieval.
  • Compression can be combined with Self-Query Retrieval.
  • Compression can be combined with Parent-Document Retrieval.
  • Compression should preserve conditions, exceptions, numbers, dates, and qualifications.
  • Source metadata must survive compression for reliable citations.
  • Compression should not be confused with chunking or summarization.
  • Excessive compression can remove answer-critical evidence.
  • Insufficient compression may provide little benefit.
  • High-stakes systems should use conservative compression strategies and stronger validation.
  • Production systems should measure compression ratio, context relevance, information retention, answer quality, latency, and cost.
  • The objective is not maximum compression.
  • The objective is maximum useful context reduction without losing important evidence.

The central pattern is:

Retrieve Broadly
      ↓
Rank Candidates
      ↓
Compress Context
      ↓
Select Evidence
      ↓
Build Prompt
      ↓
Generate Answer

Or simply:

Find More
    ↓
Keep What Matters
    ↓
Generate from Evidence

🧭 Chapter Navigation¢

Part V β€” Advanced Retrieval-Augmented GenerationΒΆ

Previous:
05. Retriever Comparison

Next:
02. Ensemble Retriever

Section:
02 β€” Enterprise Retrieval Engineering

Enterprise Retrieval Engineering PathΒΆ

01 Contextual Compression Retriever
              ↓
02 Hybrid Search
              ↓
03 Metadata Filtering
              ↓
04 Parent-Child Retrieval
              ↓
05 Multi-Vector Retrieval
              ↓
06 Re-ranking

Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β€” One Chapter at a Time.