Skip to content

Agentic RetrievalΒΆ

πŸ“– OverviewΒΆ

Agentic Retrieval extends traditional retrieval pipelines by introducing an AI-driven decision-making loop that can dynamically determine:

  • What information to retrieve
  • Which retrieval strategy to use
  • Which query to execute
  • Whether additional retrieval is required
  • Whether the retrieved evidence is sufficient
  • How to refine the search
  • When to stop retrieving

Traditional multi-stage retrieval follows a predefined pipeline:

Query
  ↓
Retriever
  ↓
Reranker
  ↓
Context

Agentic Retrieval introduces dynamic control:

Query
  ↓
Understand
  ↓
Plan
  ↓
Retrieve
  ↓
Observe
  ↓
Evaluate
  ↓
Refine
  ↓
Retrieve Again
  ↓
Sufficient Evidence?
  ↓
Generate

The key distinction is:

Multi-stage retrieval follows a predefined retrieval workflow, while Agentic Retrieval dynamically decides what to retrieve and what to do next based on the evidence it observes.

Agentic Retrieval is therefore particularly useful for complex enterprise questions where a single retrieval operation may not provide sufficient evidence.


🎯 Learning Objectives¢

After completing this chapter, you will be able to:

  • Understand Agentic Retrieval
  • Explain the difference between traditional, multi-stage, and agentic retrieval
  • Understand retrieval planning
  • Understand retrieval agents and retrieval tools
  • Design retrieval loops
  • Implement iterative retrieval
  • Implement query refinement
  • Implement retrieval evaluation
  • Implement evidence sufficiency checks
  • Combine multiple retrieval tools
  • Combine Agentic Retrieval with Router Retrieval
  • Combine Agentic Retrieval with Hybrid Search
  • Combine Agentic Retrieval with Graph RAG
  • Combine Agentic Retrieval with SQL RAG
  • Understand stopping criteria
  • Handle retrieval failures
  • Design production guardrails
  • Implement observability for retrieval agents
  • Evaluate Agentic Retrieval systems
  • Understand when Agentic Retrieval should and should not be used

1. Traditional RetrievalΒΆ

Traditional RAG usually follows:

User Query
     ↓
Embedding
     ↓
Vector Search
     ↓
Top-K Documents
     ↓
LLM
     ↓
Answer

The retrieval process is mostly predetermined.

Example:

documents = retriever.invoke(query)

answer = llm.invoke(
    build_prompt(
        query=query,
        documents=documents
    )
)

The system does not normally ask:

Are these documents sufficient?

Should I search another source?

Should I rewrite the query?

Should I search again?

2. Multi-Stage RetrievalΒΆ

Multi-stage retrieval adds multiple predefined operations:

Query
 ↓
Candidate Generation
 ↓
Filtering
 ↓
Reranking
 ↓
Compression
 ↓
Context Selection

The stages are known before execution.

Stage 1
 ↓
Stage 2
 ↓
Stage 3
 ↓
Stage 4

The pipeline itself generally does not dynamically decide that an entirely different retrieval strategy should be used.


3. Agentic RetrievalΒΆ

Agentic Retrieval introduces dynamic decision-making.

flowchart TD
    A["User Query"] --> B["Retrieval Agent"]

    B --> C["Plan"]
    C --> D["Select Retrieval Tool"]

    D --> E["Execute Retrieval"]
    E --> F["Observe Results"]

    F --> G{"Evidence Sufficient?"}

    G -->|Yes| H["Return Evidence"]
    G -->|No| I["Refine Query / Change Strategy"]

    I --> B

    H --> J["Generation LLM"]
    J --> K["Grounded Response"]

The agent can iterate until:

Sufficient Evidence

or:

Stopping Condition

is reached.


4. The Agentic Retrieval LoopΒΆ

The central loop is:

                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚    Query     β”‚
                         β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                                ↓
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚    Plan      β”‚
                         β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                                ↓
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Retrieve   β”‚
                         β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                                ↓
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Observe    β”‚
                         β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                                ↓
                         β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                         β”‚   Evaluate   β”‚
                         β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
                                ↓
                       β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”
                       ↓                 ↓
                  Sufficient?        Insufficient
                       ↓                 ↓
                    Finish        Refine / Search Again
                                         β”‚
                                         └──────→ Plan

This is the defining characteristic of Agentic Retrieval.


5. Why Agentic Retrieval?ΒΆ

Some enterprise questions cannot be answered effectively with one retrieval operation.

Consider:

"What changed in the payment service
after the migration, and did it affect
transaction processing?"

The system may need:

1. Architecture documentation
2. Migration documentation
3. Deployment history
4. Operational data
5. Transaction metrics

A fixed retriever may not know how to connect these sources.

An agent can reason about the retrieval process:

Question
 ↓
Need architecture information
 ↓
Retrieve architecture documents
 ↓
Need migration information
 ↓
Retrieve migration documents
 ↓
Need transaction metrics
 ↓
Query structured data
 ↓
Combine evidence

6. Complex QuestionsΒΆ

Agentic Retrieval is especially useful for questions involving:

Multiple documents
Multiple knowledge sources
Multiple retrieval strategies
Sequential dependencies
Query refinement
Evidence verification
Cross-domain research

For example:

"Why did API latency increase after version 3.2?"

The agent may need:

Release Notes
+
Architecture Documentation
+
Monitoring Data
+
Incident Reports

7. Agentic Retrieval vs Agentic AIΒΆ

These concepts should remain distinct.

Agentic RetrievalΒΆ

The agent's primary responsibility is:

Finding and validating information

General AI AgentΒΆ

The agent may:

Retrieve
Plan
Call APIs
Modify Systems
Execute Workflows
Send Messages
Perform Actions

Therefore:

Agentic Retrieval
β†’ Retrieval-focused agent

AI Agent
β†’ Broader autonomous system

This distinction is important for architecture and security.


8. Agentic Retrieval ComponentsΒΆ

A retrieval agent typically contains:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚      Retrieval Agent        β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Query Understanding         β”‚
β”‚ Planning                    β”‚
β”‚ Tool Selection              β”‚
β”‚ Query Generation            β”‚
β”‚ Retrieval Execution         β”‚
β”‚ Evidence Evaluation         β”‚
β”‚ Iteration                   β”‚
β”‚ Stopping Logic              β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Surrounding infrastructure includes:

Retriever Tools
Vector Stores
Search Engines
SQL Databases
Knowledge Graphs
Document Stores
Rerankers
Observability
Security

9. Retrieval ToolsΒΆ

An agent needs tools through which it can access knowledge.

Example:

tools = [
    vector_search_tool,
    hybrid_search_tool,
    sql_search_tool,
    graph_search_tool,
    document_search_tool
]

The agent decides which tool is appropriate.


10. Tool-Based RetrievalΒΆ

Conceptually:

User Query
     ↓
Agent
     ↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
↓              ↓              ↓
Vector Tool   SQL Tool      Graph Tool
↓              ↓              ↓
Documents     Data          Relationships

The important architectural principle is:

The agent should interact with retrieval capabilities through controlled interfaces rather than directly accessing infrastructure.


11. Retrieval Tool ContractΒΆ

A retrieval tool should have a clear contract.

class RetrievalTool:

    name: str
    description: str

    def execute(
        self,
        query: str,
        filters: dict | None = None
    ):
        raise NotImplementedError

Example:

class VectorSearchTool(RetrievalTool):

    name = "vector_search"

    description = (
        "Search semantic enterprise documentation."
    )

    def execute(self, query, filters=None):

        return vector_store.search(
            query=query,
            filters=filters,
            top_k=20
        )

The agent does not need to know how the vector database works internally.


12. Capability-Based Retrieval ToolsΒΆ

Prefer capabilities such as:

semantic_search
hybrid_search
structured_query
graph_search
policy_search
recent_documents

instead of exposing infrastructure directly:

pinecone_tool
postgres_tool
neo4j_tool

The capability abstraction makes the system easier to evolve.


13. Agentic Retrieval ArchitectureΒΆ

flowchart TD
    A["User Query"] --> B["Retrieval Agent"]

    B --> C["Planner"]
    C --> D["Tool Selection"]

    D --> E["Semantic Search"]
    D --> F["Hybrid Search"]
    D --> G["SQL Search"]
    D --> H["Graph Search"]

    E --> I["Observation"]
    F --> I
    G --> I
    H --> I

    I --> J["Evidence Evaluator"]

    J --> K{"Enough Evidence?"}

    K -->|No| C
    K -->|Yes| L["Context Builder"]

    L --> M["Generation LLM"]
    M --> N["Validated Response"]

14. Query PlanningΒΆ

The agent can break a complex question into retrieval tasks.

Example:

User:

"What changed in PaymentService after the
Kafka migration and did transaction latency increase?"

The agent may plan:

Task 1:
Find PaymentService migration documentation.

Task 2:
Find Kafka migration release information.

Task 3:
Find transaction latency metrics.

Task 4:
Compare timeline and evidence.

The plan creates a retrieval strategy.


15. Query DecompositionΒΆ

A complex question:

"What changed in the payment platform and
why did transaction latency increase?"

can become:

Subquery 1:
"What changes were made to the payment platform?"

Subquery 2:
"What happened to transaction latency?"

Subquery 3:
"What relationship exists between the changes
and latency?"

Each subquery may use a different retrieval source.


16. Decomposition ArchitectureΒΆ

flowchart TD
    A["Complex Query"] --> B["Retrieval Agent"]

    B --> C["Subquery 1"]
    B --> D["Subquery 2"]
    B --> E["Subquery 3"]

    C --> F["Documentation Search"]
    D --> G["Metrics / SQL Search"]
    E --> H["Graph / Documentation Search"]

    F --> I["Evidence"]
    G --> I
    H --> I

    I --> J["Evidence Synthesis"]

This is one of the most useful patterns for enterprise retrieval.


17. Sequential RetrievalΒΆ

Sometimes one retrieval result determines the next query.

Example:

Query
 ↓
Find service
 ↓
Identify dependency
 ↓
Search dependency documentation
 ↓
Find incident
 ↓
Search incident details

The retrieval path becomes:

Query
 ↓
Observation
 ↓
New Query
 ↓
Observation
 ↓
New Query

This cannot always be represented effectively as a static retrieval pipeline.


18. Iterative RetrievalΒΆ

A simplified loop:

max_iterations = 4

for iteration in range(max_iterations):

    plan = agent.plan(query, evidence)

    result = execute_tool(
        plan.tool,
        plan.query
    )

    evidence.extend(result)

    if agent.is_sufficient(
        query,
        evidence
    ):
        break

The important safeguard is:

Maximum Iterations

19. Retrieval StateΒΆ

The agent needs state.

Example:

state = {
    "original_query": query,
    "subqueries": [],
    "retrieved_documents": [],
    "tool_calls": [],
    "observations": [],
    "iteration": 0
}

The state allows the agent to understand:

What has already been searched?
What evidence has been found?
What remains unanswered?

20. Retrieval State MachineΒΆ

stateDiagram-v2
    [*] --> AnalyzeQuery
    AnalyzeQuery --> Plan
    Plan --> Retrieve
    Retrieve --> Observe
    Observe --> Evaluate

    Evaluate --> Complete: Sufficient
    Evaluate --> Refine: Insufficient
    Refine --> Plan

    Complete --> [*]

This provides a useful mental model for implementing Agentic Retrieval with workflow frameworks.


21. Evidence SufficiencyΒΆ

The agent needs to determine:

"Do I have enough evidence?"

Possible signals:

Required entities found
Required subquestions answered
Relevant documents found
Evidence from authoritative sources
Confidence above threshold
No major information gaps

Example:

def is_sufficient(evidence):

    if len(evidence) < 3:
        return False

    if not contains_required_entities(evidence):
        return False

    return relevance_score(evidence) > 0.80

Production systems should combine deterministic checks with model-based evaluation where appropriate.


22. Evidence CompletenessΒΆ

Suppose the question asks:

"What changed, when did it change,
and what was the impact?"

The evidence must cover:

What
When
Impact

If the retrieved evidence only explains:

What

the agent should continue retrieval.

Evidence Coverage:

What       βœ“
When       βœ—
Impact     βœ—

This is better than simply checking whether any documents were retrieved.


23. Evidence MatrixΒΆ

A useful representation is:

Requirement Evidence Status
Change Migration document βœ“
Timeline Release record βœ“
Impact Monitoring data βœ“
Root cause Incident report βœ—

The agent can identify:

Missing:
Root Cause

and perform another retrieval operation.


24. Query RefinementΒΆ

If retrieval fails:

Original Query
 ↓
Poor Results
 ↓
Query Refinement
 ↓
New Retrieval

Example:

Original:
"payment latency"

Refined:
"PaymentService transaction latency
after Kafka migration"

The refined query provides more retrieval-specific terminology.


25. Query ExpansionΒΆ

The agent may generate multiple search formulations.

Original Query
       ↓
 β”Œβ”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”
 ↓     ↓     ↓
Query1 Query2 Query3
 ↓     ↓     ↓
Search Search Search
 β””β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”˜
       ↓
     Fusion

This overlaps with Multi-Query Retrieval.

The distinction is:

Multi-Query Retriever
β†’ Predefined retrieval technique

Agentic Retrieval
β†’ Agent decides whether and when
  to generate additional queries

26. Agentic Retrieval + Multi-QueryΒΆ

The agent can decide:

Initial Search
 ↓
Results insufficient
 ↓
Generate 3 additional queries
 ↓
Retrieve
 ↓
Fuse

This makes Multi-Query a tool available to the agent rather than a mandatory pipeline stage.


27. Agentic Retrieval + HyDEΒΆ

The agent can also decide to use HyDE.

Query
 ↓
Initial Retrieval
 ↓
Poor Semantic Recall
 ↓
Agent selects HyDE
 ↓
Hypothetical Document
 ↓
Vector Search
 ↓
Evidence

This allows expensive retrieval techniques to be used selectively.


28. Agentic Retrieval + RouterΒΆ

The Router Retriever can be exposed as a tool.

Agent
 ↓
Router
 ↓
Selected Retriever

Or the agent can directly select retrieval capabilities.

Agent
 β”œβ”€β”€ Vector Search
 β”œβ”€β”€ Hybrid Search
 β”œβ”€β”€ SQL
 └── Graph

A router is useful when route selection follows stable rules.

An agent is useful when retrieval strategy may need to change during execution.


29. Router vs AgentΒΆ

RouterΒΆ

Query
 ↓
Select Route
 ↓
Execute

AgentΒΆ

Query
 ↓
Plan
 ↓
Select Tool
 ↓
Observe
 ↓
Evaluate
 ↓
Select Another Tool
 ↓
Observe

The key difference is:

Router
β†’ One routing decision

Agent
β†’ Potentially multiple dynamic decisions

30. Agentic Retrieval + SQLΒΆ

Consider:

"What was revenue for products affected
by the payment migration?"

The agent may need:

Step 1:
Find products affected by migration.

Step 2:
Query revenue for those products.

Step 3:
Compare results.

Architecture:

flowchart TD
    A["User Query"] --> B["Agent"]

    B --> C["Document Search"]
    C --> D["Affected Products"]

    D --> E["SQL Query"]
    E --> F["Revenue Data"]

    F --> G["Evidence Synthesis"]

This is much more powerful than treating SQL as an isolated route.


31. Agentic Retrieval + Graph RAGΒΆ

Example:

"Which services depend on PaymentService
and which incidents affected those services?"

The agent can:

1. Graph Search
2. Identify dependent services
3. Search incident documents
4. Correlate services with incidents

Architecture:

flowchart TD
    A["Query"] --> B["Agent"]

    B --> C["Graph Search"]
    C --> D["Dependent Services"]

    D --> E["Document Search"]
    E --> F["Incident Evidence"]

    F --> G["Evidence Synthesis"]

32. Cross-Source RetrievalΒΆ

Agentic Retrieval becomes particularly valuable when evidence spans:

Documents
+
SQL
+
Graph
+
Search Engine
+
APIs

Example:

User Question
       ↓
Agent
       ↓
Documentation
       ↓
Graph
       ↓
Operational Data
       ↓
SQL
       ↓
Evidence Synthesis

This is a major enterprise use case.


33. Retrieval Planning ExampleΒΆ

Query:

"Why did customer payment failures increase
after the authentication migration?"

Possible plan:

1. Search migration documentation.
2. Search incident reports.
3. Query payment failure metrics.
4. Identify authentication-related errors.
5. Correlate the timeline.
6. Determine whether evidence supports a relationship.

The agent should not assume causality merely because two events occurred around the same time.


34. Evidence CorrelationΒΆ

Agentic Retrieval can collect evidence from different sources:

Migration:
2026-04-10

Failure Increase:
2026-04-11

Incident:
Authentication timeout

Metric:
Payment failures +18%

The agent can organize the evidence:

Timeline
+
Technical Evidence
+
Operational Metrics

The final generation stage can then produce a grounded explanation.


35. Evidence Does Not Equal CausalityΒΆ

This is especially important for enterprise AI.

Finding:

Event A
+
Event B

does not prove:

A caused B

The retrieval agent should distinguish:

Observed Evidence

from:

Inference

and:

Confirmed Cause

This is particularly important in:

Finance
Healthcare
Legal
Operations
Security

36. Source AuthorityΒΆ

Not all retrieved evidence should be treated equally.

Example:

Official Policy
        > 
Internal Wiki
        >
Support Ticket
        >
User Discussion

The agent can consider source authority.

Example metadata:

{
  "source_type": "official_policy",
  "authority_score": 1.0
}

This can be used during evidence selection.


37. Source-Aware RetrievalΒΆ

The agent can prefer:

Official Documentation

for policy questions.

For operational questions:

Incident Reports
Monitoring Data
Runbooks

may be more authoritative.

The retrieval strategy should therefore consider:

Query Intent
+
Source Authority

38. Agent MemoryΒΆ

Agentic Retrieval usually requires task state, not necessarily long-term conversational memory.

Short-term retrieval memory might contain:

Original Query
Searches Performed
Documents Retrieved
Subqueries
Observations
Missing Information

This is enough for many retrieval tasks.

Do not introduce long-term memory unless the application actually requires it.


39. Retrieval ScratchpadΒΆ

A conceptual state might look like:

Goal:
Explain payment latency increase.

Completed:
βœ“ Migration documentation
βœ“ Deployment timeline

Missing:
βœ— Monitoring evidence

Next Action:
Query latency metrics.

The internal reasoning mechanism should be implementation-specific.

The application only needs structured state such as:

{
  "completed_tasks": [
    "migration_docs",
    "deployment_timeline"
  ],
  "missing_evidence": [
    "latency_metrics"
  ]
}

40. Tool SelectionΒΆ

The agent chooses tools based on their capabilities.

Example:

Question:
"What was revenue?"

β†’ SQL

Question:
"Which services depend on X?"

β†’ Graph

Question:
"How does OAuth work?"

β†’ Vector / Hybrid

Question:
"What is the current policy?"

β†’ Policy / Time-Aware Search

The tool descriptions should clearly explain:

When to use the tool
What data it accesses
What parameters it expects
What it returns

41. Tool DescriptionsΒΆ

Good:

search_engineering_docs

Search published engineering documentation,
architecture decisions, deployment guides,
and operational runbooks.

Use for questions about software systems,
architecture, deployment, APIs, and operations.

Poor:

search_docs

Search docs.

Clear descriptions improve tool selection.


42. Tool Input ValidationΒΆ

Do not allow unrestricted tool arguments.

Example:

class SearchRequest(BaseModel):

    query: str
    top_k: int = 10

    @field_validator("top_k")
    def validate_top_k(cls, value):

        if value > 50:
            raise ValueError(
                "top_k exceeds maximum allowed value"
            )

        return value

Tool boundaries are critical for production systems.


43. Retrieval GuardrailsΒΆ

Agentic Retrieval requires stronger guardrails than static retrieval.

Guard:

Maximum Iterations
Maximum Tool Calls
Maximum Token Usage
Maximum Query Length
Allowed Tools
Allowed Data Sources
Timeout
Budget

Example:

MAX_ITERATIONS = 5
MAX_TOOL_CALLS = 10
MAX_RESULTS_PER_TOOL = 50

44. Maximum IterationsΒΆ

Without an iteration limit:

Retrieve
 ↓
Insufficient
 ↓
Retrieve
 ↓
Insufficient
 ↓
Retrieve
 ↓
...

The agent can loop indefinitely.

Therefore:

Iteration Count >= MAX
        ↓
Stop

and return the best available evidence.


45. Maximum Tool CallsΒΆ

A complex query might trigger:

20 vector searches
10 SQL queries
15 graph queries

This creates:

Latency ↑
Cost ↑
Load ↑

A production agent should enforce a tool-call budget.


46. Token BudgetΒΆ

Agentic Retrieval can consume tokens through:

Planning
Tool Descriptions
Tool Results
Query Refinement
Evidence Evaluation
Final Generation

Therefore define:

Retrieval Token Budget

separately from:

Generation Token Budget

47. Cost BudgetΒΆ

A production agent can have:

budget = {
    "max_iterations": 5,
    "max_tool_calls": 10,
    "max_llm_tokens": 6000,
    "max_retrieval_cost": 0.05
}

The exact limits depend on the application.


48. Time BudgetΒΆ

Example:

Retrieval Budget = 1 second

If the agent reaches:

950 ms

it should stop expensive retrieval and use the evidence already collected.

This prevents unpredictable latency.


49. Stopping CriteriaΒΆ

The agent can stop when:

Evidence sufficient

or:

Maximum iterations reached

or:

Budget exhausted

or:

No further useful retrieval possible

or:

Required source unavailable

50. Stopping DecisionΒΆ

flowchart TD
    A["Observation"] --> B["Evidence Evaluation"]

    B --> C{"Sufficient?"}

    C -->|Yes| D["Finish"]

    C -->|No| E{"Budget Available?"}

    E -->|No| F["Finish with Best Evidence"]

    E -->|Yes| G{"Useful Next Action?"}

    G -->|Yes| H["Retrieve Again"]
    G -->|No| F

    H --> A

This makes the retrieval loop bounded and predictable.


51. Retrieval Quality GateΒΆ

A quality evaluator can score evidence.

Example:

{
  "relevance": 0.91,
  "coverage": 0.84,
  "authority": 0.95,
  "sufficient": false
}

The agent sees:

Relevance β†’ High
Coverage β†’ Medium
Authority β†’ High
Sufficient β†’ No

Therefore it searches for the missing evidence.


52. Evidence CoverageΒΆ

A useful conceptual score:

Evidence Coverage
=
Answered Requirements
---------------------
Total Requirements

For:

"What changed, when, and what was the impact?"

if only:

What
When

are answered:

Coverage = 2 / 3

The agent can continue searching.


53. Query Planning vs Query ExecutionΒΆ

Separate:

Planning

from:

Execution

Example:

plan = agent.create_plan(query)

for step in plan.steps:

    result = execute_retrieval(step)

    state.add(result)

This improves observability and testing.


54. Plan ExampleΒΆ

{
  "steps": [
    {
      "goal": "Find migration details",
      "tool": "engineering_search"
    },
    {
      "goal": "Find latency metrics",
      "tool": "metrics_search"
    },
    {
      "goal": "Find related incidents",
      "tool": "incident_search"
    }
  ]
}

The plan itself should be validated before execution.


55. Dynamic PlanningΒΆ

Unlike static pipelines, the next step can depend on the previous result.

Example:

Search migration docs
        ↓
Find migration version = 3.2
        ↓
Search incident reports for version 3.2
        ↓
Find incident INC-982
        ↓
Search incident details

The agent dynamically builds the retrieval path.


56. Retrieval GraphΒΆ

The execution can be viewed as a graph:

                    Query
                      β”‚
                      ↓
                Migration Search
                      β”‚
                      ↓
                 Version 3.2
                  /         \
                 ↓           ↓
        Incident Search   Release Search
                 ↓           ↓
             INC-982      Release Notes
                  \         /
                   ↓       ↓
                    Evidence

This is different from a simple linear retrieval pipeline.


57. Agentic Retrieval with Graph StateΒΆ

A workflow graph can represent:

START
 ↓
Analyze
 ↓
Plan
 ↓
Retrieve
 ↓
Evaluate
 β”œβ”€β”€ Complete β†’ END
 └── Continue β†’ Refine β†’ Retrieve

This pattern works well with graph-based orchestration frameworks.


58. LangGraph-Oriented ConceptΒΆ

A conceptual workflow:

workflow.add_node(
    "planner",
    planner_node
)

workflow.add_node(
    "retriever",
    retrieval_node
)

workflow.add_node(
    "evaluator",
    evaluation_node
)

workflow.add_edge(
    "planner",
    "retriever"
)

workflow.add_edge(
    "retriever",
    "evaluator"
)

Conditional routing:

workflow.add_conditional_edges(
    "evaluator",
    should_continue,
    {
        "continue": "planner",
        "finish": "answer"
    }
)

The exact implementation depends on the orchestration framework.


59. Agentic Retrieval with LlamaIndexΒΆ

A conceptual architecture can expose retrievers as tools:

Agent
 ↓
Retriever Tool
 ↓
LlamaIndex Retriever
 ↓
Documents

The important architectural boundary remains:

Agent
β†’ Tool Interface
β†’ Retrieval Engine

rather than coupling the business application directly to framework internals.


60. Agentic Retrieval with LangChainΒΆ

Similarly:

Agent
 ↓
Retriever Tool
 ↓
LangChain Retriever
 ↓
Vector Store

The application should ideally keep the agent layer independent from framework-specific retrieval implementation.


61. Framework-Agnostic ArchitectureΒΆ

A production system can define:

class RetrievalCapability:

    def search(
        self,
        query: str,
        filters: dict | None = None
    ):
        raise NotImplementedError

Adapters implement:

VectorSearchCapability
HybridSearchCapability
SQLSearchCapability
GraphSearchCapability

The agent sees:

Capability Interface

not:

Specific Framework

62. Retrieval Agent InterfaceΒΆ

class RetrievalAgent:

    def __init__(
        self,
        planner,
        tools,
        evaluator,
        budget
    ):
        self.planner = planner
        self.tools = tools
        self.evaluator = evaluator
        self.budget = budget

    def retrieve(self, query):

        state = RetrievalState(
            query=query
        )

        while not state.should_stop(
            self.budget
        ):

            plan = self.planner.plan(state)

            result = self.execute(plan)

            state.add(result)

            if self.evaluator.is_sufficient(
                state
            ):
                break

        return state.evidence

This is a simplified conceptual implementation.


63. Retrieval State ObjectΒΆ

from dataclasses import dataclass, field


@dataclass
class RetrievalState:

    query: str

    evidence: list = field(
        default_factory=list
    )

    tool_calls: list = field(
        default_factory=list
    )

    iteration: int = 0

    completed_tasks: list = field(
        default_factory=list
    )

    missing_information: list = field(
        default_factory=list
    )

Keeping state explicit makes the system easier to test.


64. Tool Execution BoundaryΒΆ

Do not allow the model to directly execute infrastructure operations.

Instead:

LLM Decision
     ↓
Validated Tool Call
     ↓
Tool Adapter
     ↓
Infrastructure

For example:

Agent
 ↓
SQL Tool
 ↓
SQL Validator
 ↓
Read-Only Database Connection

This creates an important security boundary.


65. SQL GuardrailsΒΆ

If Agentic Retrieval can call SQL:

Agent
 ↓
Generated SQL
 ↓
SQL Parser / Validator
 ↓
Allowed Operations
 ↓
Read-Only Database

Reject:

INSERT
UPDATE
DELETE
DROP
ALTER

unless the application explicitly requires them and has appropriate controls.

For retrieval-focused systems, read-only access is generally preferred.


66. Graph Query GuardrailsΒΆ

Similarly:

Agent
 ↓
Graph Query
 ↓
Query Validation
 ↓
Allowed Graph Operations
 ↓
Graph Database

The agent should not have unrestricted access to graph mutation operations.


67. Document Search GuardrailsΒΆ

Document retrieval should enforce:

Tenant
User
Role
Document ACL
Classification
Retention

before evidence enters the generation context.


68. Agentic Retrieval and Access ControlΒΆ

A critical principle:

The agent is not an authorization mechanism.

The system should never rely on:

LLM decides whether user can access document

Instead:

Application Security Layer
        ↓
Authorized Retrieval
        ↓
Agent

The agent can decide what to search, but the retrieval infrastructure decides what the user is allowed to see.


69. ObservabilityΒΆ

Agentic Retrieval requires detailed tracing.

Track:

Original Query
Plan
Tool Selected
Tool Input
Tool Output Count
Retrieval Scores
Iteration Number
Evidence State
Evaluation Result
Next Action
Stop Reason
Latency
Token Usage
Cost

Example:

{
  "query": "Why did payment failures increase?",
  "iterations": 3,
  "tool_calls": [
    "migration_search",
    "incident_search",
    "metrics_search"
  ],
  "stop_reason": "evidence_sufficient"
}

70. Retrieval TraceΒΆ

A useful trace:

Iteration 1
β”œβ”€β”€ Tool: migration_search
β”œβ”€β”€ Results: 8
└── Missing: operational impact

Iteration 2
β”œβ”€β”€ Tool: incident_search
β”œβ”€β”€ Results: 5
└── Missing: quantitative impact

Iteration 3
β”œβ”€β”€ Tool: metrics_search
β”œβ”€β”€ Results: 3
└── Evidence: sufficient

STOP

This is extremely valuable for debugging.


71. Cost ObservabilityΒΆ

Track:

Planning Tokens
Tool Selection Tokens
Query Rewrite Tokens
Tool Calls
Embedding Calls
Reranker Calls
Evaluation Calls
Final Generation Tokens

An agentic pipeline can become significantly more expensive than static retrieval.


72. Agentic Retrieval Cost FormulaΒΆ

Conceptually:

Total Cost
=
Planning Cost
+
Retrieval Cost
+
Embedding Cost
+
Reranking Cost
+
Evaluation Cost
+
Generation Cost

As iteration count increases:

Cost ↑
Latency ↑

Therefore:

Agentic Retrieval
β‰ 
Unlimited Retrieval

73. LatencyΒΆ

A static pipeline might have:

Query
 ↓
Retrieve
 ↓
Generate

Agentic retrieval can have:

Query
 ↓
Plan
 ↓
Retrieve
 ↓
Evaluate
 ↓
Retrieve
 ↓
Evaluate
 ↓
Generate

This can introduce substantial latency.

Use agentic retrieval when the additional reasoning and retrieval quality justify the cost.


74. Parallel Agentic RetrievalΒΆ

An agent may identify independent tasks:

Task A:
Search migration documents

Task B:
Search incident reports

Task C:
Query metrics

These can sometimes execute in parallel:

flowchart TD
    A["Agent Plan"] --> B["Task A"]
    A --> C["Task B"]
    A --> D["Task C"]

    B --> E["Evidence Fusion"]
    C --> E
    D --> E

    E --> F["Evidence Evaluation"]

Parallel execution reduces latency.

However, dependent tasks should remain sequential.


75. Sequential DependenciesΒΆ

Example:

Find affected service
        ↓
Find incidents for service
        ↓
Find metrics for incident period

These cannot safely be parallelized because each step depends on the previous result.

Therefore the agent should understand:

Independent Tasks
β†’ Parallel

Dependent Tasks
β†’ Sequential

76. Retrieval Planning with DependenciesΒΆ

A plan can represent dependencies:

{
  "tasks": [
    {
      "id": "find_service",
      "tool": "document_search"
    },
    {
      "id": "find_incidents",
      "tool": "incident_search",
      "depends_on": ["find_service"]
    },
    {
      "id": "find_metrics",
      "tool": "metrics_search",
      "depends_on": ["find_incidents"]
    }
  ]
}

This allows the orchestrator to execute the plan intelligently.


77. Agentic Retrieval Failure ModesΒΆ

77.1 Infinite LoopsΒΆ

Search
 ↓
Search Again
 ↓
Search Again

Prevent with:

Max Iterations
Max Tool Calls

77.2 Tool Selection ErrorsΒΆ

The agent chooses the wrong source.


77.3 Query DriftΒΆ

The agent gradually moves away from the original question.

Maintain:

Original Query

throughout the retrieval state.


77.4 Evidence AccumulationΒΆ

The agent keeps collecting documents without improving evidence quality.

Use:

Evidence Sufficiency
+
Marginal Utility

checks.


78. Marginal Retrieval ValueΒΆ

Suppose:

Iteration 1:
Evidence Quality = 0.70

Iteration 2:
Evidence Quality = 0.84

Iteration 3:
Evidence Quality = 0.85

Iteration 4:
Evidence Quality = 0.85

Iteration 4 provides almost no improvement.

The agent should stop.

Conceptually:

Marginal Gain
=
New Evidence Quality
-
Previous Evidence Quality

If:

Marginal Gain < Threshold

stop retrieving.


79. Retrieval SaturationΒΆ

A retrieval process can reach saturation:

Search 1 β†’ Significant Evidence
Search 2 β†’ Significant Evidence
Search 3 β†’ Small Improvement
Search 4 β†’ No Improvement

Continuing beyond this point wastes:

Time
Tokens
Cost
Infrastructure

80. Relevance CollapseΒΆ

More retrieval does not always mean better retrieval.

Adding many documents can introduce:

Noise
Contradictions
Redundant Evidence
Context Dilution

Therefore the agent should optimize:

Evidence Quality

not:

Evidence Quantity

81. Evidence DeduplicationΒΆ

Agentic retrieval may search the same source multiple times.

Example:

Search 1:
Document A

Search 2:
Document A

Search 3:
Document A

Deduplicate evidence before final context assembly.


82. Contradictory EvidenceΒΆ

Different sources may disagree.

Example:

Document A:
OAuth tokens expire after 60 minutes.

Document B:
OAuth tokens expire after 30 minutes.

The agent should not simply combine both statements.

It should identify:

Conflict

and retrieve authoritative or newer evidence.


83. Conflict ResolutionΒΆ

Possible strategy:

Conflict Detected
        ↓
Check Source Authority
        ↓
Check Publication Date
        ↓
Retrieve More Evidence
        ↓
Resolve / Report Uncertainty

This is particularly important for enterprise policies and changing technical documentation.


84. Source RecencyΒΆ

When information changes over time:

Current Policy

should generally prefer:

Latest Approved Version

over:

Old Archived Version

The retrieval agent can use metadata and source authority to make better retrieval decisions.


An agent may determine:

Question asks for current information

and select:

Time-Weighted Retriever

rather than:

General Vector Search

This is another example of dynamic retrieval strategy selection.


The agent may initially use:

Vector Search

If exact terminology is missing:

Switch to Hybrid Search

Example:

Initial:
Semantic Search

Result:
Low exact-term coverage

Next:
Hybrid Search

This adaptive strategy can improve retrieval without running every retriever for every query.


87. Agentic Retrieval and RerankingΒΆ

The agent can decide to rerank when:

Candidate count is high

or:

Initial scores are ambiguous

Pipeline:

Retrieve
 ↓
Evaluate
 ↓
Need precision?
 ↓
Rerank
 ↓
Evaluate

This makes reranking conditional rather than mandatory.


88. Agentic Retrieval and CompressionΒΆ

Similarly:

Large Documents
 ↓
Context Too Large
 ↓
Compression

If the retrieved documents are already concise:

Compression may be unnecessary

This can reduce cost.


89. Adaptive Retrieval PipelineΒΆ

A sophisticated retrieval agent can dynamically construct:

Query
 ↓
Vector Search
 ↓
Low Confidence
 ↓
Hybrid Search
 ↓
Many Candidates
 ↓
Reranking
 ↓
Large Context
 ↓
Compression
 ↓
Evidence Sufficient
 ↓
Finish

The pipeline is built dynamically from the agent's decisions.


90. Agentic Retrieval ArchitectureΒΆ

flowchart TD
    A["User Query"] --> B["Agent Controller"]

    B --> C["Query Understanding"]
    C --> D["Planning"]

    D --> E{"Select Capability"}

    E --> F["Vector Search"]
    E --> G["Hybrid Search"]
    E --> H["SQL"]
    E --> I["Graph"]
    E --> J["Time-Weighted Search"]
    E --> K["HyDE"]

    F --> L["Observation"]
    G --> L
    H --> L
    I --> L
    J --> L
    K --> L

    L --> M["Evidence Evaluator"]

    M --> N{"Sufficient?"}

    N -->|Yes| O["Context Assembly"]
    N -->|No| P["Query Refinement"]

    P --> D

    O --> Q["Generation"]
    Q --> R["Validation"]
    R --> S["Enterprise Response"]

91. Production Agent ControllerΒΆ

A production architecture should separate:

Agent Controller

from:

Retrieval Tools

Example:

class AgentController:

    def __init__(
        self,
        planner,
        evaluator,
        tool_registry,
        policy
    ):
        self.planner = planner
        self.evaluator = evaluator
        self.tools = tool_registry
        self.policy = policy

The controller manages:

Planning
Policy
Budgets
Tool Selection
State
Stopping

92. Retrieval PolicyΒΆ

The agent should operate within a policy.

Example:

class RetrievalPolicy:

    max_iterations = 5
    max_tool_calls = 10
    max_results = 50
    allow_sql = True
    allow_graph = True
    allow_external_search = False

The agent can select only tools permitted by policy.


93. Policy EnforcementΒΆ

Never rely solely on the LLM to follow:

"Do not use SQL."

Instead:

Agent Decision
 ↓
Policy Enforcement
 ↓
Tool Availability
 ↓
Tool Execution

If SQL is disabled:

if tool_name == "sql" and not policy.allow_sql:
    raise PermissionError(
        "SQL retrieval is disabled"
    )

94. Retrieval Budget ControllerΒΆ

A budget controller can enforce:

class RetrievalBudget:

    def __init__(
        self,
        max_iterations=5,
        max_tool_calls=10
    ):
        self.max_iterations = max_iterations
        self.max_tool_calls = max_tool_calls

The agent checks the budget before every action.

This prevents runaway execution.


95. Observability ArchitectureΒΆ

flowchart LR
    A["User Query"] --> B["Retrieval Agent"]

    B --> C["Tool Calls"]
    B --> D["Plans"]
    B --> E["Evidence Evaluation"]

    C --> F["Trace Collector"]
    D --> F
    E --> F

    F --> G["Metrics"]
    F --> H["Logs"]
    F --> I["Traces"]

    G --> J["Monitoring"]
    H --> J
    I --> J

Agentic retrieval should be observable at the decision level, not just the HTTP request level.


96. Important Agent MetricsΒΆ

Track:

Iterations per Query
Tool Calls per Query
Retrieval Success Rate
No-Result Rate
Fallback Rate
Average Evidence Coverage
Average Retrieval Latency
P95 Retrieval Latency
Token Usage
Cost per Query
Early Stop Rate
Budget Exhaustion Rate

These metrics reveal whether the agent is efficient.


97. Retrieval Quality MetricsΒΆ

Measure:

Recall@K
Precision@K
MRR
NDCG
Context Relevance
Evidence Coverage
Citation Accuracy
Answer Faithfulness
Answer Correctness

Agentic Retrieval should be compared against:

Static Retrieval Baseline

and:

Multi-Stage Retrieval Baseline

98. Evaluation DatasetΒΆ

Create questions requiring different retrieval behaviors.

SimpleΒΆ

"What is OAuth?"

Multi-documentΒΆ

"What changed across the last two releases?"

Cross-sourceΒΆ

"Did the migration affect transaction volume?"

RelationshipΒΆ

"Which services depend on PaymentService?"

StructuredΒΆ

"What was revenue in Q2?"

AmbiguousΒΆ

"How does authentication work?"

This helps determine where Agentic Retrieval actually provides value.


99. Baseline ComparisonΒΆ

Compare:

A:
Vector Search

B:
Hybrid + Reranking

C:
Multi-Stage Retrieval

D:
Agentic Retrieval

Measure:

Retrieval Quality
Answer Quality
Latency
Cost
Reliability

The goal is not automatically to choose D.

The goal is to identify:

Which architecture provides the required quality at an acceptable operational cost?


100. When Agentic Retrieval Is ValuableΒΆ

Agentic Retrieval is particularly useful when:

  • Questions require multiple retrieval steps
  • Multiple knowledge sources must be combined
  • The next search depends on previous results
  • Query decomposition is useful
  • Evidence completeness matters
  • Retrieval strategy must change dynamically
  • Knowledge sources are heterogeneous
  • Complex enterprise research is required
  • Static retrieval pipelines repeatedly fail on complex questions

Typical use cases:

Enterprise Research Assistants
Technical Investigation
Incident Analysis
Financial Research
Legal Research
Compliance Investigation
Complex Customer Support
Architecture Analysis
Cross-System Knowledge Assistants

101. When Agentic Retrieval Is Not NecessaryΒΆ

Avoid Agentic Retrieval when:

Queries are simple

or:

One retriever already provides excellent results

or:

Latency requirements are extremely strict

or:

The retrieval workflow is completely deterministic

or:

The added cost cannot be justified

For many applications:

Hybrid + Reranking

may be better than an agent.


102. Multi-Stage vs Agentic RetrievalΒΆ

Characteristic Multi-Stage Agentic
Pipeline Predefined Dynamic
Decision Making Mostly static Runtime
Query Refinement Optional Core capability
Tool Selection Usually configured Dynamic
Iteration Limited Native
Latency Predictable Variable
Cost Predictable Variable
Debugging Easier More complex
Best For Known workflows Complex research

The important design principle is:

Predictable Problem
β†’ Multi-Stage Retrieval

Dynamic Problem
β†’ Agentic Retrieval

103. Router vs Multi-Stage vs AgenticΒΆ

These three patterns form a useful progression.

RouterΒΆ

Choose

Multi-StageΒΆ

Pipeline

AgenticΒΆ

Choose
+
Execute
+
Observe
+
Adapt

Conceptually:

Router
   ↓
Select Strategy

Multi-Stage
   ↓
Execute Strategy

Agentic
   ↓
Select β†’ Execute β†’ Observe β†’ Adapt

104. Enterprise ArchitectureΒΆ

A mature enterprise retrieval platform may combine all three:

flowchart TD
    A["User Query"] --> B["Security Context"]

    B --> C["Retrieval Agent"]

    C --> D["Router"]

    D --> E["Selected Retrieval Pipeline"]

    E --> F["Candidate Generation"]
    F --> G["Filtering"]
    G --> H["Reranking"]
    H --> I["Context Selection"]

    I --> J["Evidence Evaluator"]

    J --> K{"Sufficient?"}

    K -->|Yes| L["Generation"]
    K -->|No| M["Agent Planning"]

    M --> D

    L --> N["Response Validation"]
    N --> O["Citation"]
    O --> P["Enterprise Response"]

This is a powerful production architecture.


105. Production GuardrailsΒΆ

Before allowing an Agentic Retrieval system into production:

☐ Maximum iterations configured
☐ Maximum tool calls configured
☐ Maximum token budget configured
☐ Maximum cost budget configured
☐ Tool allowlist configured
☐ Query length limits configured
☐ SQL access restricted
☐ Graph access restricted
☐ Tenant isolation enforced
☐ Document ACLs enforced
☐ Tool arguments validated
☐ Retrieval timeouts configured
☐ Fallback strategy implemented
☐ Evidence sufficiency implemented
☐ Evidence deduplication implemented
☐ Contradiction detection considered
☐ Source authority considered
☐ Original query preserved
☐ Agent state observable
☐ Tool calls traced
☐ Cost monitored
☐ Retrieval quality evaluated
☐ End-to-end answer quality evaluated
☐ Regression dataset maintained

106. Production Design PrinciplesΒΆ

Principle 1 β€” Bound the AgentΒΆ

Never allow unlimited iterations.

Principle 2 β€” Control the ToolsΒΆ

The agent should access only approved retrieval capabilities.

Principle 3 β€” Enforce Security Outside the LLMΒΆ

Authorization must be deterministic.

Principle 4 β€” Preserve EvidenceΒΆ

Every final claim should be traceable to retrieved evidence where required.

Principle 5 β€” Evaluate SufficiencyΒΆ

Do not stop simply because documents were retrieved.

Principle 6 β€” Avoid Unnecessary Agentic BehaviorΒΆ

Use static retrieval when static retrieval is sufficient.

Principle 7 β€” Measure the AgentΒΆ

Track quality, latency, iterations, tool calls, and cost.


107. Common Anti-PatternsΒΆ

Anti-Pattern 1 β€” Agent for Every QueryΒΆ

Simple Query
 ↓
Full Agent
 ↓
5 Tool Calls

This creates unnecessary cost.


Anti-Pattern 2 β€” Unlimited Tool AccessΒΆ

Agent
 ↓
Every Enterprise System

This creates major security risk.


Anti-Pattern 3 β€” No Stop ConditionΒΆ

Retrieve
 ↓
Retrieve
 ↓
Retrieve
...

This creates runaway execution.


Anti-Pattern 4 β€” No Evidence ValidationΒΆ

The agent retrieves documents and assumes they are sufficient.


Anti-Pattern 5 β€” Treating Agent Output as EvidenceΒΆ

Agent-generated reasoning is not equivalent to authoritative source evidence.


A practical architecture is:

Simple Query
      ↓
Fast Retrieval
      ↓
Answer


Complex Query
      ↓
Router / Agent
      ↓
Retrieval Planning
      ↓
Specialized Retrieval
      ↓
Evidence Evaluation
      ↓
Additional Retrieval if Needed
      ↓
Reranking
      ↓
Context Selection
      ↓
Grounded Generation

This allows the system to remain efficient for simple queries while supporting complex investigations.


109. Key TakeawaysΒΆ

  • Agentic Retrieval adds dynamic decision-making to retrieval systems.
  • The agent can plan, select tools, retrieve, observe, evaluate, and refine.
  • Agentic Retrieval is especially useful for complex multi-source questions.
  • It can dynamically select Vector, Hybrid, SQL, Graph, HyDE, or other retrieval capabilities.
  • Query decomposition allows complex questions to be split into specialized retrieval tasks.
  • Retrieval state records what has already been searched and what evidence is still missing.
  • Evidence sufficiency is a central component of Agentic Retrieval.
  • Query refinement enables iterative search.
  • Agentic Retrieval can combine multiple retrieval techniques dynamically.
  • Router Retrieval selects a route; Agentic Retrieval can repeatedly adapt the route.
  • Multi-Stage Retrieval follows a predefined workflow; Agentic Retrieval dynamically changes the workflow.
  • Tool access must be controlled through capability-based interfaces.
  • SQL and Graph tools require strict validation and authorization.
  • The agent must never act as the authorization mechanism.
  • Maximum iterations, tool calls, tokens, cost, and time should be bounded.
  • Evidence quality matters more than evidence quantity.
  • Retrieval saturation should trigger stopping.
  • Contradictory evidence should be detected and handled explicitly.
  • Source authority and recency should influence evidence selection where appropriate.
  • Agentic Retrieval introduces additional latency and cost.
  • Agentic Retrieval should be compared against strong static and multi-stage baselines.
  • The best architecture is not necessarily the most autonomous one.
  • Use Agentic Retrieval when dynamic retrieval decisions provide measurable value.

The central pattern is:

Understand
    ↓
Plan
    ↓
Retrieve
    ↓
Observe
    ↓
Evaluate
    ↓
Refine
    ↓
Retrieve Again
    ↓
Evidence Sufficient
    ↓
Generate

Or simply:

Don't just search.

Search β†’ Observe β†’ Learn β†’ Search Again

🧭 Chapter Navigation¢

Part V β€” Advanced Retrieval-Augmented GenerationΒΆ

Previous:
08. Multi-Stage Retrieval

Next:
10. Re-ranking Techniques

Section:
02 β€” Enterprise Retrieval Engineering

Enterprise Retrieval Engineering PathΒΆ

01 Contextual Compression Retriever
              ↓
02 Ensemble Retriever
              ↓
03 Multi-Vector Retriever
              ↓
04 Time-Weighted Retriever
              ↓
05 Hybrid Search Retriever
              ↓
06 HyDE Retriever
              ↓
07 Router Retriever
              ↓
08 Multi-Stage Retrieval
              ↓
09 Agentic Retrieval
              ↓
10 Re-ranking Techniques
              ↓
11 MMR & Diversity-Aware Retrieval
              ↓
12 Metadata-Aware Retrieval
              ↓
13 Advanced Query Rewriting

Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β€” One Chapter at a Time.