06. Agentic RAGΒΆ
Category: Advanced RAG Architecture
Module: Part V β Advanced Retrieval-Augmented Generation
Difficulty: Advanced
π OverviewΒΆ
Traditional RAG follows a relatively fixed pipeline:
This architecture works well for straightforward questions, but enterprise queries are often more complex.
A single question may require:
Multiple retrieval strategies
Multiple search iterations
Query rewriting
SQL queries
Knowledge Graph traversal
Vector search
Document inspection
Validation
Additional retrieval
For example:
"Which customers were affected by the payment
incident, which services were involved, and what
was the root cause?"
A fixed RAG pipeline may not know in advance whether it needs:
Agentic RAG introduces an orchestration layer capable of reasoning about the retrieval process, selecting tools, executing retrieval steps, evaluating results, and deciding whether additional retrieval is necessary.
The architecture becomes:
User Query
β
Agent / Planner
β
Decide What Information Is Needed
β
Select Retrieval Tool
β
Retrieve
β
Evaluate Evidence
β
Need More Information?
β
ββ΄ββββββββ
Yes No
β β
βΌ βΌ
Retrieve Generate
Again Answer
The important idea is:
Agentic RAG makes retrieval an adaptive process rather than a single fixed retrieval step.
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand Agentic RAG
- Understand the difference between traditional RAG and Agentic RAG
- Understand agentic retrieval
- Understand retrieval planning
- Understand query decomposition
- Understand iterative retrieval
- Understand tool-based retrieval
- Understand retrieval routing
- Build SQL retrieval tools
- Build Vector retrieval tools
- Build Knowledge Graph retrieval tools
- Combine multiple retrieval mechanisms
- Understand agent state
- Understand retrieval loops
- Implement bounded agent execution
- Understand reflection and self-evaluation
- Understand evidence verification
- Understand adaptive query rewriting
- Understand multi-hop retrieval
- Design Agentic RAG workflows
- Understand deterministic vs agentic orchestration
- Implement guardrails
- Control agent cost and latency
- Implement observability
- Evaluate Agentic RAG systems
- Design production-grade Agentic RAG architectures
π§ 1. What Is Agentic RAG?ΒΆ
Agentic RAG combines:
The agent can decide:
What should I retrieve?
Which retriever should I use?
Do I need another query?
Is the evidence sufficient?
Should I verify the answer?
Instead of:
we have:
π 2. Traditional RAG vs Agentic RAGΒΆ
Traditional RAGΒΆ
The retrieval path is mostly predetermined.
Agentic RAGΒΆ
Question
β
Agent
β
Plan
β
Select Tool
β
Retrieve
β
Evaluate
β
Re-plan
β
Retrieve Again
β
Validate
β
Answer
The retrieval path is adaptive.
π 3. ComparisonΒΆ
| Capability | Traditional RAG | Agentic RAG |
|---|---|---|
| Single retrieval | β | β |
| Query rewriting | Optional | Dynamic |
| Multiple retrieval steps | Limited | β |
| Tool selection | Fixed | Dynamic |
| SQL retrieval | Possible | Dynamic |
| Graph retrieval | Possible | Dynamic |
| Vector retrieval | β | β |
| Query decomposition | Limited | β |
| Iterative retrieval | Limited | β |
| Evidence evaluation | Basic | Dynamic |
| Planning | Limited | β |
| Adaptive routing | Limited | β |
| Complex multi-hop questions | Limited | Strong |
| Cost predictability | Higher | Lower |
| Operational complexity | Lower | Higher |
π§© 4. Why Agentic RAG?ΒΆ
Consider:
"Which customers affected by the payment incident
were using the services that experienced the highest
error rate, and what remediation was implemented?"
This may require:
Step 1:
Find payment incident.
Step 2:
Find affected services.
Step 3:
Query transaction database.
Step 4:
Identify affected customers.
Step 5:
Traverse service dependencies.
Step 6:
Retrieve remediation documentation.
Step 7:
Combine evidence.
A fixed retrieval pipeline may not know this sequence in advance.
An agent can construct it dynamically.
π§ 5. Agentic RAG Mental ModelΒΆ
USER QUERY
β
βΌ
AGENT
β
PLAN / DECIDE
β
βββββββββββββββββββββΌββββββββββββββββββββ
βΌ βΌ βΌ
Vector Search SQL Graph
β β β
βΌ βΌ βΌ
Evidence Evidence Evidence
β β β
βββββββββββββββββββββΌββββββββββββββββββββ
βΌ
EVIDENCE EVALUATION
β
βββββββββββ΄ββββββββββ
βΌ βΌ
Sufficient Insufficient
β β
βΌ βΌ
Answer Re-plan
β
βΌ
More Retrieval
ποΈ 6. Basic Agentic RAG ArchitectureΒΆ
flowchart TD
A["User Query"] --> B["Agent"]
B --> C["Planner"]
C --> D{"Select Tool"}
D --> E["Vector Retriever"]
D --> F["SQL Retriever"]
D --> G["Graph Retriever"]
E --> H["Evidence"]
F --> H
G --> H
H --> I["Evidence Evaluator"]
I --> J{"Enough Evidence?"}
J -->|No| C
J -->|Yes| K["Answer Generator"]
K --> L["Response"] π§ 7. Agent vs RAG PipelineΒΆ
A traditional RAG pipeline might be:
def rag(query):
documents = retriever.retrieve(query)
context = build_context(documents)
return llm.generate(context, query)
An Agentic RAG system is conceptually closer to:
def agentic_rag(query):
state = initialize_state(query)
while not state.completed:
action = planner.decide(state)
result = execute(action)
state = update_state(state, result)
return generate_answer(state)
The agent controls the retrieval process.
π§© 8. Agentic RetrievalΒΆ
Agentic retrieval means the system can dynamically decide:
Example:
User Query
β
Agent
β
Vector Search
β
Evidence
β
Agent
β
SQL
β
Evidence
β
Agent
β
Answer
π 9. Retrieval ToolsΒΆ
An Agentic RAG system may expose:
Example:
The agent chooses which tool to use.
π§ 10. Tool-Based RetrievalΒΆ
Each tool should have a clear contract.
Example:
SQL:
Graph:
class GraphSearchTool:
def search(
self,
entity: str,
relationship: str
):
raise NotImplementedError
ποΈ 11. Capability-Based Retrieval ToolsΒΆ
A production architecture should expose capabilities instead of infrastructure-specific implementations.
Agent
β
βββ DocumentSearch
βββ VectorSearch
βββ SQLQuery
βββ GraphQuery
βββ ImageSearch
βββ MetadataSearch
The underlying implementation can change independently.
π§ 12. Tool DescriptionsΒΆ
The agent needs to understand when a tool should be used.
Example:
Vector Search:
Search semantic enterprise documents.
SQL:
Retrieve exact structured business data.
Knowledge Graph:
Traverse entity relationships.
Image Search:
Retrieve visual enterprise evidence.
Clear tool descriptions improve tool selection.
π 13. Query ClassificationΒΆ
Before acting, the agent can classify the question.
Possible categories:
π§ 14. Query DecompositionΒΆ
Complex queries can be decomposed into sub-questions.
Example:
Decompose:
π 15. Query Decomposition PipelineΒΆ
flowchart LR
A["Complex Query"] --> B["Query Decomposer"]
B --> C["Sub-question 1"]
B --> D["Sub-question 2"]
B --> E["Sub-question 3"]
C --> F["Retriever"]
D --> G["Retriever"]
E --> H["Retriever"]
F --> I["Evidence"]
G --> I
H --> I
I --> J["Synthesis"] π§© 16. Parallel RetrievalΒΆ
Independent sub-questions can be executed in parallel.
Complex Query
β
βββββββββββΌββββββββββ
βΌ βΌ βΌ
Q1 Q2 Q3
β β β
βΌ βΌ βΌ
Vector SQL Graph
β β β
βββββββββββΌββββββββββ
βΌ
Synthesis
This can reduce latency.
β‘ 17. Parallel vs Sequential RetrievalΒΆ
Not all tasks can be parallelized.
Example:
Q2 depends on Q1.
Therefore:
Sequential execution is required.
π§ 18. Dependency-Aware PlanningΒΆ
The agent can construct a graph:
This is effectively a retrieval execution plan.
ποΈ 19. Retrieval PlanΒΆ
Example:
{
"steps": [
{
"id": "step-1",
"tool": "vector_search",
"query": "payment incident"
},
{
"id": "step-2",
"tool": "sql",
"depends_on": ["step-1"]
},
{
"id": "step-3",
"tool": "graph",
"depends_on": ["step-1"]
}
]
}
This gives the orchestration layer explicit control.
π§ 20. Agent StateΒΆ
Agentic systems need state.
Example:
{
"original_query": "...",
"sub_questions": [],
"retrieved_documents": [],
"sql_results": [],
"graph_results": [],
"observations": [],
"tool_calls": [],
"confidence": 0.0,
"iteration": 0,
"status": "running"
}
The state becomes the working memory of the retrieval workflow.
π 21. Agent State MachineΒΆ
stateDiagram-v2
[*] --> Analyze
Analyze --> Plan
Plan --> Retrieve
Retrieve --> Evaluate
Evaluate --> Plan: More Evidence Needed
Evaluate --> Generate: Evidence Sufficient
Generate --> Validate
Validate --> Retrieve: Validation Failed
Validate --> Complete: Validation Passed
Complete --> [*] π§ 22. Observation β Decision β ActionΒΆ
A useful Agentic RAG model is:
Example:
Observation:
Search results mention incident INC-123.
Decision:
Need affected customers.
Action:
Query SQL database.
Observation:
10,423 customers affected.
Decision:
Need remediation documentation.
Action:
Vector search.
Observation:
Runbook found.
Decision:
Evidence sufficient.
Action:
Generate answer.
π 23. Iterative RetrievalΒΆ
Agentic RAG can perform multiple retrieval rounds.
Round 1
β
Initial Evidence
Round 2
β
Refined Query
Round 3
β
Additional Evidence
Round 4
β
Validation
Answer
This can improve complex-query performance.
β οΈ 24. Retrieval Loops Need LimitsΒΆ
An agent should never be allowed to retrieve indefinitely.
Use:
Example:
π‘οΈ 25. Bounded Agent ExecutionΒΆ
If evidence remains insufficient:
rather than continuing indefinitely.
π§ 26. Evidence EvaluationΒΆ
The agent should evaluate:
Example:
π 27. Evidence SufficiencyΒΆ
A simple decision model:
Evidence
β
βββ Relevant?
β
βββ Complete?
β
βββ Authoritative?
β
βββ Current?
β
βββ Consistent?
β
βΌ
Sufficient?
If not:
π§ 28. ReflectionΒΆ
Reflection means the system evaluates its own intermediate result.
Example:
Reflection should be treated as a controlled verification mechanism rather than unrestricted self-reasoning.
π 29. Retrieval Reflection LoopΒΆ
flowchart TD
A["Query"] --> B["Retrieve"]
B --> C["Evidence"]
C --> D["Draft Answer"]
D --> E["Evaluate"]
E --> F{"Grounded?"}
F -->|Yes| G["Final Answer"]
F -->|No| H["Identify Missing Evidence"]
H --> I["Refine Query"]
I --> B π§© 30. Query RewritingΒΆ
The agent can rewrite a query based on retrieved evidence.
Example:
Original:
"Payment issue"
Retrieved:
Incident INC-1042
Rewritten:
"INC-1042 root cause and remediation"
This can improve subsequent retrieval.
π§ 31. Adaptive Query ExpansionΒΆ
Initial query:
Retrieved entities:
Next query:
The agent uses retrieved evidence to improve the next search.
π 32. Multi-Hop RetrievalΒΆ
Multi-hop retrieval requires several connected retrieval steps.
Example:
Question:
This is difficult for a single retrieval step.
π§ 33. Multi-Hop RAGΒΆ
flowchart LR
A["Question"] --> B["Customer"]
B --> C["Product"]
C --> D["Service"]
D --> E["Incident"]
E --> F["Evidence"]
F --> G["Answer"] Agentic orchestration can dynamically perform each hop.
π’ 34. Enterprise Multi-Hop ExampleΒΆ
Question:
"Which banking customers were impacted by
the authentication incident and what was
the business impact?"
Potential workflow:
Step 1:
Vector β Find authentication incident.
Step 2:
Graph β Find affected authentication service.
Step 3:
SQL β Find affected customers.
Step 4:
SQL β Calculate transaction impact.
Step 5:
Vector β Find incident postmortem.
Step 6:
Synthesize.
π 35. SQL + Graph + Vector AgentΒΆ
flowchart TD
A["User Query"] --> B["Agent Planner"]
B --> C["Vector Search"]
B --> D["SQL Query"]
B --> E["Graph Query"]
C --> F["Document Evidence"]
D --> G["Structured Evidence"]
E --> H["Relationship Evidence"]
F --> I["Evidence Store"]
G --> I
H --> I
I --> J["Agent Evaluator"]
J --> K{"Sufficient?"}
K -->|No| B
K -->|Yes| L["Answer Generator"]
L --> M["Validation"]
M --> N["Response"] π§ 36. Agentic RAG and Multimodal RAGΒΆ
Agents can also select visual tools.
Agent
β
βββ Vector Search
βββ SQL
βββ Graph
βββ Image Search
βββ OCR
βββ Vision Analysis
Example:
Question:
"Which service is shown communicating
with PostgreSQL in the architecture diagram?"
Agent:
β
Image Search
β
Architecture Diagram
β
Vision Analysis
β
Knowledge Graph Verification
β
Answer
π§© 37. Agentic RAG Tool RegistryΒΆ
A production system may maintain:
class ToolRegistry:
def register(self, tool):
...
def get(self, name):
...
def list_tools(self):
...
Example:
The agent can access only the tools assigned to its role.
π 38. Tool AuthorizationΒΆ
Not every agent should access every tool.
Example:
Research Agent
βββ Vector Search
βββ Graph Search
βββ Document Search
Analytics Agent
βββ SQL
βββ Metadata Search
Enterprise Agent
βββ Vector
βββ SQL
βββ Graph
βββ Document
Tool access should be governed by authorization.
π‘οΈ 39. Tool-Level SecurityΒΆ
Each tool should enforce:
Do not rely on the LLM to enforce security policies.
π¨ 40. Agent Prompt InjectionΒΆ
Retrieved documents can contain malicious instructions.
Example:
This is a retrieval-time prompt injection risk.
The agent must distinguish:
from:
π§ 41. Retrieval Content Is Data, Not AuthorityΒΆ
A critical security principle:
not:
The agent should never automatically execute instructions found inside retrieved content.
π‘οΈ 42. Tool Output ValidationΒΆ
Tool output should be treated as untrusted data.
This prevents malformed or unexpected tool output from directly influencing execution.
π§© 43. Agent GuardrailsΒΆ
Guardrails can enforce:
Allowed Tools
Allowed Domains
Allowed Databases
Allowed Tables
Maximum Iterations
Maximum Query Cost
Maximum Result Size
Maximum Token Usage
π§ 44. Deterministic vs Agentic OrchestrationΒΆ
Not every workflow needs an agent.
DeterministicΒΆ
Advantages:
AgenticΒΆ
Advantages:
π 45. When to Use Agentic RAGΒΆ
Use Agentic RAG when:
Queries require multiple retrieval steps
Queries require tool selection
Queries require multi-hop reasoning
Evidence is incomplete
Multiple knowledge sources are required
The retrieval path cannot be predetermined
π« 46. When Not to Use Agentic RAGΒΆ
Avoid agents for simple queries such as:
If:
works reliably, adding an agent may only increase:
π§ 47. Agentic RAG Decision FrameworkΒΆ
Query
β
βΌ
Is it straightforward?
/ \
Yes No
β β
βΌ βΌ
Traditional RAG Need multiple steps?
/ \
No Yes
β β
βΌ βΌ
Advanced RAG Agentic RAG
The objective is not to use agents everywhere.
The objective is to use agents where adaptive orchestration provides measurable value.
ποΈ 48. Agentic RAG Architecture LayersΒΆ
βββββββββββββββββββββββββββββββββββββββ
β User Interface β
βββββββββββββββββββββββββββββββββββββββ€
β Query Understanding β
βββββββββββββββββββββββββββββββββββββββ€
β Agent / Planner β
βββββββββββββββββββββββββββββββββββββββ€
β Tool Orchestration β
βββββββββββββββββββββββββββββββββββββββ€
β Retrieval Capabilities β
βββββββββββββββββββββββββββββββββββββββ€
β Vector β SQL β Graph β Multimodal β
βββββββββββββββββββββββββββββββββββββββ€
β Evidence / State Store β
βββββββββββββββββββββββββββββββββββββββ€
β Validation / Guardrails β
βββββββββββββββββββββββββββββββββββββββ€
β Observability / Audit β
βββββββββββββββββββββββββββββββββββββββ
π§ 49. Agent State ModelΒΆ
A production state might look like:
from dataclasses import dataclass, field
@dataclass
class AgentState:
query: str
plan: list = field(default_factory=list)
observations: list = field(default_factory=list)
evidence: list = field(default_factory=list)
tool_calls: list = field(default_factory=list)
iteration: int = 0
status: str = "RUNNING"
confidence: float = 0.0
π 50. Agent Execution LoopΒΆ
Conceptually:
def run_agent(state):
while state.status == "RUNNING":
if state.iteration >= MAX_ITERATIONS:
state.status = "STOPPED"
break
action = planner.plan(state)
observation = execute(action)
state.observations.append(observation)
state = evaluator.update(state)
state.iteration += 1
return state
Production implementations require stronger error handling, authorization, timeout, and state management.
π§© 51. Agent ActionsΒΆ
Possible actions:
SEARCH_DOCUMENTS
SEARCH_IMAGES
QUERY_SQL
QUERY_GRAPH
REWRITE_QUERY
VERIFY_EVIDENCE
GENERATE_ANSWER
ASK_CLARIFICATION
STOP
Representing actions explicitly makes orchestration easier to observe and test.
π§ 52. Action ModelΒΆ
{
"action": "QUERY_SQL",
"reason": "Need exact transaction count",
"parameters": {
"metric": "failed_transactions",
"date_range": "2026-08-10"
}
}
A structured action can be validated before execution.
π 53. Structured Tool CallsΒΆ
Prefer:
over allowing the model to produce arbitrary executable commands.
The orchestration layer can translate the structured request into a controlled operation.
π§ 54. Agent MemoryΒΆ
Agentic RAG may maintain temporary state such as:
Current Query
Retrieved Evidence
Previous Searches
Intermediate Findings
Tool Results
Open Questions
This is different from long-term user memory.
For most RAG tasks, short-lived execution state is sufficient.
ποΈ 55. Evidence StoreΒΆ
A useful internal evidence store might contain:
{
"evidence_id": "ev-102",
"source_type": "vector",
"source_id": "doc-912",
"content": "...",
"relevance": 0.91,
"timestamp": "2026-08-11T10:32:00"
}
Other evidence:
{
"evidence_id": "ev-103",
"source_type": "sql",
"source_id": "analytics-db",
"query_id": "q-819",
"result": {
"count": 12450
}
}
π§ 56. Evidence GraphΒΆ
Evidence can also be represented as relationships:
Question
β
βββ Evidence A
β β
β βββ supports β Claim 1
β
βββ Evidence B
β β
β βββ supports β Claim 2
β
βββ Evidence C
β
βββ contradicts β Claim 3
This can support sophisticated response validation.
π 57. Evidence DeduplicationΒΆ
Multiple tools may retrieve the same evidence.
Example:
Deduplicate evidence before final context construction.
This reduces:
π§ 58. Evidence RankingΒΆ
Rank evidence using:
Example:
The exact scoring strategy should be evaluated empirically.
π 59. Evidence FusionΒΆ
flowchart TD
A["Vector Evidence"] --> D["Evidence Store"]
B["SQL Evidence"] --> D
C["Graph Evidence"] --> D
E["Image Evidence"] --> D
D --> F["Deduplication"]
F --> G["Ranking"]
G --> H["Conflict Detection"]
H --> I["Evidence Fusion"]
I --> J["Context Builder"] β οΈ 60. Conflicting EvidenceΒΆ
Example:
Vector:
Payment Service uses PostgreSQL.
Graph:
Payment Service uses PostgreSQL.
SQL:
Current database = MySQL.
The agent should not blindly merge these facts.
It should consider:
and potentially retrieve more evidence.
π§ 61. Confidence-Aware RetrievalΒΆ
A useful approach:
High Confidence
β
Answer
Medium Confidence
β
Verify
Low Confidence
β
Retrieve More
No Evidence
β
Ask Clarification / Abstain
π 62. Agentic AbstentionΒΆ
A production agent should be able to say:
This is preferable to:
π§ 63. ClarificationΒΆ
Some queries cannot be safely resolved.
Example:
The system may need:
The agent can ask a clarification question instead of making assumptions.
π 64. Clarification LoopΒΆ
flowchart TD
A["User Query"] --> B["Agent"]
B --> C{"Ambiguous?"}
C -->|Yes| D["Ask Clarification"]
D --> E["User"]
E --> B
C -->|No| F["Plan Retrieval"]
F --> G["Execute"]
G --> H["Answer"] π§© 65. Agentic RAG with SQLΒΆ
Example question:
Workflow:
Vector Search
β
Find Incident ID
β
Extract Incident Time Range
β
SQL
β
Count Failed Payments
β
Validate
β
Answer
π§© 66. Agentic RAG with GraphΒΆ
Question:
Workflow:
Graph
β
Find Authentication Service
β
Traverse DEPENDS_ON
β
Find Affected Services
β
Vector Search
β
Find Outage Evidence
β
Answer
π§© 67. Agentic RAG with Multimodal RetrievalΒΆ
Question:
Workflow:
Image Search
β
Architecture Diagram
β
Vision Analysis
β
Identify Service + Database
β
Graph Verification
β
Answer
π’ 68. Enterprise Agentic RAGΒΆ
A mature enterprise agent may have access to:
Architecture:
flowchart TD
A["Enterprise User"] --> B["AI Gateway"]
B --> C["Agent Orchestrator"]
C --> D["Planner"]
D --> E["Vector"]
D --> F["SQL"]
D --> G["Graph"]
D --> H["Multimodal"]
D --> I["Enterprise APIs"]
E --> J["Evidence"]
F --> J
G --> J
H --> J
I --> J
J --> K["Evidence Evaluator"]
K --> L{"Sufficient?"}
L -->|No| D
L -->|Yes| M["Response Generator"]
M --> N["Validation"]
N --> O["Citation"]
O --> P["Enterprise Response"] π§ 69. Agentic RAG vs Agentic AIΒΆ
These terms are related but not identical.
Agentic RAGΒΆ
The agent primarily focuses on:
General Agentic AIΒΆ
An agent may:
Agentic RAG should generally maintain a narrower scope when the goal is knowledge retrieval.
π‘οΈ 70. Read vs Write ToolsΒΆ
A critical production distinction:
versus:
Agentic RAG systems should normally prioritize read-only tools.
Write capabilities require significantly stronger authorization and approval mechanisms.
π 71. Human ApprovalΒΆ
For high-impact operations:
Even though Agentic RAG is primarily retrieval-oriented, enterprise systems may eventually connect agents to operational tools.
π§ 72. Agent Planning StrategiesΒΆ
Possible approaches include:
Single-Step Planning
ReAct-Style Loop
Plan-and-Execute
Query Decomposition
Graph-Based Planning
Workflow-Based Planning
The right approach depends on:
π 73. ReAct-Style Agentic RetrievalΒΆ
Conceptually:
In production systems, internal reasoning should not be exposed to users.
The observable interface should focus on:
π§ 74. Plan-and-ExecuteΒΆ
Another approach:
Example:
1. Find incident.
2. Find affected services.
3. Query transaction impact.
4. Find remediation.
5. Synthesize.
This can provide more predictable execution than completely free-form agent loops.
βοΈ 75. Plan-and-Execute vs ReActΒΆ
| Characteristic | ReAct-Style | Plan-and-Execute |
|---|---|---|
| Dynamic adaptation | High | Medium |
| Predictability | Medium | High |
| Planning upfront | Low | High |
| Tool flexibility | High | High |
| Debugging | Medium | Easier |
| Cost control | Harder | Easier |
| Complex dependencies | Strong | Strong |
A hybrid architecture can combine both.
ποΈ 76. Hybrid Agent ArchitectureΒΆ
Query
β
Planner
β
High-Level Plan
β
Execution
β
Dynamic Re-planning
β
Validation
β
Answer
This balances:
π§ 77. Agent Failure ModesΒΆ
Common failures:
Wrong Tool Selection
Wrong Query
Retrieval Loop
Tool Overuse
Missing Evidence
Conflicting Evidence
Incorrect Planning
Hallucinated Tool Results
Prompt Injection
Unauthorized Tool Access
Excessive Cost
Excessive Latency
Incorrect Final Synthesis
π¨ 78. Tool Selection FailureΒΆ
Question:
The agent chooses:
instead of:
This produces a poor answer.
Tool selection should therefore be evaluated independently.
π¨ 79. Retrieval Loop FailureΒΆ
Example:
Use:
π¨ 80. Tool OveruseΒΆ
A simple query:
should not trigger:
if one document retrieval is enough.
Over-orchestration increases cost and latency.
π§ 81. Agent Cost ModelΒΆ
Agentic RAG cost can include:
A simple conceptual model:
β‘ 82. Latency ModelΒΆ
Similarly:
Parallel execution can reduce latency when dependencies allow it.
π§© 83. Agent BudgetΒΆ
A production agent can have:
When the budget is exhausted:
or:
depending on the confidence level.
π§ 84. Agent ObservabilityΒΆ
Every agent execution should be traceable.
Track:
Trace ID
Session ID
User
Tenant
Original Query
Plan
Tool Calls
Tool Inputs
Tool Results
Iterations
Latency
Token Usage
Cost
Errors
Evidence
Final Answer
Sensitive information should be redacted according to organizational policy.
π 85. Agent TraceΒΆ
Example:
Trace: agent-8291
00ms Query received
120ms Planner
β vector_search
350ms Vector result
β Incident INC-1042
410ms Planner
β sql_query
720ms SQL result
β 12,430 affected customers
780ms Planner
β vector_search
1040ms Retrieved remediation document
1100ms Evidence validation
1300ms Answer generated
This trace makes production debugging significantly easier.
π§ 86. Agent EvaluationΒΆ
Agentic RAG should be evaluated at multiple levels.
RetrievalΒΆ
Tool SelectionΒΆ
PlanningΒΆ
ExecutionΒΆ
GroundingΒΆ
OperationsΒΆ
π 87. Agent Evaluation MatrixΒΆ
| Metric | Purpose |
|---|---|
| Retrieval Recall | Evidence discovery |
| Tool Selection Accuracy | Correct capability selection |
| Plan Success Rate | Planning quality |
| Task Success Rate | End-to-end success |
| Evidence Sufficiency | Retrieval completeness |
| Groundedness | Answer support |
| Citation Accuracy | Source correctness |
| Iterations | Agent efficiency |
| Tool Calls | Orchestration efficiency |
| Latency | User experience |
| Cost | Operational efficiency |
π§ͺ 88. Agentic RAG Evaluation DatasetΒΆ
Include:
Simple Queries
Multi-Hop Queries
Multi-Source Queries
Ambiguous Queries
SQL Questions
Graph Questions
Visual Questions
Incomplete Evidence
Conflicting Evidence
Security Attacks
Prompt Injection
Tool Failures
π 89. Regression TestingΒΆ
Every production change should be evaluated against a regression set.
Example:
Example:
{
"query": "What caused the payment outage?",
"expected_tools": [
"vector_search"
],
"max_tool_calls": 3
}
π§ 90. Deterministic EvaluationΒΆ
Where possible, evaluate deterministic artifacts:
rather than relying only on an LLM judge.
LLM-based evaluation can supplement deterministic evaluation but should not be the only measurement mechanism.
π§© 91. Agentic RAG Security ArchitectureΒΆ
flowchart TD
A["User"] --> B["Authentication"]
B --> C["Authorization"]
C --> D["Agent Gateway"]
D --> E["Policy Engine"]
E --> F["Agent"]
F --> G["Tool Authorization"]
G --> H["Tool"]
H --> I["Tool Validation"]
I --> J["Execution"]
J --> K["Audit Log"] Security should surround the agent rather than depend on the agent.
π‘οΈ 92. Guardrail LayersΒΆ
A robust architecture can apply guardrails at multiple points:
Examples:
Input:
Prompt Injection Detection
Agent:
Allowed Tool Set
Tool:
Authorization
Data:
Tenant Isolation
Output:
Grounding / PII Validation
π§ 93. Production Agentic RAG ArchitectureΒΆ
flowchart TD
A["User"] --> B["AI Gateway"]
B --> C["Authentication"]
C --> D["Authorization"]
D --> E["Agent Orchestrator"]
E --> F["Planner"]
F --> G["Policy Engine"]
G --> H["Tool Router"]
H --> I["Vector Search"]
H --> J["SQL"]
H --> K["Knowledge Graph"]
H --> L["Multimodal Search"]
I --> M["Evidence Store"]
J --> M
K --> M
L --> M
M --> N["Evidence Evaluator"]
N --> O{"Sufficient?"}
O -->|No| F
O -->|Yes| P["Context Builder"]
P --> Q["Foundation Model"]
Q --> R["Response Validator"]
R --> S["Citation"]
S --> T["Enterprise Response"]
E --> U["Observability"]
H --> U
N --> U
Q --> U π§ 94. Enterprise Agentic RAG Reference ArchitectureΒΆ
βββββββββββββββββ
β USER β
βββββββββ¬ββββββββ
β
βΌ
βββββββββββββββββ
β AI GATEWAY β
βββββββββ¬ββββββββ
β
βΌ
βββββββββββββββββ
β AUTHORIZATION β
βββββββββ¬ββββββββ
β
βΌ
βββββββββββββββββββββ
β AGENT ORCHESTRATORβ
βββββββββββ¬ββββββββββ
β
ββββββββββββ΄βββββββββββ
βΌ βΌ
βββββββββββ ββββββββββββ
β PLANNER β β STATE β
ββββββ¬βββββ ββββββ¬ββββββ
β β
βββββββββββ¬βββββββββββ
βΌ
ββββββββββββββββββ
β TOOL ROUTER β
βββββββββ¬βββββββββ
β
ββββββββββββββββββββββββΌβββββββββββββββββββββββββ
βΌ βΌ βΌ
βββββββββββββββ ββββββββββββββ βββββββββββββββ
β Vector RAG β β SQL / Data β β Graph RAG β
ββββββββ¬βββββββ βββββββ¬βββββββ ββββββββ¬βββββββ
β β β
ββββββββββββββββββββββββΌβββββββββββββββββββββββββ
βΌ
βββββββββββββββββββ
β EVIDENCE STORE β
ββββββββββ¬βββββββββ
β
βΌ
βββββββββββββββββββ
β EVIDENCE β
β EVALUATOR β
ββββββββββ¬βββββββββ
β
ββββββββ΄βββββββ
β β
More Enough
β β
βΌ βΌ
Re-plan Context
β
βΌ
ββββββββββββββββ
β FOUNDATION β
β MODEL β
ββββββββ¬ββββββββ
β
βΌ
ββββββββββββββββ
β VALIDATION β
ββββββββ¬ββββββββ
β
βΌ
ββββββββββββββββ
β CITATIONS β
ββββββββ¬ββββββββ
β
βΌ
ββββββββββββββββ
β RESPONSE β
ββββββββββββββββ
βββββββββββββββββββββββββββββββββββββββββββββ
β OBSERVABILITY / AUDIT / COST β
βββββββββββββββββββββββββββββββββββββββββββββ
π§ͺ 95. Practical ProjectΒΆ
Build an Agentic RAG assistant for an enterprise knowledge base.
Data SourcesΒΆ
Architecture Documents
Incident Reports
Product Documentation
Customer Data
Service Dependency Graph
ToolsΒΆ
Example QueryΒΆ
"Which customers were affected by the latest
payment incident, what services were involved,
and what remediation was implemented?"
π 96. Expected Agent WorkflowΒΆ
User Query
β
Planner
β
Vector Search
β
Incident Found
β
Graph Search
β
Affected Services
β
SQL
β
Affected Customers
β
Vector Search
β
Remediation
β
Evidence Evaluation
β
Context Assembly
β
LLM
β
Validation
β
Citations
β
Response
π§ͺ 97. Advanced Project ExtensionsΒΆ
Add:
Query Decomposition
Parallel Retrieval
Adaptive Query Rewriting
Evidence Ranking
Conflict Detection
Multimodal Retrieval
SQL Validation
Graph Traversal
Agent Memory
Human Approval
Then measure:
π¨ 98. Common MistakesΒΆ
Mistake 1 β Using Agents EverywhereΒΆ
Simple RAG queries do not need an agent.
Mistake 2 β Unlimited LoopsΒΆ
Always enforce execution budgets.
Mistake 3 β Giving the Agent Every ToolΒΆ
Use capability-based authorization.
Mistake 4 β Allowing Arbitrary SQLΒΆ
Use structured query plans and SQL validation.
Mistake 5 β Trusting Retrieved InstructionsΒΆ
Retrieved content is data, not authority.
Mistake 6 β No Evidence ValidationΒΆ
The agent should verify that retrieved information actually supports the answer.
Mistake 7 β No AbstentionΒΆ
Agents must be able to stop when evidence is insufficient.
Mistake 8 β Ignoring CostΒΆ
Multiple model and tool calls can make Agentic RAG expensive.
Mistake 9 β Ignoring ObservabilityΒΆ
Without traces, debugging agent behavior becomes difficult.
Mistake 10 β Exposing Internal ReasoningΒΆ
Production systems should expose useful execution metadata rather than private chain-of-thought.
π§ 99. Design PrinciplesΒΆ
Principle 1 β Use Agents for ComplexityΒΆ
Principle 2 β Retrieval Should Be AdaptiveΒΆ
The agent should be able to change retrieval strategy based on evidence.
Principle 3 β Tools Should Be Capability-BasedΒΆ
Expose:
rather than infrastructure-specific implementation details.
Principle 4 β Keep Tools ControlledΒΆ
Every tool should have:
Principle 5 β Bound EverythingΒΆ
Control:
Principle 6 β Evidence Before AnswerΒΆ
Principle 7 β Support AbstentionΒΆ
If evidence is insufficient:
Principle 8 β Preserve ProvenanceΒΆ
Every claim should be traceable to evidence.
Principle 9 β Keep Security Outside the LLMΒΆ
The model should never be the final authorization layer.
Principle 10 β Measure the AgentΒΆ
Monitor:
π 100. Production ChecklistΒΆ
β Identify Agentic RAG use cases
β Identify queries that require adaptive retrieval
β Identify required tools
β Define tool contracts
β Define tool authorization
β Implement query understanding
β Implement query decomposition
β Implement planning
β Implement retrieval routing
β Implement agent state
β Implement tool execution
β Implement Vector Search
β Implement SQL Retrieval
β Implement Graph Retrieval
β Implement Multimodal Retrieval
β Implement Document Search
β Implement evidence collection
β Implement evidence deduplication
β Implement evidence ranking
β Implement evidence evaluation
β Implement conflict detection
β Implement query rewriting
β Implement iterative retrieval
β Implement multi-hop retrieval
β Implement parallel retrieval
β Implement dependency-aware execution
β Implement maximum iterations
β Implement maximum tool calls
β Implement latency budget
β Implement token budget
β Implement cost budget
β Implement authentication
β Implement authorization
β Implement tenant isolation
β Implement tool-level security
β Implement SQL security
β Implement prompt-injection defenses
β Treat retrieved content as untrusted data
β Validate tool inputs
β Validate tool outputs
β Restrict tool capabilities
β Implement answer grounding
β Implement response validation
β Implement citation
β Implement abstention
β Implement clarification
β Implement agent tracing
β Track tool calls
β Track iterations
β Track latency
β Track tokens
β Track cost
β Track errors
β Build evaluation dataset
β Test tool selection
β Test planning
β Test multi-hop retrieval
β Test conflicting evidence
β Test prompt injection
β Test unauthorized tool access
β Compare Agentic RAG with traditional RAG
β Measure accuracy improvement
β Measure latency overhead
β Measure cost overhead
β Establish production SLOs
π 101. Key TakeawaysΒΆ
- Agentic RAG makes retrieval adaptive.
- Traditional RAG generally follows a predetermined retrieval path.
- Agentic RAG can dynamically select tools and retrieval strategies.
- Agents are useful for complex, multi-hop, multi-source questions.
- Agentic retrieval can combine Vector Search, SQL, Knowledge Graphs, and Multimodal Retrieval.
- Query decomposition can break complex questions into manageable sub-questions.
- Dependent sub-questions should be executed sequentially.
- Independent sub-questions can often be executed in parallel.
- Agent state tracks the current retrieval process.
- Evidence evaluation helps determine whether additional retrieval is necessary.
- Query rewriting can improve subsequent retrieval rounds.
- Multi-hop retrieval allows the system to follow relationships across multiple sources.
- Agent loops must always be bounded.
- Tool calls should be controlled and authorized.
- Retrieved documents must be treated as data, not executable instructions.
- Tool outputs should be validated before influencing subsequent execution.
- SQL access should remain read-only for typical RAG workloads.
- Agentic RAG should support abstention when evidence is insufficient.
- Agents should not replace deterministic workflows where deterministic workflows are sufficient.
- Agentic RAG introduces additional latency and cost.
- Observability is essential for understanding agent behavior.
- Agent evaluation should include retrieval, tool selection, planning, grounding, task success, latency, and cost.
- Security controls must exist outside the LLM.
- Production Agentic RAG is an orchestration problem as much as an AI problem.
- The goal is not maximum autonomy.
- The goal is controlled, observable, evidence-driven adaptability.
π§ Final Mental ModelΒΆ
USER QUESTION
β
βΌ
QUERY UNDERSTANDING
β
βΌ
AGENT
β
βΌ
PLANNER
β
βΌ
TOOL SELECTION
β
βββββββββββββββββββΌββββββββββββββββββ
βΌ βΌ βΌ
VECTOR SQL GRAPH
β β β
βββββββββββββββββββΌββββββββββββββββββ
β
βΌ
TOOL RESULTS
β
βΌ
EVIDENCE STORE
β
βΌ
EVIDENCE EVALUATOR
β
ββββββββββββ΄βββββββββββ
βΌ βΌ
Insufficient Sufficient
β β
βΌ βΌ
QUERY REWRITE CONTEXT BUILD
β β
βΌ βΌ
RETRIEVE LLM
β β
ββββββββββββ βΌ
β VALIDATION
β β
β βΌ
ββββββββΊ CITATION
β
βΌ
ENTERPRISE RESPONSE
The central principle is:
Agentic RAG turns retrieval from a fixed pipeline into a controlled decision-making loop that can select tools, gather evidence, evaluate results, and adapt the retrieval strategy before generating an answer.
A production-grade Agentic RAG architecture therefore combines:
Agentic Orchestration
β
βββββββββββββββββββΌββββββββββββββββββ
βΌ βΌ βΌ
Planning Tooling State
β β β
βββββββββββββββββββΌββββββββββββββββββ
βΌ
Retrieval Fabric
β
βββββββββββββββββββΌββββββββββββββββββ
βΌ βΌ βΌ
Vector SQL Graph
β β β
βββββββββββββββββββΌββββββββββββββββββ
β
βΌ
Evidence Engineering
β
ββββββββββββΌβββββββββββ
βΌ βΌ βΌ
Ranking Validation Conflict
Detection
β β β
ββββββββββββΌβββββββββββ
βΌ
Context Engineering
β
βΌ
Foundation Model
β
βΌ
Response Validation
β
βΌ
Citation Layer
β
βΌ
Enterprise Response
The strongest enterprise design is not:
It is:
User
β
Policy
β
Agent
β
Controlled Tools
β
Trusted Retrieval
β
Evidence Evaluation
β
Context Engineering
β
LLM
β
Validation
β
Citation
β
Enterprise Response
That distinction is what separates an experimental AI agent from a production-grade enterprise Agentic RAG system.
π§ Chapter NavigationΒΆ
Part V β Advanced Retrieval-Augmented GenerationΒΆ
Previous:
05. Multimodal RAG
Next:
01. Prompt Assembly
Section:
05 β Advanced RAG Architecture
Advanced RAG Architecture PathΒΆ
01 Advanced RAG Architecture
β
02 Graph RAG
β
03 Knowledge Graphs for RAG
β
04 SQL RAG
β
05 Multimodal RAG
β
06 Agentic RAG
β
06 Production RAG Engineering
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.