Advanced Query RewritingΒΆ
π OverviewΒΆ
Advanced Query Rewriting is a retrieval-engineering technique that transforms a user's original query into one or more retrieval-optimized queries before searching the knowledge base.
A user's natural-language question is optimized for human communication, not necessarily for retrieval.
For example:
User Query:
"Why did our payment service start failing after the latest deployment,
and what should we check first?"
A single semantic search may not retrieve all the required evidence.
Query rewriting can transform it into:
Query 1:
Payment service failures after deployment
Query 2:
Payment deployment-related incidents
Query 3:
Payment service error logs after release
Query 4:
Payment service deployment rollback troubleshooting
The resulting retrieval process becomes:
User Query
β
Query Understanding
β
Query Rewriting
β
Optimized Query / Queries
β
Retrieval
β
Candidate Fusion
β
Re-ranking
β
MMR
β
Context
Advanced query rewriting is particularly valuable for enterprise RAG systems where users frequently submit:
- Ambiguous questions
- Conversational questions
- Long questions
- Multi-intent questions
- Questions containing pronouns
- Questions containing implicit context
- Questions using business terminology
- Questions requiring multiple retrieval perspectives
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand why query rewriting is required in RAG
- Distinguish query rewriting from query expansion
- Understand query normalization
- Rewrite conversational queries
- Resolve ambiguous references
- Generate multiple retrieval queries
- Implement query decomposition
- Handle multi-intent questions
- Generate hypothetical retrieval queries
- Use domain-aware query rewriting
- Combine query rewriting with hybrid search
- Combine query rewriting with re-ranking
- Combine query rewriting with MMR
- Implement query rewriting with LLMs
- Validate generated queries
- Prevent query drift
- Handle security-sensitive query rewriting
- Design adaptive query rewriting pipelines
- Evaluate query rewriting in production RAG systems
1. Why Query Rewriting MattersΒΆ
Traditional retrieval assumes:
But the original query may not be retrieval-friendly.
A user might ask:
A human understands:
and:
from conversation context.
A retriever does not automatically have the same understanding.
Query rewriting converts:
into:
2. Human Query vs Retrieval QueryΒΆ
A useful distinction is:
optimized for:
while:
should be optimized for:
Example:
Possible retrieval query:
3. Basic Query Rewriting ArchitectureΒΆ
flowchart TD
A["User Query"] --> B["Query Understanding"]
B --> C["Query Rewriter"]
C --> D["Optimized Query"]
D --> E["Retriever"]
E --> F["Candidate Documents"]
F --> G["Re-ranking"]
G --> H["MMR"]
H --> I["Context"]
I --> J["LLM"] The query rewriter sits before retrieval.
4. Query Rewriting vs Query ExpansionΒΆ
These concepts are related but not identical.
Query RewritingΒΆ
Transforms the original query:
Query ExpansionΒΆ
Adds additional terms or alternative formulations:
Example:
Expansion:
Rewriting:
5. Query Rewriting vs Multi-Query RetrievalΒΆ
Multi-query retrieval generates multiple alternative queries.
Query rewriting may generate:
or:
Therefore:
is the broader concept.
6. Query NormalizationΒΆ
The simplest form of rewriting is normalization.
Example:
becomes:
Normalization can handle:
7. Terminology NormalizationΒΆ
Suppose the enterprise uses:
but users say:
The rewriter can map these terms into canonical terminology.
Example:
User:
"How does the payment API validate transactions?"
Canonical:
"Payment Gateway transaction validation"
This is especially useful when the enterprise has a controlled vocabulary.
8. Domain-Aware RewritingΒΆ
Generic rewriting may produce:
while an enterprise-specific rewriter can produce:
Domain-aware rewriting can use:
9. Query Rewriting with Enterprise VocabularyΒΆ
flowchart LR
A["User Query"] --> B["Terminology Resolver"]
C["Enterprise Glossary"] --> B
B --> D["Canonical Query"]
D --> E["Retriever"] This reduces vocabulary mismatch between:
and:
10. Conversational Query RewritingΒΆ
Consider:
User:
How does Kafka handle payment events?
Assistant:
Kafka uses producers and consumers.
User:
What about retries?
The second query is incomplete.
The rewriter can transform:
into:
Now retrieval has sufficient context.
11. Conversational RewritingΒΆ
sequenceDiagram
participant U as User
participant R as Query Rewriter
participant V as Retriever
participant L as LLM
U->>R: What about retries?
R->>R: Resolve conversation context
R->>R: Resolve pronouns and omitted entities
R->>V: Kafka payment event retry mechanisms
V->>L: Retrieved context
L->>U: Grounded answer 12. Pronoun ResolutionΒΆ
Queries frequently contain:
Example:
The rewriter can produce:
This improves retrieval specificity.
13. Reference ResolutionΒΆ
Example:
The rewriter should identify:
from conversation context.
Result:
This is often called:
14. Query Context WindowΒΆ
A rewriting system may use:
Example:
The output should be:
rather than an answer.
15. Query Rewriting PromptΒΆ
A basic LLM prompt can be:
You are a retrieval query optimizer.
Rewrite the user's query into a concise,
search-optimized query.
Rules:
- Preserve the user's intent.
- Resolve references using conversation context.
- Do not introduce unsupported facts.
- Preserve important technical terminology.
- Return only the rewritten query.
Conversation:
{conversation}
User query:
{query}
16. Structured Query RewritingΒΆ
For production systems, structured output is often safer.
{
"original_query": "What about retries?",
"rewritten_query": "Kafka payment event retry mechanisms",
"intent": "troubleshooting",
"entities": [
"Kafka",
"payment events"
]
}
This provides additional signals for downstream retrieval.
17. Query Rewriting PipelineΒΆ
Original Query
β
Language Detection
β
Intent Detection
β
Entity Resolution
β
Terminology Normalization
β
Query Rewriting
β
Validation
β
Retrieval
Not every application requires every stage.
18. Query IntentΒΆ
A query may represent:
Fact Lookup
How-To
Troubleshooting
Comparison
Research
Policy
Definition
Historical Search
Analytical Question
The rewriter can use intent to produce a better retrieval query.
19. Intent-Aware RewritingΒΆ
Example:
Intent:
Rewrite:
Whereas:
has:
and could become:
20. Entity ExtractionΒΆ
Query rewriting benefits from identifying:
Example:
Entities:
{
"service": "payment-service-v3",
"technology": "Kafka",
"event": "upgrade",
"intent": "incident investigation"
}
21. Entity-Preserving RewritingΒΆ
The rewriter must not remove important identifiers.
Bad:
Better:
Important identifiers should survive rewriting.
22. Query DecompositionΒΆ
Some queries contain multiple questions.
Example:
"How does payment authentication work,
what happens when authentication fails,
and how are failures retried?"
One query may be insufficient.
Decompose into:
Q1:
Payment authentication mechanism
Q2:
Payment authentication failure handling
Q3:
Payment failure retry mechanism
23. Query Decomposition ArchitectureΒΆ
flowchart TD
A["Complex Query"] --> B["Intent Analysis"]
B --> C["Sub-question 1"]
B --> D["Sub-question 2"]
B --> E["Sub-question 3"]
C --> F["Retriever"]
D --> F
E --> F
F --> G["Candidate Fusion"]
G --> H["Re-ranking"]
H --> I["MMR"]
I --> J["Context"] This is particularly useful for complex enterprise questions.
24. Multi-Intent QueriesΒΆ
Example:
"Compare AWS and Azure deployment options,
explain the security differences,
and tell me which is cheaper."
Intents:
A single retrieval query may miss evidence for one or more intents.
25. Multi-Query GenerationΒΆ
A rewriter can generate:
Q1:
AWS deployment architecture
Q2:
Azure deployment architecture
Q3:
AWS vs Azure security comparison
Q4:
AWS vs Azure deployment cost
Then:
26. Query FusionΒΆ
Multiple rewritten queries can produce overlapping candidates.
Fusion produces:
Then:
can select the best evidence.
27. Query Rewriting + MMRΒΆ
This creates a powerful pipeline:
User Query
β
Query Rewriting
β
Multiple Retrieval Queries
β
Candidate Fusion
β
Re-ranking
β
MMR
β
Context
Query rewriting improves:
while MMR improves:
28. Query Rewriting + Hybrid SearchΒΆ
A rewritten query can be sent to:
Example:
flowchart TD
A["User Query"] --> B["Query Rewriter"]
B --> C["Rewritten Query"]
C --> D["Dense Search"]
C --> E["BM25"]
D --> F["Fusion"]
E --> F
F --> G["Re-ranking"]
G --> H["MMR"]
H --> I["Context"] This is useful for queries containing:
29. Query Rewriting + MetadataΒΆ
The rewriter can produce both:
and:
Example:
{
"semantic_query": "payment gateway authentication",
"filters": {
"document_type": "architecture",
"status": "approved",
"region": "EU"
}
}
This connects query rewriting directly with metadata-aware retrieval.
30. Query Rewriting + Self-Query RetrievalΒΆ
Self-query retrieval can transform:
into:
Example:
becomes:
Advanced query rewriting can act as the query-understanding layer before this process.
31. Query Rewriting + HyDEΒΆ
HyDE generates a hypothetical answer/document and uses it for retrieval.
Pipeline:
Example:
Rewritten:
Hypothetical document:
The hypothetical representation is then embedded for retrieval.
32. Query Rewriting + Re-rankingΒΆ
A strong pipeline:
The stages have different responsibilities:
33. Query Rewriting + Agentic RetrievalΒΆ
Agentic systems can dynamically rewrite queries.
flowchart TD
A["User Query"] --> B["Agent"]
B --> C["Rewrite Query"]
C --> D["Retrieve"]
D --> E["Evaluate Evidence"]
E --> F{"Enough Evidence?"}
F -->|No| G["Rewrite / Expand Query"]
G --> D
F -->|Yes| H["Context"]
H --> I["Generate Answer"] This enables iterative retrieval.
34. Iterative Query RewritingΒΆ
Example:
Retrieved evidence:
The agent recognizes:
Next query:
Then:
This creates a retrieval loop.
35. Query Reformulation LoopΒΆ
This is more powerful than static rewriting but introduces:
36. Query DriftΒΆ
One of the biggest risks of iterative rewriting is query drift.
Original:
Bad rewrite:
Further rewrite:
The system has moved away from the original intent.
37. Query Drift PreventionΒΆ
Every rewritten query should preserve:
A useful rule:
38. Query Rewriting ValidationΒΆ
Validate generated queries against:
Example:
def validate_rewrite(
original,
rewritten,
required_entities
):
for entity in required_entities:
if entity.lower() not in rewritten.lower():
return False
return True
Production validation should be more sophisticated.
39. Query Rewriting GuardrailsΒΆ
β Preserve user intent
β Preserve important entities
β Preserve identifiers
β Preserve time constraints
β Preserve tenant context
β Do not invent facts
β Do not remove security constraints
β Do not broaden authorization
β Prevent query drift
β Validate output structure
40. Security Context Must Not Be RewrittenΒΆ
Suppose the application knows:
The user asks:
The rewriter must not produce:
Security constraints should come from:
not from LLM-generated query text.
41. Query InjectionΒΆ
Users may attempt to manipulate the rewriting prompt:
The rewriter should treat the query as:
and preserve:
42. Trusted Query ContextΒΆ
A safe internal structure:
{
"user_query": "Search all company documents",
"tenant_context": {
"tenant_id": "tenant-a"
},
"authorization_context": {
"roles": [
"employee"
]
}
}
The LLM can rewrite:
but should not control:
43. Query Rewriting Output SchemaΒΆ
A production-oriented output can be:
{
"rewritten_query": "payment gateway authentication mechanism",
"intent": "technical_explanation",
"entities": [
"payment gateway",
"authentication"
],
"sub_queries": [],
"filters": {},
"must_preserve": [
"payment gateway"
]
}
Structured output makes downstream processing safer.
44. Multiple Query StrategiesΒΆ
Advanced rewriting can use different strategies.
Strategy 1 β Rewrite
Strategy 2 β Expansion
Strategy 3 β Decomposition
Strategy 4 β Multi-Query
Strategy 5 β HyDE
Strategy 6 β Query Routing
A query planner can choose the appropriate strategy.
45. Query Strategy RouterΒΆ
flowchart TD
A["User Query"] --> B["Query Classifier"]
B --> C{"Query Type"}
C -->|Simple| D["Direct Rewrite"]
C -->|Conversational| E["Contextual Rewrite"]
C -->|Complex| F["Query Decomposition"]
C -->|Research| G["Multi-Query"]
C -->|Semantic Gap| H["HyDE"]
C -->|Structured| I["Metadata Extraction"]
D --> J["Retrieval"]
E --> J
F --> J
G --> J
H --> J
I --> J This avoids applying expensive techniques to every query.
46. Simple QueryΒΆ
Example:
Do not over-process it.
Potential output:
A simple query may only need normalization.
47. Conversational QueryΒΆ
Example:
Use:
Output:
48. Complex QueryΒΆ
Example:
"How does our payment platform authenticate users,
handle failed transactions, and monitor suspicious activity?"
Decompose:
49. Research QueryΒΆ
Example:
Generate perspectives:
Kafka architecture
RabbitMQ architecture
Kafka vs RabbitMQ trade-offs
Kafka scalability
RabbitMQ routing
Kafka reliability
RabbitMQ reliability
Then fuse and rank results.
50. Comparison QueriesΒΆ
Comparison questions benefit from balanced query generation.
Example:
Generate:
AWS Lambda event-driven payment processing
Azure Functions event-driven payment processing
AWS Lambda vs Azure Functions scalability
AWS Lambda vs Azure Functions cost
AWS Lambda vs Azure Functions security
This improves evidence coverage.
51. Temporal QueriesΒΆ
Example:
Rewrite:
Metadata:
Potential sources:
52. Negative ConstraintsΒΆ
Users may specify:
The retrieval representation should preserve:
or:
depending on the enterprise schema.
Negative constraints are important because naive rewriting can accidentally remove them.
53. Exact IdentifiersΒΆ
Queries may contain:
These identifiers should generally be preserved exactly.
Example:
should not become:
The incident identifier may be the strongest retrieval signal.
54. Query Rewriting and BM25ΒΆ
Exact identifiers are particularly valuable for sparse retrieval.
Example:
Dense retrieval may interpret the concept semantically.
BM25 can match:
exactly.
Therefore:
is often stronger than dense retrieval alone for enterprise systems.
55. Query Rewriting and Hybrid SearchΒΆ
A rewritten query may contain:
Hybrid search can then exploit all three.
Rewritten Query
β
βββββββββββ΄ββββββββββ
β β
Dense Search BM25
β β
βββββββββββ¬ββββββββββ
β
Fusion
β
Re-ranking
56. Query Rewriting for AcronymsΒΆ
Enterprise environments frequently use acronyms.
Example:
The rewriter can preserve:
and potentially expand known terminology:
The expansion should only occur when the mapping is supported by trusted enterprise vocabulary.
57. Query Rewriting for TyposΒΆ
Example:
Rewrite:
This is a low-risk rewriting scenario.
58. Query Rewriting for Natural LanguageΒΆ
Example:
Rewrite:
This can bridge:
and:
59. Query Rewriting and Knowledge GraphsΒΆ
A query may contain entities and relationships:
A rewriter can identify:
The system can then route the query toward:
instead of pure vector search.
60. Query Routing by Retrieval TypeΒΆ
flowchart TD
A["Query Rewriting"] --> B["Query Planner"]
B --> C{"Retrieval Requirement"}
C -->|Semantic| D["Vector Search"]
C -->|Exact Match| E["BM25"]
C -->|Structured Data| F["SQL RAG"]
C -->|Relationships| G["Graph RAG"]
C -->|Mixed| H["Hybrid Retrieval"]
D --> I["Evidence"]
E --> I
F --> I
G --> I
H --> I Advanced query rewriting therefore becomes part of retrieval orchestration.
61. Query Rewriting for SQL RAGΒΆ
User:
The query planner may identify:
Then generate:
rather than searching documents.
62. Query Rewriting for Graph RAGΒΆ
User:
Rewrite:
This can be mapped to graph traversal.
63. Query Rewriting for Multimodal RAGΒΆ
User:
The system may detect:
Query planning:
This can trigger multimodal retrieval.
64. Query Rewriting and Agentic RAGΒΆ
Agentic RAG can use rewriting as an iterative planning operation:
Question
β
Plan
β
Rewrite
β
Retrieve
β
Evaluate
β
Rewrite Again
β
Retrieve
β
Sufficient Evidence
β
Answer
This is particularly useful for complex research questions.
65. Preventing Infinite Retrieval LoopsΒΆ
Agentic rewriting should have limits:
Example:
MAX_RETRIEVAL_ROUNDS = 3
for _ in range(MAX_RETRIEVAL_ROUNDS):
results = retrieve(query)
if sufficient_evidence(results):
break
query = rewrite(query)
Production systems should also track why each iteration occurred.
66. Query Rewriting CostΒΆ
Every rewrite may involve:
which adds:
Therefore do not rewrite every query blindly.
A practical strategy:
Simple Query
β Direct Retrieval
Complex Query
β Rewrite
Ambiguous Query
β Rewrite
Low Retrieval Confidence
β Rewrite Again
67. Adaptive Query RewritingΒΆ
flowchart TD
A["User Query"] --> B["Complexity / Ambiguity Check"]
B -->|Low| C["Direct Retrieval"]
B -->|High| D["Query Rewrite"]
C --> E["Evaluate Retrieval"]
D --> E
E --> F{"Confidence Low?"}
F -->|Yes| G["Advanced Rewrite / Expansion"]
F -->|No| H["Continue"]
G --> E This balances quality and cost.
68. Retrieval ConfidenceΒΆ
Signals may include:
If confidence is low:
may be triggered.
69. Score-Gap SignalΒΆ
Suppose:
This may indicate:
MMR can help.
But if:
the retrieval landscape may be weak or highly specific.
The system may need:
rather than simply increasing diversity.
70. Query Rewriting Decision MatrixΒΆ
| Situation | Technique |
|---|---|
| Typo | Normalization |
| Ambiguous reference | Conversational rewrite |
| Missing terminology | Domain rewrite |
| Multi-intent | Decomposition |
| Low recall | Query expansion |
| Research question | Multi-query |
| Semantic mismatch | HyDE |
| Exact identifiers | Hybrid retrieval |
| Structured filters | Metadata extraction |
| Relationship query | Graph routing |
| Database question | SQL routing |
71. Advanced Query Rewriting ServiceΒΆ
A production service can expose:
class QueryRewriter:
def rewrite(
self,
query,
conversation=None,
metadata=None
):
raise NotImplementedError
Potential implementations:
72. Rule-Based + LLM RewritingΒΆ
Not everything requires an LLM.
Use deterministic rules for:
Use LLMs for:
This can reduce cost and improve predictability.
73. Hybrid Query RewriterΒΆ
flowchart TD
A["User Query"] --> B["Rule-Based Normalization"]
B --> C["Entity / Terminology Resolution"]
C --> D["Complexity Check"]
D -->|Simple| E["Final Query"]
D -->|Complex| F["LLM Rewriter"]
F --> G["Validation"]
G --> E
E --> H["Retriever"] This is often more production-friendly than an LLM-only approach.
74. Query Rewriting CacheΒΆ
Repeated queries can be cached.
Example:
Cache invalidation must account for:
75. Query Rewriting ObservabilityΒΆ
Track:
Original Query
Rewritten Query
Rewrite Strategy
Rewrite Latency
Model
Prompt Version
Confidence
Retrieval Improvement
Query Drift
Example:
{
"strategy": "conversational_rewrite",
"original": "What about retries?",
"rewritten": "Kafka payment event retry mechanisms",
"latency_ms": 84,
"prompt_version": "v3"
}
76. Rewrite Quality MetricsΒΆ
Evaluate:
Intent Preservation
Entity Preservation
Query Specificity
Query Drift
Retrieval Recall
Retrieval Precision
Answer Quality
A rewrite should not be considered successful simply because it sounds better.
It must improve downstream retrieval.
77. Retrieval Recall ComparisonΒΆ
Compare:
vs:
Example:
This indicates improved retrieval coverage.
78. Query Drift MetricΒΆ
Conceptually measure similarity between:
and:
If semantic similarity becomes too low:
This can trigger:
79. Rewrite A/B TestingΒΆ
Compare:
Measure:
80. Query Rewriting Evaluation DatasetΒΆ
Create examples:
{
"original_query": "What about retries?",
"conversation": [
"How does Kafka handle payment events?"
],
"expected_rewrite":
"Kafka payment event retry mechanisms"
}
Include test categories:
Simple
Conversational
Ambiguous
Multi-intent
Technical
Temporal
Comparison
Identifier-heavy
Metadata-heavy
Security-sensitive
81. Offline EvaluationΒΆ
Run rewriting against a fixed dataset.
Measure:
This makes prompt/model changes measurable.
82. Online EvaluationΒΆ
Monitor production:
A useful signal is:
Repeated clarification may indicate poor rewriting or retrieval.
83. Query Rewriting and User FeedbackΒΆ
User feedback can help identify failures.
Example:
Potential causes:
Observability should allow engineers to trace the complete chain.
84. Query Rewrite TraceΒΆ
Original Query
β
Normalized Query
β
Rewritten Query
β
Metadata Filters
β
Retriever
β
Candidates
β
Re-ranker
β
MMR
β
Context
β
Answer
This trace is extremely valuable for production debugging.
85. Query Rewriting Failure ModesΒΆ
Failure 1 β Intent DriftΒΆ
Original meaning changes.
Failure 2 β Entity LossΒΆ
Important identifiers disappear.
Failure 3 β Hallucinated TermsΒΆ
The rewriter invents unsupported terminology.
Failure 4 β Over-SpecificationΒΆ
The rewriter adds assumptions.
Failure 5 β Under-SpecificationΒΆ
The rewrite remains too generic.
86. Failure Modes β ContinuedΒΆ
Failure 6 β Security ExpansionΒΆ
The query is broadened beyond authorization.
Failure 7 β Temporal LossΒΆ
"Current" or "last year" disappears.
Failure 8 β Negative Constraint LossΒΆ
"Not archived" disappears.
Failure 9 β Domain MismatchΒΆ
Technical terminology is replaced incorrectly.
Failure 10 β Excessive Query GenerationΒΆ
Too many queries increase cost without improving recall.
87. Original vs Rewritten QueryΒΆ
Example:
Good rewrite:
Bad rewrite:
The second loses:
88. Preserving ConstraintsΒΆ
A rewrite should preserve:
Example:
must retain:
89. Query Rewrite ContractΒΆ
A strong internal contract is:
Input:
- User query
- Conversation context
- Trusted metadata context
Output:
- Rewritten query
- Intent
- Entities
- Optional sub-queries
- Optional filters
- Rewrite strategy
Security context remains outside the LLM-controlled output.
90. Query Rewriting with Context EngineeringΒΆ
The rewriter itself requires context.
Useful context may include:
Recent Conversation
Entity Memory
Enterprise Glossary
User's Current Task
Current Product
Current Service
But avoid providing unnecessary information.
Too much context can introduce:
91. Context Selection for RewritingΒΆ
A useful strategy:
rather than:
This is an important connection between:
and:
92. Query Rewriting and Long ConversationsΒΆ
Long conversations can contain:
The rewriter may accidentally use context from an older topic.
Therefore:
can improve rewriting.
93. Topic-Aware Conversation ContextΒΆ
flowchart TD
A["Conversation"] --> B["Topic Segmentation"]
B --> C["Current Topic"]
C --> D["Relevant Conversation Turns"]
D --> E["Query Rewriter"]
E --> F["Retrieval Query"] This reduces accidental context contamination.
94. Query Rewriting and MemoryΒΆ
Enterprise RAG applications may maintain:
The rewriter can use these signals to resolve references.
However:
should not override:
95. Query Rewriting and Prompt AssemblyΒΆ
The rewritten query is not necessarily the final prompt.
Pipeline:
This separation is important.
The rewriter should optimize:
not generate the final answer.
96. Query Rewriting and Response ValidationΒΆ
After generation:
If the answer is poorly grounded, the system may trigger:
which can use:
again.
This creates a closed-loop RAG system.
97. Closed-Loop RetrievalΒΆ
flowchart TD
A["User Query"] --> B["Query Rewrite"]
B --> C["Retrieve"]
C --> D["Context"]
D --> E["Generate"]
E --> F["Validate"]
F --> G{"Evidence Gap?"}
G -->|Yes| H["Rewrite / Refine Query"]
H --> C
G -->|No| I["Final Response"] This is a foundation for agentic RAG.
98. Query Rewriting and CitationΒΆ
If the generated answer lacks evidence for:
the system can identify the missing claim and generate a targeted retrieval query:
Then retrieve additional supporting evidence.
This connects query rewriting with:
99. Claim-Driven RetrievalΒΆ
Example:
Validation identifies:
Query:
Retrieve:
This is an advanced production pattern.
100. Query Rewriting for Evidence GapsΒΆ
Answer Validation
β
Missing Evidence
β
Claim Extraction
β
Targeted Query Rewrite
β
Retrieval
β
Evidence
β
Citation
This moves RAG beyond:
toward:
101. Production Query Rewriting ArchitectureΒΆ
flowchart TD
A["User Query"] --> B["Query Planner"]
B --> C["Normalization"]
C --> D["Entity Resolution"]
D --> E["Intent Detection"]
E --> F{"Strategy"}
F -->|Simple| G["Direct Rewrite"]
F -->|Conversational| H["Contextual Rewrite"]
F -->|Complex| I["Query Decomposition"]
F -->|Research| J["Multi-Query"]
F -->|Semantic Gap| K["HyDE"]
F -->|Structured| L["Metadata Extraction"]
G --> M["Validation"]
H --> M
I --> M
J --> M
K --> M
L --> M
M --> N["Hybrid / Dense Retrieval"]
N --> O["Re-ranking"]
O --> P["MMR"]
P --> Q["Context Engineering"]
Q --> R["Prompt Assembly"]
R --> S["LLM"]
S --> T["Response Validation"]
T --> U{"Evidence Gap?"}
U -->|Yes| B
U -->|No| V["Enterprise Response"] 102. Recommended StrategyΒΆ
A practical production strategy is:
1. Normalize
2. Resolve conversation references
3. Preserve entities and constraints
4. Detect query complexity
5. Choose rewrite strategy
6. Generate rewrite
7. Validate rewrite
8. Retrieve
9. Re-rank
10. Apply MMR
11. Build context
12. Generate response
13. Validate response
14. Retry retrieval only when necessary
103. Production GuardrailsΒΆ
β Preserve original query
β Preserve important entities
β Preserve identifiers
β Preserve time constraints
β Preserve negative constraints
β Preserve domain terminology
β Keep security context outside LLM control
β Validate structured output
β Detect query drift
β Limit number of generated queries
β Limit retrieval iterations
β Set latency budget
β Set cost budget
β Log rewrite decisions
β Version prompts
β Version rewrite models
β Evaluate retrieval impact
104. Practical Python InterfaceΒΆ
from dataclasses import dataclass, field
@dataclass
class RewriteResult:
rewritten_query: str
intent: str | None = None
entities: list[str] = field(
default_factory=list
)
sub_queries: list[str] = field(
default_factory=list
)
filters: dict = field(
default_factory=dict
)
strategy: str | None = None
This gives downstream retrieval components a stable contract.
105. Query Rewriter InterfaceΒΆ
class QueryRewriter:
def rewrite(
self,
query: str,
conversation: list[str] | None = None
) -> RewriteResult:
raise NotImplementedError
Implementations:
RuleBasedQueryRewriter
LLMQueryRewriter
DomainQueryRewriter
HybridQueryRewriter
AgenticQueryRewriter
106. Simple Hybrid ImplementationΒΆ
class HybridQueryRewriter:
def __init__(
self,
terminology_resolver,
llm_rewriter
):
self.terminology_resolver = (
terminology_resolver
)
self.llm_rewriter = llm_rewriter
def rewrite(
self,
query,
conversation=None
):
normalized = (
self.terminology_resolver
.normalize(query)
)
if self.is_simple(normalized):
return normalized
return self.llm_rewriter.rewrite(
normalized,
conversation
)
def is_simple(self, query):
return len(query.split()) <= 6
This is intentionally simplified.
107. Query Rewrite Prompt with Structured OutputΒΆ
SYSTEM
You are an enterprise retrieval query optimizer.
Your task is to transform the user query
into retrieval-ready representations.
Rules:
1. Preserve user intent.
2. Preserve important entities.
3. Preserve exact identifiers.
4. Preserve temporal constraints.
5. Preserve negative constraints.
6. Resolve conversational references when possible.
7. Do not invent facts.
8. Do not modify authorization context.
9. Generate sub-queries only when required.
10. Prefer concise retrieval queries.
Return JSON:
{
"rewritten_query": "...",
"intent": "...",
"entities": [],
"sub_queries": [],
"filters": {}
}
108. Example InputΒΆ
Conversation:
User:
How does our payment service process failed transactions?
Assistant:
It retries some failures based on configured policies.
User:
What about the Kafka ones after deployment?
109. Example OutputΒΆ
{
"rewritten_query":
"Kafka-related payment transaction failures after deployment",
"intent":
"troubleshooting",
"entities": [
"payment service",
"Kafka",
"deployment"
],
"sub_queries": [
"Kafka payment transaction failures",
"payment service Kafka failures after deployment",
"Kafka retry handling after deployment"
],
"filters": {}
}
The retrieval system can now search multiple perspectives.
110. Query FusionΒΆ
Suppose:
Fusion:
Then:
and:
produce the final context.
111. Query Rewriting Pipeline with FusionΒΆ
Original Query
β
Query Rewriter
β
βββββββΌββββββ
β β β
Q1 Q2 Q3
β β β
R1 R2 R3
βββββββΌββββββ
β
Candidate Fusion
β
Re-ranking
β
MMR
β
Context
This is one of the most useful patterns for advanced RAG.
112. Query Rewriting and RecallΒΆ
Query rewriting primarily helps when the original query has a mismatch with the knowledge base.
Examples:
or:
or:
Rewriting attempts to reduce these mismatches.
113. Query Rewriting and PrecisionΒΆ
Rewriting can also improve precision.
Example:
may be broad.
Context-aware rewrite:
may narrow retrieval.
However, excessive specificity can reduce recall.
Therefore:
must balance:
114. Query SpecificityΒΆ
A useful conceptual spectrum:
Too Broad
β
payment security
Balanced
β
payment gateway authentication security
Too Narrow
β
payment gateway OAuth 2.0 token introspection
authorization code flow RFC...
The ideal rewrite depends on the evidence available in the corpus.
115. Query Rewriting and Recall ExpansionΒΆ
Sometimes the opposite is required.
Original:
Potential expansion:
This broadens retrieval.
Therefore advanced rewriting should support both:
and:
116. Adaptive Rewriting DirectionΒΆ
flowchart TD
A["Original Query"] --> B["Retrieval Analysis"]
B --> C{"Problem"}
C -->|Too Broad| D["Narrow Query"]
C -->|Too Narrow| E["Expand Query"]
C -->|Ambiguous| F["Clarify Query"]
C -->|Multi-Intent| G["Decompose Query"]
C -->|Low Recall| H["Generate Alternatives"]
D --> I["Retrieval"]
E --> I
F --> I
G --> I
H --> I This is more powerful than always asking an LLM to "rewrite the query."
117. Query Rewriting and Search IntentΒΆ
The same words can represent different retrieval intents.
Example:
could mean:
Context and intent classification can determine the appropriate rewrite.
118. Intent + Entity + ConstraintΒΆ
A strong retrieval representation can be modeled as:
Example:
Intent:
Troubleshooting
Entity:
payment-service
Relationship:
depends_on Kafka
Constraint:
production
Time:
after July deployment
This is much richer than the original text alone.
119. Query RepresentationΒΆ
Example:
{
"intent": "troubleshooting",
"entities": [
"payment-service",
"Kafka"
],
"relationships": [
"payment-service depends_on Kafka"
],
"constraints": [
"production"
],
"time_range": {
"from": "2026-07-01"
}
}
This representation can drive multiple retrieval mechanisms.
120. Query PlannerΒΆ
The query planner can translate the representation into:
This is the beginning of a true retrieval execution layer.
121. Enterprise Query PlanningΒΆ
flowchart TD
A["User Query"] --> B["Query Understanding"]
B --> C["Structured Query Representation"]
C --> D["Query Planner"]
D --> E["Vector Query"]
D --> F["BM25 Query"]
D --> G["Metadata Filters"]
D --> H["Graph Query"]
D --> I["SQL Query"]
E --> J["Evidence"]
F --> J
G --> J
H --> J
I --> J
J --> K["Fusion"]
K --> L["Re-ranking"]
L --> M["MMR"] This architecture allows RAG systems to retrieve from multiple knowledge representations.
122. Query Rewriting and Production RAGΒΆ
At enterprise scale, query rewriting should not be viewed as an isolated LLM prompt.
It is a capability within:
and:
The complete system becomes:
User Intent
β
Query Representation
β
Query Planning
β
Query Rewriting
β
Multi-Source Retrieval
β
Ranking
β
Context Engineering
123. Key TakeawaysΒΆ
- Query rewriting transforms user language into retrieval-optimized representations.
- Human-friendly questions are not always retrieval-friendly.
- Query rewriting can improve retrieval recall and precision.
- Conversational queries frequently require rewriting.
- Pronoun and reference resolution are important for conversational RAG.
- Domain terminology normalization can reduce vocabulary mismatch.
- Query decomposition is useful for complex multi-intent questions.
- Multi-query rewriting can improve evidence coverage.
- Query rewriting works well with hybrid retrieval.
- Rewriting can generate both semantic queries and metadata constraints.
- Query rewriting can support Graph RAG and SQL RAG routing.
- HyDE can be combined with query rewriting.
- Re-ranking improves candidate ordering after retrieval.
- MMR improves diversity among selected candidates.
- Query rewriting and MMR solve different retrieval problems.
- Iterative rewriting enables agentic retrieval.
- Iterative rewriting introduces query drift risk.
- Rewrites must preserve important entities and constraints.
- Exact identifiers should generally be preserved.
- Temporal and negative constraints must not be silently removed.
- LLM-generated filters must be validated.
- Security and authorization context must remain controlled by trusted application logic.
- Query rewriting should not broaden access permissions.
- Simple queries should not be unnecessarily processed by expensive LLM rewriting.
- Rule-based normalization can complement LLM-based rewriting.
- Adaptive rewriting can choose between narrowing, expansion, clarification, and decomposition.
- Query rewriting should be evaluated by downstream retrieval performance, not linguistic quality alone.
- Production systems should track rewrite strategy, latency, model version, prompt version, and retrieval impact.
- Query rewriting can become part of a larger query-planning and retrieval-orchestration architecture.
- The ultimate objective is not a "better-looking query."
- The objective is better evidence retrieval for the user's actual intent.
The central production pattern is:
User Query
β
Understand Intent
β
Resolve Entities & Context
β
Rewrite / Decompose
β
Generate Retrieval Queries
β
Apply Trusted Metadata & Security
β
Retrieve Broadly
β
Re-rank Precisely
β
Select Diversely
β
Build Grounded Context
β
Generate
β
Validate
β
Cite
And the core principle is:
Rewrite for retrieval without changing the user's intent.
π§ Chapter NavigationΒΆ
Part V β Advanced Retrieval-Augmented GenerationΒΆ
Previous:
12. Metadata-Aware Retrieval
Next:
01. LlamaIndex Retrievers Overview
Section:
02 β Enterprise Retrieval Engineering
Enterprise Retrieval Engineering PathΒΆ
01 Contextual Compression Retriever
β
02 Ensemble Retriever
β
03 Multi-Vector Retriever
β
04 Time-Weighted Retriever
β
05 Hybrid Search Retriever
β
06 HyDE Retriever
β
07 Router Retriever
β
08 Multi-Stage Retrieval
β
09 Agentic Retrieval
β
10 Re-ranking Techniques
β
11 MMR & Diversity-Aware Retrieval
β
12 Metadata-Aware Retrieval
β
13 Advanced Query Rewriting
β
03 LlamaIndex Retrieval Engineering
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.