18 — Building Your First RAG Pipeline¶
Build an end-to-end Retrieval-Augmented Generation (RAG) application by connecting document ingestion, chunking, embeddings, vector storage, retrieval, context construction, prompt engineering, and LLM generation into a working pipeline.
📖 Overview¶
The previous chapters introduced the individual components of a RAG system.
This chapter brings those components together into a complete working pipeline.
We will build the following flow:
Documents
↓
Document Loading
↓
Text Processing
↓
Chunking
↓
Embedding
↓
Vector Store
↓
Retriever
↓
User Query
↓
Relevant Documents
↓
Context Assembly
↓
Prompt
↓
LLM
↓
Grounded Answer
The goal is not simply to call a framework API.
The goal is to understand what happens at every stage so that the same architecture can later be implemented using:
1. What We Are Building¶
We will build a simple enterprise knowledge assistant.
Example knowledge source:
Example user question:
The system should:
1. Load the handbook.
2. Split it into chunks.
3. Generate embeddings.
4. Store the embeddings.
5. Convert the user question into an embedding.
6. Retrieve relevant chunks.
7. Build a context.
8. Send the context to an LLM.
9. Generate a grounded answer.
2. Final Architecture¶
flowchart TD
A["Enterprise Documents"] --> B["Document Loader"]
B --> C["Document Processor"]
C --> D["Text Chunker"]
D --> E["Embedding Provider"]
E --> F["Vector Store"]
G["User"] --> H["User Query"]
H --> I["Query Embedding"]
I --> J["Retriever"]
F --> J
J --> K["Retrieved Chunks"]
K --> L["Context Builder"]
L --> M["Prompt Builder"]
M --> N["LLM Provider"]
N --> O["Grounded Answer"]
This architecture contains two major paths:
and:
3. Indexing Path¶
The indexing path prepares enterprise knowledge.
This normally happens before users start asking questions.
4. Query Path¶
The query path executes when a user asks a question.
The two paths meet at the vector store.
5. Complete RAG Lifecycle¶
flowchart LR
A["Documents"] --> B["Indexing Pipeline"]
B --> C["Vector Store"]
D["User Query"] --> E["Query Pipeline"]
E --> C
C --> F["Retrieved Evidence"]
F --> G["Context"]
G --> H["LLM"]
H --> I["Answer"]
6. Project Structure¶
A simple Python implementation can use:
rag-demo/
│
├── data/
│ └── employee-handbook.txt
│
├── src/
│ ├── ingestion.py
│ ├── embeddings.py
│ ├── vector_store.py
│ ├── retriever.py
│ ├── prompt.py
│ ├── generation.py
│ └── rag_pipeline.py
│
├── tests/
│ └── test_rag_pipeline.py
│
├── requirements.txt
└── README.md
For a small learning project, everything can initially be implemented in one notebook.
For production, separate capabilities are preferable.
7. Step 1 — Prepare the Documents¶
Example document:
Company Employee Handbook
Annual Leave
Employees are entitled to 25 days of annual paid leave
per calendar year.
Employees should submit annual leave requests through
the employee portal.
Unused annual leave may be carried forward according
to company policy.
Sick Leave
Employees may take sick leave when they are unable
to work due to illness.
The document represents the knowledge source.
8. Step 2 — Load the Document¶
A simple loader:
from pathlib import Path
def load_document(path: str) -> str:
return Path(path).read_text(
encoding="utf-8"
)
document = load_document(
"data/employee-handbook.txt"
)
print(document)
The output is raw text.
9. Production Document Loading¶
Real enterprise applications rarely contain only .txt files.
Common sources include:
A production ingestion layer should hide these differences.
Possible implementations:
PdfDocumentLoader
DocxDocumentLoader
HtmlDocumentLoader
MarkdownDocumentLoader
DatabaseDocumentLoader
10. Step 3 — Normalize the Document¶
Raw documents may contain:
Extra whitespace
Headers
Footers
Page numbers
Encoding problems
Repeated content
Formatting artifacts
A simple normalization function:
def normalize_text(text: str) -> str:
lines = [
line.strip()
for line in text.splitlines()
]
lines = [
line
for line in lines
if line
]
return "\n".join(lines)
For production documents, normalization should be document-type aware.
11. Step 4 — Split the Document¶
The complete document may be too large to retrieve as one unit.
Therefore:
Example:
Chunk 1:
Company Employee Handbook
Chunk 2:
Annual Leave
Employees are entitled to 25 days...
Chunk 3:
Employees should submit annual leave requests...
Chunk 4:
Sick Leave
Employees may take sick leave...
12. Simple Chunking¶
A basic character-based chunker:
def chunk_text(
text: str,
chunk_size: int = 500
):
return [
text[i:i + chunk_size]
for i in range(
0,
len(text),
chunk_size
)
]
This is useful for understanding the pipeline.
However, production systems should generally use more meaningful boundaries.
13. Chunking with Overlap¶
A simple overlapping chunker:
def chunk_text(
text: str,
chunk_size: int = 500,
overlap: int = 50
):
chunks = []
start = 0
while start < len(text):
end = start + chunk_size
chunks.append(
text[start:end]
)
start += chunk_size - overlap
return chunks
The overlap helps preserve context across chunk boundaries.
14. Why Chunk Overlap Matters¶
Suppose a sentence crosses the boundary:
Without overlap, retrieval may lose the complete meaning.
With overlap:
Chunk 1:
Employees may carry forward unused annual leave...
Chunk 2:
...unused annual leave according to company policy.
Overlap can help preserve semantic continuity.
15. Metadata¶
Each chunk should carry metadata.
chunks = [
{
"id": "handbook-001",
"text": "...",
"metadata": {
"document_id": "employee-handbook",
"source": "employee-handbook.txt",
"section": "Annual Leave"
}
}
]
Metadata becomes important for:
16. Step 5 — Generate Embeddings¶
The next stage converts chunks into vectors.
Conceptually:
For multiple chunks:
17. Example Embedding Provider Interface¶
A framework-independent interface:
class EmbeddingProvider:
def embed_documents(
self,
texts: list[str]
) -> list[list[float]]:
raise NotImplementedError
def embed_query(
self,
text: str
) -> list[float]:
raise NotImplementedError
This keeps the RAG application independent of a specific embedding provider.
18. Embedding Provider Implementation¶
A simplified example:
class MockEmbeddingProvider(
EmbeddingProvider
):
def embed_documents(self, texts):
return [
self._embed(text)
for text in texts
]
def embed_query(self, text):
return self._embed(text)
def _embed(self, text):
# Demonstration only.
# Production systems should use
# a real embedding model.
return [0.1, 0.2, 0.3]
This demonstrates the architecture without coupling the example to a particular model provider.
19. Real Embedding Models¶
A production implementation can use:
OpenAI Embeddings
Hugging Face Models
Sentence Transformers
WatsonX Embeddings
Cloud Provider Embeddings
Self-Hosted Embedding Models
The application should interact through:
rather than embedding-provider-specific code everywhere.
20. Step 6 — Store the Vectors¶
The chunks and vectors now need to be stored.
A simple interface:
class VectorStore:
def upsert(
self,
records
):
raise NotImplementedError
def search(
self,
query_vector,
top_k=5
):
raise NotImplementedError
21. In-Memory Vector Store¶
For learning purposes, we can create a simple vector store.
import math
class InMemoryVectorStore:
def __init__(self):
self.records = []
def upsert(self, records):
self.records.extend(records)
def search(
self,
query_vector,
top_k=5
):
scored = []
for record in self.records:
score = cosine_similarity(
query_vector,
record["vector"]
)
scored.append(
(score, record)
)
scored.sort(
key=lambda item: item[0],
reverse=True
)
return scored[:top_k]
This is a teaching implementation, not a production vector database.
22. Cosine Similarity¶
For two vectors:
cosine similarity measures how closely their directions align.
def cosine_similarity(a, b):
dot = sum(
x * y
for x, y in zip(a, b)
)
magnitude_a = math.sqrt(
sum(x * x for x in a)
)
magnitude_b = math.sqrt(
sum(x * x for x in b)
)
if magnitude_a == 0 or magnitude_b == 0:
return 0.0
return dot / (
magnitude_a * magnitude_b
)
In a production system, this computation is normally handled by the vector database or search engine.
23. Step 7 — Build the Index¶
The indexing pipeline now becomes:
documents = load_documents()
chunks = chunk_documents(
documents
)
vectors = embedding_provider.embed_documents(
[
chunk["text"]
for chunk in chunks
]
)
records = []
for chunk, vector in zip(
chunks,
vectors
):
records.append(
{
"id": chunk["id"],
"text": chunk["text"],
"vector": vector,
"metadata": chunk["metadata"]
}
)
vector_store.upsert(
records
)
The knowledge base is now searchable.
24. Complete Indexing Pipeline¶
flowchart TD
A["Source Documents"] --> B["Document Loader"]
B --> C["Document Processor"]
C --> D["Chunker"]
D --> E["Metadata Enrichment"]
E --> F["Embedding Provider"]
F --> G["Vector Store"]
G --> H["Searchable Knowledge Base"]
25. Step 8 — Receive a User Query¶
Now suppose the user asks:
The query enters the runtime pipeline.
26. Step 9 — Embed the Query¶
The query must be converted into the same embedding space as the documents.
query = (
"What is the annual leave entitlement?"
)
query_vector = (
embedding_provider.embed_query(
query
)
)
The resulting vector is used for similarity search.
27. Step 10 — Retrieve Relevant Chunks¶
Conceptually:
28. Retrieval Results¶
A result might look like:
[
(
0.91,
{
"id": "handbook-002",
"text": (
"Employees are entitled to "
"25 days of annual paid leave."
),
"metadata": {
"section": "Annual Leave",
"page": 42
}
}
)
]
The similarity score is useful for:
29. Step 11 — Build the Context¶
The retrieved chunks need to be converted into context.
Example output:
Employees are entitled to 25 days
of annual paid leave.
Employees should submit annual leave
requests through the employee portal.
30. Preserve Source Metadata¶
Do not throw away metadata during context construction.
Instead:
def build_context(results):
sections = []
for score, record in results:
sections.append(
f"""
Source: {record["metadata"]["source"]}
Section: {record["metadata"]["section"]}
{record["text"]}
"""
)
return "\n\n".join(sections)
This makes citations and debugging easier.
31. Step 12 — Build the Prompt¶
A grounded prompt can be:
RAG_PROMPT = """
You are an enterprise knowledge assistant.
Answer the user's question using only
the provided context.
If the context does not contain enough
information, say that the information
is not available.
Do not invent facts.
Context:
{context}
Question:
{question}
Answer:
"""
32. Construct the Prompt¶
The resulting prompt might look like:
You are an enterprise knowledge assistant.
Answer the user's question using only
the provided context.
Context:
Employees are entitled to 25 days
of annual paid leave.
Question:
What is the annual leave entitlement?
Answer:
33. Step 13 — Invoke the LLM¶
The LLM receives the grounded prompt.
Conceptually:
34. LLM Provider Interface¶
A framework-independent interface:
Possible implementations:
35. Example LLM Provider¶
A simplified implementation:
class MockLLMProvider(
LLMProvider
):
def generate(self, prompt):
return (
"Employees are entitled to "
"25 days of annual paid leave."
)
This allows the complete pipeline to be tested without requiring an external model.
36. Step 14 — Return the Answer¶
The final response can include:
Employees are entitled to 25 days
of annual paid leave.
Source:
Employee Handbook
Section: Annual Leave
This is a grounded response because the answer comes from retrieved enterprise knowledge.
37. Complete Minimal Pipeline¶
def rag_pipeline(
query,
embedding_provider,
vector_store,
llm_provider
):
# 1. Embed query
query_vector = (
embedding_provider.embed_query(
query
)
)
# 2. Retrieve
results = vector_store.search(
query_vector,
top_k=5
)
# 3. Build context
context = "\n\n".join(
record["text"]
for score, record in results
)
# 4. Build prompt
prompt = f"""
You are an enterprise knowledge assistant.
Answer using only the provided context.
Context:
{context}
Question:
{query}
Answer:
"""
# 5. Generate
return llm_provider.generate(
prompt
)
This is the simplest representation of the runtime RAG pipeline.
38. Complete End-to-End Flow¶
flowchart TD
A["Documents"] --> B["Load"]
B --> C["Process"]
C --> D["Chunk"]
D --> E["Embed"]
E --> F["Vector Store"]
G["User Question"] --> H["Query Embedding"]
H --> I["Similarity Search"]
F --> I
I --> J["Top-K Documents"]
J --> K["Context Builder"]
K --> L["Grounded Prompt"]
L --> M["LLM"]
M --> N["Answer"]
39. Adding Metadata Filters¶
Enterprise retrieval usually requires filters.
Example:
results = vector_store.search(
query_vector=query_vector,
top_k=5,
filters={
"country": "IN",
"department": "HR",
"version": "2026"
}
)
The final search becomes:
40. Adding Similarity Thresholds¶
A production retriever may reject weak results.
The value is illustrative.
It should be determined through evaluation.
41. Handling No Results¶
Never assume retrieval will always succeed.
if not results:
return (
"I couldn't find sufficient information "
"in the available knowledge base."
)
This is safer than allowing the LLM to answer from unsupported knowledge.
42. Handling Low-Quality Results¶
Even if results exist, they may not be relevant.
relevant_results = [
record
for score, record in results
if score >= 0.72
]
if not relevant_results:
return (
"I couldn't find sufficient relevant "
"information in the knowledge base."
)
Again, the threshold is application-specific.
43. Adding Citations¶
A stronger implementation preserves source metadata.
def build_context(results):
context_parts = []
for score, record in results:
metadata = record["metadata"]
context_parts.append(
f"""
Source: {metadata["source"]}
Page: {metadata.get("page", "N/A")}
Section: {metadata.get("section", "N/A")}
{record["text"]}
"""
)
return "\n\n".join(
context_parts
)
44. Citation-Aware Response¶
The LLM can be instructed to reference source identifiers.
Use the provided source information
when answering.
Do not create sources that are not present
in the context.
The application should still maintain source metadata independently of the model.
45. Production RAG Pipeline¶
The minimal pipeline can now evolve into:
User
↓
Authentication
↓
Query Validation
↓
Query Processing
↓
Query Embedding
↓
Retriever
↓
Authorization Filtering
↓
Metadata Filtering
↓
Top-K
↓
Similarity Threshold
↓
Optional Reranking
↓
Deduplication
↓
Context Builder
↓
Token Budget
↓
Prompt Builder
↓
LLM
↓
Response Validation
↓
Citation Builder
↓
Answer
46. Production Architecture¶
flowchart TD
A["Client"] --> B["API Gateway"]
B --> C["Authentication"]
C --> D["RAG Application"]
D --> E["Query Processor"]
E --> F["Embedding Provider"]
F --> G["Retriever"]
H["Authorization Context"] --> G
I["Metadata Filters"] --> G
G --> J["Vector Store"]
J --> K["Vector Database"]
G --> L["Retrieved Candidates"]
L --> M["Ranking / Reranking"]
M --> N["Deduplication"]
N --> O["Context Builder"]
O --> P["Prompt Builder"]
P --> Q["LLM Provider"]
Q --> R["LLM"]
R --> S["Response Validator"]
S --> T["Citation Builder"]
T --> U["Final Answer"]
D --> V["Observability"]
G --> V
Q --> V
47. Framework Implementation¶
The same architecture can be implemented using a framework.
For example:
or:
The framework should simplify implementation without hiding the architectural concepts.
48. LangChain Example¶
A simplified LangChain-style implementation:
from langchain_core.prompts import ChatPromptTemplate
prompt = ChatPromptTemplate.from_template(
"""
Answer using only the provided context.
Context:
{context}
Question:
{question}
Answer:
"""
)
question = (
"What is the annual leave entitlement?"
)
documents = retriever.invoke(
question
)
context = "\n\n".join(
document.page_content
for document in documents
)
response = llm.invoke(
prompt.format_messages(
context=context,
question=question
)
)
print(response.content)
The important architectural flow remains:
49. LlamaIndex Example¶
A simplified LlamaIndex implementation:
from llama_index.core import (
VectorStoreIndex
)
index = VectorStoreIndex.from_documents(
documents
)
query_engine = index.as_query_engine(
similarity_top_k=5
)
response = query_engine.query(
"What is the annual leave entitlement?"
)
print(response)
The query engine hides several underlying operations.
Conceptually:
50. Custom Implementation vs Framework¶
Custom Implementation¶
Framework¶
For enterprise engineering, understanding both levels is valuable.
51. Java-First Enterprise RAG¶
A production backend can expose the RAG pipeline through Spring Boot.
Client
↓
Spring Boot REST API
↓
RagService
↓
Retriever
↓
VectorStore
↓
ContextBuilder
↓
PromptBuilder
↓
LLMProvider
↓
Answer
52. Spring Boot Controller¶
@RestController
@RequestMapping("/api/rag")
public class RagController {
private final RagService ragService;
public RagController(
RagService ragService
) {
this.ragService = ragService;
}
@PostMapping("/query")
public AnswerResponse query(
@RequestBody QueryRequest request
) {
return ragService.answer(
request.question()
);
}
}
The controller should remain thin.
Business orchestration belongs in the application/service layer.
53. RAG Service¶
@Service
public class RagService {
private final Retriever retriever;
private final ContextBuilder contextBuilder;
private final PromptBuilder promptBuilder;
private final LLMProvider llmProvider;
public AnswerResponse answer(
String question
) {
var documents =
retriever.retrieve(
question
);
var context =
contextBuilder.build(
documents
);
var prompt =
promptBuilder.build(
question,
context
);
var result =
llmProvider.generate(
prompt
);
return AnswerResponse.from(
result,
documents
);
}
}
54. Capability Interfaces¶
The application can use capability-based interfaces:
public interface EmbeddingProvider {
List<float[]> embedDocuments(
List<String> documents
);
float[] embedQuery(
String query
);
}
public interface VectorStore {
void upsert(
List<VectorRecord> records
);
List<RetrievedDocument> search(
float[] query,
SearchOptions options
);
}
55. Ports & Adapters Architecture¶
flowchart TD
A["REST API"] --> B["RAG Application Service"]
B --> C["Retriever"]
B --> D["ContextBuilder"]
B --> E["PromptBuilder"]
B --> F["LLMProvider"]
C --> G["EmbeddingProvider"]
C --> H["VectorStore"]
G --> I["Embedding Adapter"]
H --> J["Vector Database Adapter"]
F --> K["LLM Adapter"]
The core application does not need to know whether the infrastructure uses:
56. Configuration¶
Provider selection can be configuration-driven.
rag:
embedding:
provider: huggingface
vector-store:
provider: qdrant
llm:
provider: openai
retrieval:
top-k: 5
similarity-threshold: 0.72
This makes deployments easier to configure across environments.
57. Development vs Production¶
Development¶
Production¶
Enterprise Sources
↓
Ingestion Pipeline
↓
Enterprise Embedding Service
↓
Managed / Clustered Vector Database
↓
RAG Service
↓
Enterprise LLM
The conceptual architecture remains the same.
58. Testing Strategy¶
A production RAG pipeline should be tested at multiple levels.
59. Unit Testing¶
Test individual components.
Examples:
Example:
def test_prompt_contains_context():
prompt = build_prompt(
question="What is annual leave?",
context="Employees receive 25 days."
)
assert (
"Employees receive 25 days."
in prompt
)
60. Retrieval Testing¶
Test whether the correct chunks are retrieved.
Example:
def test_annual_leave_retrieval():
results = retriever.retrieve(
"How many annual leave days?"
)
assert any(
"25 days"
in result.content
for result in results
)
A real evaluation suite should use representative datasets rather than only one example.
61. End-to-End Testing¶
An end-to-end test can verify:
Example expected behavior:
The test should ideally validate grounding and source attribution as well.
62. Retrieval Evaluation Dataset¶
Create representative questions:
[
{
"question": "What is the annual leave entitlement?",
"expected_source": "employee-handbook",
"expected_section": "Annual Leave"
},
{
"question": "How should leave be requested?",
"expected_source": "employee-handbook",
"expected_section": "Annual Leave"
}
]
This dataset becomes valuable for regression testing.
63. RAG Regression Testing¶
Whenever you change:
rerun the evaluation dataset.
64. Observability¶
Track each pipeline stage.
Useful metrics include:
Retrieval Latency
LLM Latency
Total Latency
Top-K
Similarity Scores
Context Tokens
Input Tokens
Output Tokens
LLM Cost
Error Rate
Empty Retrieval Rate
65. Example Trace¶
{
"trace_id": "rag-12345",
"query": "What is the annual leave entitlement?",
"retrieval": {
"top_k": 5,
"results": 3,
"latency_ms": 38
},
"context": {
"chunks": 3,
"tokens": 640
},
"generation": {
"input_tokens": 710,
"output_tokens": 90,
"latency_ms": 820
}
}
Avoid logging sensitive document contents unless required and properly protected.
66. Cost Awareness¶
A RAG request can incur costs from:
Reducing unnecessary retrieved context can reduce:
67. Caching¶
Repeated queries may benefit from caching.
Caching must account for:
68. Failure Handling¶
Every stage can fail.
Production systems should define:
where appropriate.
69. Empty Knowledge Base¶
A new RAG application may start with:
The system should return a controlled response rather than attempting generation with empty context.
70. Security¶
Enterprise RAG must enforce security before generation.
Never rely on the LLM to decide whether a user is allowed to see a document.
71. Multi-Tenant RAG¶
A multi-tenant architecture may include:
in every retrieval request.
This prevents cross-tenant retrieval.
72. Document Versioning¶
If the handbook changes:
the vector index must represent the correct version.
Metadata can include:
73. Updating the Knowledge Base¶
A document update should trigger:
flowchart LR
A["Document Updated"] --> B["Change Detection"]
B --> C["Reprocess"]
C --> D["Rechunk"]
D --> E["Re-embed"]
E --> F["Upsert"]
F --> G["Updated Vector Index"]
This prevents stale enterprise knowledge from remaining indefinitely searchable.
74. RAG Pipeline with Knowledge Updates¶
┌───────────────┐
│ Source System │
└───────┬───────┘
↓
Document Update
↓
Ingestion Pipeline
↓
Vector Store
↓
RAG Query
↓
LLM
The ingestion pipeline and query pipeline are independent but connected through the knowledge index.
75. Common Beginner Mistakes¶
75.1 Sending the Entire Document to the LLM¶
This can increase:
Use retrieval instead.
75.2 Using Poor Chunking¶
Bad chunk boundaries can reduce retrieval quality.
75.3 No Metadata¶
Without metadata:
become harder.
75.4 No Retrieval Threshold¶
The system may pass weakly relevant results to the LLM.
75.5 No Empty-Result Handling¶
The LLM may generate an unsupported answer.
75.6 Mixing Embedding Models¶
Document and query vectors must belong to compatible embedding spaces.
75.7 Treating the Framework as the Architecture¶
Knowing:
is not the same as understanding:
76. Improving the First RAG Pipeline¶
The first implementation should intentionally remain simple.
Then evolve it incrementally.
Version 2
+ Metadata
Version 3
+ Filtering
Version 4
+ Citations
Version 5
+ Evaluation
Version 6
+ Observability
Version 7
+ Reranking
Version 8
+ Production Security
This is a better learning strategy than starting with a highly complex architecture.
77. RAG Evolution¶
flowchart LR
A["Basic RAG"] --> B["Metadata"]
B --> C["Filtering"]
C --> D["Citations"]
D --> E["Evaluation"]
E --> F["Observability"]
F --> G["Reranking"]
G --> H["Production RAG"]
Advanced retrieval techniques are covered in later chapters.
78. Complete Python Reference Architecture¶
class RagPipeline:
def __init__(
self,
embedding_provider,
retriever,
prompt_builder,
llm_provider
):
self.embedding_provider = (
embedding_provider
)
self.retriever = retriever
self.prompt_builder = (
prompt_builder
)
self.llm_provider = (
llm_provider
)
def answer(self, question):
documents = (
self.retriever.retrieve(
question
)
)
if not documents:
return (
"I couldn't find sufficient "
"information in the knowledge base."
)
context = self._build_context(
documents
)
prompt = (
self.prompt_builder.build(
question,
context
)
)
return self.llm_provider.generate(
prompt
)
def _build_context(
self,
documents
):
return "\n\n".join(
document.content
for document in documents
)
The purpose is to demonstrate clean separation of responsibilities.
79. Production-Oriented Python Architecture¶
rag/
│
├── domain/
│ ├── document.py
│ ├── query.py
│ ├── retrieval.py
│ └── answer.py
│
├── application/
│ └── rag_service.py
│
├── ports/
│ ├── embedding_provider.py
│ ├── vector_store.py
│ ├── retriever.py
│ └── llm_provider.py
│
├── adapters/
│ ├── embeddings/
│ ├── vectorstores/
│ └── llm/
│
└── infrastructure/
├── configuration.py
└── observability.py
This structure is closer to an enterprise architecture.
80. Production-Oriented Java Architecture¶
A Java-first implementation could use:
rag/
│
├── domain/
│ ├── model/
│ └── service/
│
├── application/
│ ├── RagService.java
│ └── RetrievalService.java
│
├── ports/
│ ├── EmbeddingProvider.java
│ ├── VectorStore.java
│ ├── Retriever.java
│ ├── LLMProvider.java
│ └── PromptBuilder.java
│
├── adapters/
│ ├── embedding/
│ ├── vectorstore/
│ └── llm/
│
└── api/
└── RagController.java
This follows the Ports & Adapters approach.
81. RAG Pipeline Contract¶
The complete pipeline can be expressed as:
and at runtime:
These explicit contracts make systems easier to test and evolve.
82. RAG Data Flow¶
flowchart LR
A["Document"] --> B["ProcessedDocument"]
B --> C["DocumentChunk"]
C --> D["VectorRecord"]
D --> E["Vector Index"]
F["User Query"] --> G["QueryVector"]
G --> H["RetrievedDocument"]
E --> H
H --> I["RetrievalContext"]
I --> J["Prompt"]
J --> K["GenerationResult"]
K --> L["Answer"]
83. What We Have Built¶
At the end of this chapter, we have a conceptual end-to-end pipeline:
INDEXING
Documents
↓
Loader
↓
Processor
↓
Chunker
↓
Embedding Provider
↓
Vector Store
QUERY
User Question
↓
Query Embedding
↓
Retriever
↓
Relevant Chunks
↓
Context Builder
↓
Prompt Builder
↓
LLM Provider
↓
Answer
This is the foundation for production RAG systems.
84. Production Readiness Checklist¶
[ ] Document ingestion implemented
[ ] Document normalization implemented
[ ] Chunking strategy defined
[ ] Chunk metadata preserved
[ ] Embedding provider abstracted
[ ] Vector store abstracted
[ ] Query embedding implemented
[ ] Retrieval implemented
[ ] Metadata filtering implemented
[ ] Authorization filtering implemented
[ ] Top-K configured
[ ] Similarity threshold evaluated
[ ] Context builder implemented
[ ] Token budget considered
[ ] Grounded prompt implemented
[ ] LLM provider abstracted
[ ] Empty retrieval handled
[ ] Citations preserved
[ ] Response validation implemented
[ ] Error handling implemented
[ ] Observability implemented
[ ] Retrieval evaluation implemented
[ ] Regression tests implemented
[ ] Document versioning considered
[ ] Knowledge update strategy defined
[ ] Security controls implemented
85. Key Takeaways¶
- A RAG pipeline combines document indexing and query-time retrieval.
- The indexing pipeline prepares knowledge before user queries arrive.
- The query pipeline retrieves knowledge and uses it to ground generation.
- Documents should be processed before chunking.
- Chunking creates manageable retrieval units.
- Metadata should be preserved throughout the pipeline.
- Embeddings convert chunks into vectors.
- Query embeddings must be compatible with document embeddings.
- Vector stores provide searchable storage for embeddings.
- Retrieval should return structured results containing content, scores, and metadata.
- Context builders should select, order, deduplicate, and format retrieved evidence.
- Prompt builders should clearly separate instructions from retrieved data.
- LLMs should generate answers from retrieved evidence rather than being treated as the enterprise source of truth.
- Empty and low-confidence retrieval must be handled explicitly.
- Citations should originate from source metadata.
- Production RAG requires authorization-aware retrieval.
- Multi-tenant systems must enforce tenant isolation.
- Document updates should trigger appropriate re-indexing.
- Embedding and chunking changes should be evaluated before production rollout.
- RAG should be evaluated at both retrieval and generation levels.
- Observability should cover the complete request path.
- Frameworks such as LangChain and LlamaIndex simplify implementation, but the underlying architecture remains the same.
- Capability-based interfaces such as
EmbeddingProvider,VectorStore,Retriever, andLLMProvidersupport provider independence. - A Java-first implementation can expose the pipeline through Spring Boot while keeping AI infrastructure behind adapters.
- The first RAG implementation should remain simple and evolve incrementally toward production readiness.
The central principle is:
A RAG application is an end-to-end knowledge pipeline: ingest authoritative information, make it searchable, retrieve the right evidence, and use that evidence to ground LLM generation.
86. Chapter Navigation¶
Part IV — Prompt Engineering & RAG Fundamentals¶
Previous Chapter: 17. Vector Databases in RAG
Current Chapter: 18 — Building Your First RAG Pipeline
Next Chapter: 19. RAG Evaluation Fundamentals
Part IV Chapters¶
- 01. Introduction to Prompt Engineering
- 02. Prompt Engineering Fundamentals
- 03. Advanced Prompt Engineering
- 04. Prompt Design Patterns
- 05. Zero-shot, One-shot & Few-shot Prompting
- 06. Chain-of-Thought Prompting
- 07. ReAct Prompting
- 08. Structured Outputs & Output Parsing
- 09. Function Calling & Tool Calling
- 10. Embeddings in Practice
- 11. Document Processing & Vectorization
- 12. Document Chunking Strategies
- 13. Vector Database Fundamentals
- 14. Similarity Search Techniques
- 15. RAG Pipeline Components
- 16. Retrieval and Generation Pipeline
- 17. Vector Databases in RAG
- 18. Building Your First RAG Pipeline
- 19. RAG Evaluation Fundamentals
- 20. Enterprise Generative AI Application Architecture
- 21. Deploying AI Applications with Gradio
References¶
- Retrieval-Augmented Generation architecture documentation
- LangChain documentation
- LlamaIndex documentation
- Hugging Face documentation
- Embedding model documentation
- Vector database documentation
- RAG retrieval documentation
- RAG evaluation documentation
- Enterprise search architecture documentation
- Spring Boot documentation
- OpenTelemetry documentation
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems — One Chapter at a Time.