12. RAG Deployment PatternsΒΆ
Category: Production RAG Engineering
Module: Part VI β Production Deployment
Difficulty: Advanced
π OverviewΒΆ
Deploying a RAG system is not simply a matter of running an API and connecting it to a vector database.
A production RAG deployment must account for:
Application Deployment
β
Retrieval Deployment
β
Index Deployment
β
Model Deployment
β
Knowledge Deployment
β
Configuration Deployment
β
Observability
β
Security
β
Scalability
β
Rollback
Different workloads require different deployment patterns.
For example:
Small Internal RAG
β Single Service
Enterprise RAG
β Microservices
High-Traffic RAG
β Horizontally Scaled Services
High-Risk RAG
β Canary / Blue-Green
Global RAG
β Multi-Region
Frequently Changing RAG
β Independent Index Deployment
The goal is not to choose the most complex deployment pattern.
The goal is to choose the simplest deployment architecture that satisfies the required quality, availability, latency, security, scalability, and cost objectives.
π― Learning ObjectivesΒΆ
After completing this chapter, you will be able to:
- Understand major RAG deployment patterns
- Design monolithic RAG deployments
- Design modular RAG deployments
- Design microservice-based RAG platforms
- Understand serverless RAG deployment
- Deploy RAG using containers
- Deploy RAG using Kubernetes
- Design blue-green deployments
- Design rolling deployments
- Design canary deployments
- Design shadow deployments
- Design A/B deployments
- Design multi-region RAG
- Design active-active architectures
- Design active-passive architectures
- Deploy retrieval and generation independently
- Deploy indexes independently from application code
- Design model rollout strategies
- Design embedding migration strategies
- Design zero-downtime RAG deployments
- Design rollback mechanisms
- Design disaster recovery
- Design deployment pipelines
- Build CI/CD quality gates
- Design environment strategies
- Understand infrastructure-as-code for RAG
- Design deployment observability
- Choose appropriate deployment patterns based on workload requirements
π§ 1. Why RAG Deployment Is DifferentΒΆ
Traditional backend deployment often looks like:
RAG introduces additional deployable components:
Application
Retriever
Embedding Model
Vector Index
Keyword Index
Reranker
Prompt
LLM
Knowledge Base
Configuration
Evaluation Dataset
Therefore:
It is a coordinated deployment of multiple versioned artifacts.
π§ 2. RAG Deployment SurfaceΒΆ
flowchart TD
A["RAG System"] --> B["Application"]
A --> C["Retrieval"]
A --> D["Indexes"]
A --> E["Models"]
A --> F["Prompts"]
A --> G["Knowledge"]
A --> H["Configuration"]
B --> I["Deployment"]
C --> I
D --> I
E --> I
F --> I
G --> I
H --> I π§ 3. Version EverythingΒΆ
A production RAG deployment should identify:
Application Version
Retriever Version
Embedding Version
Index Version
Reranker Version
Prompt Version
LLM Version
Configuration Version
Knowledge Version
Example:
{
"application": "v12",
"retriever": "v8",
"embedding": "v4",
"index": "v17",
"reranker": "v3",
"prompt": "v9",
"model": "model-x",
"configuration": "v11"
}
This enables:
π§ 4. Deployment UnitsΒΆ
A mature RAG platform may have:
API Service
Query Service
Retrieval Service
Reranking Service
Generation Service
Ingestion Worker
Indexing Worker
Evaluation Service
Not every system needs all of them as separate services.
π§ 5. Deployment GranularityΒΆ
There are several options:
Choose based on:
π§ 6. Pattern 1 β Monolithic RAGΒΆ
The simplest deployment:
RAG Application
β
βββββββββββββββββΌββββββββββββββββ
βΌ βΌ βΌ
Retrieval Context LLM
β β β
βββββββββββββββββΌββββββββββββββββ
βΌ
Response
Everything runs inside one application.
π§ 7. Monolithic DeploymentΒΆ
Docker Container
β
βββ API
βββ Retrieval
βββ Prompt
βββ Validation
βββ Generation
External:
π§ 8. Advantages of Monolithic RAGΒΆ
Simple
Easy to Develop
Easy to Debug
Low Operational Overhead
Low Network Overhead
Easy Local Deployment
π§ 9. LimitationsΒΆ
Independent Scaling is Difficult
Large Deployment Unit
Higher Blast Radius
Retrieval and Generation Coupled
Harder Provider Isolation
π§ 10. When to UseΒΆ
Good for:
Avoid unnecessary microservices at this stage.
π§ 11. Pattern 2 β Modular MonolithΒΆ
A stronger intermediate architecture:
RAG Application
β
βββββββββββββββββΌββββββββββββββββ
βΌ βΌ βΌ
Query Module Retrieval Module Generation
β β β
βββββββββββββββββΌββββββββββββββββ
βΌ
Response
The modules have explicit interfaces but run in one process.
π§ 12. Why Modular Monolith?ΒΆ
It provides:
without immediately introducing distributed-system complexity.
π§ 13. Pattern 3 β Microservice RAGΒΆ
At larger scale:
flowchart TD
A["API Gateway"] --> B["RAG Orchestrator"]
B --> C["Query Service"]
B --> D["Retrieval Service"]
B --> E["Generation Service"]
D --> F["Vector Search"]
D --> G["Keyword Search"]
E --> H["Model Gateway"]
I["Ingestion Service"] --> J["Indexing"] π§ 14. Microservice BoundariesΒΆ
Potential services:
Query Service
Retrieval Service
Reranking Service
Context Service
Generation Service
Ingestion Service
Indexing Service
Evaluation Service
But do not split services purely because components exist.
π§ 15. Independent ScalingΒΆ
One of the strongest reasons for service separation:
They have very different scaling characteristics.
π§ 16. Microservice Trade-OffΒΆ
Benefits:
Costs:
Network Latency
Distributed Tracing
Operational Complexity
Service Discovery
Failure Handling
Deployment Complexity
π§ 17. Pattern 4 β Retrieval as a PlatformΒΆ
Multiple applications can consume a shared retrieval platform.
Retrieval Platform
β
βββββββββββββΌββββββββββββ
βΌ βΌ βΌ
Chat Copilot Agent
β β β
βββββββββββββΌββββββββββββ
βΌ
Evidence
π§ 18. Why Retrieval Platform?ΒΆ
Centralize:
Applications focus on business capabilities.
π§ 19. Pattern 5 β Serverless RAGΒΆ
A serverless deployment may use:
Ingestion:
π§ 20. Serverless AdvantagesΒΆ
No Server Management
Automatic Scaling
Pay Per Use
Fast Initial Deployment
Good for Variable Traffic
π§ 21. Serverless LimitationsΒΆ
Potential issues:
Cold Starts
Execution Limits
Concurrency Limits
Long-Running Processing
Connection Management
Vendor Coupling
π§ 22. Good Serverless WorkloadsΒΆ
Low / Variable Traffic
Event-Driven Ingestion
Document Processing
Lightweight APIs
Scheduled Evaluation
π§ 23. Pattern 6 β Containerized RAGΒΆ
A common production pattern:
Possible platforms:
π§ 24. Container ArchitectureΒΆ
flowchart LR
A["Container Registry"] --> B["Deployment Platform"]
B --> C["RAG API"]
B --> D["Retrieval Worker"]
B --> E["Ingestion Worker"]
B --> F["Evaluation Worker"]
C --> G["Vector DB"]
D --> G π§ 25. Why Containers?ΒΆ
Containers provide:
π§ 26. Pattern 7 β Kubernetes RAGΒΆ
For complex enterprise environments:
Kubernetes Cluster
β
βββ RAG API
βββ Retrieval Service
βββ Reranker
βββ Ingestion Workers
βββ Indexing Workers
βββ Evaluation Workers
External:
π§ 27. Kubernetes ScalingΒΆ
Load Balancer
β
βββββββββββΌββββββββββ
βΌ βΌ βΌ
Pod 1 Pod 2 Pod 3
β β β
βββββββββββΌββββββββββ
βΌ
Retrieval
Autoscaling can respond to:
π§ 28. Kubernetes AdvantagesΒΆ
Horizontal Scaling
Self-Healing
Rolling Deployments
Service Discovery
Resource Isolation
Declarative Configuration
π§ 29. Kubernetes ChallengesΒΆ
Do not use Kubernetes merely because it is available.
π§ 30. Pattern 8 β Rolling DeploymentΒΆ
Replace instances gradually.
Version 1:
Pod A
Pod B
Pod C
Pod D
β
Update A
Pod A = V2
Pod B = V1
Pod C = V1
Pod D = V1
β
Update B
Pod A = V2
Pod B = V2
Pod C = V1
Pod D = V1
Eventually:
π§ 31. Rolling Deployment AdvantagesΒΆ
π§ 32. Rolling Deployment RiskΒΆ
During rollout:
may run simultaneously.
Therefore:
must remain compatible.
π§ 33. Pattern 9 β Blue-Green DeploymentΒΆ
Maintain two environments:
Traffic initially:
After validation:
π§ 34. Blue-Green AdvantagesΒΆ
π§ 35. Blue-Green RAGΒΆ
Blue and Green may contain:
But index deployment requires additional planning.
π§ 36. Index Blue-GreenΒΆ
This allows index rollback independently from application rollback.
π§ 37. Pattern 10 β Canary DeploymentΒΆ
Send a small percentage of traffic to the new version.
Monitor:
π§ 38. Canary for RAGΒΆ
Canary changes can include:
π§ 39. Quality-Aware CanaryΒΆ
Traditional canary:
RAG canary should additionally monitor:
π§ 40. Canary PromotionΒΆ
Promotion should stop if quality or operational metrics degrade.
π§ 41. Pattern 11 β Shadow DeploymentΒΆ
The new version receives copied traffic but does not affect the user response.
flowchart LR
A["Production Request"] --> B["Current RAG"]
A --> C["Shadow RAG"]
B --> D["User Response"]
C --> E["Evaluation Only"] π§ 42. Shadow Deployment BenefitsΒΆ
Useful for testing:
against real production queries.
π§ 43. Shadow Deployment RiskΒΆ
Shadow systems still consume:
Therefore cost must be controlled.
π§ 44. Pattern 12 β A/B DeploymentΒΆ
Different users receive different versions.
Compare:
π§ 45. A/B Testing in RAGΒΆ
Possible experiments:
Prompt A vs Prompt B
Retriever A vs Retriever B
Chunking A vs Chunking B
Reranker A vs Reranker B
Model A vs Model B
π§ 46. Pattern 13 β Multi-RegionΒΆ
Global systems may deploy RAG into multiple regions.
Global Router
β
ββββββββββββ΄βββββββββββ
βΌ βΌ
Region A Region B
β β
RAG Stack RAG Stack
β β
Index A Index B
π§ 47. Why Multi-Region?ΒΆ
Reasons include:
π§ 48. Active-ActiveΒΆ
Both regions serve traffic:
Advantages:
Challenges:
π§ 49. Active-PassiveΒΆ
One region serves traffic.
During failure:
π§ 50. Active-Active vs Active-PassiveΒΆ
| Pattern | Availability | Complexity | Cost |
|---|---|---|---|
| Active-Active | Very High | High | High |
| Active-Passive | High | Medium | Medium |
| Single Region | Lower | Low | Lower |
Choose based on business requirements.
π§ 51. Multi-Region Index StrategyΒΆ
Possible approaches:
or:
or:
Selection depends on:
π§ 52. Pattern 14 β Edge / Regional RetrievalΒΆ
For latency-sensitive workloads:
Useful when:
π§ 53. Pattern 15 β Dedicated Tenant DeploymentΒΆ
For high-value enterprise tenants:
Tenant A
βββ RAG API
βββ Index
βββ Storage
Tenant B
βββ RAG API
βββ Index
βββ Storage
Benefits:
Cost:
π§ 54. Pattern 16 β Shared RAG PlatformΒΆ
Multiple tenants share infrastructure:
RAG Platform
β
βββββββββββββββΌββββββββββββββ
βΌ βΌ βΌ
Tenant A Tenant B Tenant C
Isolation occurs through:
π§ 55. Hybrid Tenant DeploymentΒΆ
A mature enterprise platform can use:
This balances:
π§ 56. Pattern 17 β Independent Index DeploymentΒΆ
Do not necessarily deploy the index together with application code.
The two can evolve independently.
π§ 57. Why Independent Index Deployment?ΒΆ
Useful when:
π§ 58. Index Deployment PipelineΒΆ
flowchart LR
A["Source Data"] --> B["Index Builder"]
B --> C["Evaluation"]
C --> D["Index V18"]
D --> E["Canary"]
E --> F["Production"] π§ 59. Embedding MigrationΒΆ
Changing embeddings can require re-indexing.
Embedding V1
β
New Embedding V2
β
Re-Embed Documents
β
Build New Index
β
Evaluate
β
Deploy
Do not blindly replace the existing index.
π§ 60. Dual-Index MigrationΒΆ
During migration:
Compare:
Then switch traffic.
π§ 61. Pattern 18 β Dual ReadΒΆ
Both systems are queried:
Useful for:
π§ 62. Pattern 19 β Dual WriteΒΆ
During migration:
This helps keep both indexes current.
Use carefully because it increases:
π§ 63. Migration StrategyΒΆ
A safer migration:
Build New
β
Backfill
β
Dual Write
β
Dual Read
β
Compare
β
Canary
β
Switch
β
Retire Old
π§ 64. Pattern 20 β Immutable DeploymentΒΆ
Treat deployment artifacts as immutable:
Do not modify deployed artifacts in place.
π§ 65. Reproducible DeploymentΒΆ
Given:
you should be able to reconstruct the deployment.
π§ 66. Deployment ManifestΒΆ
Example:
application:
version: v12
retriever:
version: v8
embedding:
version: v4
index:
version: v17
reranker:
version: v3
prompt:
version: v9
model:
version: model-x
π§ 67. Deployment MetadataΒΆ
Expose deployment information through:
Example:
π§ 68. Environment StrategyΒΆ
Use:
Each environment should have appropriate:
π§ 69. Development EnvironmentΒΆ
Optimize for:
Possible:
π§ 70. Staging EnvironmentΒΆ
Should resemble production enough to validate:
π§ 71. Production EnvironmentΒΆ
Requires:
π§ 72. Configuration PromotionΒΆ
Do not copy configuration manually.
Use:
π§ 73. Secrets PromotionΒΆ
Never move secrets through Git.
Use:
π§ 74. CI/CD PipelineΒΆ
flowchart LR
A["Git Commit"] --> B["Build"]
B --> C["Unit Tests"]
C --> D["Integration Tests"]
D --> E["RAG Evaluation"]
E --> F["Security Tests"]
F --> G["Performance Tests"]
G --> H["Build Artifact"]
H --> I["Staging"]
I --> J["Canary"]
J --> K["Production"] π§ 75. RAG-Specific Quality GateΒΆ
Traditional deployment:
RAG deployment:
Tests
β
Retrieval Evaluation
β
Generation Evaluation
β
Citation Evaluation
β
Security
β
Performance
β
Deploy
π§ 76. Deployment Gate ExampleΒΆ
Recall@10 >= target
AND
Groundedness >= target
AND
Citation Accuracy >= target
AND
p95 <= target
AND
Cost/request <= target
If any critical condition fails:
π§ 77. Deployment ObservabilityΒΆ
Monitor deployment impact:
π§ 78. Deployment DashboardΒΆ
Track:
π§ 79. Rollback StrategyΒΆ
Rollback may involve:
These should not necessarily be rolled back together.
π§ 80. Application RollbackΒΆ
Simple if deployments are immutable.
π§ 81. Index RollbackΒΆ
Traffic can be switched back.
π§ 82. Prompt RollbackΒΆ
Prompt versioning makes this possible.
π§ 83. Model RollbackΒΆ
Use model gateways where possible to simplify routing.
π§ 84. Partial RollbackΒΆ
A powerful production capability:
The system does not need to roll back every component.
π§ 85. Zero-Downtime DeploymentΒΆ
A production RAG deployment should ideally maintain:
while replacing components.
Use:
depending on risk.
π§ 86. Deployment CompatibilityΒΆ
During rollout:
may coexist.
Therefore ensure compatibility between:
π§ 87. Schema EvolutionΒΆ
Example:
New fields should ideally be introduced compatibly before old fields are removed.
π§ 88. Expand-and-ContractΒΆ
A safer migration pattern:
Example:
Add New Metadata Field
β
Deploy Consumers
β
Populate Field
β
Switch Retrieval
β
Remove Old Field
π§ 89. Deployment Blast RadiusΒΆ
Not every change should affect:
Use:
π§ 90. Tenant-Based RolloutΒΆ
Useful for enterprise platforms.
π§ 91. Region-Based RolloutΒΆ
Useful for global systems.
π§ 92. Feature Flag RolloutΒΆ
Roll out gradually:
π§ 93. Model Deployment PatternsΒΆ
Models can be deployed through:
π§ 94. Managed Model DeploymentΒΆ
Advantages:
π§ 95. Self-Hosted Model DeploymentΒΆ
Benefits:
Challenges:
π§ 96. Model CanaryΒΆ
Compare:
π§ 97. Prompt DeploymentΒΆ
Prompts should be treated as versioned artifacts.
Deploy independently when architecture permits.
π§ 98. Prompt CanaryΒΆ
Evaluate:
π§ 99. Retrieval DeploymentΒΆ
Retrieval components can also be independently deployed:
Use:
before full rollout.
π§ 100. Deployment Pattern SelectionΒΆ
Use the following mental model:
Small System
β
Monolith
Growing System
β
Modular Monolith
Independent Scaling Required
β
Microservices
Variable / Event-Driven Workload
β
Serverless
Complex Enterprise Platform
β
Containers / Kubernetes
High Deployment Risk
β
Canary / Blue-Green
Global Availability
β
Multi-Region
Migration
β
Dual Read / Dual Write
π§ 101. Deployment Decision MatrixΒΆ
| Requirement | Recommended Pattern |
|---|---|
| Simple application | Monolith |
| Strong modularity | Modular Monolith |
| Independent scaling | Microservices |
| Event-driven workload | Serverless / Workers |
| Enterprise platform | Containers / Kubernetes |
| Low deployment risk | Blue-Green |
| Gradual rollout | Canary |
| Production comparison | Shadow |
| Experimentation | A/B |
| Global availability | Multi-Region |
| High isolation tenant | Dedicated |
| Migration | Dual Read / Dual Write |
| Independent index lifecycle | Separate Index Deployment |
π§ 102. Deployment Architecture by MaturityΒΆ
Stage 1ΒΆ
Stage 2ΒΆ
Stage 3ΒΆ
Stage 4ΒΆ
Stage 5ΒΆ
π§ 103. Production Deployment ArchitectureΒΆ
flowchart TD
A["Users"] --> B["Global Load Balancer"]
B --> C["Region A"]
B --> D["Region B"]
C --> E["RAG Gateway"]
D --> F["RAG Gateway"]
E --> G["Retrieval Platform"]
F --> H["Retrieval Platform"]
G --> I["Index A"]
H --> J["Index B"]
E --> K["Model Gateway"]
F --> L["Model Gateway"]
K --> M["LLM"]
L --> N["LLM"]
O["CI/CD"] --> P["Deployment Controller"]
P --> E
P --> F
Q["Index Pipeline"] --> I
Q --> J
R["Observability"] --> E
R --> F
R --> G
R --> H π§ 104. Production Deployment WorkflowΒΆ
Developer
β
Git
β
CI
β
Unit Tests
β
Integration Tests
β
RAG Evaluation
β
Security Tests
β
Performance Tests
β
Artifact
β
Staging
β
Canary
β
Production
β
Observe
β
Promote / Rollback
π§ 105. Deployment SafetyΒΆ
Every production deployment should answer:
What changed?
Who receives the change?
How do we measure impact?
How quickly can we rollback?
What happens to existing requests?
What happens to indexes?
What happens to cached results?
What happens if the new version fails?
π§ 106. Cache During DeploymentΒΆ
Deployment can create stale caches.
Example:
Cache keys should include version information where appropriate:
π§ 107. Deployment and Index CompatibilityΒΆ
Avoid:
This creates runtime failures.
Use:
during transitions.
π§ 108. Deployment and FreshnessΒΆ
Application deployment does not automatically mean:
Keep separate lifecycles:
π§ 109. Knowledge DeploymentΒΆ
A document update may follow:
This is a deployment of knowledge rather than application code.
π§ 110. Knowledge CanaryΒΆ
For high-risk knowledge changes:
Useful for:
π§ 111. Regulated RAG DeploymentΒΆ
Regulated systems may require:
π§ 112. Approval WorkflowΒΆ
π§ 113. GitOpsΒΆ
For Kubernetes-based environments:
Benefits:
π§ 114. Infrastructure as CodeΒΆ
Infrastructure should be versioned:
Example:
π§ 115. Deployment as CodeΒΆ
The same principle applies to:
π§ 116. Immutable ArtifactsΒΆ
Examples:
Avoid mutable production artifacts.
π§ 117. Disaster Recovery DeploymentΒΆ
A recovery architecture should define:
π§ 118. RAG Disaster RecoveryΒΆ
flowchart LR
A["Primary Region"] --> B["Replication"]
B --> C["Secondary Region"]
A --> D["Backup"]
D --> E["Restore"]
C --> F["Failover"]
E --> F π§ 119. RPO / RTOΒΆ
Define:
Example:
These are illustrative.
π§ 120. Deployment Observability ChecklistΒΆ
β Deployment version tracked
β Traffic split visible
β Error rate monitored
β Latency monitored
β Retrieval quality monitored
β Groundedness monitored
β Citation quality monitored
β Token usage monitored
β Cost monitored
β Index version tracked
β Rollback available
π§ͺ 121. Practical ProjectΒΆ
Build a deployment platform for:
Production RAG Application
Support:
Docker
CI/CD
Staging
Canary
Blue-Green
Index Versioning
Prompt Versioning
Model Routing
Rollback
Observability
π§ͺ 122. Suggested RepositoryΒΆ
production-rag-deployment/
β
βββ application/
β βββ rag-api/
β βββ retrieval/
β βββ generation/
β
βββ deployment/
β βββ docker/
β βββ kubernetes/
β βββ helm/
β βββ manifests/
β
βββ infrastructure/
β βββ terraform/
β
βββ ci/
β βββ build.yaml
β βββ test.yaml
β βββ deploy.yaml
β
βββ evaluation/
β βββ datasets/
β βββ gates/
β
βββ indexes/
β βββ v1/
β βββ v2/
β
βββ docs/
βββ architecture/
βββ deployment/
βββ runbooks/
π§ͺ 123. Deployment ExerciseΒΆ
Implement:
Version 1
β
Production
Version 2
β
Staging
β
Evaluation
β
Canary 5%
β
Canary 25%
β
Canary 50%
β
100%
Then deliberately introduce:
and verify:
π§ͺ 124. Index Deployment ExerciseΒΆ
Create:
Then:
with an improved chunking strategy.
Perform:
π§ͺ 125. Model Deployment ExerciseΒΆ
Compare:
using:
Measure:
π§ͺ 126. Multi-Region ExerciseΒΆ
Deploy:
Test:
Verify:
π§ 127. Deployment Anti-PatternsΒΆ
Anti-Pattern 1ΒΆ
without evaluation or canary for high-risk changes.
Anti-Pattern 2ΒΆ
all changed simultaneously without version tracking.
Anti-Pattern 3ΒΆ
with no version or rollback capability.
Anti-Pattern 4ΒΆ
Anti-Pattern 5ΒΆ
Anti-Pattern 6ΒΆ
Anti-Pattern 7ΒΆ
Anti-Pattern 8ΒΆ
with no audit trail.
π§ 128. Deployment Design PrinciplesΒΆ
Principle 1 β Deploy Small ChangesΒΆ
Principle 2 β Separate LifecyclesΒΆ
Treat these independently:
Principle 3 β Automate Quality GatesΒΆ
Principle 4 β Make Rollback EasyΒΆ
Rollback should be:
Principle 5 β Prefer Progressive DeliveryΒΆ
Principle 6 β Observe QualityΒΆ
Traditional deployment metrics are not enough.
Track:
π§ 129. Deployment Pattern SummaryΒΆ
MONOLITH
β
Simple
MODULAR MONOLITH
β
Structured
MICROSERVICES
β
Independent Scaling
SERVERLESS
β
Variable / Event-Driven
CONTAINERS
β
Portable Production
KUBERNETES
β
Complex Enterprise Platform
ROLLING
β
Incremental Replacement
BLUE-GREEN
β
Fast Rollback
CANARY
β
Controlled Risk
SHADOW
β
Production Comparison
A/B
β
Experimentation
ACTIVE-ACTIVE
β
High Availability
ACTIVE-PASSIVE
β
Disaster Recovery
DUAL READ
β
Migration
DUAL WRITE
β
Migration Synchronization
π§ 130. Final Mental ModelΒΆ
Production RAG deployment should be viewed as:
RAG DEPLOYMENT
β
βββββββββββββββββββΌββββββββββββββββββ
βΌ βΌ βΌ
APPLICATION INDEX MODEL
β β β
Version Version Version
β β β
βββββββββββββββββββΌββββββββββββββββββ
βΌ
CONFIGURATION
β
βΌ
CI / CD PIPELINE
β
βΌ
STAGING
β
βΌ
EVALUATION
β
βΌ
CANARY
β
βββββββ΄ββββββ
βΌ βΌ
PASS FAIL
β β
βΌ βΌ
PROMOTE ROLLBACK
β
βΌ
PRODUCTION
β
βΌ
OBSERVABILITY
β
βΌ
CONTINUOUS IMPROVEMENT
π§ 131. Deployment FormulaΒΆ
A useful conceptual model:
π§ 132. What Makes RAG Deployment Production-Grade?ΒΆ
A production deployment should answer:
What changed?
Which version is running?
Which index is active?
Which model is active?
Which users receive the change?
How was the change evaluated?
What is the blast radius?
How do we monitor quality?
How do we rollback?
Can we reproduce the deployment?
Can we recover from regional failure?
If these questions cannot be answered, the deployment architecture is not mature enough.
π 133. Key TakeawaysΒΆ
- RAG deployment involves more than deploying an application.
- Application, retrieval, index, model, prompt, knowledge, and configuration have different lifecycles.
- Version every important RAG artifact.
- Monolithic RAG is appropriate for simple systems.
- Modular monoliths provide strong boundaries without distributed-system overhead.
- Microservices become useful when components require independent scaling or ownership.
- Retrieval can be exposed as a shared platform capability.
- Serverless is useful for variable and event-driven workloads.
- Containers provide portability and predictable runtime environments.
- Kubernetes is useful for complex enterprise platforms but introduces significant operational complexity.
- Rolling deployments provide incremental replacement.
- Blue-green deployments provide clean environments and fast rollback.
- Canary deployments reduce blast radius.
- Shadow deployments allow production comparison without affecting users.
- A/B deployments enable controlled experimentation.
- Multi-region deployments improve availability and latency but increase complexity.
- Active-active architectures provide high availability at higher operational cost.
- Active-passive architectures simplify disaster recovery.
- Dedicated tenant deployment provides strong isolation at higher cost.
- Shared tenant deployment improves infrastructure efficiency but requires strong isolation controls.
- Hybrid tenant deployment balances cost and isolation.
- Indexes should have independent versioning and deployment lifecycles.
- Embedding migrations should use controlled index migration strategies.
- Dual-read and dual-write patterns can help during large migrations.
- RAG deployments should use immutable artifacts where practical.
- CI/CD pipelines should include retrieval and generation evaluation, not just unit tests.
- Quality gates should include retrieval quality, groundedness, citation quality, latency, and cost.
- Progressive delivery is particularly valuable for high-risk RAG changes.
- Deployment observability must measure AI-specific quality signals.
- Rollback should be possible for application, retriever, index, prompt, model, and configuration independently where architecture permits.
- Schema evolution must maintain compatibility during rolling deployments.
- Knowledge deployment and application deployment should be treated as separate lifecycles.
- Infrastructure should be managed through infrastructure-as-code.
- Secrets should never be stored in source control.
- Production deployments should be observable, reproducible, auditable, and reversible.
- The correct deployment pattern depends on scale, availability, security, latency, team capability, cost, and business risk.
- The objective is not the most sophisticated deployment architecture.
- The objective is safe, measurable, repeatable, scalable, and reversible RAG delivery.
π§ 134. Chapter NavigationΒΆ
Part VI β Production RAG Deployment & OperationsΒΆ
Previous:
11. Building Production RAG Systems
Next:
13. RAG Caching Strategies
Production RAG Engineering PathΒΆ
01 Prompt Assembly
β
02 Context Selection & Context Engineering
β
03 Response Validation
β
04 Citation & Source Attribution
β
05 Enterprise Response
β
06 RAG Evaluation & Benchmarking
β
07 RAG Observability
β
08 RAG Performance Optimization
β
09 RAG Cost Optimization
β
10 Production Retrieval Architecture
β
11 Building Production RAG Systems
β
12 RAG Deployment Patterns
β
13 RAG Caching Strategies
β
14 Multi-Tenant RAG
β
15 RAG Testing Frameworks
β
16 RAG Failure Patterns
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems β One Chapter at a Time.