01. Agent Communication Overview¶
Category: Agent Communication
Module: AI Agents
Prerequisites: AI Agent Fundamentals, Agent Memory
Difficulty: IntermediateNote: Modern AI systems rarely consist of a single intelligent agent. Instead, multiple specialized agents collaborate to solve complex problems by exchanging tasks, information, decisions, and results. Agent Communication defines how AI agents interact with each other, external tools, APIs, humans, and enterprise systems in a reliable, scalable, and production-ready manner.
Overview¶
Imagine building an AI Software Engineering Assistant.
Instead of one large agent doing everything, you create specialized agents.
Software Engineer
↓
AI Supervisor
↓
Planner Agent
↓
Code Agent
↓
Testing Agent
↓
Documentation Agent
Each agent has a specific responsibility.
However, specialization alone is not enough.
The agents must communicate efficiently.
For example:
Planner Agent
↓
Generate Implementation Plan
↓
Code Agent
↓
Write Source Code
↓
Testing Agent
↓
Execute Tests
↓
Documentation Agent
↓
Generate README
Without communication, every agent works independently and cannot collaborate.
Agent Communication enables multiple AI agents to exchange information, coordinate tasks, and achieve a common objective.
🎯 Learning Objectives¶
After completing this note, you will understand:
- What Agent Communication is.
- Why communication is important in multi-agent systems.
- How router agents select agents, tools, and workflows.
- The major communication participants in enterprise AI systems.
- Direct, message-based, event-driven, and shared memory communication.
- Common communication patterns and when to use them.
- How agent communication is implemented using Python, LangGraph, and Kafka.
- Production architecture considerations for distributed AI systems.
- Communication best practices and common mistakes.
- How major AI agent frameworks support communication and coordination.
1. Why Agent Communication Matters¶
Without communication:
Problems include:
- Duplicate work
- No coordination
- Inconsistent decisions
- Workflow failures
- Poor scalability
With Communication¶
Supervisor Agent
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Planner Developer Tester
│ │ │
└──────────────┼──────────────┘
▼
Shared Information
Benefits include:
- Better collaboration
- Parallel execution
- Task delegation
- Faster workflows
- Scalable AI systems
Agent Communication allows specialized agents to operate as part of a coordinated system rather than as isolated components.
2. Communication Participants¶
AI agents communicate with multiple systems.
AI Agent
│
┌───────────────┼─────────────────┐
▼ ▼ ▼
Other Agents Humans External APIs
│ │
▼ ▼
Message Queue Enterprise Systems
Communication is not limited to agent-to-agent interactions.
Agents frequently communicate with:
- Other AI agents
- Human users
- APIs
- Databases
- Enterprise applications
- Workflow engines
- Messaging infrastructure
This means Agent Communication is broader than simply:
It represents the coordination layer connecting AI reasoning with enterprise systems and workflows.
3. High-Level Agent Communication Architecture¶
A multi-agent system may use the following architecture:
User
│
▼
Supervisor Agent
│
┌───────────────────────────┼────────────────────────────┐
▼ ▼ ▼
Planner Agent Coding Agent Testing Agent
│ │ │
└──────────────┬────────────┴──────────────┬─────────────┘
▼ ▼
Communication Layer
│
┌──────────────┼───────────────┐
▼ ▼ ▼
Message Bus Event Queue Shared Memory
│
▼
External Services
The Communication Layer enables agents to exchange information without requiring every agent to maintain direct knowledge of every other agent.
This supports:
- Loose coupling
- Independent deployment
- Better scalability
- Easier orchestration
4. Communication Lifecycle¶
Agent communication generally follows a structured lifecycle.
Create Message
↓
Send Message
↓
Receive Message
↓
Process Message
↓
Generate Response
↓
Continue Workflow
Each communication step should be:
- Reliable
- Traceable
- Fault tolerant
- Observable
In production systems, communication often also includes:
The communication mechanism depends on the architecture and delivery requirements.
5. 🔀 Router Agents¶
A router agent is responsible for deciding which specialized agent, tool, workflow, or system should handle a request.
User Request
↓
Router Agent
│
┌─────┼─────────────┐
▼ ▼ ▼
RAG SQL Support
Agent Agent Agent
│
└──────┬──────┘
↓
Final Response
The router analyzes the request and selects the most appropriate execution path.
For example:
User Request:
"Why did customer revenue drop last quarter?"
↓
Router Agent
│
┌────┴─────┐
▼ ▼
Database Knowledge
Agent Agent
│ │
└────┬─────┘
↓
Response
Router Responsibilities¶
A router agent may decide:
- Which specialized agent to invoke
- Which tool to use
- Which data source to query
- Whether multiple agents are required
- Which workflow should handle the request
Router vs Specialized Agent¶
This separation helps prevent routing logic from becoming mixed with domain-specific execution logic.
💼 Backend Architecture Parallel¶
A router agent is similar to a request-routing layer in distributed systems.
Similarly:
Key Principle¶
A router agent should focus on selecting the right execution path rather than solving every task itself.
This enables specialized agents to handle specific responsibilities while keeping the overall system modular and easier to scale.
6. Communication Models¶
Enterprise AI systems commonly use several communication models.
6.1 Direct Communication¶
One agent directly invokes another.
Characteristics¶
- Simple
- Low latency
- Tight coupling
Typical Uses¶
- Small agent systems
- Local workflows
- Simple orchestration
Direct communication is easy to implement but becomes harder to manage as the number of agents grows.
6.2 Message-Based Communication¶
Messages are exchanged through a broker.
Characteristics¶
- Loose coupling
- Reliable delivery
- Scalable
Typical Uses¶
- Enterprise AI platforms
- Distributed agents
- Independent services
The sender does not need to directly know the implementation details of the receiving agent.
6.3 Event-Driven Communication¶
Agents react to events.
Characteristics¶
- Asynchronous
- Highly scalable
- Decoupled
Typical Uses¶
- Business workflows
- Enterprise automation
- Long-running processes
Event-driven communication is useful when systems react to business events rather than directly invoking each other.
6.4 Shared Memory Communication¶
Agents exchange information through a common memory or state store.
Characteristics¶
- Shared context
- Easy collaboration
- Simple coordination
Typical Uses¶
- Multi-agent reasoning
- Shared planning
- Workflow state
Shared memory allows agents to collaborate through common context rather than sending every piece of information directly between agents.
7. Communication Patterns¶
Different workflows require different communication strategies.
| Pattern | Typical Use Case |
|---|---|
| Direct Calls | Small workflows |
| Message Queue | Distributed systems |
| Publish-Subscribe | Event processing |
| Shared Memory | Collaborative reasoning |
| Request-Response | Tool invocation |
| Broadcast | Multi-agent notifications |
8. Choosing the Right Communication Pattern¶
| Scenario | Recommended Pattern |
|---|---|
| Single workflow | Direct Communication |
| Multi-agent collaboration | Shared Memory |
| Enterprise automation | Event-Driven |
| Distributed AI platform | Message Queue |
| External APIs | Request-Response |
| Notifications | Publish-Subscribe |
There is no universal communication model.
The correct choice depends on:
- Workflow complexity
- Number of agents
- Latency requirements
- Scalability requirements
- Failure handling requirements
- Deployment architecture
9. Implementation¶
Example 1 — Core Python¶
A simple direct communication example.
class PlannerAgent:
def create_task(self):
return "Generate project documentation"
class DocumentationAgent:
def execute(self, task):
print(f"Executing: {task}")
planner = PlannerAgent()
documentation = DocumentationAgent()
task = planner.create_task()
documentation.execute(task)
Output:
This demonstrates synchronous communication between two agents.
Example 2 — LangGraph¶
LangGraph enables communication through shared workflow state.
from typing import TypedDict
from langgraph.graph import StateGraph
class AgentState(TypedDict):
task: str
result: str
workflow = StateGraph(AgentState)
workflow.add_node("planner", planner_node)
workflow.add_node("developer", developer_node)
workflow.add_node("tester", tester_node)
Each node communicates by reading from and writing to the shared workflow state rather than directly invoking other agents.
Conceptually:
This approach is particularly useful for structured multi-step workflows.
Example 3 — Production Example: Kafka¶
Enterprise AI systems commonly use Kafka for asynchronous communication.
from kafka import KafkaProducer
import json
producer = KafkaProducer(
bootstrap_servers="localhost:9092",
value_serializer=lambda v: json.dumps(v).encode("utf-8")
)
producer.send(
"agent-tasks",
{
"agent": "documentation",
"task": "Generate API documentation"
}
)
producer.flush()
Instead of directly invoking another agent, the task is published to a Kafka topic.
An authorized agent subscribed to the topic can consume and process the message.
Architecture:
This enables:
- Loose coupling
- Asynchronous processing
- Independent scaling
- Distributed deployment
10. Enterprise Use Cases¶
Software Development Assistant¶
Multiple specialized agents collaborate to complete software development tasks.
Examples:
- Requirement Analysis Agent
- Architecture Agent
- Code Generation Agent
- Testing Agent
- Documentation Agent
Developer
↓
Supervisor Agent
↓
Planner
↓
Code Generator
↓
Tester
↓
Documentation Agent
↓
Final Solution
Each agent communicates its output to the next stage of the workflow.
Customer Support Platform¶
Customer support systems commonly consist of several specialized agents.
Examples:
- Intent Detection Agent
- Knowledge Retrieval Agent
- Ticket Management Agent
- Escalation Agent
- Feedback Agent
Instead of a single agent handling everything, each agent performs one specialized task and communicates the results.
Financial Services¶
Enterprise banking systems may orchestrate multiple AI agents.
Examples:
- Fraud Detection Agent
- Risk Assessment Agent
- Compliance Agent
- Recommendation Agent
- Customer Notification Agent
Each agent exchanges structured messages while maintaining auditability.
Healthcare Assistant¶
Healthcare AI systems may require collaboration among multiple specialized agents.
Examples:
- Patient Intake Agent
- Diagnosis Assistant
- Medical Knowledge Agent
- Treatment Recommendation Agent
- Appointment Scheduling Agent
Communication enables each agent to contribute to the overall workflow without unnecessarily duplicating responsibilities.
Enterprise Workflow Automation¶
Large organizations automate business workflows using communicating agents.
Invoice Received
↓
Validation Agent
↓
Approval Agent
↓
Payment Agent
↓
Notification Agent
↓
ERP System
Each agent performs a specific business operation and communicates completion before the workflow proceeds.
11. Production Communication Architecture¶
Enterprise AI agents often communicate through messaging infrastructure rather than relying exclusively on direct method calls.
Supervisor Agent
│
▼
Communication Layer
│
┌─────────────────┼──────────────────┐
▼ ▼ ▼
Kafka RabbitMQ Redis Streams
│ │ │
▼ ▼ ▼
Planner Agent Coding Agent Testing Agent
This architecture can provide:
- Loose coupling
- Independent deployment
- Horizontal scalability
- Fault tolerance
- Reliable message delivery
A key architectural principle is:
Separate business logic from communication infrastructure.
Agents should focus on their business capability while messaging infrastructure handles:
- Routing
- Delivery
- Queuing
- Retry mechanisms
- Event distribution
12. Architecture Decision Guide¶
| Scenario | Recommended Communication Pattern |
|---|---|
| Single application | Direct Communication |
| Microservices | Message Queue |
| Event-driven workflows | Publish-Subscribe |
| Shared planning | Shared Memory |
| External API integration | Request-Response |
| Long-running workflows | Event Bus |
| Enterprise AI Platform | Kafka + Workflow Engine |
| Multi-Agent Systems | Message Queue + Shared Memory |
13. Advantages¶
Agent Communication:
- Enables collaboration between specialized agents
- Supports distributed AI architectures
- Improves scalability
- Enables asynchronous execution
- Reduces coupling between agents
- Improves fault tolerance
- Simplifies workflow orchestration
14. Limitations¶
Distributed communication also introduces challenges.
Examples:
- Additional infrastructure requirements
- Increased architectural complexity
- Message ordering challenges
- Network latency
- More complex error handling
- Distributed debugging difficulties
As the system becomes more distributed, observability and communication design become increasingly important.
15. Best Practices¶
- Keep messages small and self-contained.
- Use structured message formats such as JSON, Protobuf, or Avro.
- Design communication to be asynchronous whenever possible.
- Avoid unnecessary direct dependencies between specialized agents.
- Implement retry strategies.
- Use dead-letter queues where appropriate.
- Use correlation IDs for request tracing.
- Version message schemas.
- Monitor communication latency and failures.
- Separate business logic from messaging logic.
- Define clear ownership for messages and workflows.
16. Common Mistakes¶
❌ Creating tightly coupled agents
❌ Using synchronous communication everywhere
❌ Sending large payloads between agents
❌ Ignoring message versioning
❌ No retry strategy
❌ Missing correlation IDs
❌ No communication monitoring
❌ Mixing business logic with messaging logic
These problems become significantly more difficult to manage as the number of agents and distributed services increases.
17. Framework Comparison¶
| Framework | Communication Model |
|---|---|
| LangChain | Chains, Tool Calling, Runnable Pipelines |
| LangGraph | Shared Graph State, Directed Workflow Edges |
| CrewAI | Agent-to-Agent Collaboration |
| AutoGen | Conversational Multi-Agent Messaging |
| OpenAI Agents SDK | Tool Invocation & Session Context |
| Google ADK | Workflow & Agent Coordination |
Different frameworks provide different abstractions, but the underlying architectural concerns remain similar:
- Information exchange
- Task delegation
- State management
- Coordination
- Workflow control
18. 💼 Backend Architecture Parallel¶
Agent Communication closely resembles communication patterns used in distributed backend systems.
Traditional distributed architecture:
Multi-agent architecture:
The same distributed systems concepts apply:
- Synchronous vs asynchronous communication
- Loose coupling
- Message schemas
- Retry strategies
- Dead-letter queues
- Correlation IDs
- Event-driven architecture
- Observability
- Independent deployment
A useful mental model is:
The major difference is that agents may make probabilistic decisions about:
- Which agent to contact
- Which workflow to execute
- Which tool to invoke
- What information should be shared
19. Interview Questions¶
What is Agent Communication?¶
Why is communication important in multi-agent systems?¶
What is the difference between direct communication and message-based communication?¶
When should event-driven communication be preferred?¶
Why are message queues commonly used in enterprise AI systems?¶
What is shared memory communication?¶
What is the responsibility of a router agent?¶
How does asynchronous communication improve scalability?¶
What challenges arise in distributed agent communication?¶
How would you design communication for a large enterprise multi-agent system?¶
20. 🚀 Quick Revision¶
AI Agents
│
▼
Communication Layer
│
┌──────────────┼──────────────┐
▼ ▼ ▼
Direct Call Message Queue Event Bus
│ │ │
▼ ▼ ▼
Shared Memory Request-Response Publish-Subscribe
│
▼
External Systems
Router Agent¶
Router vs Specialized Agent¶
Communication Lifecycle¶
Create Message
↓
Send Message
↓
Receive Message
↓
Process Message
↓
Generate Response
↓
Continue Workflow
Production Principle¶
Keep them separated.
21. Key Takeaways¶
- Agent Communication enables AI agents to collaborate, coordinate tasks, and exchange information efficiently.
- Multi-agent systems require communication to avoid isolated execution, duplicate work, and inconsistent workflows.
- Router agents determine the appropriate execution path by selecting agents, tools, workflows, or data sources.
- Enterprise AI platforms use multiple communication models, including direct communication, message queues, event-driven communication, publish-subscribe, and shared memory.
- Modern production systems often use messaging infrastructure such as Kafka, RabbitMQ, or Redis Streams to improve scalability, reliability, and loose coupling.
- Choosing the appropriate communication pattern depends on workflow complexity, latency requirements, scalability goals, reliability requirements, and deployment architecture.
- Distributed agent communication introduces familiar backend engineering concerns such as retries, message versioning, correlation IDs, dead-letter queues, and observability.
- Effective communication is a foundation for multi-agent systems, distributed AI platforms, and enterprise workflow automation.
22. References¶
- LangGraph Documentation — Multi-Agent Workflows
- CrewAI Documentation — Agent Collaboration
- AutoGen Documentation — Multi-Agent Conversations
- Apache Kafka Documentation
- RabbitMQ Documentation
- Redis Streams Documentation
23. Next Note¶
02-message-passing.md
In the next note, we'll explore Message Passing, the most fundamental communication mechanism in multi-agent systems.
You will learn:
- Synchronous vs asynchronous messaging
- Message structure
- Delivery guarantees
- Serialization formats
- Routing strategies
- Acknowledgments
- Retries
- Production implementations using Kafka, RabbitMQ, Redis Streams, and cloud messaging services
Enterprise AI Engineering Handbook
Building Production-Grade Enterprise AI Systems — One Chapter at a Time.