🎯 By the end of this module, you will:
Single-Agent vs. Multi-Agent Systems
Single agents minimize latency, cost, and debugging headaches for focused tasks. Multi-agent teams partition complexity across specialized subagents with isolated context sandboxes.
One prompt, one LLM loop, and a focused tool catalog (<15 tools). Easiest to test, debug, and monitor.
# Single-Agent ReAct Loop
agent = ReActAgent(
model="gemini-2.0-flash",
tools=[search_db, run_sql, format_table],
system_prompt="You are a SQL data analyst."
)
response = agent.run("Show last month's churn rate by region.")Mental Model: Family Clinic Doctor vs. Hospital Surgical Team
For a common fever, visiting a single doctor takes 10 minutes and gives immediate relief. Summoning a committee of 6 surgeons, an anesthesiologist, and a radiologist for a mild cold would be absurdly slow and expensive.
However, for open-heart surgery, that solo doctor cannot do it alone. You need a specialized hospital team (surgeon, cardiologist, anesthesiologist, scrub nurse) following strict surgical protocols.
The 3 Canonical Multi-Agent Topologies
Explore how multi-agent teams communicate: Hierarchical Supervisors (Claude Code), Deterministic Pipelines (Deep Research), and Swarm Networks (Customer Service).
Interactive Multi-Agent Topology Simulator
Explore the 3 canonical coordination patterns: Hierarchical Supervisor, Sequential Pipeline, and Peer Network.
Supervisor Orchestrator Agent
Decomposes task & routes to subagents
Claude Code (Anthropic)
Claude Code employs a master orchestrator that spawns isolated subagents (e.g. for codebase exploration). The subagents search and test code in clean separate sandboxes, returning only their verified conclusions back to the main agent.
Multi-agent systems add coordination lag, latency, and debugging complexity. Start with a single well-designed agent! Only split into multiple agents when you hit clear limitations: tool confusion (>20 tools) or conflicting role requirements.
Single-Agent vs. Multi-Agent Decision Engine
Adjust your project parameters (tool count, concurrency needs, and persona divergence) to receive an immediate architecture recommendation.
Architectural Compass: Single-Agent vs. Multi-Agent Decision Studio
Configure your technical requirements to discover the ideal architectural pattern and avoid over-engineering.
✓ Within safe single-agent capacity (low risk of tool hallucination).
Build a Single Well-Architected ReAct Agent
Your workload has fewer than 15 tools and no conflicting role personas. Introducing multi-agent orchestration right now would needlessly inflate token costs, introduce coordination latency, and make debugging much harder. Keep it lean and simple!
# LangGraph Multi-Agent Supervisor Pattern
from langchain_core.messages import HumanMessage
from langgraph.graph import StateGraph, START, END
# 1. Define Specialized Subagents
def researcher_agent(state):
# Runs search tools in isolation; returns factual summary
return {"messages": ["Researcher: Retrieved 5 relevant financial reports."]}
def coder_agent(state):
# Runs code interpreter in clean sandbox
return {"messages": ["Coder: Executed pandas script; churn rate = 3.2%."]}
# 2. Supervisor Orchestrator Router
def supervisor_node(state):
# Decides whether to route to researcher, coder, or finish
last_msg = state["messages"][-1]
if "Retrieved" in last_msg:
return "coder"
return END
# 3. Compile Graph with Shallow Hierarchy
workflow = StateGraph(dict)
workflow.add_node("researcher", researcher_agent)
workflow.add_node("coder", coder_agent)
workflow.add_conditional_edges("researcher", supervisor_node)
app = workflow.compile()Architectural Traps to Avoid
Splitting a simple CRUD or search task into 5 debating agents creates massive token latency, high bills, and non-deterministic loops. If a single prompt with 4 tools can do the job, keep it single-agent!
Research demonstrates model accuracy plunges when one agent is loaded with >15-20 tools at once (tool confusion). When you cross this threshold, partition into specialized subagents holding 3-5 tools each!
Key Architectural Takeaways
- 1.Default to Single-Agent: It delivers the lowest latency, lowest cost, and easiest observability.
- 2.Partition on 3 Triggers: Split only when tool count >15, strict parallel concurrency is required, or roles directly conflict.
- 3.Use Shallow Hierarchies: In supervisor setups (like Claude Code), worker subagents run in isolated sandboxes and discard noise before reporting back.
Enhancing Agents with Retrieval-Augmented Generation (RAG)
Learn how autonomous agents convert static RAG pipelines into dynamic, agentic search tools with self-correction, query re-writing, and iterative chunk retrieval.