🎯 By the end of this module, you will:
Configuring Async and Sync Agent Execution
LLM applications are almost entirely I/O-bound. While waiting 2s for a model response, your CPU is completely idle. Async execution lets your agent handle dozens of concurrent user tasks simultaneously with zero extra hardware.
# SYNC: Sequential — blocks on every LLM call
# 6 tasks × 2s each = 12s total wall-clock time
from langchain_openai import ChatOpenAI
model = ChatOpenAI(model="gpt-4o-mini")
def process_sync(prompts: list[str]) -> list[str]:
results = []
for prompt in prompts:
# .invoke() BLOCKS the thread for ~2s each!
result = model.invoke(prompt)
results.append(result.content)
return results
# 6 prompts → executed one by one → 12s total
answers = process_sync([
"Summarize quantum computing",
"Explain transformers",
"What is RAG?",
"Define agent loops",
"What is LangGraph?",
"Explain vector DBs",
])👨🍳 Mental Model: Chef Waiting vs. Chef Multitasking
Chef boils pasta, then stares at the pot for 10 minutes until it's done. Only then do they start chopping vegetables. Staring = your CPU waiting for the LLM response. Completely idle. 6 dishes = 60 minutes.
Chef starts all 6 dishes, sets timers, and bounces between tasks during each other's wait periods. 6 dishes in the time it takes to cook 1 — the exact same kitchen, same chef, just no idle waiting. That's asyncio.
🔬 Async vs. Sync Benchmarker
Async vs. Sync Concurrency BenchmarkerI/O-Bound Optimization
Experience why asynchronous execution (ainvoke / asyncio.gather) is mandatory for multi-user AI agents
import asyncio
# 1. Async invocation with 'a' prefix:
async def run_concurrently():
tasks = [
agent.ainvoke({"input": q1}),
agent.ainvoke({"input": q2}),
agent.ainvoke({"input": q3}),
agent.ainvoke({"input": q4}),
]
# Executes all 4 in parallel!
results = await asyncio.gather(*tasks)
# 2. Wrapping blocking sync tool:
import asyncio
result = await asyncio.to_thread(sync_tool)Async Execution Traps
Calling time.sleep() or requests.get() inside an async function blocks the entire event loop, making all concurrent tasks wait in line. Use await asyncio.sleep() and await httpx.AsyncClient.get() instead. One sync call can kill all concurrency.
If one task in asyncio.gather() raises an exception, all other tasks are cancelled by default. Use asyncio.gather(*tasks, return_exceptions=True) to capture individual failures and continue processing the remaining successful tasks.
Key Takeaways
- 1.LLM Calls are I/O-Bound: While waiting for the API response, your CPU is idle. asyncio lets you fill that idle time by running other requests concurrently — same CPU, 5x+ throughput.
- 2.asyncio.gather() is the Multiplier: Dispatching N tasks with asyncio.gather() makes them run concurrently. N × 2s of sequential work becomes ~2s of concurrent work. Use this for batch processing and multi-user agent servers.
- 3.asyncio.to_thread() for Legacy Tools: Any blocking sync function (DB driver, requests, file I/O) that you can't swap must be wrapped in asyncio.to_thread() to prevent event loop starvation.
Creating Reusable Dynamic Prompt Templates
Stop embedding prompts as raw f-strings. Build composable ChatPromptTemplates with role isolation, history placeholders, and runtime variable binding.