Mod 2.13Configuring Async and Sync Agent Execution
Level 2›Module 2.13
Level 2: Core Implementation & WorkflowsModule 2.13

Configuring Async and Sync Agent Execution

Async and Sync

Level 2 • Core Implementation & Workflows
Est. ~27 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Configuring Async and Sync Agent Execution? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone

🎯 By the end of this module, you will:

Explain why LLM calls are I/O-bound and why sync agents waste CPU waiting
Use asyncio.gather() to run 6 concurrent LLM calls in the time of 1 sequential call
Wrap legacy blocking tools with asyncio.to_thread() to prevent event loop starvation
Wire LangGraph's ainvoke() and astream() into a FastAPI async endpoint
Async Execution • Concurrency Patterns

Configuring Async and Sync Agent Execution

LLM applications are almost entirely I/O-bound. While waiting 2s for a model response, your CPU is completely idle. Async execution lets your agent handle dozens of concurrent user tasks simultaneously with zero extra hardware.

Sync (Sequential): Blocking Execution
# SYNC: Sequential — blocks on every LLM call
# 6 tasks × 2s each = 12s total wall-clock time

from langchain_openai import ChatOpenAI

model = ChatOpenAI(model="gpt-4o-mini")

def process_sync(prompts: list[str]) -> list[str]:
    results = []
    for prompt in prompts:
        # .invoke() BLOCKS the thread for ~2s each!
        result = model.invoke(prompt)
        results.append(result.content)
    return results

# 6 prompts → executed one by one → 12s total
answers = process_sync([
    "Summarize quantum computing",
    "Explain transformers",
    "What is RAG?",
    "Define agent loops",
    "What is LangGraph?",
    "Explain vector DBs",
])

👨‍🍳 Mental Model: Chef Waiting vs. Chef Multitasking

❌ Sync Agent (Inefficient Chef)

Chef boils pasta, then stares at the pot for 10 minutes until it's done. Only then do they start chopping vegetables. Staring = your CPU waiting for the LLM response. Completely idle. 6 dishes = 60 minutes.

✅ Async Agent (Expert Chef)

Chef starts all 6 dishes, sets timers, and bounces between tasks during each other's wait periods. 6 dishes in the time it takes to cook 1 — the exact same kitchen, same chef, just no idle waiting. That's asyncio.

🔬 Async vs. Sync Benchmarker

Async vs. Sync Concurrency BenchmarkerI/O-Bound Optimization

Experience why asynchronous execution (ainvoke / asyncio.gather) is mandatory for multi-user AI agents

4 Concurrent Tool Calls (Search / DB / Calculations)
Task #1: Query API EndpointQueued
Task #2: Query API EndpointQueued
Task #3: Query API EndpointQueued
Task #4: Query API EndpointQueued
Python Async Method Patterns
import asyncio

# 1. Async invocation with 'a' prefix:
async def run_concurrently():
    tasks = [
        agent.ainvoke({"input": q1}),
        agent.ainvoke({"input": q2}),
        agent.ainvoke({"input": q3}),
        agent.ainvoke({"input": q4}),
    ]
    # Executes all 4 in parallel!
    results = await asyncio.gather(*tasks)

# 2. Wrapping blocking sync tool:
import asyncio
result = await asyncio.to_thread(sync_tool)
💡 Why async rules AI: LLM apps spend 95% of execution waiting for cloud sockets. While waiting, Python's async event loop can process dozens of other user requests simultaneously!
✍️ Instructor Note: "Don't mix sync and async carelessly! Calling a blocking sync function (like requests.get()) inside an async function will FREEZE your entire event loop — blocking ALL concurrent users, not just one. Always use await httpx.AsyncClient() instead of requests in async code."
📌 Core Rule: CPU-bound work (numpy calculations, image processing) does NOT benefit from asyncio — use multiprocessing instead. asyncio shines for I/O-bound work only: network calls, database queries, file I/O. LLM API calls are 100% I/O-bound.

Async Execution Traps

TRAP #1: Blocking the Event Loop

Calling time.sleep() or requests.get() inside an async function blocks the entire event loop, making all concurrent tasks wait in line. Use await asyncio.sleep() and await httpx.AsyncClient.get() instead. One sync call can kill all concurrency.

TRAP #2: asyncio.gather() Without Error Handling

If one task in asyncio.gather() raises an exception, all other tasks are cancelled by default. Use asyncio.gather(*tasks, return_exceptions=True) to capture individual failures and continue processing the remaining successful tasks.

Key Takeaways

  • 1.LLM Calls are I/O-Bound: While waiting for the API response, your CPU is idle. asyncio lets you fill that idle time by running other requests concurrently — same CPU, 5x+ throughput.
  • 2.asyncio.gather() is the Multiplier: Dispatching N tasks with asyncio.gather() makes them run concurrently. N × 2s of sequential work becomes ~2s of concurrent work. Use this for batch processing and multi-user agent servers.
  • 3.asyncio.to_thread() for Legacy Tools: Any blocking sync function (DB driver, requests, file I/O) that you can't swap must be wrapped in asyncio.to_thread() to prevent event loop starvation.
Up Next • Module 2.14

Creating Reusable Dynamic Prompt Templates

Stop embedding prompts as raw f-strings. Build composable ChatPromptTemplates with role isolation, history placeholders, and runtime variable binding.

Continue to Module 2.14