🎯 By the end of this module, you will:
Implementing Streaming Output for Real-Time Responses
Nobody likes a 8-second spinner. Streaming delivers tokens as they're generated, slashing TTFT from seconds to milliseconds and giving users the instant typing experience they expect from modern AI apps.
# stream_mode="messages" — raw token chunks for chat UIs
# Each chunk arrives as AIMessageChunk with .content str
async def stream_to_ui(user_query: str):
async for chunk, metadata in app.astream(
{"messages": [("user", user_query)]},
stream_mode="messages"
):
if chunk.content and not isinstance(chunk, HumanMessage):
# Send token immediately via SSE / WebSocket
yield f"data: {chunk.content}\n\n"
# React UI: setResponse(prev => prev + chunk.content)
# Result: Users see "Q", "ua", "ntum", " comput", "ing..."
# TTFT (Time to First Token) ~150ms vs 8s for batch response⚡ Mental Model: Restaurant vs. Streaming
Like a restaurant that serves nothing until the entire 5-course meal is cooked. You wait 25 minutes staring at an empty table before eating your first bite — even if the bread was ready in 30 seconds.
Like a restaurant that brings bread immediately, then soup, then salad as each course finishes. You're eating within 30 seconds of ordering — even if the full meal takes 25 minutes total to prepare.
🔬 Streaming Output Simulator
Real-Time Token Streaming StudioTTFT & Stream Modes
Inspect how LangGraph streams messages, node updates, and state snapshots to eliminate user wait time
for chunk in app.stream(inputs, stream_mode="messages"):
print(chunk)Streaming Output Traps
Nginx, AWS ALB, and most reverse proxies buffer responses by default. Unless you explicitly set X-Accel-Buffering: no, all tokens accumulate server-side and deliver in one batch — making streaming appear broken to the client.
The most common mistake: developers add streaming code to the UI but forget to change the backend from agent.invoke() to agent.stream(). invoke() waits for the full response before returning — your streaming frontend never gets any chunks to display.
Key Takeaways
- 1.Choose Mode by Use Case: messages for chat UIs (token chunks), updates for progress indicators (node steps), values for telemetry and debugging (full state snapshots).
- 2.TTFT Is Your UX Metric: Users perceive streaming as dramatically faster even when total latency is identical. Measure and optimize Time to First Token for user-facing features.
- 3.Fix Your Proxy Headers: Always set X-Accel-Buffering: no and Cache-Control: no-cache on your streaming endpoints. This is the #1 reason streaming appears broken in production deployments.
Configuring Async and Sync Agent Execution
LLM calls are I/O-bound. Learn how async execution lets your agent handle 6 concurrent users in the same time it would take to handle 1 synchronously.