Mod 2.12Implementing Streaming Output for Real-Time Responses
Level 2›Module 2.12
Level 2: Core Implementation & WorkflowsModule 2.12

Implementing Streaming Output for Real-Time Responses

Streaming Output

Level 2 • Core Implementation & Workflows
Est. ~15 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Implementing Streaming Output for Real-Time Responses? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone

🎯 By the end of this module, you will:

Distinguish the 3 LangGraph stream modes and choose the right one for your UI
Stream raw token chunks using stream_mode="messages" for chat interfaces
Build a FastAPI SSE endpoint that streams LangGraph output to browsers
Measure Time to First Token (TTFT) and understand its UX impact
Streaming Output • Real-Time UX

Implementing Streaming Output for Real-Time Responses

Nobody likes a 8-second spinner. Streaming delivers tokens as they're generated, slashing TTFT from seconds to milliseconds and giving users the instant typing experience they expect from modern AI apps.

stream_mode="messages"Chat UIs
# stream_mode="messages" — raw token chunks for chat UIs
# Each chunk arrives as AIMessageChunk with .content str

async def stream_to_ui(user_query: str):
    async for chunk, metadata in app.astream(
        {"messages": [("user", user_query)]},
        stream_mode="messages"
    ):
        if chunk.content and not isinstance(chunk, HumanMessage):
            # Send token immediately via SSE / WebSocket
            yield f"data: {chunk.content}\n\n"
            # React UI: setResponse(prev => prev + chunk.content)

# Result: Users see "Q", "ua", "ntum", " comput", "ing..."
# TTFT (Time to First Token) ~150ms vs 8s for batch response

⚡ Mental Model: Restaurant vs. Streaming

❌ Batch (No Streaming)

Like a restaurant that serves nothing until the entire 5-course meal is cooked. You wait 25 minutes staring at an empty table before eating your first bite — even if the bread was ready in 30 seconds.

✅ Streaming (Token by Token)

Like a restaurant that brings bread immediately, then soup, then salad as each course finishes. You're eating within 30 seconds of ordering — even if the full meal takes 25 minutes total to prepare.

🔬 Streaming Output Simulator

Real-Time Token Streaming StudioTTFT & Stream Modes

Inspect how LangGraph streams messages, node updates, and state snapshots to eliminate user wait time

Client Render Output
Tokens: 0
Awaiting streaming dispatch... Click “Start Token Stream” below.
Emitted Chunk Stream0 Chunks
Chunk payloads will appear here during stream...
# LangGraph Streaming:
for chunk in app.stream(inputs, stream_mode="messages"):
    print(chunk)
✍️ Instructor Note: "Streaming only reduces perceived latency (TTFT), NOT total generation time. A 5-second generation takes 5 seconds regardless. Streaming just means your user sees the first word at 150ms instead of waiting all 5 seconds for the full response."
📌 Core Rule: Always set headers X-Accel-Buffering: no and Cache-Control: no-cache on your SSE streaming endpoint. Without these, nginx and cloud load balancers will buffer your entire response before sending it, completely defeating the purpose of streaming!

Streaming Output Traps

TRAP #1: Buffering Middleware

Nginx, AWS ALB, and most reverse proxies buffer responses by default. Unless you explicitly set X-Accel-Buffering: no, all tokens accumulate server-side and deliver in one batch — making streaming appear broken to the client.

TRAP #2: Calling invoke() Instead of stream()

The most common mistake: developers add streaming code to the UI but forget to change the backend from agent.invoke() to agent.stream(). invoke() waits for the full response before returning — your streaming frontend never gets any chunks to display.

Key Takeaways

  • 1.Choose Mode by Use Case: messages for chat UIs (token chunks), updates for progress indicators (node steps), values for telemetry and debugging (full state snapshots).
  • 2.TTFT Is Your UX Metric: Users perceive streaming as dramatically faster even when total latency is identical. Measure and optimize Time to First Token for user-facing features.
  • 3.Fix Your Proxy Headers: Always set X-Accel-Buffering: no and Cache-Control: no-cache on your streaming endpoints. This is the #1 reason streaming appears broken in production deployments.
Up Next • Module 2.13

Configuring Async and Sync Agent Execution

LLM calls are I/O-bound. Learn how async execution lets your agent handle 6 concurrent users in the same time it would take to handle 1 synchronously.

Continue to Module 2.13