Mod 2.10Troubleshooting Common LLM API
Level 2›Module 2.10
Level 2: Core Implementation & WorkflowsModule 2.10

Troubleshooting Common LLM API

Common LLM API

Level 2 • Core Implementation & Workflows
Est. ~39 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Troubleshooting Common LLM API? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone

🎯 By the end of this module, you will:

Categorize 401, 429, 500, and timeout errors with correct retry vs. halt decisions
Implement tenacity exponential backoff with jitter for 429 rate limit handling
Build multi-provider fallback with LangChain's with_fallbacks() API
Enforce strict async timeout deadlines to prevent agent hang on slow APIs
Error Handling • Resilience Patterns

Troubleshooting Common LLM API Issues

Networks flake, rate limits hit during traffic spikes, and providers go down. A self-healing agent handles these predictably — no 3am pages required.

401/403 Auth: Fatal Config ErrorNo Retry
# 401/403: Fatal auth error — alert, never retry
from openai import AuthenticationError

try:
    response = client.chat.completions.create(...)
except AuthenticationError as e:
    # NEVER retry auth errors — the key is broken
    logger.critical(json.dumps({
        "event": "auth_error", "http_status": 401,
        "action": "alert_developer",
        "message": "Revoke and rotate API key immediately"
    }))
    # Raise to caller — this needs human intervention
    raise SystemExit("Invalid API key — shutting down agent")

🔀 Error Decision Tree: Retry vs. Halt vs. Fallback

401 / 403 → HALT

Invalid API key or insufficient permissions. Fix config, don't retry.

429 → RETRY

Exponential backoff + jitter. Max 5 attempts, then escalate.

500 / 503 → FALLBACK

Route to secondary provider. If all providers fail, queue for retry.

🔬 API Error Troubleshooter

Production API Error & Resilience LabHTTP & Retries

Simulate common API errors (401, 429, 503) and test self-healing mitigations in real time

Resilience Defenses
Exponential Backoff (Tenacity)
Jittered delay between 429 retries
Multi-Provider Fallback Router
Auto-routes to Claude/Gemini on 503
# LangChain Fallback Pattern:
primary = ChatOpenAI(model="gpt-4o")
fallback = ChatAnthropic(model="claude-3-5-haiku")

# Self-healing fallback model!
resilient_model = primary.with_fallbacks([fallback])
Runtime Network TelemetryScenario: HTTP 429

Your application exceeded the Tokens Per Minute limit during a burst of concurrent tool calls.

HTTP REQUEST / RESPONSE LIFECYCLESTATUS LOG
Click “Simulate HTTP 429” to observe network telemetry and mitigation logic.
✍️ Instructor Note: "Always add jitter (randomized delay) to your backoff! Without jitter, 100 concurrent requests that all get rate-limited will retry at exactly the same time, causing a thundering herd that worsens the 429 problem."
💡 Mental Model: Think of API errors like traffic signals. Red (401) = stop forever, call a mechanic. Yellow (429) = slow down, wait, then proceed. Red-then-detour (500) = road is closed, take the alternate route (fallback provider).

API Error Handling Traps

TRAP #1: Retrying Auth Errors

Retrying a 401 error is like trying to swipe a cancelled credit card 5 times. Each attempt wastes latency and fails identically. Detect auth errors immediately, log them as critical, and halt the agent pending human intervention.

TRAP #2: Linear Retry Without Jitter

Retrying every 2 seconds with a fixed delay during a 429 storm causes synchronized thundering herds. Always use exponential backoff (2s, 4s, 8s...) plus a random ±30% jitter to stagger retries across concurrent requests.

Key Takeaways

  • 1.Classify Before Responding: Not all errors are retryable. 401=halt, 429=backoff, 500=fallback, timeout=retry. Build your error handler as a decision tree, not a catch-all retry loop.
  • 2.Use tenacity for Retries: tenacity's @retry decorator with wait_exponential handles the entire backoff/jitter/max-attempt logic in 4 lines. Don't roll your own retry loop.
  • 3.with_fallbacks() is Free Insurance: LangChain's with_fallbacks() API gives you multi-provider failover with zero custom code. Primary → backup provider routing happens transparently on any server error.
Up Next • Module 2.11

Building a Chatbot Agent in LangGraph

Build a stateful multi-turn conversational agent using LangGraph checkpointers and thread_id — so your agent remembers what users said across sessions.

Continue to Module 2.11