🎯 By the end of this module, you will:
Troubleshooting Common LLM API Issues
Networks flake, rate limits hit during traffic spikes, and providers go down. A self-healing agent handles these predictably — no 3am pages required.
# 401/403: Fatal auth error — alert, never retry
from openai import AuthenticationError
try:
response = client.chat.completions.create(...)
except AuthenticationError as e:
# NEVER retry auth errors — the key is broken
logger.critical(json.dumps({
"event": "auth_error", "http_status": 401,
"action": "alert_developer",
"message": "Revoke and rotate API key immediately"
}))
# Raise to caller — this needs human intervention
raise SystemExit("Invalid API key — shutting down agent")🔀 Error Decision Tree: Retry vs. Halt vs. Fallback
Invalid API key or insufficient permissions. Fix config, don't retry.
Exponential backoff + jitter. Max 5 attempts, then escalate.
Route to secondary provider. If all providers fail, queue for retry.
🔬 API Error Troubleshooter
Production API Error & Resilience LabHTTP & Retries
Simulate common API errors (401, 429, 503) and test self-healing mitigations in real time
primary = ChatOpenAI(model="gpt-4o") fallback = ChatAnthropic(model="claude-3-5-haiku") # Self-healing fallback model! resilient_model = primary.with_fallbacks([fallback])
Your application exceeded the Tokens Per Minute limit during a burst of concurrent tool calls.
API Error Handling Traps
Retrying a 401 error is like trying to swipe a cancelled credit card 5 times. Each attempt wastes latency and fails identically. Detect auth errors immediately, log them as critical, and halt the agent pending human intervention.
Retrying every 2 seconds with a fixed delay during a 429 storm causes synchronized thundering herds. Always use exponential backoff (2s, 4s, 8s...) plus a random ±30% jitter to stagger retries across concurrent requests.
Key Takeaways
- 1.Classify Before Responding: Not all errors are retryable. 401=halt, 429=backoff, 500=fallback, timeout=retry. Build your error handler as a decision tree, not a catch-all retry loop.
- 2.Use tenacity for Retries: tenacity's @retry decorator with wait_exponential handles the entire backoff/jitter/max-attempt logic in 4 lines. Don't roll your own retry loop.
- 3.with_fallbacks() is Free Insurance: LangChain's with_fallbacks() API gives you multi-provider failover with zero custom code. Primary → backup provider routing happens transparently on any server error.
Building a Chatbot Agent in LangGraph
Build a stateful multi-turn conversational agent using LangGraph checkpointers and thread_id — so your agent remembers what users said across sessions.