Developing Error Handling and Recovery Pathways
In production, systems fail constantly. In this lesson, we build self-healing agent recovery pathways: classifying failure modes, tripping graph circuit breakers, and cascading to backup model providers.
1. Core Mechanics of Resilient Error Handling
Select a resilience strategyFault Tolerance Inspector
# 1. CLASSIFYING AGENT EXCEPTIONS
class AgentErrorClassifier:
@staticmethod
def classify(status_code: int) -> str:
if status_code in (429, 502, 503, 504):
return "TRANSIENT" # Retry with exponential backoff
elif status_code in (400, 401, 403, 404):
return "PERMANENT" # Terminate, do NOT waste retry tokens
return "UNKNOWN"2. Interactive Error Recovery Simulator
Outage & Circuit Breaker SandboxScenario
Search API returns 503 (server overloaded). This is temporary — it'll recover.
Common Engineering Traps
When an API key expires, naive agents retry 10 times with exponential backoff, delaying user error notification by 60 seconds. Only retry transient 5xx/429 errors. Fail permanent errors immediately.
If an error recovery node attempts to prompt the LLM to fix a bug, but the LLM repeats the same bug, the agent loops infinitely. Enforce a strict max_recovery_attempts = 2 counter in State.
Key Architectural Takeaways
- 1.Classify Before Action: Categorize errors into Transient, Permanent, Partial, and Cascading before selecting a recovery strategy.
- 2.Circuit Breakers Save Dependencies: Trip the breaker after 3 failures to protect downstream databases from cascading thundering herds.
- 3.Multi-Provider Redundancy: Deploy with_fallbacks across diverse model providers to achieve five-nines availability.
Using Time Travel for State Branching
What if you could rewind an agent run, change a single tool result, and fork reality? Master LangGraph Time Travel debugging, checkpoint diffing, and speculative branching.