🎯 By the end of this module, you will:
Debugging Agent Executions with Logging
Agents fail silently. Without structured logs, you can't tell which tool failed, what the LLM decided, or why it looped. Four observability pillars transform your black box into an open window.
# Pillar 1: Log all LLM interactions with full context
import json, logging
from datetime import datetime
logger = logging.getLogger("agent.llm")
def log_llm_call(prompt: str, response: str, token_usage: dict, latency_ms: float):
logger.info(json.dumps({
"timestamp": datetime.utcnow().isoformat(),
"event": "llm_call",
"run_id": current_run_id,
"model": "gpt-4o-mini",
"prompt_tokens": token_usage["prompt_tokens"],
"completion_tokens": token_usage["completion_tokens"],
"latency_ms": round(latency_ms, 2),
"response_preview": response[:200] # First 200 chars only
}))✈️ Mental Model: Agent Logs = Flight Data Recorder
Like an airplane crash with no black box. You know it crashed, you have the wreckage — but no idea what happened in the last 30 seconds before impact. Impossible to diagnose or prevent.
Like a flight data recorder that captures altitude, speed, and every control input every 10ms. When the agent fails, you replay its exact decisions, tool calls, and LLM responses step by step.
🔬 Structured Logging Debugger
Structured Logging & Observability StudioTrace Diagnostics
Compare noisy print statements against structured JSON traces with token, latency, and cost telemetry
run_88fa1{
"timestamp": "2026-09-28T13:20:01.102Z",
"level": "INFO",
"run_id": "run_88fa1",
"node": "call_model",
"event": "llm_invoke",
"tokens": {
"prompt": 142,
"completion": 24,
"total": 166
},
"latency_ms": 310,
"cost_usd": 0.00028
}{
"timestamp": "2026-09-28T13:20:01.415Z",
"level": "DEBUG",
"run_id": "run_88fa1",
"node": "call_tools",
"event": "tool_dispatch",
"tool": "check_inventory",
"args": {
"sku": "SKU-901"
},
"latency_ms": 45
}{
"timestamp": "2026-09-28T13:20:01.462Z",
"level": "WARNING",
"run_id": "run_88fa1",
"node": "call_tools",
"event": "db_slow_query",
"detail": "Warehouse DB replica latency above 40ms threshold"
}{
"timestamp": "2026-09-28T13:20:01.780Z",
"level": "INFO",
"run_id": "run_88fa1",
"node": "call_model",
"event": "final_synthesis",
"tokens": {
"prompt": 198,
"completion": 32,
"total": 230
},
"latency_ms": 285,
"cost_usd": 0.00036
}Agent Logging Traps
Logging the complete system prompt + user message + history to stdout will leak PII into server logs, violate GDPR, and fill your disk in hours. Always truncate, redact, or hash sensitive fields before logging.
Without monotonic sequence numbers, concurrent async agent runs produce interleaved logs that are impossible to reconstruct in order. Always add a seq: int counter per run to enable deterministic replay ordering.
Key Takeaways
- 1.Structured JSON Logs: Every log event should be machine-parsable JSON, not human-readable text blobs. This enables filter, aggregate, and alert tooling in Datadog, Grafana, or custom dashboards.
- 2.run_id is Sacred: Every log event for one agent execution must share the same run_id. This is your forensic thread for reconstructing exactly what happened in any failure.
- 3.Log Cost in Real-Time: Calculate and log token cost per run using your model's pricing table. Alert when a single run exceeds a cost ceiling — runaway loops can burn $50 in minutes without guardrails.
Troubleshooting Common LLM API Issues
Build self-healing agents with exponential backoff, multi-provider fallback routing, and structured error categorization for 401, 429, and 500 errors.