Mod 2.9Debugging Agent Executions with Logging
Level 2›Module 2.9
Level 2: Core Implementation & WorkflowsModule 2.9

Debugging Agent Executions with Logging

Agent Executions

Level 2 • Core Implementation & Workflows
Est. ~30 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Debugging Agent Executions with Logging? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone

🎯 By the end of this module, you will:

Build a structured JSON logging system for every LLM call and tool dispatch
Attach run_id, session_id, and sequence numbers to enable full run replay
Calculate per-run token cost in real-time using model pricing tables
Distinguish debug vs. info vs. error log levels for production observability
Observability • Structured Logging

Debugging Agent Executions with Logging

Agents fail silently. Without structured logs, you can't tell which tool failed, what the LLM decided, or why it looped. Four observability pillars transform your black box into an open window.

1. LLM InteractionsPillar 1
# Pillar 1: Log all LLM interactions with full context
import json, logging
from datetime import datetime

logger = logging.getLogger("agent.llm")

def log_llm_call(prompt: str, response: str, token_usage: dict, latency_ms: float):
    logger.info(json.dumps({
        "timestamp": datetime.utcnow().isoformat(),
        "event": "llm_call",
        "run_id": current_run_id,
        "model": "gpt-4o-mini",
        "prompt_tokens": token_usage["prompt_tokens"],
        "completion_tokens": token_usage["completion_tokens"],
        "latency_ms": round(latency_ms, 2),
        "response_preview": response[:200]  # First 200 chars only
    }))

✈️ Mental Model: Agent Logs = Flight Data Recorder

❌ No Structured Logs

Like an airplane crash with no black box. You know it crashed, you have the wreckage — but no idea what happened in the last 30 seconds before impact. Impossible to diagnose or prevent.

✅ Structured JSON Logs

Like a flight data recorder that captures altitude, speed, and every control input every 10ms. When the agent fails, you replay its exact decisions, tool calls, and LLM responses step by step.

🔬 Structured Logging Debugger

Structured Logging & Observability StudioTrace Diagnostics

Compare noisy print statements against structured JSON traces with token, latency, and cost telemetry

Run ID: run_88fa1
Level:
INFO[call_model] llm_invoke
2026-09-28T13:20:01.102Z
{
  "timestamp": "2026-09-28T13:20:01.102Z",
  "level": "INFO",
  "run_id": "run_88fa1",
  "node": "call_model",
  "event": "llm_invoke",
  "tokens": {
    "prompt": 142,
    "completion": 24,
    "total": 166
  },
  "latency_ms": 310,
  "cost_usd": 0.00028
}
DEBUG[call_tools] tool_dispatch
2026-09-28T13:20:01.415Z
{
  "timestamp": "2026-09-28T13:20:01.415Z",
  "level": "DEBUG",
  "run_id": "run_88fa1",
  "node": "call_tools",
  "event": "tool_dispatch",
  "tool": "check_inventory",
  "args": {
    "sku": "SKU-901"
  },
  "latency_ms": 45
}
WARNING[call_tools] db_slow_query
2026-09-28T13:20:01.462Z
{
  "timestamp": "2026-09-28T13:20:01.462Z",
  "level": "WARNING",
  "run_id": "run_88fa1",
  "node": "call_tools",
  "event": "db_slow_query",
  "detail": "Warehouse DB replica latency above 40ms threshold"
}
INFO[call_model] final_synthesis
2026-09-28T13:20:01.780Z
{
  "timestamp": "2026-09-28T13:20:01.780Z",
  "level": "INFO",
  "run_id": "run_88fa1",
  "node": "call_model",
  "event": "final_synthesis",
  "tokens": {
    "prompt": 198,
    "completion": 32,
    "total": 230
  },
  "latency_ms": 285,
  "cost_usd": 0.00036
}
✍️ Instructor Note: "Never use print() for production observability. print() is synchronous, has no log level, and can't be filtered by severity. Use Python's logging module with a JSON formatter from day one."
💡 Mental Model: Think of run_id like a tracking number on a courier package. Every log event for one agent run carries that same tracking number so you can reconstruct the full journey from start to delivery (or failure!).
📌 Core Rule: Log the first 200–300 chars of every prompt and response. Logging the full text is expensive and often hits PII regulations. Log IDs, token counts, and truncated previews — then retrieve full text from LangSmith only when debugging a specific failure.

Agent Logging Traps

TRAP #1: Logging Full Prompts to stdout

Logging the complete system prompt + user message + history to stdout will leak PII into server logs, violate GDPR, and fill your disk in hours. Always truncate, redact, or hash sensitive fields before logging.

TRAP #2: No Sequence Numbers

Without monotonic sequence numbers, concurrent async agent runs produce interleaved logs that are impossible to reconstruct in order. Always add a seq: int counter per run to enable deterministic replay ordering.

Key Takeaways

  • 1.Structured JSON Logs: Every log event should be machine-parsable JSON, not human-readable text blobs. This enables filter, aggregate, and alert tooling in Datadog, Grafana, or custom dashboards.
  • 2.run_id is Sacred: Every log event for one agent execution must share the same run_id. This is your forensic thread for reconstructing exactly what happened in any failure.
  • 3.Log Cost in Real-Time: Calculate and log token cost per run using your model's pricing table. Alert when a single run exceeds a cost ceiling — runaway loops can burn $50 in minutes without guardrails.
Up Next • Module 2.10

Troubleshooting Common LLM API Issues

Build self-healing agents with exponential backoff, multi-provider fallback routing, and structured error categorization for 401, 429, and 500 errors.

Continue to Module 2.10