Measuring Agent Performance and Cost
Unmonitored agents are an open checkbook and a latency gamble. Production readiness requires tracking 4 telemetry pillars — latency, cost, resolution rate, and system efficiency.
🎯 By the end of this module, you will:
The 4 Production Telemetry Pillars
Moving from local scripts to production means answering 4 questions every run: How fast was it? How much did it cost? Did it succeed? Was it efficient?
# PILLAR 1: Latency profiling with callbacks
from langchain_core.callbacks import BaseCallbackHandler
import time
class LatencyTracker(BaseCallbackHandler):
def __init__(self):
self.node_timings = {}
self.ttft_recorded = False
def on_llm_start(self, serialized, prompts, **kwargs):
self.node_timings["llm_start"] = time.perf_counter()
def on_llm_new_token(self, token, **kwargs):
if not self.ttft_recorded:
# First token = TTFT
ttft_ms = (time.perf_counter() - self.node_timings["llm_start"]) * 1000
metrics.record("agent.ttft_ms", ttft_ms) # Target: < 400ms
self.ttft_recorded = True
def on_llm_end(self, response, **kwargs):
total_ms = (time.perf_counter() - self.node_timings["llm_start"]) * 1000
metrics.record("agent.llm.duration_ms", total_ms)
metrics.record("agent.llm.p95", total_ms, percentile=95)📊 Production Dashboard: What to Alert On
🔬 Telemetry & Unit Economics Studio
Production Agent Telemetry & Cost Studio
Model real-world agent operating expenses, node latencies, and token unit economics
Anthropic/OpenAI prompt cache reuse (cuts input token costs by up to 50%).
$6.438 per 1,000 tasks
1900 tokens consumed
User perceives instant response
Across all 4 graph nodes
Key Takeaways
- 1.4 Pillars = 4 Dashboards: Every production agent deployment needs 4 metrics dashboards: Latency (TTFT, P95), Cost (per-run, daily budget), Resolution Rate (success %, escalation %), Efficiency (tokens/action, loop drift count).
- 2.Use Callbacks, Not Manual Timing: LangChain's BaseCallbackHandler gives you on_llm_start / on_llm_end / on_tool_start hooks that fire automatically. Never add manual timing with time.time() scattered across business logic.
- 3.Alert Aggressively, Early: Set cost alerts at 2x your expected per-run cost, not at your monthly budget limit. A single runaway loop caught in 5 minutes costs $5. Caught after 1 day = catastrophic. Alerting is your circuit breaker.
You Have Mastered Core Implementation & Workflows
You now possess the full engineering toolkit to build production agents: LangGraph state graphs, Pydantic schemas, tool calling, streaming telemetry, async concurrency, prompt templates, defense-in-depth guardrails, automated pytest suites, and financial cost monitoring.
MCP, Deep Planning, Sub-graphs, Human-in-the-Loop