Mod 2.17Measuring Agent Performance and Cost
Level 2›Module 2.17
Level 2: Core Implementation & WorkflowsModule 2.17

Measuring Agent Performance and Cost

Agent Performance

Level 2 • Core Implementation & Workflows
Est. ~27 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Measuring Agent Performance and Cost? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone
Level 2 Capstone • Module 2.17

Measuring Agent Performance and Cost

Unmonitored agents are an open checkbook and a latency gamble. Production readiness requires tracking 4 telemetry pillars — latency, cost, resolution rate, and system efficiency.

🎯 By the end of this module, you will:

Implement latency callbacks that track TTFT and P95 tail latency per LangGraph node
Calculate real-time token cost using model pricing tables and alert on cost spikes
Define and track task resolution rate vs. escalation rate for business-level SLAs
Detect loop drift by flagging agents that repeat the same tool call 3+ times
Production Telemetry • Unit Economics

The 4 Production Telemetry Pillars

Moving from local scripts to production means answering 4 questions every run: How fast was it? How much did it cost? Did it succeed? Was it efficient?

1. Latency Profile: TTFT & P95 TailTarget: TTFT < 400ms
# PILLAR 1: Latency profiling with callbacks
from langchain_core.callbacks import BaseCallbackHandler
import time

class LatencyTracker(BaseCallbackHandler):
    def __init__(self):
        self.node_timings = {}
        self.ttft_recorded = False
    
    def on_llm_start(self, serialized, prompts, **kwargs):
        self.node_timings["llm_start"] = time.perf_counter()
    
    def on_llm_new_token(self, token, **kwargs):
        if not self.ttft_recorded:
            # First token = TTFT
            ttft_ms = (time.perf_counter() - self.node_timings["llm_start"]) * 1000
            metrics.record("agent.ttft_ms", ttft_ms)  # Target: < 400ms
            self.ttft_recorded = True
    
    def on_llm_end(self, response, **kwargs):
        total_ms = (time.perf_counter() - self.node_timings["llm_start"]) * 1000
        metrics.record("agent.llm.duration_ms", total_ms)
        metrics.record("agent.llm.p95", total_ms, percentile=95)

📊 Production Dashboard: What to Alert On

TTFT
287ms
target: < 400ms
Cost/Run
$0.0003
target: < $0.05
Resolution
87.3%
target: > 85%
Loop Drift
0 runs
target: 0 detected

🔬 Telemetry & Unit Economics Studio

Production Agent Telemetry & Cost Studio

Model real-world agent operating expenses, node latencies, and token unit economics

Live Financial Model
Daily Requests5,000 / day
150,000 requests/month
Input Tokens1500
Output Tokens400
Prompt Caching

Anthropic/OpenAI prompt cache reuse (cuts input token costs by up to 50%).

Monthly LLM API Bill
$965.63/mo

$6.438 per 1,000 tasks

Unit Cost Per Resolution
$0.644¢/task

1900 tokens consumed

Time to First Token (TTFT)
380msP50

User perceives instant response

End-to-End Latency
1.73stotal

Across all 4 graph nodes

Graph Node Latency Waterfall BreakdownTotal Execution: 1734ms
guardrail_verification
14ms1% of total
planner_reasoning_node
810ms47% of total
tool_execution_api
280ms16% of total
synthesis_and_formatting
630ms36% of total
✍️ Instructor Note: "Don't wait until your bill arrives to discover a cost problem. Set up automated alerts when agent.cost_usd exceeds $0.05 per run AND when 7-day average crosses your budget threshold. One runaway loop can burn $50 in under a minute."
📌 Core Rule: Track input_tokens vs. output_tokens separately. Prompt caching reduces input token costs by 50–80%. If your cache hit rate is low, your prompts are changing too frequently between requests — optimize for cache stability to cut costs dramatically.

Key Takeaways

  • 1.4 Pillars = 4 Dashboards: Every production agent deployment needs 4 metrics dashboards: Latency (TTFT, P95), Cost (per-run, daily budget), Resolution Rate (success %, escalation %), Efficiency (tokens/action, loop drift count).
  • 2.Use Callbacks, Not Manual Timing: LangChain's BaseCallbackHandler gives you on_llm_start / on_llm_end / on_tool_start hooks that fire automatically. Never add manual timing with time.time() scattered across business logic.
  • 3.Alert Aggressively, Early: Set cost alerts at 2x your expected per-run cost, not at your monthly budget limit. A single runaway loop caught in 5 minutes costs $5. Caught after 1 day = catastrophic. Alerting is your circuit breaker.
🎉 Level 2 Completed! (17 of 17 Modules)

You Have Mastered Core Implementation & Workflows

You now possess the full engineering toolkit to build production agents: LangGraph state graphs, Pydantic schemas, tool calling, streaming telemetry, async concurrency, prompt templates, defense-in-depth guardrails, automated pytest suites, and financial cost monitoring.

LangGraph State Graphs
Pydantic Schemas
Streaming Output
Async Concurrency
Prompt Templates
Input Guardrails
Pytest Test Suites
Cost Telemetry
Next Milestone: Level 3 • Advanced Patterns & System Design
MCP, Deep Planning, Sub-graphs, Human-in-the-Loop
Explore Level 3
🔒 Certificate Locked • 0/17 Lessons DoneComplete All Lessons to Unlock

Level 2: Core Implementation & Workflows

To generate the Level 2 certificate, you must complete all 17 lessons in this level. You currently have completed 0 of 17 lessons (17 remaining).

Level 2 Progress0%