🎯 By the end of this module, you will:
Writing Test Cases for Agent Actions
Traditional tests assert strict equality. Agent tests mix deterministic components (state reducers, schemas, tools) with stochastic LLM reasoning. Effective suites verify tool selection and parameter extraction under real-world conditions.
# DETERMINISTIC: Exact assertions on pure functions & state
import pytest
from langchain_core.messages import HumanMessage, AIMessage
def test_add_messages_reducer_appends_correctly():
"""State reducer must APPEND, not overwrite, messages"""
from langgraph.graph.message import add_messages
# Arrange: Initial state with existing messages
existing = [HumanMessage("Hello"), AIMessage("Hi there!")]
new_messages = [HumanMessage("What is RAG?")]
# Act: Apply reducer
result = add_messages(existing, new_messages)
# Assert: Deterministic — must have exactly 3 messages
assert len(result) == 3, "Reducer should append, not replace"
assert result[-1].content == "What is RAG?"
def test_pydantic_schema_rejects_invalid_input():
"""Schema validation must catch bad data before LLM call"""
from pydantic import ValidationError
with pytest.raises(ValidationError) as exc_info:
AgentInput(query="", max_tokens=-1) # Both fields invalid
errors = exc_info.value.errors()
assert len(errors) == 2 # Exactly 2 validation failures🔺 The Agent Test Pyramid
Test tool functions, state reducers, schemas in isolation. Fast, cheap, deterministic. No LLM calls.
Test agent graph with mocked LLMs. Verify routing, error handling, and tool dispatch logic.
Full agent run with real LLM calls. Cover critical user journeys only — expensive but catches real regressions.
🔬 Pytest Suite Runner
Pytest Agent Action Test Runner
Run deterministic & stochastic test assertions against agent tools and states
test_weather_tool_parameter_extraction
"What is the current temperature in San Francisco?"
state = {"messages": [HumanMessage("What is the current temperature in San Francisco?")]}response = agent.invoke(state)
assert len(response["tool_calls"]) == 1 assert response["tool_calls"][0]["name"] == "get_weather" assert response["tool_calls"][0]["args"]["location"] == "San Francisco"
Agent Testing Traps
Never assert that result['messages'][-1].content == "The weather in Tokyo is 22°C". LLM text is non-deterministic. Assert behavior: tool_calls[0]['name'] == 'get_weather'. Test what the agent DID, not what it SAID.
Test suites filled only with pristine, well-formatted queries give false confidence. Add noisy inputs ("aapl$$ price?"), empty strings, extremely long queries, and compound questions. Production agents meet all of these — your test suite must too.
Key Takeaways
- 1.Dual Testing Strategy: Deterministic tests for all pure functions (reducers, schemas, tools) + stochastic behavioral tests for LLM reasoning. Together they give comprehensive agent coverage.
- 2.AAA Pattern for Every Agent Test: Arrange (build realistic state), Act (invoke graph), Assert (verify 4 behavioral checkpoints: tool selected, args extracted, count correct, error recovery). Never skip a phase.
- 3.Mock for Speed, E2E for Confidence: Run mocked integration tests on every commit (fast). Run real LLM E2E tests nightly or on release branches (slow, expensive). This balances development speed with production confidence.
Measuring Agent Performance and Cost
The Level 2 capstone — track the 4 production telemetry pillars: TTFT, per-node latency waterfalls, token unit economics, and task completion rates.