Mod 2.16Writing Test Cases for Agent Actions
Level 2›Module 2.16
Level 2: Core Implementation & WorkflowsModule 2.16

Writing Test Cases for Agent Actions

Cases for Agent

Level 2 • Core Implementation & Workflows
Est. ~27 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Writing Test Cases for Agent Actions? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone

🎯 By the end of this module, you will:

Distinguish deterministic tests (exact assertions) from stochastic tests (behavioral checks)
Apply the Arrange-Act-Assert pattern to write clean, reproducible agent test cases
Mock LLMs and external APIs with pytest fixtures for fast, offline testing
Write 4 critical behavioral checkpoints: tool name, args, count, and error recovery
Agent Testing • Pytest Patterns

Writing Test Cases for Agent Actions

Traditional tests assert strict equality. Agent tests mix deterministic components (state reducers, schemas, tools) with stochastic LLM reasoning. Effective suites verify tool selection and parameter extraction under real-world conditions.

⚙️ Deterministic Testingassert x == y
# DETERMINISTIC: Exact assertions on pure functions & state
import pytest
from langchain_core.messages import HumanMessage, AIMessage

def test_add_messages_reducer_appends_correctly():
    """State reducer must APPEND, not overwrite, messages"""
    from langgraph.graph.message import add_messages
    
    # Arrange: Initial state with existing messages
    existing = [HumanMessage("Hello"), AIMessage("Hi there!")]
    new_messages = [HumanMessage("What is RAG?")]
    
    # Act: Apply reducer
    result = add_messages(existing, new_messages)
    
    # Assert: Deterministic — must have exactly 3 messages
    assert len(result) == 3, "Reducer should append, not replace"
    assert result[-1].content == "What is RAG?"

def test_pydantic_schema_rejects_invalid_input():
    """Schema validation must catch bad data before LLM call"""
    from pydantic import ValidationError
    
    with pytest.raises(ValidationError) as exc_info:
        AgentInput(query="", max_tokens=-1)  # Both fields invalid
    
    errors = exc_info.value.errors()
    assert len(errors) == 2  # Exactly 2 validation failures

🔺 The Agent Test Pyramid

Unit (70%)

Test tool functions, state reducers, schemas in isolation. Fast, cheap, deterministic. No LLM calls.

Integration (20%)

Test agent graph with mocked LLMs. Verify routing, error handling, and tool dispatch logic.

E2E (10%)

Full agent run with real LLM calls. Cover critical user journeys only — expensive but catches real regressions.

🔬 Pytest Suite Runner

Pytest Agent Action Test Runner

Run deterministic & stochastic test assertions against agent tools and states

Test Files & Cases0/4 Passed
Inspecting: Deterministic Test

test_weather_tool_parameter_extraction

~24ms
Input User Query:

"What is the current temperature in San Francisco?"

1. Arrange (Setup State & Mocks)
state = {"messages": [HumanMessage("What is the current temperature in San Francisco?")]}
2. Act (Invoke Agent Node)
response = agent.invoke(state)
3. Assert (Verify Tool Call & Arguments)
assert len(response["tool_calls"]) == 1
assert response["tool_calls"][0]["name"] == "get_weather"
assert response["tool_calls"][0]["args"]["location"] == "San Francisco"
Deterministic Tool AssertionsNoisy Prompt Resilience
Status: Ready
✍️ Instructor Note: "For stochastic LLM tests, run each test case 3–5 times and require it to pass ≥80% of runs. A test that occasionally fails due to model randomness is not a flaky test — it reveals a prompt that needs more few-shot examples to be reliable."
📌 Core Rule: Always include noisy, incomplete, and adversarial inputs in your test suite. In staging, developers test "Please fetch the stock price of AAPL". Real users send "aapl price rn pls", compound questions, and typos. Your agent must handle real-world messiness, not lab-clean inputs.

Agent Testing Traps

TRAP #1: Asserting LLM Response Text

Never assert that result['messages'][-1].content == "The weather in Tokyo is 22°C". LLM text is non-deterministic. Assert behavior: tool_calls[0]['name'] == 'get_weather'. Test what the agent DID, not what it SAID.

TRAP #2: Testing Clean Inputs Only

Test suites filled only with pristine, well-formatted queries give false confidence. Add noisy inputs ("aapl$$ price?"), empty strings, extremely long queries, and compound questions. Production agents meet all of these — your test suite must too.

Key Takeaways

  • 1.Dual Testing Strategy: Deterministic tests for all pure functions (reducers, schemas, tools) + stochastic behavioral tests for LLM reasoning. Together they give comprehensive agent coverage.
  • 2.AAA Pattern for Every Agent Test: Arrange (build realistic state), Act (invoke graph), Assert (verify 4 behavioral checkpoints: tool selected, args extracted, count correct, error recovery). Never skip a phase.
  • 3.Mock for Speed, E2E for Confidence: Run mocked integration tests on every commit (fast). Run real LLM E2E tests nightly or on release branches (slow, expensive). This balances development speed with production confidence.
Up Next • Module 2.17 (Capstone)

Measuring Agent Performance and Cost

The Level 2 capstone — track the 4 production telemetry pillars: TTFT, per-node latency waterfalls, token unit economics, and task completion rates.

Continue to Module 2.17