Building Robust Tools with Validation and Logging
Most agent failures in production aren't LLM failures — they are tool failures. Master enterprise tool engineering: Pydantic pre-validation, structured error contracts, exponential backoff retries, and correlation telemetry.
1. The 4 Pillars of a Production Tool
Select a pillar to inspectRobust Tool Implementation
# 1. PYDANTIC INPUT VALIDATION SCHEMA
from pydantic import BaseModel, Field
class SearchInput(BaseModel):
query: str = Field(..., min_length=2, max_length=200, description="Search query string")
max_results: int = Field(default=5, ge=1, le=20, description="Number of results between 1 and 20")
domain_filter: str | None = Field(default=None, description="Optional domain restriction")
# If agent passes max_results=999, Pydantic intercepts and returns a clean error
# before any network call occurs!2. Interactive Robust Tool Studio
Live Tool Validation SandboxSelect Scenario
Tool Input
{
"query": "Apple revenue 2024",
"max_results": 5
}Structured Log Output
Common Engineering Traps
Retrying on 400 Bad Request or 401 Unauthorized is futile — the server will never accept the request without changed credentials. Only retry on transient 5xx or network connection timeouts.
Returning None or an empty string "" when a database tool fails leaves the LLM blind. The agent assumes the database is empty rather than realizing the query syntax was malformed. Always return explicit error objects.
Key Architectural Takeaways
- 1.Pydantic Pre-Validation: Intercept bad inputs locally to prevent wasting external API quota and tripping rate limits.
- 2.Actionable Error Contracts: Format error outputs as JSON with explicit guidance to trigger model self-correction.
- 3.Correlation Tracing: Pass request IDs through every tool call to debug multi-step autonomous chains in production.
Building Multi-Agent Supervisor Systems
When a single agent has too many tools, its reasoning degrades. Learn how to architect a Multi-Agent Supervisor that routes tasks to specialized worker agents.