Mod 4.9Optimizing Prompts and Tool Selection
Level 4›Module 4.9
Level 4: Production, Scaling & OptimizationModule 4.9

Optimizing Prompts and Tool Selection

Prompts and Tool

Level 4 • Production, Scaling & Optimization
Est. ~33 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Optimizing Prompts and Tool Selection? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone
Module 4.9 • Production, Scaling & Optimization~25 min interactive

Optimizing Prompts and Tool Selection

When an agent picks the wrong tool or stops halfway, engineers reflexively blame the LLM. In reality, 90% of agent routing failures stem from vague tool docstrings and ambiguous system prompts. Master the art of deterministic prompt engineering and tool boundary definition.

1. Core Pillars of Tool Prompt Optimization

Select a design pillar

Prompt & Tool Routing Inspector

prompt_tuning.py • role_framing
# 1. PRODUCTION SYSTEM PROMPT WITH PERSISTENCE DIRECTIVE
SYSTEM_PROMPT = """You are an expert equity research intelligence agent.
OPERATIONAL SCOPE:
- Synthesize verifiable financial facts from SEC filings and live feeds.
- CRITICAL PERSISTENCE RULE: Do not terminate with partial insights. If an initial
  search query produces ambiguous numbers, inspect secondary sources until completely verified.
- Do not stop until all user questions have citation-backed financial metrics."""

2. Interactive Prompt Optimizer Lab

A/B Test Prompt Revisions Live

System Prompt

You are a helpful financial assistant. Answer questions and use tools when needed.

Problems with this prompt:

No tool selection criteria — agent guesses which tool to call
No output format specified — inconsistent responses
No persistence instruction — agent stops too early
No guardrails — agent may hallucinate rather than use tools

Resulting Tool Calls

web_search()

Called for a price query that needed live_price() — wrong tool

calculator()

Skipped — agent guessed the answer instead

END (premature)

Stopped after 1 step even though task was incomplete

Agent Performance Score
38%
📌 Production Insight: The LLM does not execute your Python function — it reads your docstring and tries to guess how to populate JSON arguments. If your docstring is "Searches data", the agent will hallucinate parameters every time. Write docstrings with the exact same rigor you give to customer-facing REST API specs!

Common Engineering Traps

TRAP #1: The 2-3 Query Testing Fallacy

A developer tweaks a prompt, runs 2 test queries in terminal, sees good output, and pushes to production. Two days later, 15% of edge-case user queries break. Never deploy a prompt change without benchmarking against a golden dataset of at least 30 representative test questions.

TRAP #2: Overlapping Tool Semantic Space

Giving an agent both search_web and fetch_online_articles without clear boundary criteria leads to erratic coin-flip routing. When tool scopes overlap, merge them into a single parameterized tool or explicitly state mutually exclusive routing rules.

Key Architectural Takeaways

  • 1.Explicit Negative Boundaries: Include 'WHEN NOT TO USE' clauses in tool docstrings to actively reject incorrect tools.
  • 2.Strict Pydantic Validation: Constrain parameters with regex, bounds, and enums so syntax errors are caught before runtime.
  • 3.Persistence Directives: Mandate in the system prompt that the agent must not terminate until tasks are thoroughly verified.
Up Next • Module 4.10

Managing Context Windows Effectively

Large context windows are not infinite trash cans. Learn how to prevent the 'Lost in the Middle' degradation with windowing, summarization nodes, and selective state pruning.

Continue to Module 4.10