🎯 By the end of this module, you will:
Implementing Input Validation & Guardrails
Direct user input is inherently untrusted. Without safeguards, adversarial prompts can hijack your agent's persona, leak private API keys, or induce dangerous hallucinations. Defense-in-Depth quarantines inputs across three escalating validation layers.
# TIER 1: Regex rules — zero-cost, sub-millisecond circuit breakers
import re
# Adversarial keywords that bypass role instructions
BLOCKED_PATTERNS = [
r"ignore\s+(previous|all|prior)\s+instructions?",
r"system\s*prompt",
r"DAN\s*mode",
r"jailbreak",
r"forget\s+you\s+are",
r"act\s+as\s+if\s+you\s+have\s+no\s+rules",
]
MAX_INPUT_TOKENS = 2000 # ~8000 chars — enforce at tier 1
def tier1_regex_check(user_input: str) -> tuple[bool, str]:
# Length ceiling
if len(user_input) > MAX_INPUT_TOKENS * 4:
return False, "Input exceeds maximum allowed length"
# Keyword scan
input_lower = user_input.lower()
for pattern in BLOCKED_PATTERNS:
if re.search(pattern, input_lower):
return False, f"Blocked pattern detected: {pattern}"
return True, "APPROVED"🏰 Defense-in-Depth Pipeline Flow
Each stage either blocks the request (returns a safe rejection) or passes it to the next layer. Any stage can short-circuit the pipeline.
🔬 Defense Pipeline Studio
3-Layer Guardrails Pipeline Studio
Inspect how Defense-in-Depth blocks prompt attacks before they reach your primary agent
Token & Keyword Filter
Checks banned overrides ("ignore instructions", "chaosgpt") & length ceilings.
XML Delimiter Quarantine
Encapsulates user payload in <user_input> to prevent prompt escaping.
LLM Intent Classifier
Evaluates malicious intent, jailbreaks, and sensitive data leakage.
Guardrails Traps
Overly aggressive regex that blocks phrases like "ignore" or "system" in isolation will block legitimate requests like "Can you ignore the formatting and just give me the raw JSON?". Always use context-aware patterns that require surrounding adversarial context.
Input guardrails only prevent malicious inputs from reaching the LLM. They don't prevent the LLM from accidentally hallucinating or echoing sensitive data in its output. Always scan responses for API keys, internal paths, and PII before returning them to users.
Key Takeaways
- 1.Defense-in-Depth, Not Single Layer: No single guardrail catches everything. Layer fast regex (0.2ms), XML quarantine (0ms), and semantic judgment (50ms) so each layer catches what the previous one misses.
- 2.Guardrail Node Before Agent Node: In LangGraph, add your guardrail_node with a conditional router before your main agent_node. The conditional edge routes APPROVED inputs to the agent and FLAGGED inputs to a safe_reject_node.
- 3.Dual Direction — Input AND Output: Input guardrails prevent jailbreaks. Output guardrails prevent accidental secret leakage. You need both. An agent that never gets jailbroken but still echoes API keys in responses is a serious security risk.
Writing Test Cases for Agent Actions
Agents mix deterministic tools with stochastic LLM reasoning. Learn the Arrange-Act-Assert pattern for testing tool selection, parameter extraction, and error recovery.