Mod 2.15Implementing Input Validation and Guardrails
Level 2›Module 2.15
Level 2: Core Implementation & WorkflowsModule 2.15

Implementing Input Validation and Guardrails

Input Validation

Level 2 • Core Implementation & Workflows
Est. ~33 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Implementing Input Validation and Guardrails? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone

🎯 By the end of this module, you will:

Build a 3-tier Defense-in-Depth pipeline: Regex → XML Delimiters → Semantic Judge
Implement regex circuit breakers to catch common prompt injection patterns instantly
Wrap untrusted input in XML tags to quarantine it structurally from system instructions
Add output guardrails that scan LLM responses for leaked secrets and hallucinated fields
Input Validation • Defense-in-Depth Guardrails

Implementing Input Validation & Guardrails

Direct user input is inherently untrusted. Without safeguards, adversarial prompts can hijack your agent's persona, leak private API keys, or induce dangerous hallucinations. Defense-in-Depth quarantines inputs across three escalating validation layers.

Tier 1: Regex Rules: Instant Circuit Breakers<0.2ms • $0.00
# TIER 1: Regex rules — zero-cost, sub-millisecond circuit breakers
import re

# Adversarial keywords that bypass role instructions
BLOCKED_PATTERNS = [
    r"ignore\s+(previous|all|prior)\s+instructions?",
    r"system\s*prompt",
    r"DAN\s*mode",
    r"jailbreak",
    r"forget\s+you\s+are",
    r"act\s+as\s+if\s+you\s+have\s+no\s+rules",
]

MAX_INPUT_TOKENS = 2000  # ~8000 chars — enforce at tier 1

def tier1_regex_check(user_input: str) -> tuple[bool, str]:
    # Length ceiling
    if len(user_input) > MAX_INPUT_TOKENS * 4:
        return False, "Input exceeds maximum allowed length"
    
    # Keyword scan
    input_lower = user_input.lower()
    for pattern in BLOCKED_PATTERNS:
        if re.search(pattern, input_lower):
            return False, f"Blocked pattern detected: {pattern}"
    
    return True, "APPROVED"

🏰 Defense-in-Depth Pipeline Flow

User Input
→
Tier 1 Regex
→
Tier 2 XML Wrap
→
Tier 3 Judge LLM
→
Main Agent
→
Output Guardrail
→
Safe Response

Each stage either blocks the request (returns a safe rejection) or passes it to the next layer. Any stage can short-circuit the pipeline.

🔬 Defense Pipeline Studio

3-Layer Guardrails Pipeline Studio

Inspect how Defense-in-Depth blocks prompt attacks before they reach your primary agent

66 chars
Tier 1: Rules & Regex

Token & Keyword Filter

Checks banned overrides ("ignore instructions", "chaosgpt") & length ceilings.

Tier 2: Delimiters

XML Delimiter Quarantine

Encapsulates user payload in <user_input> to prevent prompt escaping.

Tier 3: Semantic Judge

LLM Intent Classifier

Evaluates malicious intent, jailbreaks, and sensitive data leakage.

✍️ Instructor Note: "Always place your guardrail node BEFORE the main agent node in LangGraph — not after. A guardrail that runs after the LLM has already processed adversarial input is useless. The damage is done before your safety check fires."
📌 Core Rule: Use cheap models (GPT-4o-mini, Claude Haiku) as your Tier 3 semantic judge — NOT your main model. The judge runs on EVERY request and must be cost-effective. Reserve expensive models for the main agent that only runs on APPROVED inputs.

Guardrails Traps

TRAP #1: Over-Blocking with Regex

Overly aggressive regex that blocks phrases like "ignore" or "system" in isolation will block legitimate requests like "Can you ignore the formatting and just give me the raw JSON?". Always use context-aware patterns that require surrounding adversarial context.

TRAP #2: Forgetting Output Guardrails

Input guardrails only prevent malicious inputs from reaching the LLM. They don't prevent the LLM from accidentally hallucinating or echoing sensitive data in its output. Always scan responses for API keys, internal paths, and PII before returning them to users.

Key Takeaways

  • 1.Defense-in-Depth, Not Single Layer: No single guardrail catches everything. Layer fast regex (0.2ms), XML quarantine (0ms), and semantic judgment (50ms) so each layer catches what the previous one misses.
  • 2.Guardrail Node Before Agent Node: In LangGraph, add your guardrail_node with a conditional router before your main agent_node. The conditional edge routes APPROVED inputs to the agent and FLAGGED inputs to a safe_reject_node.
  • 3.Dual Direction — Input AND Output: Input guardrails prevent jailbreaks. Output guardrails prevent accidental secret leakage. You need both. An agent that never gets jailbroken but still echoes API keys in responses is a serious security risk.
Up Next • Module 2.16

Writing Test Cases for Agent Actions

Agents mix deterministic tools with stochastic LLM reasoning. Learn the Arrange-Act-Assert pattern for testing tool selection, parameter extraction, and error recovery.

Continue to Module 2.16