Mod 1.10Enhancing Agents with Retrieval Augmented
Level 1›Module 1.10
Level 1: Foundations & ArchitectureModule 1.10

Enhancing Agents with Retrieval Augmented

Agents with

Level 1 • Foundations & Architecture
Est. ~33 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Enhancing Agents with Retrieval Augmented? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone

🎯 By the end of this module, you will:

Understand the 2-phase RAG lifecycle: Offline Ingestion vs Online Inference
Implement optimal chunking (256-512 tokens with overlap) to avoid boundary slicing
Architect RAG as an autonomous agent tool vs a rigid deterministic workflow node
Debunk the "RAG eliminates hallucinations" myth and apply context validation guards
Architecture Patterns • Knowledge Grounding

Retrieval-Augmented Generation (RAG)

RAG bridges the gap between frozen model weights and dynamic, private enterprise documents without costly fine-tuning.

Phase Mechanics: 1. Chunking
Semantic Segmentation

Raw docs are segmented into 256-512 token chunks with 10-20% overlap so crucial conditions are never sliced in half.

# 1. Semantic Chunking with Overlap
from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=400,
    chunk_overlap=60,
    separators=["\n\n", "\n", ". ", " "]
)
chunks = splitter.split_text(raw_enterprise_policy)
💡

Mental Model: Closed-Book Cramming vs. The Open-Book Exam

• Closed-Book Exam (Standalone LLM): A student tries to memorize 10,000 pages of corporate bylaws. When asked about a specific 2026 refund formula, their memory blurs and they invent a convincing fake answer (Hallucination).
• Open-Book Exam with a Library (RAG): The student doesn't memorize every nuance. When asked the question, they walk to the catalog, pull the exact SLA handbook, and quote the exact sentence with page citations.

✍️ Instructor Note: "RAG transforms every query into an open-book exam, grounding the agent in verified corporate evidence!"
Agentic Evolution • Node vs. Tool

RAG as a Workflow Node vs. Autonomous Agent Tool

Compare fixed pipeline execution (where retrieval is mandatory on every step) against agentic tool calling (where the agent autonomously decides if, when, and what to query).

Interactive Architectural Paradigm: RAG as a Node vs. RAG as a Tool

Compare deterministic pipeline retrieval (Node) against dynamic agentic self-directed search (Tool).

Deterministic Pipeline (Retrieval Always Happens)
1. User Query Ingestion Node"What is our refund SLA?"
↓ Mandatory Transition
Fixed Graph Node
Vector Knowledge Base Retrieval

Query vector DB ➔ Ingest Top-3 policy chunks

↓ Mandatory Transition with Injected Context
3. Final Response Generator NodeStrictly Grounded Answer
Architectural Analysis

When to use RAG as a Node:

• Mandatory Domain Pipelines: Customer support chatbots where every single inquiry MUST check the official FAQ before replying.

• Predictable SLAs & Cost: Exactly one retrieval vector search per request. Zero risk of the agent looping or skipping search.

Critical Truth: RAG Does NOT Solve Hallucinations

RAG merely shifts the problem from "making things up from weights" to "retrieval quality and context adherence". If your vector database returns irrelevant or truncated chunks, the model will hallucinate around the noise!

Interactive Laboratory • Ingestion & Noise

RAG Lifecycle & Noise Impact Studio

Experiment with document chunk size in the Ingestion phase, and observe how noisy vector retrieval triggers hallucinations in the Inference phase.

RAG Lifecycle & Retrieval Quality Laboratory

Test the Ingestion Phase (Chunking & Embedding) and the Inference Phase (Vector Search & Noise Impact).

Raw Enterprise Policy Document

Raw unstructured documents must be segmented into chunks before they can be converted into numerical vectors.

"FinCorp Enterprise SLA (Rev 2026.2): Tier 1 enterprise clients receive 99.99% uptime guarantees. Scheduled maintenance occurs on the first Sunday of each month between 02:00 and 04:00 UTC. In the event of an unplanned outage exceeding 15 minutes, enterprise accounts are eligible for a 15% billing credit upon submitting a formal claim within 30 days."
Chunk Window Size:256 tokens
💡 Chunking Rule: Too small (e.g. 50 tokens) loses sentence context; too large (e.g. 2000 tokens) dilutes semantic embedding specificity.
Compiled Vector Database Payloads (Pinecone Index)
Chunk_ID: chunk_sla_0011536-dim vector embedded

"Tier 1 enterprise clients receive 99.99% uptime guarantees. Scheduled maintenance occurs on the first Sunday..."

Chunk_ID: chunk_sla_0021536-dim vector embedded

"...unplanned outage exceeding 15 minutes, enterprise accounts are eligible for a 15% billing credit upon formal claim..."

agentic_rag_tool.py
# Agentic RAG: Giving the Agent a Knowledge Search Tool
from langchain_core.tools import tool

@tool
def search_corporate_kb(query: str, domain: str = "all") -> str:
    """Search internal enterprise knowledge base for policies, SLAs, and technical manuals."""
    # 1. Embed incoming search query
    query_vector = embed_model.embed_query(query)
    
    # 2. Query Vector DB with metadata filtering
    results = vector_db.query(
        vector=query_vector,
        filter={"domain": domain} if domain != "all" else None,
        top_k=2
    )
    
    # 3. Format grounded passages with citation IDs
    return "\n".join([f"[{r.id}] {r.metadata['title']}: {r.text}" for r in results])

# Agent can now autonomously decide IF and WHEN to call search_corporate_kb!
agent = initialize_agent(tools=[search_corporate_kb, calculate_refund])

Architectural Traps in RAG Systems

TRAP #1: The Hallucination Immunity Myth

RAG does NOT eliminate hallucinations; it shifts the problem to retrieval quality. If your vector database returns irrelevant chunks, the model will hallucinate around the noise to please the prompt!

TRAP #2: Zero-Overlap Chunk Slicing

Splitting text strictly by token count without sentence overlap risks slicing critical conditions in half (e.g. "Except in cases of..." moves to chunk 2, inverting the legal meaning in chunk 1). Always include 10-20% chunk overlap!

Key Architectural Takeaways

  • 1.Decouple Ingestion from Inference: Ingestion (chunking, embedding, indexing) happens offline; Inference (vector similarity, grounding) happens live.
  • 2.Use RAG as a Tool for Dynamic Agents: Give the agent an explicit search tool so it can query multiple databases conditionally rather than forcing rigid pipeline retrievals.
  • 3.Enforce Strict Relevance Filters: Discard retrieved chunks with low cosine similarity scores (<0.80) to prevent the model from hallucinating plausible nonsense.
Up Next • Module 1.11

Evaluating AI Agent Frameworks

Compare LangGraph, CrewAI, AutoGen, LlamaIndex, and the OpenAI Agents SDK to choose the exact right framework for your production architecture.

Continue to Module 1.11