Mod 3.5Building a Basic RAG System for Agents
Level 3›Module 3.5
Level 3: Advanced Patterns & System DesignModule 3.5

Building a Basic RAG System for Agents

RAG System for

Level 3 • Advanced Patterns & System Design
Est. ~33 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Building a Basic RAG System for Agents? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone
Module 3.5 • Retrieval-Augmented Generation~20 min interactive

Building a Basic RAG System for Agents

Give your agents factual grounding. In this lesson, we build an Agentic RAG pipeline: transforming proprietary documents into vector embeddings and wrapping retrieval as a model-controlled tool with citation synthesis.

1. Core Architecture of Agentic RAG

Select a stage to inspect

RAG Pipeline Inspector

agentic_rag.py • ingestion
# 1. INGESTION & SEMANTIC CHUNKING
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter

# Load documents
loader = PyPDFLoader("company_policies.pdf")
raw_docs = loader.load()

# Split into semantically sound chunks with overlap
splitter = RecursiveCharacterTextSplitter(
    chunk_size=800,       # ~200 tokens per chunk
    chunk_overlap=120,     # Preserves context across chunk seams
    separators=["\n\n", "\n", " ", ""]
)
chunks = splitter.split_documents(raw_docs)
print(f"Divided {len(raw_docs)} pages into {len(chunks)} searchable chunks.")

2. Interactive RAG Pipeline Workbench

Chunking & Retrieval Inspector

Interactive Two-Phase RAG Pipeline Studio

Visualize how vector embeddings retrieve grounded facts to augment agent prompts

Top-K Chunks2 chunks
Controls how many matching chunks enter LLM prompt
Phase 1: Ingestion (Offline / Batch)
PDF / DocsChunksEmbeddingsVector Store

Documents are pre-chunked and embedded once into high-dimensional vectors.

Phase 2: Inference (Runtime Query)
User QuerySimilarity SearchAugmented Prompt

Query is converted into a vector with the same model to find nearest neighbor chunks.

📌 Agentic RAG vs Standard RAG: In traditional RAG, you retrieve on every single message whether needed or not. In Agentic RAG, the LLM decides WHEN to query the knowledge base and what query to search for. If the first search query returns poor results, an agent can rephrase and try again!

Common Engineering Traps

TRAP #1: Zero Chunk Overlap

Splitting text strictly on character boundaries without overlap chops critical sentences in half. If an important clause starts at character 790 and finishes at 820, neither chunk contains the full meaning, causing similarity scores to plummet. Always set chunk_overlap to 10-15% of chunk size.

TRAP #2: The 'Lost in the Middle' Syndrome

Retrieving 15 chunks (k=15) overwhelms the model with noisy context. Research shows LLMs pay attention to the beginning and end of retrieved passages, frequently ignoring chunks buried in the middle. Retrieve fewer, higher-quality chunks (k=3-5) and use a re-ranker.

Key Architectural Takeaways

  • 1.Two-Phase Lifecycle: RAG consists of offline Ingestion (load, chunk, embed, index) and online Retrieval (query, vector search, tool execution, grounded answer).
  • 2.Tool-Based Autonomy: Treating the vector retriever as a tool gives agents full autonomy to decide whether domain knowledge is required for a specific turn.
  • 3.Explicit Citation Grounding: System prompts should strictly instruct the model to ground answers in provided documents and refuse speculation.
Up Next • Module 3.6

Constructing Plan-and-Execute Agent Systems

Pure ReAct agents wander off track on complex 10-step workflows. Discover the Plan-and-Execute architecture: separating high-level strategic decomposition from tactical tool execution.

Continue to Module 3.6