Building a Basic RAG System for Agents
Give your agents factual grounding. In this lesson, we build an Agentic RAG pipeline: transforming proprietary documents into vector embeddings and wrapping retrieval as a model-controlled tool with citation synthesis.
1. Core Architecture of Agentic RAG
Select a stage to inspectRAG Pipeline Inspector
# 1. INGESTION & SEMANTIC CHUNKING
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
# Load documents
loader = PyPDFLoader("company_policies.pdf")
raw_docs = loader.load()
# Split into semantically sound chunks with overlap
splitter = RecursiveCharacterTextSplitter(
chunk_size=800, # ~200 tokens per chunk
chunk_overlap=120, # Preserves context across chunk seams
separators=["\n\n", "\n", " ", ""]
)
chunks = splitter.split_documents(raw_docs)
print(f"Divided {len(raw_docs)} pages into {len(chunks)} searchable chunks.")2. Interactive RAG Pipeline Workbench
Chunking & Retrieval InspectorInteractive Two-Phase RAG Pipeline Studio
Visualize how vector embeddings retrieve grounded facts to augment agent prompts
Documents are pre-chunked and embedded once into high-dimensional vectors.
Query is converted into a vector with the same model to find nearest neighbor chunks.
Common Engineering Traps
Splitting text strictly on character boundaries without overlap chops critical sentences in half. If an important clause starts at character 790 and finishes at 820, neither chunk contains the full meaning, causing similarity scores to plummet. Always set chunk_overlap to 10-15% of chunk size.
Retrieving 15 chunks (k=15) overwhelms the model with noisy context. Research shows LLMs pay attention to the beginning and end of retrieved passages, frequently ignoring chunks buried in the middle. Retrieve fewer, higher-quality chunks (k=3-5) and use a re-ranker.
Key Architectural Takeaways
- 1.Two-Phase Lifecycle: RAG consists of offline Ingestion (load, chunk, embed, index) and online Retrieval (query, vector search, tool execution, grounded answer).
- 2.Tool-Based Autonomy: Treating the vector retriever as a tool gives agents full autonomy to decide whether domain knowledge is required for a specific turn.
- 3.Explicit Citation Grounding: System prompts should strictly instruct the model to ground answers in provided documents and refuse speculation.
Constructing Plan-and-Execute Agent Systems
Pure ReAct agents wander off track on complex 10-step workflows. Discover the Plan-and-Execute architecture: separating high-level strategic decomposition from tactical tool execution.