Mod 3.11Implementing Semantic Memory with Vector Stores
Level 3›Module 3.11
Level 3: Advanced Patterns & System DesignModule 3.11

Implementing Semantic Memory with Vector Stores

Semantic Memory

Level 3 • Advanced Patterns & System Design
Est. ~24 mins
5 Key Topics
🎁 Free Learner Perk

Unlock Verified Certificate & Daily Streak Tracker

Ready to master Implementing Semantic Memory with Vector Stores? Enable cloud sync to record your daily streak 🔥 and earn your Informational Completion Badge for your study milestones.

Day 1 Streak ActiveFree Completion BadgeSync Laptop & Phone
Module 3.11 • Long-Term Cognitive Memory~20 min interactive

Implementing Semantic Memory with Vector Stores

Short-term memory tracks the current session; Semantic Memory preserves user context across months. Learn how to extract episodic facts, store them in vector databases, and inject personalized preferences dynamically into agent prompts.

1. Core Mechanics of Semantic Memory

Select a memory mechanism to inspect

Semantic Memory Inspector

semantic_memory.py • concept
# 1. DUAL MEMORY ARCHITECTURE IN LANGGRAPH
# Short-Term Memory: Checkpointer (Session-bound, raw messages)
# Long-Term Memory: Vector Store (Cross-session, distilled facts)

from langchain_openai import OpenAIEmbeddings
from langchain_community.vectorstores import Chroma

# Long-term semantic memory store
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
long_term_memory = Chroma(
    collection_name="user_semantic_memories",
    embedding_function=embeddings
)

2. Interactive Semantic Memory Studio

Vector Preference Sandbox

Try a preset query:

📌 Browser History vs Human Brain: Checkpointers are like your browser history — 500 lines of exact logs. Semantic vector memory is like your brain — you don't remember the exact sentence a colleague said 3 months ago, but you remember they prefer dark mode and TypeScript!

Common Engineering Traps

TRAP #1: Unfiltered Cross-Tenant Vector Bleed

Querying a vector database without a strict metadata filter (filter={ "user_id": uid }) allows User A to retrieve private memories and API keys stored by User B if semantic similarity is high. Always enforce user-scoped metadata partitioning.

TRAP #2: Storing Ephemeral Noise as Permanent Facts

Embedding raw chat messages like "I will be back in 5 minutes" clutters vector indices with useless junk. Always use an extraction model to filter for enduring facts (preferences, roles, long-term project parameters).

Key Architectural Takeaways

  • 1.Dual Memory Stratum: Checkpointers provide session auditability, while vector stores provide cross-session personalization.
  • 2.Fact Distillation: Extract only high-signal permanent preferences, ignoring transient chit-chat.
  • 3.Multi-Tenant Privacy: Mandatory metadata filtering guarantees user memory boundaries are completely airtight.
Up Next • Module 3.12

Deploying Agents with FastAPI

Take your LangGraph agents to production. Learn how to wrap agents in async FastAPI endpoints with Server-Sent Events (SSE) streaming and background task queues.

Continue to Module 3.12