Implementing Semantic Memory with Vector Stores
Short-term memory tracks the current session; Semantic Memory preserves user context across months. Learn how to extract episodic facts, store them in vector databases, and inject personalized preferences dynamically into agent prompts.
1. Core Mechanics of Semantic Memory
Select a memory mechanism to inspectSemantic Memory Inspector
# 1. DUAL MEMORY ARCHITECTURE IN LANGGRAPH
# Short-Term Memory: Checkpointer (Session-bound, raw messages)
# Long-Term Memory: Vector Store (Cross-session, distilled facts)
from langchain_openai import OpenAIEmbeddings
from langchain_community.vectorstores import Chroma
# Long-term semantic memory store
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
long_term_memory = Chroma(
collection_name="user_semantic_memories",
embedding_function=embeddings
)2. Interactive Semantic Memory Studio
Vector Preference SandboxTry a preset query:
Common Engineering Traps
Querying a vector database without a strict metadata filter (filter={ "user_id": uid }) allows User A to retrieve private memories and API keys stored by User B if semantic similarity is high. Always enforce user-scoped metadata partitioning.
Embedding raw chat messages like "I will be back in 5 minutes" clutters vector indices with useless junk. Always use an extraction model to filter for enduring facts (preferences, roles, long-term project parameters).
Key Architectural Takeaways
- 1.Dual Memory Stratum: Checkpointers provide session auditability, while vector stores provide cross-session personalization.
- 2.Fact Distillation: Extract only high-signal permanent preferences, ignoring transient chit-chat.
- 3.Multi-Tenant Privacy: Mandatory metadata filtering guarantees user memory boundaries are completely airtight.
Deploying Agents with FastAPI
Take your LangGraph agents to production. Learn how to wrap agents in async FastAPI endpoints with Server-Sent Events (SSE) streaming and background task queues.