Concept
Long-term memory is the part of AI agent memory
that lasts across sessions: the facts, preferences, and past events an agent can recall
later. Think of it like a well-kept customer file, where each fact is written once, dated,
tied to where it came from, and crossed out rather than erased when it changes. Built well,
it lets a support assistant remember that a customer moved to annual billing, forget what
it is asked to forget, and keep every customer’s memories separate. The recipe: store each
fact as a versioned, user-scoped record linked to its source, then search only the current
memories that user may see.
Learning objectives
After reading this article you will be able to:
- Name the parts of an agent memory system and where each one runs
- Design a memory data model that is scoped to users and keeps versions
- Explain how new memories are deduplicated, updated, and forgotten
- Describe how an agent recalls memories, from scoping to building the prompt
What are the parts of an agent memory architecture?
A memory system has two halves: application code that interprets content, usually with AI models, and a database that stores, links, and searches it. For example, a support assistant’s extractor might pull “the customer’s plan renews in March” from a chat, and the retrieval service finds that fact the next time the customer writes in.What data model does agent memory need?
Store each memory as one small, self-contained fact with a clear owner, fields that say whether it is still current, and links to where it came from. The records are nodes and the links are edges, as in any property graph. Adapt the generic names below to your domain.- User owns memories; in multi-tenant products, where many customers share one system, every record also has a tenant key.
- Memory is one atomic fact: its text, embedding, kind (fact, preference, episode, or
procedure), salience (how important it is), confidence, and lifecycle fields such as
createdAt,isLatest,validFrom,validTo,expiresAt, anddeletedAt. - Source document and chunk hold raw context for citations and document search, as in RAG.
- Profile is the summarizer’s output, loaded as always-on context.
UPDATES, plus
EXTENDS for added detail), and associations with entities and categories, much like a
knowledge graph. Every memory, current
or superseded, has its own provenance and association edges; the diagram draws them once.
Here, validFrom and validTo record when a memory was current in the store
(system time), not when the fact was true in the world (the valid time that bitemporal
databases track). Keep real-world dates, such as a trip date, in a property like eventAt.
Try HelixDB
Keep memories, their sources, and their versions in one open-source graph database, and
run traversals, vector search, and BM25 search in the same transaction.
How does the memory write path work?
When the agent learns something, it pulls out the fact, checks for similar memories, decides whether the fact is new, a repeat, a correction, or added detail, and saves the result in one step.1
Extract with context
Give the extractor the current message, recent turns including the previous
assistant message, related memories, and the current date, so short answers resolve
in context. Ask for structured output: content, kind, confidence, entities, source,
and scope.
2
Find deduplication candidates
Search the user’s current memories with vector search for similar facts and keyword
search for shared names. Check for an exact content match inside the write
transaction, or use a scoped content hash as an idempotency key, so a retried write
never stores the same fact twice.
3
Classify the relationship
A similarity threshold alone cannot tell a restatement from a correction, so
adjudicate with rules or a model.
4
Write atomically
Write the new memory, its ownership and provenance edges, its entity and category
links, and any invalidation of an older version together, so recall never sees a
half-applied change. An update writes a new memory with its own embedding and never
edits an existing memory’s text. This is harder when records and indexes live in
separate systems; see
one database for graph, vector, and text.
How do you implement forgetting in agent memory?
Agents do not forget on their own. Forgetting is explicit writes when a fact changes or is removed, plus filters on every read that hide anything no longer current.- Supersession. An updated fact gets
isLatestfalse and avalidTotime but stays available for audits and “what changed” questions. - Soft deletion. Set
deletedAtwhen a user removes a memory. Recall excludes it, and the change is reversible. - Expiry. Time-bound facts, such as “the user is traveling this week”, get an
expiresAttime. A sweeper hides or removes them once it passes. - Decay. A sweeper can hide episodic memories that are old, low in salience, and rarely recalled.
- Hard deletion. When a user or policy requires physical removal, follow provenance
edges to every memory derived from the deleted source and delete its chunks too. Check
earlier versions through
UPDATESedges, which can hold the same data, and regenerate any profile built from removed memories.
How does the memory read path work?
To recall, the agent narrows to current memories this user may see, searches only inside that set, follows links to related facts, and packs the best results into the prompt.- Scope. Start from the tenant and user, or from the project or access-control boundary that decides visibility.
- Filter lifecycle. Keep memories where
deletedAtis empty,isLatestis true,validTois empty, andexpiresAtis empty or in the future. - Recall inside the candidate set. Run vector search for paraphrases and BM25 keyword search for names, IDs, and error codes over only the scoped, current memories. Filtering after a global top-k search can return too few results or leak data; see filtered vector search.
- Expand. From the top hits, follow edges to entities, related memories, and source chunks, as in GraphRAG. Apply the same tenant, user, and lifecycle filters to every memory reached, include earlier versions only when the question asks for history, and keep traversal depth bounded.
- Fuse and rerank. Merge the vector and keyword lists, for example with reciprocal rank fusion as in hybrid search, then adjust by salience, recency, and relationship type.
- Assemble context. Include the profile, the top memories with their sources, and supporting chunks within a token budget, without embeddings.
What are common pitfalls when building agent memory?
Common pitfalls are relying on vector search alone, leaving out tenancy scope or provenance, and overwriting memories in place.- Vector-only memory. Similarity search can miss or misrank exact identifiers such as IDs and error codes, can return a fact next to its correction, and has no concept of ownership or recency.
- No tenancy scope. Recall not bounded by tenant and user can return another user’s memories. Mark shared memories explicitly instead of treating a missing user ID as shared.
- No provenance. Without links to sources, the agent cannot cite, you cannot debug a wrong memory, and deleting a source document leaves its derived facts behind.
- Overwriting in place. Editing a memory’s text destroys the history needed to explain a change and, without re-embedding, leaves the vector describing the old fact.
How does HelixDB fit into this architecture?
HelixDB can serve as the memory store in this design, holding these nodes and directed edges in one labeled property graph.- Tenant scope, identity, and lifecycle fields are top-level properties with equality and range secondary indexes, because nested properties cannot be indexed.
- Vector indexes on the memory and chunk embedding properties, one per label, serve deduplication and paraphrase recall. Text indexes on their text properties rank exact names, IDs, and rare tokens with BM25. Both can be partitioned by tenant.
- Prefiltered search ranks only the candidate set a traversal builds, so scope and lifecycle rules apply before ranking and no search result falls outside the set. Graph expansion after the search needs the same filters.
- All entries in a write batch commit or roll back together, keeping a new version, its edges, and the old version’s invalidation consistent. See guarantees.
Frequently asked questions
What happens when you change the embedding model?
Vectors from different models are not comparable, so re-embed every memory and chunk, embed queries with the new model, and retune deduplication thresholds, which depend on the model and the distance metric. See vector embeddings.How do you handle shared memory for teams?
Add a scope key beyond the user, such as a project or workspace ID, and mark shared memories explicitly. Build the candidate set from every scope the requesting user can access.Should the database run memory extraction?
In most architectures, no. Extraction, embedding, and summarization depend on models and prompts that change often, so they run in the application, where you can change them without migrating stored data.Related topics
What is AI agent memory?
Memory types, the memory lifecycle, and what a memory layer needs.
What is GraphRAG?
Retrieval that follows relationships between entities and sources.
What is hybrid search?
Combine vector similarity with BM25 keyword ranking.
What is filtered vector search?
Why search inside a candidate set instead of filtering afterward.
What are vector embeddings?
How models turn content into vectors, and when to re-embed.
What is BM25?
The keyword ranking behind exact recall of names and IDs.