Concept
AI agent memory is the system that lets an AI agent, an assistant that carries out tasks,
remember facts, past conversations, and routines across sessions. It works like a notebook
the agent keeps: the underlying model retains nothing between calls, so the agent writes
down what matters in a separate store and reads the relevant pages back, typically before
it replies. Without memory, a support assistant asks for your account details every time,
and a music assistant forgets that you skip live recordings. With it, the agent can pick
up where it left off, keep up with facts that change, and stop asking the same questions.
Learning objectives
After reading this article you will be able to:
- Explain why a model’s context window cannot serve as an agent’s long-term memory
- Describe the main types of agent memory, from working memory to the user profile
- Outline how memories are captured, updated, forgotten, and recalled
- Identify the storage and search features a memory layer needs
Why is a context window not memory?
A context window is the text a model reads in a single call, and it is cleared when that call ends. It can hold the conversation in progress, but it cannot remember last month. Under the hood, the window is a maximum amount of text, measured in tokens (words or pieces of words); in many models, the reply the model generates counts against the same limit. It works like working memory: everything the model knows about the current task has to fit inside it. Using it as long-term memory breaks down for several reasons:- It resets. A new session starts empty unless something outside the model puts earlier information back.
- It is bounded and typically priced per token. Replaying every past conversation grows cost and latency with each turn, and eventually stops fitting.
- It has no notion of truth over time. If a user said one thing last month and the opposite today, a transcript contains both, and the model has to guess which is current.
- It is not scoped. Nothing in a raw prompt records who owns a fact, where it came from, or whether this user may see it.
- More text is not always better context. Models do not use every part of a long prompt equally well, so a focused set of relevant facts often produces better answers than a full history.
What types of memory do AI agents use?
Agents use short-term memory for the task at hand and long-term memory for anything that should survive to the next session. The terms are borrowed from cognitive psychology, the study of how people think and remember.
The difference between episodic and semantic memory matters in practice. Episodes
answer “what happened”; semantic facts answer “what is true now.” Agents often derive
semantic facts from episodes, for example inferring a preference from several
conversations, and keep a link back to those episodes as evidence.
The user profile is not a separate source of truth. It is a summary maintained from
long-term memories, so the agent can load broad context cheaply on every request.
Try HelixDB
Store an agent’s memories, their sources, and their embeddings in one open-source graph
database, with vector and BM25 search over the same records.
How does the memory lifecycle work?
Memory works as a loop. Each interaction can add, change, or remove memories, and each new request reads from the result. The loop has six steps: ingest, extract, deduplicate, update, forget, and recall.- Ingest. Collect raw input: conversation turns, uploaded documents, tool results, and application events.
- Extract. Turn raw input into small, self-contained facts. A reply such as “next Tuesday” only means something together with the question before it, so the extractor needs recent conversation context and the current date.
- Deduplicate. Compare each candidate with existing memories, both by meaning with vector search and by exact terms with full-text search, so the same fact is not stored many times.
- Update and version. When a fact changes, write a new version and mark the old one as superseded instead of overwriting it. History stays available for audits and for questions about what changed.
- Forget. Expire time-bound facts, let low-value episodes decay, and delete what a user or policy asks to remove.
- Recall. Retrieve the memories this user may see, that are still current, and that are relevant to the request, then assemble them into context. Recall can also follow links to related facts, much like GraphRAG.
What infrastructure does a memory layer need?
A memory layer has to know who owns each memory and how memories connect. It also has to find them by meaning and by exact words, while skipping anything outdated or deleted.
Many teams assemble these from several systems, such as a relational store for records, a
vector store for embeddings, and a search engine for keywords. That works, but the
application then has to keep them consistent.
For example, a memory deleted in one system can still surface from another. A permission
filter applied after a similarity search can remove every result or, if it is forgotten,
leak one; see filtered vector search. For
the trade-offs of combining these, see
one database for graph, vector, and text.
How does HelixDB support agent memory?
HelixDB is an open-source graph database with native vector search and BM25 full-text search, so the requirements above can be met in one labeled property graph:- Properties with indexes hold tenant scope, identity, and lifecycle fields such as
isLatest,validTo,deletedAt, andexpiresAt. Equality and range secondary indexes make them indexed filters. - Edges record provenance, categories, entities, version updates, and associations.
- Vector search handles deduplication and paraphrase recall, and BM25 handles exact names, IDs, and rare tokens. Both index types can be partitioned by tenant.
- A profile node holds stable user context that the agent loads on every request.
Frequently asked questions
How does an agent decide which memories to recall?
It narrows by scope first: only memories owned by or shared with the current user or team. It then keeps only current memories, excluding anything deleted, superseded, or expired, and ranks what remains by relevance with semantic and keyword search, as in hybrid search. These filters should apply before or during ranking, not after, and the agent keeps only a small budget of top results so the context stays focused.Is agent memory the same as RAG?
They share retrieval techniques but differ in what they retrieve. Retrieval-augmented generation (RAG) typically reads from a document collection that changes independently of the agent, such as a company wiki. Memory is written by the agent from its own interactions, changes as facts are corrected, and is scoped to a user or team. Many systems combine the two and link each memory to the source passage it came from.Can a vector database alone serve as agent memory?
Not on its own. A vector database covers semantic recall, and many also filter on metadata copied onto each vector, which can handle ownership and lifecycle flags. That copy has to be kept in sync, and it cannot represent provenance, versions, or links between entities as relationships; many vector databases also lack keyword ranking for exact identifiers. A complete memory layer needs properties, lifecycle filters, keyword search, and relationships alongside the vectors.What should an agent store in memory?
Store information that is likely to matter again: stable facts and preferences, notable events, and workflows the agent has learned. Skip small talk, passing details, and anything a user or policy says must not be kept. Each stored fact should be atomic and self-contained, and should record its owner, its source, and when it was learned so it can be superseded or expired later.How should an AI agent forget?
Many systems use soft deletion: mark a memory as deleted, superseded, or expired and exclude it from every recall path, and reserve physical deletion for user requests and retention policies. The guide to building long-term memory covers expiry, decay, and deletion in detail.Related topics
Building long-term memory
Data model, write path, forgetting, and recall for agent memory.
What is RAG?
Ground a model’s answers in data retrieved at query time.
What is GraphRAG?
Retrieval over entities and relationships, not only text chunks.
What is hybrid search?
Combine vector similarity with BM25 keyword ranking.
What is filtered vector search?
Restrict similarity search to the records a user may see.
What is a vector database?
Store embeddings and find the nearest ones to a query.