Concept
Often you don’t need separate graph, vector, and text databases: one database can store
relationships and search the same records by meaning and by keyword. Think of it as
keeping one address book instead of three copies you must update by hand whenever a
friend moves. When the search indexes cover the same records as the graph and update in
the same transaction (an all-or-nothing change), there are no sync pipelines to maintain
and no copies that disagree. Separate systems still make sense in some cases, such as
when search is a standalone product.
Learning objectives
After reading this article you will be able to:
- Describe what a stack of separate graph, vector, and search stores looks like
- Explain the hidden costs of separate stores, from sync pipelines to permission drift
- Compare separate stores with one database that indexes the same records
- Recognize workloads where separate systems are the better choice
What does a stack of separate stores look like?
A common setup uses one system per job: a primary or graph database for entities and relationships, a vector database for embeddings, and a search engine for keyword queries. Pipelines copy data between them, and application code queries each one and merges the answers. For example, a company wiki might keep pages and team memberships in one database, page embeddings in a vector store, and a keyword index in a search engine. Each system can do its own job well. The cost is in the seams between them.What does it cost to run separate stores?
Mostly, it costs the work of keeping several copies of the same data in agreement, plus the work of running and joining several systems:
Permission drift deserves attention in
retrieval-augmented generation (RAG), where a chatbot
answers from retrieved documents. If the vector store still returns a chunk from a
document the user lost access to, the model can quote it.
Cross-system joins also cause a subtle recall problem, meaning relevant results go
missing. A common workaround is to take the top results from the vector store, then
drop the ones the graph says the user cannot see. If most of the top results are
dropped, the user gets too few results even though relevant, permitted documents exist
further down. See filtered vector search.
What changes with one transactional store?
When the database updates its vector and text indexes in the same transaction as the records, most of those seams typically go away:- No copy to keep in sync. The embedding and the text are properties of the record itself, so there is no second copy to replicate. If the application writes the new text and its recomputed embedding in the same transaction, the two cannot disagree.
- One transaction, one snapshot. Graph and search results within one request come from the same committed snapshot (a consistent view of the data at one moment), so the graph neighborhood and the search hits agree.
- Exact filtering before ranking. If the database supports prefiltered search, search can run inside the set of records a traversal reaches, such as documents a user can read, instead of filtering a truncated result list afterward.
- One system to operate, secure, and back up.
Try HelixDB
Store your graph, embeddings, and text in open-source HelixDB, and run traversals,
vector search, and BM25 search in one transaction.
When do separate systems make sense?
Separate systems make sense when a specialized need or an organizational constraint outweighs the cost of the seams:- Search is its own product. A large public search experience, such as an online store’s main search box, may need features that many multi-purpose databases do not offer, such as faceting (result counts by category), highlighting, typo tolerance, or custom language analysis.
- The workload has no relationships or filters. Pure similarity search over a very large embedding collection may be served well by a dedicated vector index.
- An existing system of record cannot move. If another database owns your transactions, you will be syncing data anyway, and the question becomes where the search and graph copies should live.
- Workloads need strict isolation. Different teams, scaling profiles, or failure domains can justify separate systems.
- The job is analytics. Large scans and aggregations over history typically belong in a data warehouse, not in an online graph or search store.
How do you decide between one database and separate systems?
Choose one database when your queries mix relationships with similarity or keyword relevance, access depends on relationships, and stale results would cause real problems. That pattern is common in permission-aware RAG and in AI agent memory. Otherwise, separate systems may fit. These questions help you decide:- Do your queries combine relationships with similarity or keyword relevance?
- Does access control depend on relationships, such as teams, shares, or ownership?
- How stale can search results be after a write or a permission change?
- Do you need exact identifiers as well as semantic matches?
- How many data systems can your team operate well?
How does HelixDB keep graph, vector, and text together?
HelixDB stores data in one labeled property graph, treats vector and text indexes as access paths over that graph, and runs traversals and searches in one ACID transaction per request:- Indexes are access paths over the same data. Secondary, vector, and text indexes are optional access paths over a label and a top-level property on nodes or edges, so relationships can be indexed and searched too.
- One ACID transaction per request. Each request runs over a committed snapshot with serializable snapshot isolation, and conflicting writes are caught at commit. Graph traversals, vector search, text search, and index lookups run in the same transaction, and all entries in a write batch commit or roll back together.
- Prefiltered search. Vector and BM25 search can run inside an exact, traversal-defined candidate set. The order is graph traversal, exact candidate membership, ranking, then top k, so a result outside the candidate set is never returned.
- Freshness. On Helix Cloud, readers see new commits after a snapshot refresh; writer-only reads give read-after-write. A newly created index backfills existing data asynchronously and becomes visible only after validation and atomic activation.
- Application-side pieces. The application computes embeddings and fuses vector and BM25 results. HelixDB has no built-in rank fusion or reranking.
Frequently asked questions
Can a graph database do vector search?
Some can. Graph databases differ: some include vector indexes natively, and others rely on an external vector store. Check whether vector search can be restricted to the results of a traversal and whether it runs in the same transaction as graph reads.Do I need a knowledge graph and a vector database for RAG?
You need both kinds of retrieval if your questions depend on relationships as well as meaning, but not necessarily two databases. A graph database with vector and text indexes can hold the knowledge graph and serve similarity search over the same records. See What is GraphRAG?Does one database become a bottleneck?
It can if its architecture does not scale the part of the workload you stress. Look at how it scales reads and storage, how writes are coordinated, and what isolation it offers. See Databases on object storage for one approach.Can I still add a search engine later?
Yes. Starting with one store does not prevent adding a specialized system for a specific need later. It means you add the seams only when a workload justifies them.Related topics
What is RAG?
Grounding model answers in data retrieved at query time.
Filtered vector search
Pre-filtering, post-filtering, and returning the right top k.
Hybrid search
Combining keyword and vector results with rank fusion.
What is a vector database?
How vector stores index embeddings, and where they fall short.
What is a graph database?
Nodes, edges, and traversals for connected data.
Databases on object storage
Separating storage from compute, and what it means for cost.