Concept
Hybrid search runs a keyword search and a meaning-based search (vector search) for the
same query, then merges the two result lists into one. Think of it as asking two
librarians at once: one finds every book that uses your exact words, the other finds
books about the same idea in different words, and you combine their shortlists. Keyword
search catches exact names, IDs, and error codes, while vector search catches
paraphrases. Together they hold up whether a shopper types a model number or
“comfortable shoes for standing all day.”
Learning objectives
After reading this article you will be able to:
- Explain why keyword search and vector search fail in complementary ways
- Describe the hybrid pipeline: search both indexes, merge the lists, optionally rerank
- Calculate reciprocal rank fusion scores and compare them with weighted blending
- Recognize when hybrid search is worth its extra cost
How do keyword search and vector search differ?
Keyword search matches the words you typed, usually ranked with BM25. Vector search matches what you meant. Each succeeds where the other fails, and that is the whole case for combining them.
A support-ticket search shows the gap. A user typing “E1042” needs the tickets that
contain that exact code, which an embedding may blur into generic “error” text. A user
typing “I can’t log in after resetting my password” needs tickets about authentication
failures, even when they share few words with the query.
How does hybrid search work?
Hybrid search asks both indexes the same question, merges the two ranked lists, and optionally re-scores the top of the merged list before returning results.1
Retrieve from both indexes
Run BM25 and vector search for the same query, usually in parallel. Embed the query
with the same model that embedded the documents. Retrieve more candidates from each
search than you plan to return, so good results from one list are not cut off
before fusion.
2
Fuse the lists
Combine the two rankings into one with a fusion method such as reciprocal rank
fusion.
3
Rerank, optionally
Score the fused top results with a slower, more accurate model, then return the
final top results.
Try HelixDB
Run prefiltered vector search and BM25 search over the same graph records, in one
transaction, with open-source HelixDB.
How do you combine BM25 and vector search results?
Put simply, you either merge by each document’s position in each list, usually with reciprocal rank fusion (RRF), or rescale both sets of scores and blend them with a weight. Either can be followed by a reranker that rescores the fused top candidates.Reciprocal rank fusion
Reciprocal rank fusion ignores raw scores and uses only each document’s rank in each list. A document earns credit for sitting near the top of any list, and more credit for appearing in several:k dampens the advantage of the very top ranks. The value 60 comes from
the paper that introduced RRF and is a common convention, not a requirement. Suppose
BM25 returns A, B, C and vector search returns D, C, A:
In this example, documents that appear in both lists rise to the top. In general, RRF
rewards agreement between lists, but a high rank in one list can still outscore low
ranks in both. RRF needs no score normalization and no training data, which makes it a
robust default. Its limitation is that it discards score magnitude: a document that is
far better than the rest in one list gets no extra credit. A weighted variant
multiplies each list’s terms by a weight to favor one method.
Weighted score blending
The other common approach normalizes each list’s scores and blends them:- BM25 scores have no fixed range and shift with the query’s terms and the collection’s statistics.
- Vector scores depend on the embedding model and the distance metric. Some are distances, where lower is better, and some are similarities, where higher is better.
- Both distributions change from query to query, so a fixed threshold on either is unreliable.
w is then tuned on labeled queries (test searches with known right
answers). This can beat RRF when tuned well, but it is sensitive to outliers and to how
many candidates each search returns.
Reranking
A reranker, often a cross-encoder model (one that reads the query and a document together) or a large language model (LLM), produces a new relevance score for each candidate. It is usually more accurate than either first-stage search but far more expensive per document, so it runs only on the fused top candidates.Why does hybrid search matter for RAG?
A chatbot that answers from documents is only as good as the passages it retrieves. In retrieval-augmented generation (RAG), many questions mix a concept with an identifier, and vector retrieval alone can miss the one chunk that names it. Adding BM25 reduces that risk. For example, in “what changed in the refund policy for plan B-7,” BM25 finds the chunks that contain “B-7” even when their embeddings are not the closest to the question. For following relationships from retrieved hits, see GraphRAG.When is hybrid search worth it?
Hybrid search pays off when queries mix natural language with exact terms: support tickets, product catalogs, code and logs, legal and policy documents, and AI agent memory. It adds less when every query is conversational and the corpus has no identifiers, or when every lookup is an exact key. The cost is a second index to build and keep current, a query embedding per request, and a fusion step to tune. Measure whether it pays for itself: compare BM25 alone, vector alone, and the fused result on a set of real queries with known relevant documents.How does HelixDB support hybrid search?
In HelixDB, vector and BM25 text indexes are optional access paths over a label and a top-level property, on nodes or edges of its property graph. When you define both on the same label, both searches rank records of that label:- One request is one ACID transaction over a committed snapshot. A single request can build a candidate set with a graph traversal, then run a prefiltered vector search and a prefiltered BM25 search over that same candidate set.
- Prefiltering guarantees that neither search returns a result outside the candidate set, such as documents the current user cannot read. Each prefiltered search returns at most 800 results. Prefiltered BM25 takes its term statistics from the full tenant partition, not only from the candidate set.
- Vector results are ordered by distance, closest first, then ID. BM25 results are ordered by score, then ID. These orders give the ranks that rank-based fusion uses.
- The application computes embeddings; HelixDB stores and indexes the vectors.
- HelixDB has no built-in rank-fusion operator. The application fuses the two result lists, for example with RRF, and applies any reranking.
Frequently asked questions
Is hybrid search better than vector search?
On workloads that mix exact terms with natural language, it often improves recall of exact terms while keeping most semantic matches, but it is not automatically better. Fusion reorders the combined list, so some hits that only vector search found can fall below the cutoff.Can hybrid search combine more than two searches?
Yes. RRF sums over any number of ranked lists, so one query can fuse, for example, BM25 over titles, BM25 over body text, and vector search. Weighted blending also extends to more lists, but each list needs its own normalization and weight, which means more tuning.How many results should each search return before fusion?
More than the final number you need, because a document ranked modestly in both lists can outrank one ranked highly in only one. The right depth depends on your corpus and latency budget, so tune it along with the fusion method.Does hybrid search need two databases?
No. It needs a keyword index and a vector index over the same content, not necessarily a search engine plus a separate vector database. When both indexes live in one database, the two searches can read the same snapshot of the data and apply the same filters, for example within one transaction. See Do you need separate graph, vector, and text databases?Related topics
What is BM25?
The ranking function behind most keyword search, explained.
What is vector search?
Nearest neighbors, approximate search, and recall.
What are vector embeddings?
How models turn text and other inputs into comparable vectors.
What is retrieval-augmented generation (RAG)?
Grounding model answers in retrieved context.
What is filtered vector search?
Pre-filtering, post-filtering, and returning the right top k.
Prefiltered search guide
Rank only the records a traversal reaches in HelixDB.