Concept
Full-text search is a way to find documents by the words inside them and list the best
matches first. Think of it as the index at the back of a book: instead of reading every
page, you look up a word and jump straight to the pages that mention it. That is how a
support team can type “login error” and get the most relevant tickets back without
scanning every ticket. The same idea sits behind the search box of an online store, a
company wiki, and many chatbots that answer questions from documents.
Learning objectives
After reading this article you will be able to:
- Explain how a search engine turns text into ranked results
- Describe an inverted index and how cleanup steps such as stemming change matches
- Explain why full-text search beats a SQL
LIKEquery for free-form searches - Decide when full-text search fits and when vector search fits better
How does full-text search work?
Full-text search does much of its work ahead of time. When a document is saved, the engine records which documents contain each word. When someone searches, it looks up only the search words in that record, instead of scanning every document, and ranks what it finds. Under the hood, most systems follow the same four stages:1
Analyze
An analyzer, the component that cleans up text, splits raw text into tokens
(usually individual words) and then normalizes them, for example by lowercasing or
stemming.
2
Index
Each resulting term is added to an inverted index that maps the term to the
documents that contain it.
3
Query
The search string goes through the same analyzer, and the engine looks up each
resulting term in the index to collect candidate documents.
4
Rank
A scoring function gives every candidate a relevance score, and the engine returns
the top results.
What is an inverted index?
An inverted index is a lookup table from each word to the documents that contain it, like a book index that maps a word to page numbers. The list of documents stored for each word is called a postings list. For example, take three short support tickets:- Ticket 1: “Reset your password from the login page”
- Ticket 2: “The login page shows an error”
- Ticket 3: “Password rules for new accounts”
login -> 1, 2 and
error -> 2. Together they tell the engine that only tickets 1 and 2 contain any of
the words, and only ticket 2 contains both, so ranking can put it first.
Postings typically also store how often each word appears in each document, which
ranking needs. Many systems can store term positions (where each word sits in the
document) as well, which make phrase searches such as "login page" and proximity
scoring possible.
How do tokenization, stemming, and stop words affect matches?
These cleanup steps decide what counts as “the same word,” and therefore which documents match. They are why a search for “connecting” can find a ticket that says “connected.”
Put simply, normalization and stemming usually raise recall (finding more of the
relevant documents) at some cost to precision (keeping irrelevant ones out). Stemming
lets “connected” match “connection issues,” but an aggressive stemmer can also merge
unrelated words.
Removing stop words shrinks the index but can break phrase searches where those words
matter. That is why many modern systems keep stop words and let ranking give them
little weight.
Try HelixDB
Run BM25 full-text search and graph traversals in the same transaction with
open-source HelixDB.
How is full-text search different from a SQL LIKE query?
A SQLLIKE '%password%' condition scans rows for a raw substring (any run of
characters) and returns matches in no particular order. Full-text search looks up whole,
cleaned-up words in an inverted index and ranks the results by relevance. In other
words, LIKE asks “does this text contain these characters?” while full-text search
asks “which documents are most about these words?”
LIKE is fine for small tables, prefix lookups, and structured codes in a
relational database. Once people
type free-form searches over a large collection, you need both an index and a ranking.
How are full-text results ranked?
Documents typically rank higher when they contain more of your search words, use them more often, and especially when they contain the rare ones. A relevance function turns those signals into one score per document. Under the hood, the two classic signals are:- Term frequency (TF): a document that mentions a word more often is probably more about it.
- Inverse document frequency (IDF): a word that appears in few documents, like “timeout,” tells you more than one that appears everywhere, like “issue.”
When should you use full-text search?
Use full-text search when the exact words matter. For example:- A support agent looking up an exact error message or order ID, or a shopper typing a part number.
- Rare or in-house terms, such as project names on a company wiki, that an embedding model may not represent well.
- An AI agent recalling an exact name or ID from its memory.
- Results that need to be explainable (“this matched because it contains
timeout”).
status = "open", use an equality or
range index instead of text search.
How does HelixDB do full-text search?
HelixDB provides full-text search through text indexes on its property graph, where data lives as nodes (entities) and edges (relationships):- A text index provides BM25-ranked search over a top-level string or string-array property of a node label or an edge label. Results are ordered by BM25 score, then ID.
- Each index uses one of three analyzers:
standard,standard_stem_en, orwhitespace_lowercase. Term positions are optional. - Text indexes can be partitioned by tenant, which keeps a separate index per tenant value.
- Prefiltered search ranks only the nodes or edges a graph traversal reaches, such as “documents this user can read.” A result outside that candidate set is never returned, and BM25 statistics still come from the full tenant partition.
- Within one request, text search runs in the same ACID transaction (an all-or-nothing, isolated unit of work) as graph traversals, vector search, and index lookups. There is no built-in rank fusion, so the application combines BM25 and vector results.
Frequently asked questions
Is full-text search the same as keyword search?
In everyday usage, yes. Both describe search that matches the words in a query against the words in documents. “Full-text” emphasizes that the whole body of each document is indexed, rather than only a title or a set of tags.Does full-text search understand synonyms?
Not by default. It matches analyzed terms, so “car” does not match “automobile” unless you configure a synonym list or expand the query. Vector search captures this kind of similarity without a hand-built list, which is one reason to use hybrid search.What is the difference between stemming and lemmatization?
Stemming applies suffix-stripping rules and can produce stems that are not real words. Lemmatization uses vocabulary and grammar to map each word to its dictionary form, so “better” can map to “good.” Lemmatization is usually more accurate and more expensive.Do I need a separate search engine for full-text search?
Not necessarily. Some databases include inverted indexes and BM25 ranking, which avoids copying data into a second system and keeping it in sync. A dedicated search engine can still make sense when search is its own product with specialized requirements. For the trade-offs, see Do you need separate graph, vector, and text databases?.Related topics
What is BM25?
A widely used ranking function for keyword search, explained.
What is hybrid search?
Combining keyword and vector results into one ranking, often with rank fusion.
What is vector search?
Embeddings, nearest neighbors, and approximate search.
Text indexes guide
Create BM25 indexes and run ranked search in HelixDB.