Dmitriy Kononov.
Let’s talkContact

News and analysis

Hybrid search: why tokenization matters alongside embeddings

Why hybrid search misses identifiers, versions and accented words. Test BM25 tokenization and embeddings against representative queries.

AIPublished:

An employee searches for the instructions for AC-42, but the search ranks documents for the similar AC-24 device. Another query fails to match the stored name café with the typed cafe. If search supplies context to AI support, the model may confidently explain the wrong instructions. Replacing the model does not resolve the difference between the required object and a similar text.

Hybrid search combines a semantic embedding signal with lexical ranking such as BM25. The first helps with meaning and paraphrases; the second helps find particular terms. However, lexical retrieval operates on tokens produced during text analysis. Check that stage when search misses identifiers, version names or words in another language.

Changes in text-analysis tooling

Weaviate's May 14, 2026 article describes features in version 1.37: property-level text analysis, accent folding and an API for inspecting tokenization. The provider shows two distinct token sets: indexed tokens and tokens used to process a query. This makes it easier to see where normalisation or stopwords change the lexical signal.

The account concerns a particular database version. It does not establish an equivalent contract in every search engine or an existing Weaviate deployment. Check your configuration, client library and index-change requirements. The wider engineering principle is to inspect the text representation before comparing the model's final response.

Different fields need different rules

A product description and its identifier serve different purposes. In a description, normalising case and separating words from punctuation may help. In an identifier, a hyphen, underscore or letter case may carry meaning. One shared analyser can improve descriptions while undermining exact object distinctions.

Weaviate lists word, lowercase, whitespace and field methods. In its example, word splits User_42 - café into words and a number, while field keeps the entire string as one token. Choose based on the field's actual search behaviour. Preserving a whole value can help with exact matching without making it suitable for a free-form customer question.

A hypothetical catalogue might separate identifier, name, description and document language. A query with an explicit identifier first checks the exact field; “How do I clean the filter?” searches textual instructions. This is a proposed architecture that needs testing on your data. The custom API project outline illustrates why resources, access and response contracts need their own documentation; the historical outline does not establish adoption of this search design.

A search bar above result cards.
K. Limpitsouni / unDraw · License

Accents and stopwords are not universally noise

Accent folding can match café with cafe. Weaviate describes applying the analysis at index and query time, with selected characters exempted from transformation. That matters when a distinction separates names or catalogue items. Normalisation that helps one language can merge values that the business considers different.

Stopwords present a similar issue. A common function word may be unimportant in a long guide but meaningful in a brand name. The provider describes property-level stopword settings for properties using word tokenization and clarifies that this mechanism keeps stopwords in the index while excluding them from the BM25 query. Changing the list is therefore not the same as removing words from stored documents.

For a Russian-language catalogue, handling Latin-script accents does not establish correct Russian morphology. Test inflected words, mixed Russian and English queries and transliteration. Assess each language analyser separately when the system supports multiple languages.

Evaluate retrieved documents first

In the May 6 retrieval-quality overview, the author discusses research observations including stale material, semantically close but insufficient passages and low-relevance top-k results. These are the author's published observations rather than a universal assessment of all RAG systems. A useful practical response is to separate finding evidence from generating an explanation.

Build a small reference set: an exact identifier, a neighbouring identifier, a version name, a typo, an accented word, an inflected Russian word and a semantic paraphrase. Each query needs expected documents and clearly unsuitable matches. Include questions absent from the knowledge base: always returning several passages does not establish that useful information was found.

Compare lexical search, vector search and their combination on the same set. Record results before passing them to the model: whether the correct document appeared, its position and whether dangerously similar versions also appeared. Adjusting hybrid-ranking weights makes sense after checking tokens; otherwise it may simply conceal a defect in one signal.

When the cause is beyond tokenization

Correct tokens cannot repair a document whose identifier is in one passage while the applicable instructions are in another without connecting context. They also cannot make old material current or authorise access. Document and data processing includes input preparation and transformation checks, which directly affect what reaches the index.

For PDFs containing tables and charts, check whether extraction preserved their structure. That is the separate subject of retrieval from PDF tables and charts. If an answer needs a chain of relationships between a contract, supplier and asset, exact words will not resolve the whole task either. Consider choosing knowledge graphs or vector RAG.

For an initial improvement, choose one class of missed queries, inspect its tokens and compare an alternative configuration on the reference set. Then evaluate final answers to the same questions. Hybrid search becomes useful when both signals fit the task; its advantage over alternatives needs evidence rather than an architectural label.

Sources