Yandex B2B Tech has launched hybrid search in its database YDB: the full-text index with BM25 ranking handles exact matches, the vector index handles semantic search, and the results of both can be combined in a single SQL query. This compresses the classic stack of “main database plus separate search engine plus vector storage” into a single stack without separate search infrastructure. The feature is already available in open source, in the Yandex Cloud managed service, and in local on-premises installations.


What happened
Hybrid search was included in the YDB 26.3 release. Both indexes — the full-text index with BM25 ranking and the vector index — are implemented as distributed YDB service tables and are updated in a single transaction along with the main data. The release also introduced the SQL function HybridRank, which combines BM25 search and vector search results in a single query, and an improved filterable vector index with the adaptive_clusters parameter, which automatically selects the number of clusters. The company made the feature available in three delivery formats: open source on ydb.tech, Managed Service for YDB from September 23, including Serverless, and on-premises.
Context
The demand for combining two search modes has been around for a long time: BM25 ranking reliably finds SKUs, names, and exact terms, while vector search catches paraphrased formulations and semantically similar fragments. Hybrid search itself is an established industry approach with limited scientific novelty; the main point here is different — where the indexes physically reside. Teams that need both modes (RAG assistants, knowledge bases, support, catalogs) usually have to assemble and synchronize three separate systems — the main DB, a search engine, and a vector storage; a synchronization window arises, where a row written a second ago is not yet visible in the search index, which directly hits the quality of RAG on fresh data. This is compounded by the double transition “index → identifier → row” during retrieval and licenses for each storage. The market responds to such tasks with growth: according to Grand View Research, the vector and hybrid search segment could reach $2–4 billion by 2028, while the vendor itself positions YDB as the first Russian database with built-in hybrid search.
Why this matters for the industry
For building teams and companies, the main win is architectural: since the search indexes are updated in a single transaction with the main data, two separate components along with their synchronization infrastructure disappear from the RAG application stack, and with them the synchronization window that directly affected the freshness of assistant responses. Teams get the ability to build search and RAG products on a single stack: for startups, this is a reduction in operating expenses and faster iterations, for existing systems — a scenario for consolidating three storages into one database. For Yandex Cloud, this is a classic step in expanding the value of a managed database and retaining workloads within its own ecosystem. If transactional indexes prove themselves under production load, hybrid search inside an SQL database could become an expected default for RAG infrastructure, and competition will shift from checking “is the feature there” to measurable characteristics — latency, cost, and ranking quality.
Why this matters for users
The feature can be tried immediately in any of the three delivery formats. A prototype of search over your own documents or tickets is assembled with standard SQL: the full-text index is created with a regular ALTER TABLE operator with the USING fulltext_relevance option and BM25 ranking, and the combined ranking of BM25 and vector search results is performed by the HybridRank function in a single query — a separate search engine and vector storage are not needed for the prototype. For those who want to see the assembly in full, Yandex Cloud will hold a webinar on October 15, “Find what you need quickly with YDB”: in one hour, they cover building search over documents and tickets with an AI assistant connection, registration is open on the webinar page.
What is still unknown / limitations
The key promises cannot yet be verified with numbers: the available materials contain neither recall@k, nor latency, nor throughput, nor comparisons with pgvector, Elasticsearch, or ClickHouse, so claims about speed and ranking quality remain declarations. The transactionality of search indexes as a product feature is claimed, but not demonstrated under live loads, and the methodology for merging two rankings in HybridRank is not disclosed. All three main sources — the Yandex Cloud blog, an article by a YDB team member on Habr, and the webinar page — are vendor sources, there are no independent reviews; the formulation about “the first Russian database with built-in hybrid search” is a marketing framework, and the market assessment methodology from Grand View Research is not disclosed. How transactional hybrid indexes will behave under production scale will be shown by operation and third-party measurements.
Sources
- Launched hybrid search in YDB — Yandex Cloud blog
- Launched full-text and hybrid search in YDB: we tell you what's under the hood — Habr, Alexander Zevaykin (YDB team)
- Find what you need quickly with YDB — Yandex Cloud webinar, October 15 (registration)
Author
Look at AI, editorial team
