Let’s Connect
( ← Back )

Hybrid Retrieval Isn’t BM25 Plus Vectors. It’s Knowing Where Queries Go

Hybrid retrieval router sending queries to a keyword index, a trace log and vector search

A pharmacy counter on a busy Saturday evening. The first customer says, “Dolo 650, one strip.” The pharmacist turns, picks the exact box and hands it over in seconds. The next customer says, “Something for fever that won’t upset my stomach.” Now the pharmacist has to understand a need, not match a name. A good pharmacist never mixes the two up.

Most hybrid retrieval systems skip that moment of judgement. They send every query to a keyword index and a vector index, merge the two result lists and trust the blend. It usually works. When it doesn’t, the miss is silent: the language model receives plausible context and writes a confident, wrong answer.

This matters more now because agents call retrieval as a tool, and many of those calls carry exact identifiers pulled from tickets, logs or records. The two retrievers are the easy part of the system. The routing layer in front of them is the real design work.

By the end of this piece you will have a three lane taxonomy for sorting queries, a router that runs the cheapest reliable check first, and a trace span that records every routing decision so you can debug it later.

The Short Answer

Hybrid retrieval combines lexical search, which matches exact words, with semantic search, which matches meaning. Engines such as Azure AI Search run both on every hybrid query and fuse the results. Fusion rewards documents that both retrievers partly agree on, so an exact match found by only one retriever can lose to a vaguer document. The fix is a routing layer before retrieval: identifiers and codes go to lexical search, intent and paraphrase go to semantic search, and anything uncertain goes to both with weighted fusion. Every decision is logged as a trace span. The retrievers are configuration. The router is the design.

Two good retrievers can still return the wrong page

Both retrievers can be working perfectly and the user still gets the wrong document, because each one was handed a question it was never built to answer.

Picture a support engineer at a broadband provider typing ERR-4032 after firmware update into an internal assistant. Keyword search finds the one page that mentions that exact code. Vector search returns a general guide on error handling during firmware updates, which is close in meaning and useless in practice.

Ten minutes later the same engineer types customer can’t log in after password reset. This time keyword search returns the generic password reset page, because it shares the most words. Vector search finds the article that actually solves the problem, titled around account access issues after a credential change. The words differ; the meaning matches.

Two support queries showing keyword search winning one and vector search winning the other
Fig 1. Neither retriever is broken. Each is answering a different kind of question from the one it was given. Illustrative support desk example.

Neither result is a bug. The failure sits upstream, in a system that never asked which kind of question it had received. A retrieval pipeline makes it only if someone designs the step.

Consider a multimodal assistant working over repair manuals and product catalogues, with images retrieved alongside text. A query as plain as how to check tyre pressure can return general tyre photos above the actual procedure images from the manual, and occasionally miss the procedure image altogether. Every one of those tyre photos is close to the query in meaning. None of them shows the reader how to do the job. Images make the gap easy to see, because a wrong picture is obvious in a way a wrong paragraph is not.

To test your own pipeline, run a handful of real identifiers, such as error codes or part numbers, through vector search alone. If the document or image that contains them keeps missing the top results, you have found the gap a router closes.

Lexical search vs semantic search: what each one is for

Lexical search and semantic search solve different problems, and treating them as interchangeable scorers is how hybrid retrieval quietly goes wrong.

Lexical search ranks documents by the words they share with the query. Its standard scoring function is BM25 (Best Matching 25), which gives rare terms more weight, stops rewarding a word after it repeats enough times and adjusts for document length. It is precise, cheap and blind to synonyms.

Semantic search turns the query and every chunk of text into embedding vectors, then returns the chunks whose vectors sit closest to the query. It handles paraphrase and vague descriptions well. It struggles with strings that carry little learnable meaning, such as codes, part numbers and version strings.

Anthropic’s engineering team made the same point in its 2024 write-up on contextual retrieval. It noted that embedding models can miss crucial exact matches, and that BM25 is particularly effective for queries containing unique identifiers or technical terms. Its example was an error code in a technical support database, where embeddings surfaced content about error codes in general.

There is an older warning too. The BEIR benchmark, published at NeurIPS in 2021, evaluated retrieval systems across 18 datasets. It found BM25 to be a robust baseline, while dense retrievers often underperformed on unfamiliar data. Embedding models have improved a great deal since, so read that as history rather than a current ranking. The lesson holds: on unfamiliar data, the old lexical baseline is harder to beat than it looks.

What is hybrid retrieval, and why does fusion alone fall short?

Hybrid retrieval is a search design that combines lexical and semantic retrieval over the same corpus so each covers the other’s blind spots. The common implementation runs both retrievers on every query and merges the ranked lists with Reciprocal Rank Fusion. That merge is exactly where exact answers get lost.

Reciprocal Rank Fusion (RRF) scores each document by adding up 1 / (60 + rank) in every list it appears in. It was introduced by Cormack, Clarke and Büttcher at SIGIR in 2009, and their original paper fixed the constant at 60 during a pilot study. Microsoft's Azure AI Search documentation describes the same method in current use. Hybrid queries run full text and vector search in parallel, and RRF gives higher importance to items ranked higher in multiple lists.

That design choice is the catch. Agreement is a strong signal for an ambiguous question and a weak one for an exact lookup.

Reciprocal Rank Fusion scores ranking a general document above the exact error code page
Fig 2. Worked arithmetic with illustrative documents. The one page containing the code ranks first in one list and still loses to a page both lists half agree on.

None of this is an argument against running both retrievers. In the same 2024 study, Anthropic reported that combining contextual embeddings with contextual BM25 reduced the top 20 retrieval failure rate from 5.7% to 2.9%. That is a 49% reduction averaged across the domains it tested. The evidence says keep both retrievers. It does not say send every query to both with equal say.

Modern retrieval engines already expose per-query controls such as vector weighting and caps on how many BM25 results enter fusion. Those controls are exactly the hook a router needs.

Query routing with the Three Lanes model

Query routing is the step that decides, before any retrieval runs, which retriever should answer a query and how much each one’s opinion counts. The Three Lanes model keeps that decision small enough to test.

The exact lane takes queries built around an identifier, a code, a quoted phrase or a named product, and sends them to lexical search only. The meaning lane takes descriptions of a need, a symptom or an intent, and sends them to semantic search only. The judgement lane takes everything else: mixed queries, short ambiguous ones and anything the router is unsure about. Those run through both retrievers with weights chosen at routing time.

Router sending queries into exact, judgement and meaning lanes before reranking, with trace logging
Fig 3. Three lanes, one router, one span. Uncertainty is not an error state; it is a lane.

The taxonomy below is an illustrative example for an enterprise support knowledge base. Build yours from your own query logs.

Table 1. An illustrative query taxonomy. The last column is the reason the lane exists.
Query class Example query Detection signal Lane What breaks if misrouted
Identifier or code ERR-4032, INV-88213 Pattern match on known ID formats Exact Vector search returns generic error content
Quoted phrase "force majeure" clause Quote marks in the query Exact Paraphrases replace the wording the user asked for
Named product or policy Unlimited 599 plan Lookup against the product catalogue Exact Similar plans outrank the one named
Symptom or intent internet drops every evening No identifier, descriptive language Meaning Keyword search misses articles worded differently
Paraphrase how do I get my money back Classifier, high confidence Meaning The refund policy page shares no words with the query
Mixed ERR-4032 after password reset Identifier plus intent Judgement Either single lane drops half the question
Short or ambiguous P2 SLA Classifier, low confidence Judgement A confident miss in the wrong lane

Keep the lane list short. Every new lane is a new class to label, test and monitor, and most of the value comes from separating exact lookups from everything else.

How should the router decide?

Use the cheapest check that is reliable for each class. Deterministic rules catch identifiers and quoted text. A small trained classifier separates lookup from intent and reports a confidence. A language model router, if you use one at all, handles only the residue the first two cannot place. Anything below your threshold goes to the judgement lane.

Table 2. Three ways to build the router, stacked from cheapest to most expensive.
Router approach Good at Weak at What the trace can explain Relative cost and latency
Rules and patterns Identifiers, codes, quoted text Anything phrased freely The exact rule that fired Lowest
Small classifier Lookup versus intent at volume Rare classes it has not seen A label and a confidence score Low
Language model router Messy, mixed or novel queries Consistency and speed A stated reason, which needs checking Highest

Order matters because each layer only sees what the one before it could not place. That keeps the expensive router small and the explainable one busy.

Where this approach does not apply

Routing is overhead you should not pay everywhere. If your knowledge base is under 200,000 tokens (about 500 pages), Anthropic’s write-up suggests skipping retrieval altogether and placing it all in the prompt context. If your queries are nearly all one kind, a single well-tuned retriever may be enough. Routing also cannot rescue an index that never stored the identifier field in the first place.

This usually fails when a rule fires on something that only looks like a code. Protect the exact lane with a fallback: if lexical search returns nothing above a minimum score, retry the query in the judgement lane and log that the fallback fired.

Why is the routing decision a trace span?

Because a routing decision is a runtime choice that changes what the model sees, and a choice you cannot see is a choice you cannot debug. Recording it as a span inside the request trace places the lane, the rule or model that decided and its confidence right beside the retrieval results it caused.

The OpenTelemetry semantic conventions for generative AI already define a retrieval span. Its operation name is set to retrieval, and the span is named after the data source it queried. The conventions do not define routing, so the route span carries custom attributes alongside the standard ones. They are still marked as in development, so pin the version you emit.

Trace timeline and route span attributes showing lane, rule, confidence and skipped vector retrieval
Fig 4. The route span sits above the retrievals it caused. A skipped retriever is recorded as skipped, not left out.

Once routes are spans, four questions become answerable with a query instead of a guess:

  • Which lane produces the most wrong answers?
  • Which rule fires on queries it should not?
  • Did the new router version shift traffic between lanes?
  • Is the share of judgement lane traffic creeping up, a sign that users have started asking something new?

Here is how that plays out in an illustrative case. A team notices that answers about one product line got worse after a catalogue update. Without route spans, the first instinct is to retune the embeddings. With them, one filter shows the drop is confined to the exact lane, and the rule_id attribute points to a single pattern: the new product codes use a format the rules did not recognise, so those queries fell through to the meaning lane. The fix is one new rule, not a new model.

Where to start: building a query router in six steps

Start with the queries you already have, not with a routing framework.

  1. Label a sample of real queries from your production logs with the lane each should take. Output: a labelled routing set you can test against.
  2. Write rules for identifier formats you already know, such as ticket IDs, SKUs and error codes. Output: a rule file with a test case for every pattern.
  3. Measure each retriever alone on each lane using the labelled set. Output: a per-lane recall table that shows where each retriever wins.
  4. Train or prompt a classifier for the queries rules cannot place, and set a confidence threshold. Output: a classifier, its threshold and its error pattern on the labelled set.
  5. Emit a route span on every request with the lane, the decider, the confidence and the router version. Output: answer quality broken down by lane.
  6. Add the exact lane fallback and review misroutes on a fixed cadence. Output: a fallback rate you watch and a versioned router changelog.

Questions about query routing

Is hybrid retrieval always better than vector search alone?

Usually, but not by default. Anthropic’s 2024 study concluded that embeddings plus BM25 outperformed embeddings alone in its tests. The gain depends on your query mix. A corpus searched mostly in natural language may gain little, while one full of codes and part numbers can gain a lot. Measure per query class before deciding.

What is the difference between BM25 and vector search?

BM25 ranks documents by shared words, weighting rare terms higher and adjusting for document length. Vector search ranks documents by how close their embeddings sit to the query’s embedding, which captures meaning rather than wording. BM25 is strong on identifiers and exact phrases. Vector search is strong on paraphrase and loosely described needs.

Does query routing add latency?

A rules layer adds almost nothing, and a small classifier adds little. A language model router adds a full model call, which is why it belongs at the end of the stack for the few queries nothing else can place. Routing can also save time, because exact lane queries skip the vector search and fusion steps entirely.

Can a language model do the query routing?

It can, but it should rarely be the first check. Language models handle messy and novel queries well, yet they are slower, costlier and less consistent than rules or a trained classifier. Use one for the residue, log its stated reason in the route span and audit that reason against the retrieval outcome.

How do I know whether my router is working?

Compare answer quality by lane using the route span, and track three signals over time: the misroute rate on a labelled sample, the exact lane fallback rate and the share of traffic landing in the judgement lane. A rising judgement share often signals that user behaviour has changed and your taxonomy needs a review.

Route first, then fuse

Hybrid retrieval earns its name only when the system knows which kind of question it is answering. Running BM25 and vectors side by side is table stakes, and fusion does a fine job on ambiguous queries. Exact lookups need a different path, and the router is what provides it.

The pharmacist at the counter does not run every request through both the shelf and a conversation. They decide first, and they would be able to tell you why. Your retrieval layer should be able to do the same, and the trace span is where it tells you.

If you take one action this week, pull a sample of real queries from your logs and label each one exact, meaning or judgement. The share that lands in each lane tells you whether routing will pay for itself, and it gives you the test set for everything that follows.

The Agneya perspective

If you’re building an AI application, an internal knowledge engine or an agent with retrieval tools, the two retrievers are the easy part. The real engineering is the routing layer in front of them: deciding which queries need exact precision, which need semantic nuance, and recording every decision as a trace span so you can inspect and improve it.

If you want to review your retrieval architecture or build a production RAG pipeline that doesn’t fail quietly, we can walk through how your queries flow. Tell us what you’re building.

Sources & References

  1. Anthropic, “Contextual Retrieval”, 2024.
  2. Thakur et al., “BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models”, NeurIPS 2021.
  3. Cormack, Clarke, and Büttcher, “Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods”, SIGIR 2009.
  4. Microsoft Azure, “Hybrid search ranking using Reciprocal Rank Fusion (RRF)”, Azure AI Search Documentation.
  5. OpenTelemetry, “Semantic Conventions for Generative AI Operations”.

More from the blog