My experiments· 2 min read
When hybrid search loses: a lesson from Reciprocal Rank Fusion
In my hybrid RAG benchmark, a chunk ranked #1 by dense search lost to a mediocre one. Here's why RRF rewards agreement over confidence.
#RAG#Retrieval#Evaluation
Hybrid retrieval, keyword search and semantic search fused together, is often recommended for RAG by default. In my benchmark it did retrieve the most evidence. But while reading the per-question results, I found a case where it clearly made things worse, and the reason is built into how the fusion works.
How Reciprocal Rank Fusion works
Reciprocal Rank Fusion (RRF) ignores the raw scores of each retriever and only looks at ranks. Every document gets a score from each list it appears in:
RRF(d) = Σ over lists 1 / (k + rank(d))With the usual k = 60, a document only gets points from lists it actually appears in. In my setup, each retriever contributed its top 50.
The failure case
On one question, dense search ranked the right chunk #1, but BM25 didn't have it in its top 50 at all. Another chunk was ranked around the middle by both retrievers. Here's what that does to the scores (ranks are illustrative):
k = 60
confident_in_one = 1 / (k + 1) # #1 in dense, missing from BM25
agreed_by_both = 1 / (k + 10) + 1 / (k + 10) # #10 in both lists
print(f"{confident_in_one:.4f}") # 0.0164
print(f"{agreed_by_both:.4f}") # 0.0286 -> winsThe mediocre-but-agreed chunk scores almost twice as much. RRF rewards agreement between lists, not the confidence of a single list.
What the benchmark said overall
Across 40 answerable questions, hybrid still had the best context recall:
| Mode | Context recall |
|---|---|
| dense | 0.525 |
| sparse (BM25) | 0.550 |
| hybrid (RRF) | 0.625 |
But a paired bootstrap put the gain over dense at +0.10, 95% CI [0.00, +0.23]. That's suggestive, not conclusive, at this sample size. The bigger finding was elsewhere: the top 5 results held only 17–26% of the evidence the questions needed, so retrieval, not the LLM, was the bottleneck.
What I'd test next
- Retrieve more (larger k) and measure recall@k directly, instead of only through RAGAS.
- A reranker after fusion, so a confident single-list hit can win back its place.
- Diversity (MMR or per-article caps), since several multi-hop questions need chunks from different articles.
- Query decomposition for multi-hop questions, retrieving for each sub-question separately.
The full numbers, charts and every recorded answer are in the report and the live explorer.