Is Hybrid Retrieval Worth It? An Evaluation-Driven Comparison of Dense, Sparse and Hybrid RAG on Multi-Hop News Questions
Ahmad Tawil · Sep 2026
The question
Hybrid search is recommended for RAG by default. Does it actually help, and by how much?
01
TL;DR
- 1
Built one RAG pipeline three ways: dense (Qdrant + bge-small), sparse (BM25) and hybrid (Reciprocal Rank Fusion), over 609 news articles.
- 2
Hybrid retrieved the most evidence (context recall 0.525 → 0.625), but a paired bootstrap puts the gain at +0.10, 95% CI [0.00, +0.23]: suggestive, not conclusive.
- 3
The real bottleneck is retrieval: the top 5 held only 17–26% of the needed evidence, and retrieval takes 38–83 ms while the LLM takes 10–12 s.
02
Key results
- context recall, hybrid (vs 0.525 dense)
- 0.625context recall, hybrid (vs 0.525 dense)
- trap questions refused
- 9 / 10trap questions refused
- of needed evidence in top 5
- 17–26%of needed evidence in top 5
- retrieval latency, no LLM
- 38–83 msretrieval latency, no LLM
03
In plain words
Instead of judging RAG by how good one answer looks, this study builds the same pipeline three ways and measures each one on a fixed benchmark: 50 MultiHop-RAG questions (20 direct, 20 multi-hop and 10 unanswerable traps), scored with RAGAS, refusal checks, latency and paired bootstrap confidence intervals.
Generation is grounded: the model must cite every claim by chunk id or reply exactly “Insufficient information.” Trap questions are scored separately, because RAGAS scores a correct refusal as zero.
The study also documents a weakness of Reciprocal Rank Fusion: a chunk ranked #1 by one retriever but missing from the other's list can lose to a mediocre chunk both lists agree on.
Honest limitations
- !50 questions is noisy; the direct set is mostly yes/no.
- !Retrieval parameters (RRF k, chunk size, embedding model) were not tuned.
- !Generator and judge come from the same vendor.
- !The exact-chunk evidence check is a strict lower bound.
- 1
Corpus
609 articles → 17,653 chunks (512 chars, 64 overlap) with stable hash ids
- 2
Index
Qdrant + BAAI/bge-small-en-v1.5 (dense) and BM25 (sparse), one shared id space
- 3
Retrieve
dense, sparse, or hybrid via RRF (k = 60, top 50 from each list), top 5 to the model
- 4
Generate
grounded answer with [chunk-id] citations, JSON-schema validated
- 5
Evaluate
RAGAS with a separate judge model, trap-refusal rate, latency, bootstrap CIs
04
Read the full paper
05
Cite this work
@techreport{tawil2026hybrid,
title = {Is Hybrid Retrieval Worth It? An Evaluation-Driven Comparison of Dense, Sparse and Hybrid RAG on Multi-Hop News Questions},
author = {Tawil, Ahmad},
year = {2026},
month = sep,
type = {Technical report},
url = {https://ahmadtawil.dev/research/hybrid-retrieval}
}06



