Technical report · Sep 2026RAGInformation retrievalLLM evaluation

Is Hybrid Retrieval Worth It? An Evaluation-Driven Comparison of Dense, Sparse and Hybrid RAG on Multi-Hop News Questions

Ahmad Tawil · Sep 2026

The question

Hybrid search is recommended for RAG by default. Does it actually help, and by how much?

01

TL;DR

  1. 1

    Built one RAG pipeline three ways: dense (Qdrant + bge-small), sparse (BM25) and hybrid (Reciprocal Rank Fusion), over 609 news articles.

  2. 2

    Hybrid retrieved the most evidence (context recall 0.525 → 0.625), but a paired bootstrap puts the gain at +0.10, 95% CI [0.00, +0.23]: suggestive, not conclusive.

  3. 3

    The real bottleneck is retrieval: the top 5 held only 17–26% of the needed evidence, and retrieval takes 38–83 ms while the LLM takes 10–12 s.

02

Key results

context recall, hybrid (vs 0.525 dense)
0.625context recall, hybrid (vs 0.525 dense)
trap questions refused
9 / 10trap questions refused
of needed evidence in top 5
17–26%of needed evidence in top 5
retrieval latency, no LLM
38–83 msretrieval latency, no LLM
RAGAS metrics by retrieval mode
RAGAS metrics by retrieval mode
Share of required evidence retrieved
Share of required evidence retrieved
Refusal behaviour on trap and answerable questions
Refusal behaviour on trap and answerable questions
Latency: retrieval vs generation
Latency: retrieval vs generation

03

In plain words

Instead of judging RAG by how good one answer looks, this study builds the same pipeline three ways and measures each one on a fixed benchmark: 50 MultiHop-RAG questions (20 direct, 20 multi-hop and 10 unanswerable traps), scored with RAGAS, refusal checks, latency and paired bootstrap confidence intervals.

Generation is grounded: the model must cite every claim by chunk id or reply exactly “Insufficient information.” Trap questions are scored separately, because RAGAS scores a correct refusal as zero.

The study also documents a weakness of Reciprocal Rank Fusion: a chunk ranked #1 by one retriever but missing from the other's list can lose to a mediocre chunk both lists agree on.

Honest limitations

  • !50 questions is noisy; the direct set is mostly yes/no.
  • !Retrieval parameters (RRF k, chunk size, embedding model) were not tuned.
  • !Generator and judge come from the same vendor.
  • !The exact-chunk evidence check is a strict lower bound.
  1. 1

    Corpus

    609 articles → 17,653 chunks (512 chars, 64 overlap) with stable hash ids

  2. 2

    Index

    Qdrant + BAAI/bge-small-en-v1.5 (dense) and BM25 (sparse), one shared id space

  3. 3

    Retrieve

    dense, sparse, or hybrid via RRF (k = 60, top 50 from each list), top 5 to the model

  4. 4

    Generate

    grounded answer with [chunk-id] citations, JSON-schema validated

  5. 5

    Evaluate

    RAGAS with a separate judge model, trap-refusal rate, latency, bootstrap CIs

04

Read the full paper

05

Cite this work

BibTeX
@techreport{tawil2026hybrid,
  title       = {Is Hybrid Retrieval Worth It? An Evaluation-Driven Comparison of Dense, Sparse and Hybrid RAG on Multi-Hop News Questions},
  author      = {Tawil, Ahmad},
  year        = {2026},
  month       = sep,
  type        = {Technical report},
  url         = {https://ahmadtawil.dev/research/hybrid-retrieval}
}

06

Keep reading