eval-hybrid-rag

An evaluation-driven comparison of dense, sparse and hybrid retrieval for RAG on multi-hop news questions.

This is a static page: there is no server behind it, so you cannot type a new question here. What you can browse are the real answers recorded during the benchmark, for all 50 test questions in all three search modes.

Explore the benchmark

Pick a question to see how dense, sparse and hybrid search did on it: what the model answered, which sources it cited, how much of the needed evidence search found, and the exact passages the model saw. Answers were written by gpt-5-mini from the top 5 passages.

Start with:

Loading the recorded answers...

How to read the numbers: "needed passages" is a strict check against the exact evidence chunks the dataset lists, so it is a lower bound. RAGAS context recall is judged by a language model against the short reference answer, so the two can disagree: a short answer such as a company name can appear in a passage that is not one of the needed ones.

Results

Loading the benchmark results...

Grouped bar chart of the four RAGAS scores for dense, sparse and hybrid search
RAGAS scores per search mode (40 answerable questions).
Stacked bars showing, for each mode, how many answerable questions were answered or refused, and how much of the needed evidence search found
Most refusals happened because search had not found the needed evidence.
Bar chart of the share of needed evidence found in the top 5 results for each mode
How much of the needed evidence the top 5 results contained.
Bar chart of mean and 95th percentile latency for each mode
Latency is almost all the language model call.

Key finding. Hybrid search found the most evidence, but with 40 scored questions the gain over dense is suggestive, not conclusive. The bigger finding is that search, not the language model, is the bottleneck: the top 5 results held only 17 to 26% of the evidence a question needs, so the model correctly refused more than half of the answerable questions. The technical report has the details and the limitations.

How it works

Architecture diagram: articles are chunked into a dense Qdrant index and a BM25 index, retrieved and fused with RRF, then answered with citations by gpt-5-mini; a separate evaluation flow scores the results

Run it yourself

The live system (FastAPI, Qdrant, the same three search modes and a try-it box) starts with one command. It needs Docker; the first start takes about 20 minutes because it builds the search indexes. Asking new questions needs an OpenAI key in a .env file; browsing retrieved passages does not.

docker compose up --build
# then open http://localhost:8000