RAG retrieval on FiQA, from a BM25 baseline to hybrid search and a reranker
Reference build · in progress. This is not client work. VenianAI is building and measuring this system on a public dataset to document the method. The method below is the one the run follows. The results are not in yet, and each one says so where it will appear.
Summary
We are measuring retrieval on a public financial question-answering benchmark, starting from a plain BM25 baseline and changing one component at a time: dense retrieval, hybrid search, a reranker, then chunk size.
Every variant is scored on the same offline eval set, with retrieval and generation measured separately so a wrong answer can be traced to the stage that caused it, and the targets are committed before the first variant runs.
This is a reference build on public data, not client work, and every result on this page reads as measuring until a logged run produces it.
- Failure mode
- Poor data access
- Maturity
- L1 → L2
- Results measured
- 0 of 52
- Updated
Context
Teams at L1 who have grounded an assistant in their own documents usually cannot say whether a wrong answer came from retrieval or from generation. Both look the same from the outside: a confident answer that is not supported by the source.
This build separates the two and measures each. Retrieval is scored against human relevance judgements, generation against the passages actually retrieved. The method is the one we use on client corpora; the corpus here is public so anyone can rerun it.
When the assistant gives a wrong answer, was it retrieval or generation, and which change fixes it?
Failure mode
Poor data access. The model is only as good as the passages it is handed, and the gap between a demo and a useful system is almost always retrieval: chunking, embedding, hybrid search and reranking.
The pattern this targets is an assistant that answers fluently from a passage that is on topic but does not contain the answer. Without retrieval measured on its own, the fix people reach for is a bigger model, which makes the answer more fluent and no more correct.
The fault, as we define it
Poor data access: The model can't reach the system that holds the answer.
The test: When your assistant answers a question, can it point at the document or record it came from?
Read the full entry on poor data accessMetric agreed
Two retrieval metrics on the benchmark's own relevance judgements, two generation metrics on a sampled subset, and the latency and cost every variant pays. Targets are committed in the repository's first commit, before any variant runs, so the results cannot be framed after the fact.
Recall@10
Share of the judged-relevant passages for a query that appear in the top 10 retrieved.
Higher is better
- Baseline
- MeasuringBM25 Recall@10 on the test split
- Target
- MeasuringCommitted before any variant runs
- Result
- MeasuringRecall@10 of the recommended configuration
nDCG@10
The standard BEIR ranking metric: rewards relevant passages near the top of the list more than lower down.
Higher is better
- Baseline
- MeasuringBM25 nDCG@10 on the test split
- Target
- MeasuringCommitted before any variant runs
- Result
- MeasuringnDCG@10 of the recommended configuration
Faithfulness
Share of the claims in a generated answer that are supported by the retrieved context.
Higher is better
- Baseline
- MeasuringFaithfulness with BM25 retrieval, on the generation subset
- Target
- MeasuringCommitted before any variant runs
- Result
- MeasuringFaithfulness of the recommended configuration
Citation correctness
Share of citations where the cited passage actually supports the claim it is attached to. Judged, and calibrated.
Higher is better
- Baseline
- MeasuringCitation correctness with BM25 retrieval
- Target
- MeasuringCommitted before any variant runs
- Result
- MeasuringCitation correctness of the recommended configuration
p95 retrieval latency
95th-percentile time from query to ranked passages, including reranking, on recorded hardware.
Lower is better
- Baseline
- MeasuringBM25 p95 latency
- Target
- MeasuringLatency budget committed up front
- Result
- Measuringp95 latency of the recommended configuration
Cost per 1,000 queries
Embedding, reranking and generation spend for 1,000 queries, from the gateway's billing records.
Lower is better
- Baseline
- MeasuringBM25 cost per 1,000 queries
- Target
- MeasuringCost ceiling committed up front
- Result
- MeasuringCost of the recommended configuration
Eval design (offline)
Retrieval metrics use the benchmark's relevance judgements on its test split, computed with standard IR evaluation tooling so the numbers are comparable with published BEIR results. Generation metrics use a sampled subset, because each item costs a judge call.
Each variant changes exactly one component from the one before it, and every run records its commit hash, model versions and index configuration. A variant that changes two things at once cannot say which one helped.
Faithfulness and citation correctness are judged by an LLM calibrated against human labels. Retrieval metrics need no judge at all, which is why they carry the gate.
- Eval set
- MeasuringQueries in the test split with relevance judgements
- Generation subset
- MeasuringQueries sampled for faithfulness and citation scoring
- Judge calibration sample
- MeasuringItems labelled by hand to calibrate the judge
- Source
- Public dataset
- Split
- The benchmark's own test split for retrieval. The generation subset is drawn from it by a seeded script committed with the code.
- Stratification
- The generation subset is stratified by how many judged-relevant passages a query has, so queries with a single relevant passage, the hardest case, are not under-sampled.
- Versioning
- Index configurations, the generation subset and the judge prompt are all versioned with the code. Every result row cites the commit it came from.
Graders
Programmatic
Recall@10 and nDCG@10 from the benchmark's relevance judgements, with standard IR evaluation tooling.
LLM judge
Faithfulness and citation correctness. Judge model and version recorded per run.
Agreement with humans
MeasuringCohen's κ between the judge and human labels on the calibration sampleHuman labels
Claim-level support labels on the calibration sample, used only to calibrate the judge.
Online monitoring
There is no live traffic on a reference build, so online monitoring is simulated: held-out queries are replayed as a stream. We say so wherever the numbers appear.
Mid-stream we inject drift: relevant documents are removed from the index to simulate a stale corpus, and off-distribution queries are mixed in. The number reported is the delay between the drift going in and an alert firing.
Simulated: replayed traffic, not live users
Groundedness rate, judge-scored
AlertRolling-window rate drops past a threshold fixed up front
Empty or low-score retrievals
AlertShare rises past a threshold fixed up front
Top-1 retrieval score
AlertDistribution shifts from the offline baseline
p95 retrieval latency
AlertRises past the latency budget
- Judge sampling
- MeasuringShare of replayed queries sent to the judge
- Detection delay
- MeasuringTime from injected drift to the alert firing
Regression gate
Once a configuration is chosen, it becomes the bar. Any pull request that changes chunking, the embedding model or the index configuration must hold Recall@10 and faithfulness within a tolerance of the current best, or the check fails.
Retrieval metrics run on the full test split on every such PR, because they are cheap and deterministic. Faithfulness runs on the sampled subset.
- Where it runs
- A GitHub Actions job on every pull request that changes chunking, the embedding model or the index configuration.
- Noise floor
Retrieval is deterministic for a fixed index, so its tolerance can be tight. Generation is not, so the faithfulness tolerance sits above its measured run-to-run spread.
MeasuringRepeated generation runs the faithfulness tolerance is set from
What fails the check
Recall@10
Fails if it falls outside the tolerance of the current best
Faithfulness
Fails if it falls outside the tolerance of the current best
p95 retrieval latency
Fails if it exceeds the latency budget
Each limit's value is in the metric table above.
Result
One table: each variant against each metric, with its latency and cost. The recommended configuration is the cheapest variant that meets every target, not the one with the highest single score.
Variants
Each variant changes one component from the one before it. Scored on the same eval set, with the cost and latency each one pays.
Variant 1
BM25
Lexical baseline. The implementation is recorded with the run.
- Recall@10
- Measuring: Recall@10 for this variant
- nDCG@10
- Measuring: nDCG@10 for this variant
- Faithfulness
- Measuring: Faithfulness for this variant
- p95 latency
- Measuring: p95 retrieval latency for this variant
- Cost / 1k
- Measuring: Cost per 1,000 queries for this variant
Variant 2
Dense retrieval
One embedding model, chosen and recorded at run time.
- Recall@10
- Measuring: Recall@10 for this variant
- nDCG@10
- Measuring: nDCG@10 for this variant
- Faithfulness
- Measuring: Faithfulness for this variant
- p95 latency
- Measuring: p95 retrieval latency for this variant
- Cost / 1k
- Measuring: Cost per 1,000 queries for this variant
Variant 3
Hybrid
BM25 and dense results merged by reciprocal rank fusion.
- Recall@10
- Measuring: Recall@10 for this variant
- nDCG@10
- Measuring: nDCG@10 for this variant
- Faithfulness
- Measuring: Faithfulness for this variant
- p95 latency
- Measuring: p95 retrieval latency for this variant
- Cost / 1k
- Measuring: Cost per 1,000 queries for this variant
Variant 4
Hybrid plus reranker
A cross-encoder reranks the hybrid candidates. The first-stage depth it reranks is recorded, because the reranker's gain depends on it.
- Recall@10
- Measuring: Recall@10 for this variant
- nDCG@10
- Measuring: nDCG@10 for this variant
- Faithfulness
- Measuring: Faithfulness for this variant
- p95 latency
- Measuring: p95 retrieval latency for this variant
- Cost / 1k
- Measuring: Cost per 1,000 queries for this variant
Variant 5
Best variant, different chunk size
The best of the above with passages re-chunked. Chunks are mapped back to their source document before scoring, so the benchmark's judgements still apply.
- Recall@10
- Measuring: Recall@10 for this variant
- nDCG@10
- Measuring: nDCG@10 for this variant
- Faithfulness
- Measuring: Faithfulness for this variant
- p95 latency
- Measuring: p95 retrieval latency for this variant
- Cost / 1k
- Measuring: Cost per 1,000 queries for this variant
What we'd do differently
Written after the run, from what actually went wrong. Every missed regression or failed target gets an entry with its root cause.
Dataset and licence
Financial opinion questions with human relevance judgements over a passage corpus, as packaged in the BEIR benchmark.
- Dataset
- FiQA-2018 (BEIR)
- Licence
- Not yet confirmed
To be confirmed on the dataset card before any derived eval set is redistributed. Until then the repository ships the download and sampling scripts, not the data.
- Version
- Recorded with the first run, along with the download date.
- Corpus passages
- MeasuringPassages in the corpus at download time
- Test queries
- MeasuringQueries in the test split
Cite as
- Thakur, Reimers, Rücklé, Srivastava and Gurevych. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. NeurIPS Datasets and Benchmarks, 2021.
- Maia et al. WWW'18 Open Challenge: Financial Opinion Mining and Question Answering. Companion Proceedings of The Web Conference, 2018.
Stack
- Vector store
- Qdrant or pgvector, recorded per run
- Lexical
- BM25, implementation recorded per run
- Embedding and reranking
- Chosen and recorded at run time
- Grading
- Ragas, Standard IR evaluation tooling
- Gateway and tracing
- LiteLLM, Langfuse
- CI
- GitHub Actions
Models
Embedding
Chosen and recorded at run time, with its exact version
Reranker
Chosen and recorded at run time, with its exact version
Generator
Chosen and recorded at run time, with its exact version
Judge
Chosen and recorded at run time, with its exact version
What would make this fail
The ways this build could break, or produce numbers that mislead, and what we do about each. A method you cannot argue with is a method nobody checked.
Relevance judgements on FiQA are sparse. A passage nobody judged counts as irrelevant, which penalises a retriever that finds relevant passages the annotators never saw.
We hand-check a sample of top-ranked unjudged passages for the leading variants and report how often they were in fact relevant, next to the headline metrics.
Embedding and reranking models may have been trained on BEIR data, including FiQA, which flatters the dense variants.
The training-data disclosure of each model is checked and recorded with the run. Where a model is known to have seen FiQA, its row says so.
FiQA is opinion-style financial Q&A from forums. Your corpus of contracts, SOPs or tickets is different, and the ranking of variants may not transfer.
The page claims the method transfers, not the numbers. On a client corpus the same ladder is rerun on a set built from that corpus.
Changing one component at a time misses interactions, such as a reranker that only helps when the first stage retrieves deeply.
The reranker variant records its first-stage depth, and any interaction we find is reported as its own row rather than folded into the best result.
Latency depends on hardware and hosting, so a p95 from our machines is not a p95 from yours.
Hardware and hosting are recorded with every latency figure, and latency is compared between variants on the same setup, not across setups.
Faithfulness and citation correctness rest on an LLM judge calibrated on a small sample.
Agreement is reported with its n. If it does not clear the threshold fixed up front, those metrics are reported but do not carry the gate.
Evidence and publication
No client and no permission involved: we are the client, and the data is public. This page drops its in-progress label only when every item below is done, and the build refuses to publish it before then.
Still to do before publishing
- 52 values not yet measured
- “What we'd do differently” not yet written
- Repository with code and run logs not yet public
- Dataset licence not yet confirmed for redistribution
- Dataset version not yet recorded
- Model versions and run dates not yet recorded
Code, eval manifest and run logs are published with the first results. Every result row will cite the commit or run it came from.
Want this method run on your system?
A 30-minute call. We look at one system you are trying to get into production, name the metric that matters, and say what an eval gate on it would take.
hello@venian.aiReply within one business day, from an engineer.
What happens next
- 01
A reply within one business day
From an engineer who would work on it, not a sales team.
- 02
A 30-minute call
We map one process, name the metric it should move, and tell you plainly whether AI is the right tool for it.
- 03
A written scope
If it is, a one-page scope: the metric, how we baseline it, the target, and which engagement fits.
- For
- Teams with an LLM feature in or near production, or a pilot that stalled, and a number it should move.
- Not for
- Chatbot-on-a-website projects, or strategy decks with no build behind them.