Reference build · public dataIn progress

RAG retrieval on FiQA, from a BM25 baseline to hybrid search and a reranker

Reference build · in progress. This is not client work. VenianAI is building and measuring this system on a public dataset to document the method. The method below is the one the run follows. The results are not in yet, and each one says so where it will appear.

Summary

We are measuring retrieval on a public financial question-answering benchmark, starting from a plain BM25 baseline and changing one component at a time: dense retrieval, hybrid search, a reranker, then chunk size.

Every variant is scored on the same offline eval set, with retrieval and generation measured separately so a wrong answer can be traced to the stage that caused it, and the targets are committed before the first variant runs.

This is a reference build on public data, not client work, and every result on this page reads as measuring until a logged run produces it.

Failure mode
Poor data access
Maturity
L1 → L2
Results measured
0 of 52
Updated

Context

Teams at L1 who have grounded an assistant in their own documents usually cannot say whether a wrong answer came from retrieval or from generation. Both look the same from the outside: a confident answer that is not supported by the source.

This build separates the two and measures each. Retrieval is scored against human relevance judgements, generation against the passages actually retrieved. The method is the one we use on client corpora; the corpus here is public so anyone can rerun it.

The question it answers

When the assistant gives a wrong answer, was it retrieval or generation, and which change fixes it?

Failure mode

Poor data access. The model is only as good as the passages it is handed, and the gap between a demo and a useful system is almost always retrieval: chunking, embedding, hybrid search and reranking.

The pattern this targets is an assistant that answers fluently from a passage that is on topic but does not contain the answer. Without retrieval measured on its own, the fix people reach for is a bigger model, which makes the answer more fluent and no more correct.

The fault, as we define it

Poor data access: The model can't reach the system that holds the answer.

The test: When your assistant answers a question, can it point at the document or record it came from?

Read the full entry on poor data access

Metric agreed

Two retrieval metrics on the benchmark's own relevance judgements, two generation metrics on a sampled subset, and the latency and cost every variant pays. Targets are committed in the repository's first commit, before any variant runs, so the results cannot be framed after the fact.

  • Recall@10

    Share of the judged-relevant passages for a query that appear in the top 10 retrieved.

    Higher is better

    Baseline
    MeasuringBM25 Recall@10 on the test split
    Target
    MeasuringCommitted before any variant runs
    Result
    MeasuringRecall@10 of the recommended configuration
  • nDCG@10

    The standard BEIR ranking metric: rewards relevant passages near the top of the list more than lower down.

    Higher is better

    Baseline
    MeasuringBM25 nDCG@10 on the test split
    Target
    MeasuringCommitted before any variant runs
    Result
    MeasuringnDCG@10 of the recommended configuration
  • Faithfulness

    Share of the claims in a generated answer that are supported by the retrieved context.

    Higher is better

    Baseline
    MeasuringFaithfulness with BM25 retrieval, on the generation subset
    Target
    MeasuringCommitted before any variant runs
    Result
    MeasuringFaithfulness of the recommended configuration
  • Citation correctness

    Share of citations where the cited passage actually supports the claim it is attached to. Judged, and calibrated.

    Higher is better

    Baseline
    MeasuringCitation correctness with BM25 retrieval
    Target
    MeasuringCommitted before any variant runs
    Result
    MeasuringCitation correctness of the recommended configuration
  • p95 retrieval latency

    95th-percentile time from query to ranked passages, including reranking, on recorded hardware.

    Lower is better

    Baseline
    MeasuringBM25 p95 latency
    Target
    MeasuringLatency budget committed up front
    Result
    Measuringp95 latency of the recommended configuration
  • Cost per 1,000 queries

    Embedding, reranking and generation spend for 1,000 queries, from the gateway's billing records.

    Lower is better

    Baseline
    MeasuringBM25 cost per 1,000 queries
    Target
    MeasuringCost ceiling committed up front
    Result
    MeasuringCost of the recommended configuration

Eval design (offline)

Retrieval metrics use the benchmark's relevance judgements on its test split, computed with standard IR evaluation tooling so the numbers are comparable with published BEIR results. Generation metrics use a sampled subset, because each item costs a judge call.

Each variant changes exactly one component from the one before it, and every run records its commit hash, model versions and index configuration. A variant that changes two things at once cannot say which one helped.

Faithfulness and citation correctness are judged by an LLM calibrated against human labels. Retrieval metrics need no judge at all, which is why they carry the gate.

Eval set
MeasuringQueries in the test split with relevance judgements
Generation subset
MeasuringQueries sampled for faithfulness and citation scoring
Judge calibration sample
MeasuringItems labelled by hand to calibrate the judge
Source
Public dataset
Split
The benchmark's own test split for retrieval. The generation subset is drawn from it by a seeded script committed with the code.
Stratification
The generation subset is stratified by how many judged-relevant passages a query has, so queries with a single relevant passage, the hardest case, are not under-sampled.
Versioning
Index configurations, the generation subset and the judge prompt are all versioned with the code. Every result row cites the commit it came from.

Graders

  • Programmatic

    Recall@10 and nDCG@10 from the benchmark's relevance judgements, with standard IR evaluation tooling.

  • LLM judge

    Faithfulness and citation correctness. Judge model and version recorded per run.

    Agreement with humans

    MeasuringCohen's κ between the judge and human labels on the calibration sample
  • Human labels

    Claim-level support labels on the calibration sample, used only to calibrate the judge.

Online monitoring

There is no live traffic on a reference build, so online monitoring is simulated: held-out queries are replayed as a stream. We say so wherever the numbers appear.

Mid-stream we inject drift: relevant documents are removed from the index to simulate a stale corpus, and off-distribution queries are mixed in. The number reported is the delay between the drift going in and an alert firing.

Simulated: replayed traffic, not live users

  • Groundedness rate, judge-scored

    AlertRolling-window rate drops past a threshold fixed up front

  • Empty or low-score retrievals

    AlertShare rises past a threshold fixed up front

  • Top-1 retrieval score

    AlertDistribution shifts from the offline baseline

  • p95 retrieval latency

    AlertRises past the latency budget

Judge sampling
MeasuringShare of replayed queries sent to the judge
Detection delay
MeasuringTime from injected drift to the alert firing

Regression gate

Once a configuration is chosen, it becomes the bar. Any pull request that changes chunking, the embedding model or the index configuration must hold Recall@10 and faithfulness within a tolerance of the current best, or the check fails.

Retrieval metrics run on the full test split on every such PR, because they are cheap and deterministic. Faithfulness runs on the sampled subset.

Where it runs
A GitHub Actions job on every pull request that changes chunking, the embedding model or the index configuration.
Noise floor

Retrieval is deterministic for a fixed index, so its tolerance can be tight. Generation is not, so the faithfulness tolerance sits above its measured run-to-run spread.

MeasuringRepeated generation runs the faithfulness tolerance is set from

What fails the check

  • Recall@10

    Fails if it falls outside the tolerance of the current best

  • Faithfulness

    Fails if it falls outside the tolerance of the current best

  • p95 retrieval latency

    Fails if it exceeds the latency budget

Each limit's value is in the metric table above.

Result

One table: each variant against each metric, with its latency and cost. The recommended configuration is the cheapest variant that meets every target, not the one with the highest single score.

Variants

Each variant changes one component from the one before it. Scored on the same eval set, with the cost and latency each one pays.

  • Variant 1

    BM25

    Lexical baseline. The implementation is recorded with the run.

    Recall@10
    Measuring: Recall@10 for this variant
    nDCG@10
    Measuring: nDCG@10 for this variant
    Faithfulness
    Measuring: Faithfulness for this variant
    p95 latency
    Measuring: p95 retrieval latency for this variant
    Cost / 1k
    Measuring: Cost per 1,000 queries for this variant
  • Variant 2

    Dense retrieval

    One embedding model, chosen and recorded at run time.

    Recall@10
    Measuring: Recall@10 for this variant
    nDCG@10
    Measuring: nDCG@10 for this variant
    Faithfulness
    Measuring: Faithfulness for this variant
    p95 latency
    Measuring: p95 retrieval latency for this variant
    Cost / 1k
    Measuring: Cost per 1,000 queries for this variant
  • Variant 3

    Hybrid

    BM25 and dense results merged by reciprocal rank fusion.

    Recall@10
    Measuring: Recall@10 for this variant
    nDCG@10
    Measuring: nDCG@10 for this variant
    Faithfulness
    Measuring: Faithfulness for this variant
    p95 latency
    Measuring: p95 retrieval latency for this variant
    Cost / 1k
    Measuring: Cost per 1,000 queries for this variant
  • Variant 4

    Hybrid plus reranker

    A cross-encoder reranks the hybrid candidates. The first-stage depth it reranks is recorded, because the reranker's gain depends on it.

    Recall@10
    Measuring: Recall@10 for this variant
    nDCG@10
    Measuring: nDCG@10 for this variant
    Faithfulness
    Measuring: Faithfulness for this variant
    p95 latency
    Measuring: p95 retrieval latency for this variant
    Cost / 1k
    Measuring: Cost per 1,000 queries for this variant
  • Variant 5

    Best variant, different chunk size

    The best of the above with passages re-chunked. Chunks are mapped back to their source document before scoring, so the benchmark's judgements still apply.

    Recall@10
    Measuring: Recall@10 for this variant
    nDCG@10
    Measuring: nDCG@10 for this variant
    Faithfulness
    Measuring: Faithfulness for this variant
    p95 latency
    Measuring: p95 retrieval latency for this variant
    Cost / 1k
    Measuring: Cost per 1,000 queries for this variant

What we'd do differently

Not written yet

Written after the run, from what actually went wrong. Every missed regression or failed target gets an entry with its root cause.

Dataset and licence

Financial opinion questions with human relevance judgements over a passage corpus, as packaged in the BEIR benchmark.

Licence
Not yet confirmed

To be confirmed on the dataset card before any derived eval set is redistributed. Until then the repository ships the download and sampling scripts, not the data.

Version
Recorded with the first run, along with the download date.
Corpus passages
MeasuringPassages in the corpus at download time
Test queries
MeasuringQueries in the test split

Cite as

  • Thakur, Reimers, Rücklé, Srivastava and Gurevych. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. NeurIPS Datasets and Benchmarks, 2021.
  • Maia et al. WWW'18 Open Challenge: Financial Opinion Mining and Question Answering. Companion Proceedings of The Web Conference, 2018.

Stack

Vector store
Qdrant or pgvector, recorded per run
Lexical
BM25, implementation recorded per run
Embedding and reranking
Chosen and recorded at run time
Grading
Ragas, Standard IR evaluation tooling
Gateway and tracing
LiteLLM, Langfuse
CI
GitHub Actions

Models

  • Embedding

    Chosen and recorded at run time, with its exact version

  • Reranker

    Chosen and recorded at run time, with its exact version

  • Generator

    Chosen and recorded at run time, with its exact version

  • Judge

    Chosen and recorded at run time, with its exact version

What would make this fail

The ways this build could break, or produce numbers that mislead, and what we do about each. A method you cannot argue with is a method nobody checked.

  • Relevance judgements on FiQA are sparse. A passage nobody judged counts as irrelevant, which penalises a retriever that finds relevant passages the annotators never saw.

    We hand-check a sample of top-ranked unjudged passages for the leading variants and report how often they were in fact relevant, next to the headline metrics.

  • Embedding and reranking models may have been trained on BEIR data, including FiQA, which flatters the dense variants.

    The training-data disclosure of each model is checked and recorded with the run. Where a model is known to have seen FiQA, its row says so.

  • FiQA is opinion-style financial Q&A from forums. Your corpus of contracts, SOPs or tickets is different, and the ranking of variants may not transfer.

    The page claims the method transfers, not the numbers. On a client corpus the same ladder is rerun on a set built from that corpus.

  • Changing one component at a time misses interactions, such as a reranker that only helps when the first stage retrieves deeply.

    The reranker variant records its first-stage depth, and any interaction we find is reported as its own row rather than folded into the best result.

  • Latency depends on hardware and hosting, so a p95 from our machines is not a p95 from yours.

    Hardware and hosting are recorded with every latency figure, and latency is compared between variants on the same setup, not across setups.

  • Faithfulness and citation correctness rest on an LLM judge calibrated on a small sample.

    Agreement is reported with its n. If it does not clear the threshold fixed up front, those metrics are reported but do not carry the gate.

Evidence and publication

No client and no permission involved: we are the client, and the data is public. This page drops its in-progress label only when every item below is done, and the build refuses to publish it before then.

Still to do before publishing

  • 52 values not yet measured
  • “What we'd do differently” not yet written
  • Repository with code and run logs not yet public
  • Dataset licence not yet confirmed for redistribution
  • Dataset version not yet recorded
  • Model versions and run dates not yet recorded

Code, eval manifest and run logs are published with the first results. Every result row will cite the commit or run it came from.

Want this method run on your system?

A 30-minute call. We look at one system you are trying to get into production, name the metric that matters, and say what an eval gate on it would take.

hello@venian.aiReply within one business day, from an engineer.

What happens next

  1. 01

    A reply within one business day

    From an engineer who would work on it, not a sales team.

  2. 02

    A 30-minute call

    We map one process, name the metric it should move, and tell you plainly whether AI is the right tool for it.

  3. 03

    A written scope

    If it is, a one-page scope: the metric, how we baseline it, the target, and which engagement fits.

For
Teams with an LLM feature in or near production, or a pilot that stalled, and a number it should move.
Not for
Chatbot-on-a-website projects, or strategy decks with no build behind them.