Case studies and reference builds
Short answer
Every write-up here follows one structure: the context, the failure mode, the metric agreed before work started, the offline eval design, online monitoring, the regression gate, the measured result, and what we would do differently.
Client case studies are published only with written permission and only once every number in them is measured. None meets that bar yet, so none is shown.
Reference builds are systems we build and measure on public data, run in the open so the method can be checked before it is trusted. Each is labelled as a reference build, never as client work, and results read as measuring until a logged run produces them.
What gets published, and when
Every number on these pages has a source we can send to your engineers on request: a client report, a signed-off dashboard export, or a public repository and its run logs. If we cannot send it, we do not print it.
Client studies need permission and real numbers
A client engagement is published only with the client's written permission, an anonymised one included, and only once every number in it is measured, with its window and sample size, and approved by the client. We do not write composite or representative clients, and we do not round in the flattering direction. Until one meets that bar, there is nothing here to show.
Reference builds run in the open
A reference build is a system we build and measure ourselves on a public dataset, so you can check the method before you trust it with your data. It is always labelled as such and never presented as client work. While it is in progress the method is published and every result reads as measuring. Code, eval manifest and run logs go public with the first results.
Reference builds
Systems we built and measured ourselves on public data, so you can check the method before you trust it with yours.
- Reference build · public dataIn progress
A regression gate for a support agent, tested against regressions we plant ourselves
If someone changes the system tonight, will anything tell you tomorrow whether it got worse?
- Failure mode
- No evaluation
- Maturity
- L2 → L2
- Results measured
- 0 of 36
- Updated
- Reference build · public dataIn progress
RAG retrieval on FiQA, from a BM25 baseline to hybrid search and a reranker
When the assistant gives a wrong answer, was it retrieval or generation, and which change fixes it?
- Failure mode
- Poor data access
- Maturity
- L1 → L2
- Results measured
- 0 of 52
- Updated
The structure every study follows
The headings are fixed, in this order, for client studies and reference builds alike. A study that cannot fill one of them is not ready to publish.
- 01
Context
Who it is, what they had, and where they sat on the maturity ladder.
- 02
Failure mode
What was actually breaking, mapped to one of the seven faults that keep pilots out of production.
- 03
Metric agreed
The number, its baseline with method and window, and the target, fixed before any work started.
- 04
Offline eval design
The eval set: its size, source, stratification and labels, the graders, and how the judge was calibrated.
- 05
Online monitoring
What is watched on live traffic, how it is sampled, and what fires an alert.
- 06
Regression gate
Where it runs, the limit per metric, and what happened the first time it blocked a deploy.
- 07
Result
Measured values against baseline and target, each with its window and n. A table, not prose.
- 08
What we'd do differently
At least one real item, written after the run. Never a humble-brag.
Want this method on your system?
A 30-minute call. We look at one system you are trying to get into production, name the metric that matters, and say what an eval gate on it would take.
hello@venian.aiReply within one business day, from an engineer.
What happens next
- 01
A reply within one business day
From an engineer who would work on it, not a sales team.
- 02
A 30-minute call
We map one process, name the metric it should move, and tell you plainly whether AI is the right tool for it.
- 03
A written scope
If it is, a one-page scope: the metric, how we baseline it, the target, and which engagement fits.
- For
- Teams with an LLM feature in or near production, or a pilot that stalled, and a number it should move.
- Not for
- Chatbot-on-a-website projects, or strategy decks with no build behind them.