Case studies and reference builds

Short answer

Every write-up here follows one structure: the context, the failure mode, the metric agreed before work started, the offline eval design, online monitoring, the regression gate, the measured result, and what we would do differently.

Client case studies are published only with written permission and only once every number in them is measured. None meets that bar yet, so none is shown.

Reference builds are systems we build and measure on public data, run in the open so the method can be checked before it is trusted. Each is labelled as a reference build, never as client work, and results read as measuring until a logged run produces them.

What gets published, and when

Every number on these pages has a source we can send to your engineers on request: a client report, a signed-off dashboard export, or a public repository and its run logs. If we cannot send it, we do not print it.

Client studies need permission and real numbers

A client engagement is published only with the client's written permission, an anonymised one included, and only once every number in it is measured, with its window and sample size, and approved by the client. We do not write composite or representative clients, and we do not round in the flattering direction. Until one meets that bar, there is nothing here to show.

Reference builds run in the open

A reference build is a system we build and measure ourselves on a public dataset, so you can check the method before you trust it with your data. It is always labelled as such and never presented as client work. While it is in progress the method is published and every result reads as measuring. Code, eval manifest and run logs go public with the first results.

The structure every study follows

The headings are fixed, in this order, for client studies and reference builds alike. A study that cannot fill one of them is not ready to publish.

  1. 01

    Context

    Who it is, what they had, and where they sat on the maturity ladder.

  2. 02

    Failure mode

    What was actually breaking, mapped to one of the seven faults that keep pilots out of production.

  3. 03

    Metric agreed

    The number, its baseline with method and window, and the target, fixed before any work started.

  4. 04

    Offline eval design

    The eval set: its size, source, stratification and labels, the graders, and how the judge was calibrated.

  5. 05

    Online monitoring

    What is watched on live traffic, how it is sampled, and what fires an alert.

  6. 06

    Regression gate

    Where it runs, the limit per metric, and what happened the first time it blocked a deploy.

  7. 07

    Result

    Measured values against baseline and target, each with its window and n. A table, not prose.

  8. 08

    What we'd do differently

    At least one real item, written after the run. Never a humble-brag.

Want this method on your system?

A 30-minute call. We look at one system you are trying to get into production, name the metric that matters, and say what an eval gate on it would take.

hello@venian.aiReply within one business day, from an engineer.

What happens next

  1. 01

    A reply within one business day

    From an engineer who would work on it, not a sales team.

  2. 02

    A 30-minute call

    We map one process, name the metric it should move, and tell you plainly whether AI is the right tool for it.

  3. 03

    A written scope

    If it is, a one-page scope: the metric, how we baseline it, the target, and which engagement fits.

For
Teams with an LLM feature in or near production, or a pilot that stalled, and a number it should move.
Not for
Chatbot-on-a-website projects, or strategy decks with no build behind them.