LLM evaluation that gates every deploy

Short answer

LLM evaluation consulting builds the system that tells you whether an LLM feature is good enough to ship, and whether the next change made it worse. An eval programme has four parts. An offline eval set: real inputs from your traffic, each graded against a written rubric. Evaluators: code checks where the answer is checkable, and an LLM-as-judge where it is not, calibrated against human labels before anyone trusts its score. Online evals: the same evaluators run on a sample of live traffic, so drift shows up on a dashboard rather than in a support ticket. And a regression gate in CI that fails the build when a prompt edit, a model swap or a retrieval change drops a score past an agreed threshold.

VenianAI starts with a two-week Eval Audit: error analysis on your real traces, a graded golden set you own, and a baseline scorecard. The Regression Gate then wires that suite into CI.

What is an LLM eval programme?

Four parts, each catching a failure the others cannot. An eval set with no gate is a report nobody reruns; a gate with an uncalibrated judge blocks good changes and passes bad ones.

01

An offline eval set

Real inputs taken from your traces, each with a written rubric for what a correct answer must do. It runs before a change ships, on demand, against the current system and the candidate, so the two can be compared case by case.

Offline and online evals compared
02

Evaluators you can trust

Code checks wherever the answer is checkable: schema, citations, tool arguments, forbidden content. An LLM-as-judge where it is not, scored against human labels first and used only once its agreement is known and written down.

03

Online evals on live traffic

The same evaluators run on a sample of production requests after the fact. That is what catches the inputs nobody thought to put in the offline set, and the slow drift that no single deploy caused.

04

A regression gate in CI

Every pull request that touches a prompt, a model setting, retrieval or a tool definition runs the suite. The build fails when a score drops below its floor or past the agreed tolerance, and the diff is posted on the PR.

How to regression-test an LLM system

How do you grade open-ended output?

With the cheapest method that can decide the question. Many criteria people reach for a judge to grade can be checked in code, and the ones that cannot need a judge whose error rate is known.

MethodUse it forWhy you can trust itCost
Code checksOutput that is checkable: valid JSON against a schema, required fields present, cited document IDs that exist, a tool called with valid arguments, banned strings absent.Deterministic. Unit-tested like the rest of the codebase.Effectively free. Run on every case and every live request.
Binary rubricOne pass-or-fail question per failure mode: “Does every claim in the answer appear in a retrieved passage?” Written from error analysis, not from a generic quality list.A pass/fail question has one boundary for two graders to agree on. A five-point scale has four.The rubric is written once. Applying it is the judge's or the reviewer's job.
LLM-as-judgeApplying a rubric at scale where code cannot: faithfulness, tone, whether the answer addressed the question that was asked.Only after its verdicts are compared with human labels on a held-out set, reporting how often it catches real failures and how often it passes real successes. Re-checked when the judge model or the rubric changes.One model call per case per criterion. Usually run on a sample of live traffic, not all of it.
Human reviewLabelling the calibration set, settling disagreements between graders, reading the production traces the evaluators flag.A domain expert working from written guidelines. Disagreements go back into the guidelines.The most expensive per case, so it is spent where it calibrates everything else.

How many examples does an eval set need?

Enough to make the decision you are gating on. A pass rate measured on a small set has a wide margin of error, and a gate tighter than that margin fails on noise.

Start with 50 to 200 cases chosen from error analysis, not generated in bulk. A case earns its place by representing a failure you have seen or a path the product cannot afford to break. A thousand synthetic happy-path questions measure very little.

Coverage by slice matters more than the total. If refund requests are a tenth of your traffic and the set has ten of them, the set can tell you whether refunds broke completely, not whether they got slightly worse. Slices you gate on separately need enough cases to be measured separately.

The set then grows from production. Every failure an online eval or a user report surfaces becomes a case, so the set ends up weighted toward the inputs that actually go wrong.

Margin of error on a 90% pass rate

Cases95% interval
50±8.3 pts
100±5.9 pts
200±4.2 pts
400±2.9 pts

Arithmetic, not a measurement: 1.96 × √(p(1−p)/n)

What does an eval gate in CI look like?

A required check on the pull request, the same as tests and lint. It runs the suite against the candidate, compares it with the baseline from main, and fails the build on any of three conditions.

  1. 01

    Below the floor

    An evaluator's pass rate falls under the absolute minimum agreed for it, whatever the baseline was.

  2. 02

    Past the tolerance

    The drop against the baseline is larger than the run-to-run noise measured on that suite. Smaller drops are reported, not blocked.

  3. 03

    A critical case fails

    Cases marked critical, such as a refusal that must hold or a figure that must be exact, have to pass on every sample.

Each case runs several times at production temperature, so a single unlucky sample cannot fail the build and a single lucky one cannot pass it. The score diff is posted as a comment on the pull request, case by case, so the reviewer sees what changed, not only that something did.

The full guide, with a working CI step

What do we hand over?

Two engagements cover evaluation. Both are fixed scope, and both leave everything in your repository: the set, the evaluators and the gate are yours, not a tenant in our account. Fixed price, quoted after the scoping call.

2 weeks

Eval Audit

For a team with an LLM feature in or near production and no quality signal it trusts.

  • Error analysis on 100 to 300 of your real traces
  • A failure taxonomy for your system
  • A graded golden set of 50 to 200 cases, owned by you
  • 3 to 5 binary evaluators, with their agreement against human labels reported
  • A baseline scorecard and a written regression and monitoring plan

You know how good the system is today, where it fails, and what to gate on.

Eval Audit in full

1 to 2 weeks

Regression Gate

For teams that change prompts or models every week.

  • Your eval suite wired into CI
  • Pass and fail thresholds per evaluator
  • A comment on every pull request with the score diff
  • Playbooks for model swaps and prompt changes

A bad prompt or a model swap fails the build instead of reaching users.

Regression Gate in full

How does an evals engagement run?

From the first call to a gate on your pull requests. The Eval Audit is steps three and four; the rest is what happens either side of it.

  1. 01

    Scoping call

    Thirty minutes. Which feature, which failure you are worried about, and which decision the evals have to support: ship or not, model A or model B, keep the prompt or revert it.

  2. 02

    Access

    Read access to traces or logs, and a few hours from someone who knows what a correct answer looks like. Nothing is labelled without a domain expert.

  3. 03

    Error analysis, week one

    We read your real traces, group what goes wrong into a failure taxonomy, and choose what is worth measuring. Most of the value of an eval set is decided here.

  4. 04

    Build and calibrate, week two

    The golden set, the evaluators, the judge's agreement with human labels, and a baseline scorecard. Handed over in your repository.

  5. 05

    Gate, if you change things weekly

    The Regression Gate wires the suite into CI with a threshold per evaluator, so the next prompt edit is scored before it merges.

  6. 06

    Keep it current

    Production failures go back into the set. Your team can own that loop, or the Reliability Retainer runs it monthly.

All four engagements, and how they connect

Who is LLM evaluation consulting for?

Teams that already have something to measure. An eval set is built from real traces, so it comes after the first working version, not before it.

A good fit

  • An LLM feature in or near production, changed every week or two, where nobody can say with evidence whether the last change helped.
  • A pilot that stalled because nobody could show it worked well enough to ship.
  • A team about to swap model providers, or upgrade a model version, and needing to know what breaks first.
Why missing evaluation stalls pilots

Not a fit

  • No working version yet. There are no traces to learn from; start with the Pilot-to-Production Sprint, which builds the evals alongside the system.
  • A ranking of public models on public benchmarks. That measures someone else’s task.
  • Choosing an eval vendor with no system to evaluate. The tool matters far less than the set and the rubric.

Questions about LLM evaluation

What is the difference between an LLM eval and a test?

A unit test asserts one exact output and either passes or fails. An LLM eval scores a set of cases against a rubric and reports a pass rate, because the same input can produce different acceptable outputs and the question is how often the system meets the bar, not whether one run did. A regression gate is where the two meet: the eval runs in CI and the build fails when the pass rate drops past a threshold.

Can an LLM reliably grade another LLM's output?

Only once it has been checked. An LLM-as-judge is a classifier, and like any classifier it has a false-pass rate and a false-fail rate. We measure both against human labels on a held-out set before the judge's scores are used for anything, keep its prompt to one binary question, and re-measure whenever the judge model or the rubric changes.

Do we need to buy an eval platform?

No. An eval programme is a golden set, a set of evaluators and a gate, and all three can live as plain files and code in your own repository. If you already run a tracing or eval tool such as Langfuse or Ragas we build on it; if you do not, we do not make one a condition. What you own at the end is the same either way.

How is this different from a public model benchmark?

A public benchmark measures a model on someone else's task. An eval set measures your system, including its prompts, retrieval and tools, on your task, with your inputs and your definition of correct. A model can gain on a leaderboard and still regress on your eval set, which is exactly the case a regression gate exists to catch.

How long until we have a quality baseline?

The Eval Audit runs for two weeks and ends with a baseline scorecard: the pass rate of the current system on a graded golden set, broken down by failure mode. That baseline is what every later change, and every target in a statement of work, is measured against.

Find out how good it is today

A 30-minute call. We look at one LLM feature, the failure you are worried about, and the decision the evals have to support, then tell you whether an Eval Audit is the right first step.

hello@venian.aiReply within one business day, from an engineer.

What happens next

  1. 01

    A reply within one business day

    From an engineer who would work on it, not a sales team.

  2. 02

    A 30-minute call

    We map one process, name the metric it should move, and tell you plainly whether AI is the right tool for it.

  3. 03

    A written scope

    If it is, a one-page scope: the metric, how we baseline it, the target, and which engagement fits.

For
Teams with an LLM feature in or near production, or a pilot that stalled, and a number it should move.
Not for
Chatbot-on-a-website projects, or strategy decks with no build behind them.