Reference build · public dataIn progress

A regression gate for a support agent, tested against regressions we plant ourselves

Reference build · in progress. This is not client work. VenianAI is building and measuring this system on a public dataset to document the method. The method below is the one the run follows. The results are not in yet, and each one says so where it will appear.

Summary

We are building a ticket-triage-and-answer agent on a public customer-support dataset, wrapping it in an offline eval suite, and wiring that suite into CI as a deploy gate.

We then plant the regressions that reach production in real systems, a prompt edit, a model swap and a stale knowledge base, and record which ones the gate blocks, on which metric, and which it misses.

This is a reference build on public data, not client work, and every result on this page reads as measuring until a logged run produces it.

Failure mode
No evaluation
Maturity
L2 → L2L2 with deploy gates, the prerequisite for L3
Results measured
0 of 36
Updated

Context

Most teams at L2 change prompts weekly, swap model versions when a provider ships one, and edit the knowledge base whenever a policy changes. Each of those changes behaviour without a code diff, and none of them fails a unit test.

This build is the smallest honest version of the system that catches that: an agent that does a real support job, a graded eval set, and a CI job that refuses a merge when quality drops. It is the method we install on client systems, run in the open on data anyone can download.

The question it answers

If someone changes the system tonight, will anything tell you tomorrow whether it got worse?

Failure mode

No evaluation. Quality is assessed by whoever last looked at the output, so a change that makes the agent worse looks exactly like a change that makes it better until a customer notices.

The regressions that matter are rarely dramatic. A removed sentence in a system prompt can stop escalations without changing the tone of a single answer. An outdated policy article produces answers that are fluent, confident and wrong. Neither shows up in a demo.

The fault, as we define it

No evaluation: Nobody can say whether it got better or worse this week.

The test: If someone changed the system prompt tonight, what would tell you tomorrow whether it got worse?

Read the full entry on no evaluation

Metric agreed

Four quality metrics and one cost metric, each with a definition precise enough that two readers compute the same number. Baselines come from the unchanged agent on the held-out split. Gate limits are fixed in the repository's first commit, before the first planted regression runs, so the thresholds cannot be tuned to the result.

  • Intent accuracy

    Share of held-out tickets where the predicted intent exactly matches the dataset label. Also reported per intent.

    Higher is better

    Baseline
    MeasuringAccuracy of the unchanged agent on the held-out split
    Gate limit
    MeasuringLargest drop, in percentage points, a PR may cause
    Result
    MeasuringAccuracy under each planted regression
  • Answer quality

    Mean rubric score from 1 to 5, given by an LLM judge calibrated against human labels.

    Higher is better

    Baseline
    MeasuringMean judge score of the unchanged agent
    Gate limit
    MeasuringLargest drop in mean score a PR may cause
    Result
    MeasuringMean judge score under each planted regression
  • Escalation precision

    Of the tickets the agent escalated to a human, the share our rubric says should have been escalated.

    Higher is better

    Baseline
    MeasuringPrecision of the unchanged agent
    Gate limit
    MeasuringLargest drop in precision a PR may cause
    Result
    MeasuringPrecision under each planted regression
  • Escalation recall

    Of the tickets our rubric says should reach a human, the share the agent escalated. Catches the agent that quietly stops escalating.

    Higher is better

    Baseline
    MeasuringRecall of the unchanged agent
    Gate limit
    MeasuringLargest drop in recall a PR may cause
    Result
    MeasuringRecall under each planted regression
  • Cost per 1,000 tickets

    Spend reported by the gateway for the held-out split, normalised to 1,000 tickets.

    Lower is better

    Baseline
    MeasuringGateway-reported cost of the unchanged agent
    Gate limit
    MeasuringLargest cost increase a PR may cause
    Result
    MeasuringCost under each planted regression

Eval design (offline)

The eval set is a stratified sample of the dataset with a held-out split that is never used for prompt iteration. Prompt work happens on a separate development split, so the gate is never grading the examples the prompt was tuned on.

Intent is graded by exact match against the dataset label. Answer quality is graded by an LLM judge against a rubric published in the repository. Escalation is graded programmatically against labels we write ourselves, because the dataset does not say which tickets should reach a human.

The judge only gates once it has been calibrated against human labels and its agreement clears a threshold we set before seeing the calibration result. If it does not clear, answer quality is reported but does not block a merge.

Eval set
MeasuringItems in the held-out split, stratified by intent
Judge calibration sample
MeasuringItems labelled by hand to calibrate the judge
Source
Public dataset
Split
A seeded script in the first commit fixes a development split and a held-out split. Prompt iteration only ever sees the development split.
Stratification
Stratified by intent label, so a rare intent cannot drop out of the set and a regression on it cannot hide inside the average.
Versioning
The eval set is a manifest of item IDs plus the sampling script, versioned with the code. Changing the set is a reviewed change of its own, never folded into a PR that also changes the agent.

Graders

  • Exact match

    Predicted intent against the dataset's intent label.

  • LLM judge

    Answer quality against a published 1-to-5 rubric. Judge model and version recorded per run, from a different model family than the agent so it is not grading its own style.

    Agreement with humans

    MeasuringCohen's κ between the judge and human labels on the calibration sample
  • Programmatic

    Escalation decision against our own should-escalate labels, written to a rubric in the repository.

  • Human labels

    Labels on the calibration sample, written against the same rubric the judge uses. Used to calibrate, never to tune the agent.

Online monitoring

There is no live traffic on a reference build, so online monitoring is simulated: a held-out ticket stream is replayed through the deployed agent with full traces. We say so wherever the numbers appear.

A sample of replayed traffic is scored by the same judge, and alerts fire on a rolling-window drop rather than a single bad answer. Drift is planted mid-stream, and the number reported is the time from the drift going in to the alert firing.

Simulated: replayed traffic, not live users

  • Answer quality, judge-scored

    AlertRolling-window mean drops past the gate limit

  • Escalation rate

    AlertDeparts from the offline baseline rate in either direction

  • Cost per 1,000 tickets

    AlertRises past the gate limit

  • p95 latency

    AlertRises past a budget fixed in the first commit

Judge sampling
MeasuringShare of replayed tickets sent to the judge
Detection delay
MeasuringTime from planted drift to the alert firing

Regression gate

The gate is a CI job, not a dashboard someone has to remember to check. It runs the suite on every pull request that touches prompts, model configuration or the knowledge base, and fails the check if any metric breaches its limit.

LLM output is non-deterministic, so the unchanged system is run across several seeds first and every limit sits above that noise floor. A gate that flaps on benign changes gets switched off within a month; that is a failure mode of its own, and the build measures it.

Where it runs
A GitHub Actions job on every pull request that touches prompts, model configuration or the knowledge base.
Noise floor

Every limit sits above the run-to-run spread of the unchanged system, so the gate does not flap.

MeasuringSeeds the unchanged system is run across

What fails the check

  • Intent accuracy

    Fails on a drop past its limit

  • Answer quality

    Fails on a drop past its limit, once the judge is calibrated

  • Escalation precision and recall

    Fails on a drop in either past its limit

  • Cost per 1,000 tickets

    Fails on a rise past its limit

Each limit's value is in the metric table above.

Result

Each planted regression is one row: caught or missed, and on which metric. Benign changes are a control group, and any block on one counts as a false block. The table fills in from logged runs only.

Planted regressions

Three classes of change that alter behaviour without a code diff, plus a control group of changes that should not. The metric we expect each one to move is written down before it runs.

  • Prompt edit

    Escalation instruction removed

    The sentence telling the agent when to hand a ticket to a human is deleted from the system prompt.

    Expected to move: Escalation recall

    Caught
    Measuring: Whether the gate blocked it
    Metric
    Measuring: Which metric breached first
  • Prompt edit

    Answer format loosened

    The instruction to cite the policy article an answer relies on is softened to a suggestion.

    Expected to move: Answer quality

    Caught
    Measuring: Whether the gate blocked it
    Metric
    Measuring: Which metric breached first
  • Model swap

    Smaller model, same family

    The agent's model is replaced by a cheaper, smaller one from the same provider.

    Expected to move: Intent accuracy, answer quality

    Caught
    Measuring: Whether the gate blocked it
    Metric
    Measuring: Which metric breached first
  • Model swap

    Older version pinned

    The model alias is pinned to an earlier dated version, the way a stale config file does it.

    Expected to move: Answer quality

    Caught
    Measuring: Whether the gate blocked it
    Metric
    Measuring: Which metric breached first
  • Stale corpus

    Policy article deleted

    One article the agent depends on is removed from the knowledge base.

    Expected to move: Answer quality

    Caught
    Measuring: Whether the gate blocked it
    Metric
    Measuring: Which metric breached first
  • Stale corpus

    Policy article outdated

    One article is replaced with an older version that contradicts the current policy. The hard case: answers stay fluent.

    Expected to move: Answer quality

    Caught
    Measuring: Whether the gate blocked it
    Metric
    Measuring: Which metric breached first
  • Control

    Benign changes

    Rewordings and refactors with no intended change in behaviour. Any block here is a false block.

    Expected to move: None

    False blocks
    Measuring: Benign changes the gate blocked, out of those run

What we'd do differently

Not written yet

Written after the run, from what actually went wrong. Every missed regression or failed target gets an entry with its root cause.

Dataset and licence

Customer-support requests labelled with intents and categories, across retail-style support scenarios.

Licence
Not yet confirmed

To be confirmed on the dataset card before any derived eval set is redistributed. Until then the repository ships the sampling script, not the sampled data.

Version
Recorded with the first run, along with the download date.
Rows
MeasuringRows in the dataset at download time
Intents
MeasuringDistinct intent labels

Cite as

  • Bitext Customer Support LLM Chatbot Training Dataset, published by Bitext on Hugging Face.

Stack

Gateway
LiteLLM
Orchestration
LangGraph
Tracing
Langfuse
Grading
Custom LLM judge with a published rubric, Ragas where its metrics fit
CI
GitHub Actions

Models

  • Agent

    Chosen and recorded at run time, with its exact version

  • Smaller swap-in model

    Chosen and recorded at run time, with its exact version

  • Judge

    Chosen and recorded at run time, with its exact version

What would make this fail

The ways this build could break, or produce numbers that mislead, and what we do about each. A method you cannot argue with is a method nobody checked.

  • We wrote the regressions and the gate, so the gate can end up tuned to the regressions.

    The regression list, the limits and the expected metric for each are committed before the first gated run. A control group of benign changes measures false blocks alongside catches.

  • The dataset's requests are single-turn and carry one intent each. Production tickets are often multi-turn, multi-intent and messier.

    The build demonstrates the method, not coverage of your traffic. On a client system the eval set is sampled from production, and this page does not claim otherwise.

  • A public dataset may be in a model's training data, which inflates absolute accuracy.

    The gate acts on changes between runs, not on absolute scores, and deltas are much less affected by contamination. Absolute numbers are reported with that caveat.

  • An LLM judge can prefer its own family's style, or drift when its provider updates it.

    The judge comes from a different model family than the agent, is pinned to a dated version, and is recalibrated whenever that version changes.

  • The should-escalate labels are ours, so both escalation metrics inherit our rubric's judgement.

    The rubric and the labels are published with the code, so a reader can disagree with a specific label rather than with the result as a whole.

  • Limits set too tight make the gate flap; set too loose, it passes real regressions.

    Both failure directions are measured: misses on the planted regressions, false blocks on the control group. Neither is hidden behind the other.

Evidence and publication

No client and no permission involved: we are the client, and the data is public. This page drops its in-progress label only when every item below is done, and the build refuses to publish it before then.

Still to do before publishing

  • 36 values not yet measured
  • “What we'd do differently” not yet written
  • Repository with code and run logs not yet public
  • Dataset licence not yet confirmed for redistribution
  • Dataset version not yet recorded
  • Model versions and run dates not yet recorded

Code, eval manifest and run logs are published with the first results. Every result row will cite the commit or run it came from.

Want this method run on your system?

A 30-minute call. We look at one system you are trying to get into production, name the metric that matters, and say what an eval gate on it would take.

hello@venian.aiReply within one business day, from an engineer.

What happens next

  1. 01

    A reply within one business day

    From an engineer who would work on it, not a sales team.

  2. 02

    A 30-minute call

    We map one process, name the metric it should move, and tell you plainly whether AI is the right tool for it.

  3. 03

    A written scope

    If it is, a one-page scope: the metric, how we baseline it, the target, and which engagement fits.

For
Teams with an LLM feature in or near production, or a pilot that stalled, and a number it should move.
Not for
Chatbot-on-a-website projects, or strategy decks with no build behind them.