How to regression-test an LLM system

Short answer

LLM regression testing checks that a change to an LLM system did not make it worse on cases it used to get right. These systems regress without a code diff: a prompt edit, a model version swap, a reindexed or rechunked corpus, or a changed tool schema can each shift behaviour on inputs nobody touched.

A regression suite is a golden set of real cases, each with an evaluator that checks a property of the answer rather than its exact wording, plus a stored baseline from the version in production. Because the output is non-deterministic, each case runs several times and the gate compares pass rates, not single outputs. A change fails if an evaluator falls below its floor, drops further than the measured run-to-run noise, or fails a case marked critical. The suite runs in CI on every pull request that touches prompts, model configuration, retrieval or tools, and on a schedule against whatever model version the provider is serving.

Why do LLM systems regress without a code change?

Because most of the behaviour lives outside the code: in the prompt, the model, the index and the tool definitions. Each one changes on its own schedule, and each needs its own kind of case in the suite.

ChangeWhat actually changesWhat tends to breakWhat the suite needs
Prompt editInstruction wording or order, few-shot examples, the system prompt, a template variable.Instructions that used to be followed get dropped. Output format drifts. The fix for one complaint quietly breaks cases the old prompt was tuned on.The cases the current prompt was tuned against, plus format checks in code.
Model swapA new provider, a new version, or a version alias that starts pointing at a new snapshot with no change on your side.Refusal behaviour, verbosity, JSON validity, tool-call formatting, and cost and latency per task.The full suite run on the pinned and the new version side by side, with cost and latency recorded next to quality.
Retrieval changeA reindex, new chunking, a new embedding model or reranker, or a corpus refresh.The right passage stops being retrieved, a stale document outranks the current one, citations point at the wrong chunk.Retrieval-level checks (is the gold passage in the top k) scored separately from answer checks, so a failure points at its layer.
Tool schema changeA renamed field, a new required argument, an edited tool description, a tool added to the set.The model calls the wrong tool, or the right one with the old argument shape. Agents loop or give up.Tool-call cases that assert the tool chosen and validate its arguments against the live schema.

What belongs in a regression suite?

Cases that represent a failure you have seen or a path you cannot afford to break. Each one checks a property of the answer, not a string, so a better-worded correct answer still passes.

Production failures

Every bug that reached a user becomes a case when it is fixed, so it cannot come back unnoticed. Over time this is most of the suite, and the most valuable part of it.

Core paths

The handful of things the product must always do. Few cases, each marked critical, each required to pass on every sample.

Slices and edge cases

Inputs grouped by what makes them different: intent, language, length, customer tier, missing data. Tagged, so a broken slice cannot hide inside a healthy average.

Refusals and safety

What the system must decline, redact or escalate. A regression here is usually the expensive kind, so these are critical by default.

What makes a case golden

A domain expert has agreed what correct means for it, and that agreement is written down as the evaluator. The set is versioned in git and changes to it are reviewed like code.

A case is never edited to make a change pass. If the new answer is better and the rubric rejects it, the rubric is wrong, and fixing it is its own reviewed change.

evals/golden/refunds.yamlyaml
# evals/golden/refunds.yaml
- id: refund-outside-window-017
  slice: refunds
  critical: true
  source: production-trace        # where the case came from, for audit
  input:
    messages:
      - role: user
        content: "I bought the annual plan 45 days ago. Can I still get a refund?"
  context:
    plan: annual
    days_since_purchase: 45
  expect:
    - check: cites_document        # code: the cited ID exists and was retrieved
      doc_id: refund-policy-v3
    - check: no_refund_promised    # code: pattern and structured-output check
    - judge: states_window_correctly
      rubric: >
        Does the answer say refunds close 30 days after purchase,
        without offering an exception? Answer pass or fail.

How do you set pass thresholds on non-deterministic output?

By measuring the noise before setting the bar. A tolerance tighter than the suite's own run-to-run variation fails good changes on luck, and the team learns to rerun the job until it goes green.

  1. 01

    Measure the noise floor

    Run the unchanged production version against the suite several times. The spread of pass rates between those runs is the noise. Setting temperature to zero narrows it but does not remove it: providers do not guarantee identical outputs for identical requests.

  2. 02

    Sample each case more than once

    Run every case several times at production settings and score each sample. A case that passes two runs in three is information a single run throws away, and it is exactly the kind of case a change tips over.

  3. 03

    Compare paired, not pooled

    Baseline and candidate run on the same cases, so compare them case by case. A paired bootstrap over cases, or McNemar's test on per-case pass and fail, says whether a drop is larger than chance and gives an interval to report instead of a single number.

  4. 04

    Set three kinds of threshold

    A floor per evaluator that no change may cross. A tolerance on the drop against main, set from the measured noise. And critical cases that must pass on every sample, whatever the averages do.

  5. 05

    Hold each slice to its own bar

    An overall pass rate can hold steady while one slice collapses. Gate the slices that matter separately, and give each enough cases that its own margin of error is narrower than the drop you want to catch.

How eval set size sets the margin of error

How do you wire the suite into CI?

As a required check on pull requests that touch the parts of the system that change behaviour, plus a nightly run that catches the changes nobody opened a pull request for. The two files below are the whole gate.

.github/workflows/llm-regression.ymlyaml
# .github/workflows/llm-regression.yml
name: llm-regression
on:
  pull_request:
    paths:
      - "prompts/**"
      - "src/llm/**"
      - "config/models.yaml"
      - "evals/**"
  schedule:
    - cron: "0 5 * * *"   # nightly: catches a provider-side model change with no PR

permissions:
  contents: read
  pull-requests: write

jobs:
  gate:
    runs-on: ubuntu-latest
    timeout-minutes: 30
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install -r evals/requirements.txt

      - name: Run the suite, 3 samples per case
        run: python evals/run.py --suite regression --samples 3 --out evals/out/candidate.json
        env:
          MODEL_API_KEY: ${{ secrets.MODEL_API_KEY }}

      - name: Gate against the baseline from main
        run: pytest evals/test_regression_gate.py -q

      - name: Post the score diff on the PR
        if: always() && github.event_name == 'pull_request'
        run: |
          python evals/report.py evals/baseline.json evals/out/candidate.json > diff.md
          gh pr comment ${{ github.event.pull_request.number }} --body-file diff.md
        env:
          GH_TOKEN: ${{ github.token }}
evals/test_regression_gate.pypython
# evals/test_regression_gate.py
import json
from pathlib import Path

import pytest

BASELINE = json.loads(Path("evals/baseline.json").read_text())
CANDIDATE = json.loads(Path("evals/out/candidate.json").read_text())
# Per evaluator: {"floor": 0.90, "tolerance": 0.03}. Set the tolerance from
# measured run-to-run noise on this suite, never by copying a number.
THRESHOLDS = json.loads(Path("evals/thresholds.json").read_text())


def pass_rate(run: dict, evaluator: str) -> float:
    # Every sample of every case is one verdict, so 3 samples give 3 votes.
    verdicts = [v for case in run["cases"] for v in case["verdicts"][evaluator]]
    return sum(verdicts) / len(verdicts)


@pytest.mark.parametrize("evaluator", sorted(THRESHOLDS))
def test_evaluator_holds(evaluator: str) -> None:
    t = THRESHOLDS[evaluator]
    base, cand = pass_rate(BASELINE, evaluator), pass_rate(CANDIDATE, evaluator)
    assert cand >= t["floor"], f"{evaluator}: {cand:.3f} is under its floor {t['floor']}"
    assert base - cand <= t["tolerance"], (
        f"{evaluator}: fell {base - cand:.3f} against main (tolerance {t['tolerance']})"
    )


def test_critical_cases_pass_every_sample() -> None:
    failed = [
        case["id"]
        for case in CANDIDATE["cases"]
        if case.get("critical") and not all(all(v) for v in case["verdicts"].values())
    ]
    assert not failed, f"critical cases failed: {failed}"

The harness is yours

run.py stands for whatever runs your system and its evaluators. The gate depends only on the JSON it writes: cases, each with a list of verdicts per evaluator, one per sample.

The baseline moves on merge

baseline.json is regenerated from main after each merge, never from a branch. A pull request is always compared with what is actually in production.

The nightly run has no diff

A failure on the scheduled run with no merged change means something upstream moved: a model alias, a provider update, the corpus. That is the run that earns its place.

What should happen when the gate fails?

Someone reads the failing cases before anyone touches a threshold. A red gate has three possible causes, and only one of them is a regression.

A real regression

The candidate is worse on cases it used to pass. Fix the change, or ship it knowingly with the trade-off written in the pull request.

A better answer, rejected

The new output is correct and the evaluator is wrong. Fix the evaluator in its own reviewed change, then rerun.

Noise

The drop is inside the measured spread. Add samples or cases until the suite can tell the difference. Never lower a threshold in the pull request that tripped it.

Make the next model swap boring

A 30-minute call. We look at what changes in your system and how often, and tell you what a regression gate would need to cover before the next prompt edit or model upgrade ships.

hello@venian.aiReply within one business day, from an engineer.

What happens next

  1. 01

    A reply within one business day

    From an engineer who would work on it, not a sales team.

  2. 02

    A 30-minute call

    We map one process, name the metric it should move, and tell you plainly whether AI is the right tool for it.

  3. 03

    A written scope

    If it is, a one-page scope: the metric, how we baseline it, the target, and which engagement fits.

For
Teams with an LLM feature in or near production, or a pilot that stalled, and a number it should move.
Not for
Chatbot-on-a-website projects, or strategy decks with no build behind them.