Online vs offline evaluation, and why you need both

Short answer

Offline evaluation scores an LLM system on a fixed, labelled eval set before a change ships. Online evaluation scores the live system on a sample of real production traffic after it ships.

Offline evals are controlled and repeatable: the same cases run against the current version and the candidate, so a drop can be pinned on the change, and the result can gate a deploy in CI. Their blind spot is everything not in the set. Online evals see real inputs, real users and real upstream data, so they catch new intents, distribution shift and slow drift that no single deploy caused. But they report after users have seen the output, and live traffic has no ground-truth labels, so they rely on code checks, a calibrated LLM-as-judge and user signals.

A reliable system runs both and connects them: failures that online evals surface are labelled and added to the offline set, so the next change is gated on them.

What is an offline eval?

Offline evaluation scores an LLM system on a fixed, labelled eval set before a change ships, so the current version and the candidate can be compared case by case and the result can gate a deploy.

The set is a few dozen to a few hundred real cases, each with an evaluator: a code check where the answer is checkable, a rubric applied by a judge where it is not. Because the cases are fixed, two runs differ only in the system under test, which is what lets a score difference be read as the effect of a change.

Wired into CI, an offline eval becomes a regression gate: the build fails when a prompt edit or a model swap drops a score past its threshold. It is the only kind of evaluation that can stop a bad change before a user sees it.

Building the offline suite into a CI gate

What is an online eval?

Online evaluation scores the live LLM system on a sample of real production traffic after it ships, using code checks, a calibrated LLM-as-judge and user signals, because live traffic has no ground-truth labels.

The evaluators are, as far as possible, the same ones the offline set uses, so a pass rate in production and a pass rate in CI mean the same thing. They run out of the request path, on traces collected by your tracing layer, so scoring never adds latency for the user.

What comes out is a rate over time, per evaluator and per slice, alongside cost and latency per task. Its job is to notice what the offline set does not contain yet, and to say so before a support queue does.

How do online and offline evals compare?

Each one's blind spot is the other's strength. That is the case for running both rather than choosing between them.

 Offline evalOnline eval
When it runsBefore a change merges or deploys: on the pull request, on a model upgrade, nightly.Continuously after deploy, on requests as they arrive, scored asynchronously.
DataA fixed golden set of real cases, versioned in git.A sample of live traffic: real inputs, real users, real upstream data.
Ground truthYes. Every case has an agreed definition of correct.No. Verdicts come from code checks, a judge and user signals such as edits, retries and escalations.
Time to a verdictMinutes, inside CI, before anyone sees the output.Minutes to hours after the response, and a trend needs a window of traffic to be visible.
What it catchesRegressions caused by a specific change: a prompt edit, a model swap, a retrieval or tool change.New intents, distribution shift, slow drift, upstream data changes, failures under real load.
What it missesAnything not in the set. A set that passes is only as good as the cases it holds.It cannot stop a bad change: users have already seen the output by the time it is scored.
CostBounded and predictable: set size times samples times evaluators, per run.Scales with traffic, which is why judges run on a sample and code checks run on everything.
What it can blockA merge or a deploy, as a required CI check.Nothing directly. It raises alerts and feeds new cases to the offline set.

When does an offline pass hide an online failure?

Whenever production stops looking like the eval set. A green CI run says the change did not break the cases you have; it says nothing about the cases you do not.

Users asked something new

The eval set was built from last quarter's traffic. A launch, a pricing change or a new customer segment brings intents nobody has written a case for.

Data upstream moved

The corpus was refreshed, an API started returning a new field, a document was retired. The prompt and model are unchanged and the answers are now wrong.

Conversations run longer

Offline cases are often single-turn. Real sessions carry history, and failures appear at turn eight that no one-turn case can reach.

Load changes behaviour

Timeouts, rate limits, retries and truncated context happen under production traffic and not in a CI job running one case at a time.

The set was tuned on

A prompt iterated against the same fifty cases for a month learns those fifty cases. The score holds; the behaviour on everything else does not.

How do production failures get back into the eval set?

Through a loop someone owns. Without it, online evals produce a dashboard and the offline set goes stale; with it, the set ends up weighted toward the inputs that actually go wrong.

  1. 01

    Score a sample of live traffic

    Code checks on every request, the judge on a sample, user signals logged alongside each trace.

  2. 02

    Triage what the scores flag

    Someone reads the failing traces, not just the chart, and groups them by failure mode. New modes go into the taxonomy.

  3. 03

    Label the ones worth keeping

    A domain expert agrees what the correct answer was. Inputs are redacted of personal data before they leave the trace store.

  4. 04

    Add them to the offline set

    As reviewed cases with a slice tag, an evaluator and, where the failure was costly, a critical flag.

  5. 05

    Gate the next change on them

    The fix ships only when the new cases pass, and the regression suite keeps them passing from then on.

What is drift, and how do online evals catch it?

Drift is a change in behaviour that no single deploy caused. It is the failure offline evals are structurally unable to see, because nothing changed in the pull request.

Input drift

What users ask changes: the intent mix, languages, lengths. Watch the distribution of incoming traffic by slice, not only the scores.

Output drift

What the system says changes: answer length, refusal rate, tool-call rate, citation rate. Often the first visible sign of a provider-side model change.

Quality drift

The pass rate falls. Compare a rolling window per evaluator and per slice with the baseline window, and alert on a sustained drop rather than one bad hour.

How much live traffic should you sample?

As much as each evaluator can afford, weighted toward where failures are. Cheap checks see everything; expensive ones see a sample sized for the margin of error you need per window.

Every request

Code checks: schema validity, citation resolution, tool-argument validation, PII patterns. They cost almost nothing, so there is no reason to sample them.

A uniform random sample

The LLM-as-judge. The rate is set by how many verdicts per slice per window you need for a pass rate narrow enough to alert on, not by a round percentage.

Rare slices, oversampled

A slice that is a small share of traffic gets too few verdicts at the uniform rate. Sample it at a higher rate and weight it back when reporting the overall number.

Everything flagged

Negative feedback, retries, escalations to a human, guardrail triggers, tool errors. These traces are where new cases come from, so all of them are scored.

A small human-reviewed slice

Judged traces read by a person on a schedule. It is how you notice the judge itself drifting, and it keeps the calibration current.

How sample size sets the margin of error

Are guardrails the same as online evals?

No. A guardrail acts on one response, inline. An online eval measures many responses, after the fact. A system needs both, and each one's output is an input to the other.

Guardrail

  • Runs in the request path, on every request.
  • Acts on a single response: block, redact, rewrite, escalate.
  • Has to fit the latency budget, so it is cheap and narrow.
  • Answers “may this response go out?”

Online eval

  • Runs out of the request path, on a sample.
  • Acts on nothing directly: it measures and alerts.
  • Can afford a slower, more thorough judge.
  • Answers “how often is the system getting this right?”

A guardrail’s trigger rate is itself an online metric. One that never fires may not be checking anything; one whose rate climbs is usually the earliest sign that inputs or the model have moved.

Close the loop between CI and production

A 30-minute call. We look at what you score today, before and after deploy, and name the gap most likely to let the next failure reach users.

hello@venian.aiReply within one business day, from an engineer.

What happens next

  1. 01

    A reply within one business day

    From an engineer who would work on it, not a sales team.

  2. 02

    A 30-minute call

    We map one process, name the metric it should move, and tell you plainly whether AI is the right tool for it.

  3. 03

    A written scope

    If it is, a one-page scope: the metric, how we baseline it, the target, and which engagement fits.

For
Teams with an LLM feature in or near production, or a pilot that stalled, and a number it should move.
Not for
Chatbot-on-a-website projects, or strategy decks with no build behind them.