Skip to main content
Evaluations close Veri’s self-improvement loop: serve a model, capture its traffic, turn traffic into data, train, and only promote the new version when it measurably holds up. You write the evaluators (a judge prompt, a code function, or a human rubric); Veri runs them against your own deployments, stores every score, and can refuse a production promotion that does not meet your bar. Target and judge calls are ordinary requests to your own deployments, through the same serving path your users hit. They cost what those deployments cost; there is no second inference stack.

The pieces

The loop

1

Capture traffic

Deploy with request traces on. Send an X-Veri-Thread-Id header on every request of a conversation so its turns group into a thread.
2

Turn traffic into data

Import a window of traces into a dataset stream, optionally only the ones scored as failing (by reviewers or your own harness), and cut a snapshot.
3

Train and register

Train on the snapshot and register the job as the lineage’s next version.
4

Evaluate

Run the gate’s experiment against the new version: veri models evaluate acme-bot@v7 --wait.
5

Promote

veri models promote acme-bot v7 -m "...". With the gate on, the move succeeds only if the evidence meets every rule; otherwise it is refused with the reasons.

Try it in five commands

A golden set, a code evaluator that runs on your machine, and an experiment against a serving deployment:
exact.py defines one function:
veri experiments run waits, prints a per-evaluator table, and exits 1 when the run fails, which makes it a CI check.

Limitations

What evaluations do not do today:
  • No span or tool-span targets. Evaluators score a whole request (a trace), a whole conversation (a thread), or an experiment item. Individual tool calls or steps inside a request are not scoring targets.
  • No pairwise or composite evaluators. Each evaluator scores one output on its own; there is no A-vs-B judge and no evaluator built from other evaluators. Compare versions with an experiment’s baseline diff instead.
  • Only Veri-served models. Experiment targets and LLM judges are deployments in your workspace. The exceptions: target: none scores the outputs already recorded in your dataset, and local code evaluators run on your machine.
  • No templates. Veri ships no built-in evaluators, rubrics, or benchmark suites; you write every prompt and function.
  • Text only. No multimodal inputs or outputs.
  • Cloud code evaluators are JavaScript or TypeScript only. Python code evaluators run locally through the SDK or CLI.
  • No online monitors yet. Evaluators do not run automatically on sampled live traffic; score traffic through experiments, the scores API, or annotation queues.

Where to go next

Evaluators

Judge prompts, code functions, human rubrics, and how variables bind.

Experiments

Datasets, targets, thread replay, the local flow, and baseline diffs.

Promotion gate

Require passing evidence before production moves.

Human review

Annotation queues, corrections, and judge calibration.