production promotion that does not meet your bar.
Target and judge calls are ordinary requests to your own deployments, through the same serving path your users hit. They cost what those deployments cost; there is no second inference stack.
The pieces
The loop
1
Capture traffic
Deploy with request traces on. Send an
X-Veri-Thread-Id header on every request of a conversation so its turns group into a thread.2
Turn traffic into data
Import a window of traces into a dataset stream, optionally only the ones scored as failing (by reviewers or your own harness), and cut a snapshot.
3
Train and register
Train on the snapshot and register the job as the lineage’s next version.
4
Evaluate
Run the gate’s experiment against the new version:
veri models evaluate acme-bot@v7 --wait.5
Promote
veri models promote acme-bot v7 -m "...". With the gate on, the move succeeds only if the evidence meets every rule; otherwise it is refused with the reasons.Try it in five commands
A golden set, a code evaluator that runs on your machine, and an experiment against a serving deployment:exact.py defines one function:
veri experiments run waits, prints a per-evaluator table, and exits 1 when the run fails, which makes it a CI check.
Limitations
What evaluations do not do today:- No span or tool-span targets. Evaluators score a whole request (a trace), a whole conversation (a thread), or an experiment item. Individual tool calls or steps inside a request are not scoring targets.
- No pairwise or composite evaluators. Each evaluator scores one output on its own; there is no A-vs-B judge and no evaluator built from other evaluators. Compare versions with an experiment’s baseline diff instead.
- Only Veri-served models. Experiment targets and LLM judges are deployments in your workspace. The exceptions:
target: nonescores the outputs already recorded in your dataset, and local code evaluators run on your machine. - No templates. Veri ships no built-in evaluators, rubrics, or benchmark suites; you write every prompt and function.
- Text only. No multimodal inputs or outputs.
- Cloud code evaluators are JavaScript or TypeScript only. Python code evaluators run locally through the SDK or CLI.
- No online monitors yet. Evaluators do not run automatically on sampled live traffic; score traffic through experiments, the scores API, or annotation queues.
Where to go next
Evaluators
Judge prompts, code functions, human rubrics, and how variables bind.
Experiments
Datasets, targets, thread replay, the local flow, and baseline diffs.
Promotion gate
Require passing evidence before
production moves.Human review
Annotation queues, corrections, and judge calibration.

