Datasets and the eval row format
Experiments read a dataset stream snapshot: golden@snap-3, or a stream name (which pins @latest, cutting a snapshot if rows arrived since the last one). Create a stream for eval rows with veri datasets create golden --format eval. An eval row is one of two shapes:
inputis a string (sent as one user message) or a message list. An optionaloutputrecords an answer to score without calling a target (seetarget: none).turnsis a list of strings (user messages) or message objects; every user turn gets a reply from the target.expectedmay be a list with one entry per user turn.metadatais free-form.metadata.critical: truemarks a row that must pass: any failure on it is listed undercritical_failuresand fails the promotion gate.
prompt row is the input; a completion row is the input plus expected; a chat row ending in an assistant turn is the conversation before it as input, and that turn as the recorded output. preference rows are rejected.
Targets
Targets are deployments in your workspace.
--temperature and --max-tokens set the target’s sampling; --trials N (1-10) runs each row N times.
Thread replay
Aturns row replays a conversation: the target answers the first user turn, its reply joins the history, the next user turn is sent with that history, and so on. All calls of one replay carry the thread id exp-<experiment>-<row>-<trial>, so they group like any other thread in your request log.
- A
threadevaluator scores the whole transcript once per replay. - A
single_turnevaluator scores every turn separately (item ids<row>.<trial>.t<turn>), each with that turn’s user message asinput, its reply asoutput, and the conversation so far astranscript.
Run it
veri experiments run creates the experiment, waits, prints a table per evaluator (and the baseline diff), and sets its exit code for CI:
--json prints the finished experiment instead of the table; --no-wait only creates it (as does veri experiments create). Follow a run with veri experiments get <id>, list its items with veri experiments items <id>, and stop it with veri experiments cancel <id>.
Local code evaluators
Evaluators withruntime: local run on your machine, not in Veri. The experiment calls the target and runs every cloud evaluator itself, then waits in status awaiting_local until the local scores arrive. Supply them while it runs:
GET /v1/experiments/{id}/items?pending_local=true), runs each evaluator on the pinned version’s source, and posts each result with the source’s SHA-256. An evaluator that raises, times out, or returns a value its output type rejects is recorded as an error score, so the experiment still completes. It completes when the last owed score lands.
Results
When an experiment completes,aggregates.evaluators holds one entry per pinned evaluator: n (items), ok, errors, skipped, pass_rate (passes over all items, so errors and skips count against it), mean (numeric outputs), label_counts (categorical), and critical_failures.
passed is the experiment’s verdict: true when every evaluator with a pass rule passed every item with no errors, false otherwise, and null when no evaluator has a pass rule. The promotion gate applies its own thresholds instead of this all-or-nothing verdict.
Each item carries its input, output, transcript, latency, token counts, the request ids it made, and its scores. Subscribe to the experiment.completed webhook instead of polling.
Baseline diff
Set a baseline (--baseline exp_... at create) and the completed experiment carries a diff; compare any two runs later with veri experiments diff <experiment> <baseline>. Per evaluator it lists the items that newly fail, newly pass, and still fail, plus the pass-rate, mean, and error deltas.
Two runs are comparable only on the same dataset snapshot with identical evaluator versions, judge models, and bindings. Otherwise the diff says comparable: false with the reason, rather than comparing different questions.
Limits
- Two live (queued, running, or awaiting-local) experiments per workspace; a third is refused with
429 CONCURRENT_EXPERIMENT_LIMIT. Wait for one or cancel it. - Up to 20,000 scores per experiment (rows x trials x evaluators) and 16 evaluators.
- A target that is stopped is refused at create; one that scales to zero is woken.
Where to go next
Promotion gate
Make an experiment the evidence a production promotion needs.
Human review
Queue failing items for reviewers and collect corrections.

