Names are 1-64 lowercase letters, digits,
., _ or -, unique per workspace, and usable in place of the evl_... id everywhere.
Output types and pass rules
The output config says what a score may be and, optionally, when it passes. Experiment verdicts and gate thresholds count passes. An evaluator without a pass rule is informational in an experiment’s verdict (its mean or label counts are reported); the gate’s non-regression rule still compares a numeric evaluator’s mean with production’s.Variables and bindings
Evaluators read five built-in variables from the item being scored:
Bindings map a variable to a JSONPath inside the item, so an evaluator can read a nested field or define a new variable for its prompt:
$ followed by .name, ['name'], or [index]. Wildcards, slices, and filters are rejected: a binding selects exactly one value. An experiment can layer its own bindings over every evaluator’s.
Set requires_reference: true when the evaluator compares against expected: items without an expected value are then recorded as skipped instead of being scored.
LLM judge
{{input}}, {{output}}, and any other variable name (built-in or bound) are replaced with the item’s values. Strings are inserted as-is and other JSON values as compact JSON. A placeholder that names no variable is an error at scoring time; a known variable with no value renders empty and is reported.
How the judge runs:
- The judge is a deployment in your workspace (
--judge-deployment;--judge-modelpicks a model or adapter on it). An experiment pins the judge’s served model when it starts. - Veri adds a fixed system instruction (grade strictly on the evidence shown), asks for a JSON verdict
{"reasoning", "score"}with structured outputs at temperature 0, and validates the reply against your output type. - A reply that does not parse or is out of range becomes an error score, never a pass.
Code evaluators
A code evaluator is one function that receives{input, output, expected, metadata, transcript} (after bindings) and returns {"score": ..., "reasoning": ...}. The score is a boolean, a finite number, or a label string; reasoning is optional.
.py, .js/.mjs, .ts/.mts) unless you pass --language. Where the code runs is the runtime:
Both runtimes enforce the same contract: a 5 second wall clock, source up to 256 KB, arguments up to 5 MB, a result up to 256 KB, labels up to 256 characters, and reasoning truncated to 8 KB. Use the standard library only: a local evaluator that imports a package you have installed would score differently (or fail) anywhere it is missing.
A local evaluator’s score must carry the SHA-256 of the evaluator version’s source; the SDK computes it, and a score from different code is refused (
409 source_mismatch). Python code evaluators are local-only: Python in the cloud sandbox would not match local results.
Human evaluators
llm_judge and code evaluators.
Spec files
Keep evaluators in version control as JSON (or YAML, with PyYAML installed) and create them with--file. method_config.source_file is read relative to the spec file:
evaluator.json
Versions
Every edit is a new immutable version; nothing that pinned an earlier version changes meaning. Send only what changes, and method flags are merged into the latest version’s config:helpful@2 in experiments, gates, and queues; a bare name means the latest.
Scores
A score is one evaluator result on one target. Scores are immutable: a second score for the same evaluator version and target is refused with409 score_exists.
Your own harness can post verdicts on production traffic, and they show up on the trace in the dashboard and in
veri datasets import-traces --score-evaluator:
source (api, sdk, human, or experiment), its author, passed (the version’s pass rule applied to the value), and the model_version_id that answered the scored request, when known. status is ok, error (with error), or skipped.
Threads
A thread is the requests of one conversation. Send the sameX-Veri-Thread-Id header (1-128 characters of A-Z a-z 0-9 . _ : -) on every request of the conversation; the SDK’s session_id argument sends the equivalent X-Veri-Session-Id, which is accepted as an alias. The thread id is stored on each request and trace, and thread evaluators score the whole conversation up to its latest request.
Where to go next
Experiments
Run evaluators over a dataset snapshot against a target.
Human review
Score human evaluators and calibrate judges.

