Skip to main content
An evaluator is a named rule that scores a model output. Every evaluator has three independent choices: Names are 1-64 lowercase letters, digits, ., _ or -, unique per workspace, and usable in place of the evl_... id everywhere.

Output types and pass rules

The output config says what a score may be and, optionally, when it passes. Experiment verdicts and gate thresholds count passes. An evaluator without a pass rule is informational in an experiment’s verdict (its mean or label counts are reported); the gate’s non-regression rule still compares a numeric evaluator’s mean with production’s.

Variables and bindings

Evaluators read five built-in variables from the item being scored: Bindings map a variable to a JSONPath inside the item, so an evaluator can read a nested field or define a new variable for its prompt:
The path grammar is $ followed by .name, ['name'], or [index]. Wildcards, slices, and filters are rejected: a binding selects exactly one value. An experiment can layer its own bindings over every evaluator’s. Set requires_reference: true when the evaluator compares against expected: items without an expected value are then recorded as skipped instead of being scored.

LLM judge

The prompt is a template: {{input}}, {{output}}, and any other variable name (built-in or bound) are replaced with the item’s values. Strings are inserted as-is and other JSON values as compact JSON. A placeholder that names no variable is an error at scoring time; a known variable with no value renders empty and is reported. How the judge runs:
  • The judge is a deployment in your workspace (--judge-deployment; --judge-model picks a model or adapter on it). An experiment pins the judge’s served model when it starts.
  • Veri adds a fixed system instruction (grade strictly on the evidence shown), asks for a JSON verdict {"reasoning", "score"} with structured outputs at temperature 0, and validates the reply against your output type.
  • A reply that does not parse or is out of range becomes an error score, never a pass.

Code evaluators

A code evaluator is one function that receives {input, output, expected, metadata, transcript} (after bindings) and returns {"score": ..., "reasoning": ...}. The score is a boolean, a finite number, or a label string; reasoning is optional.
The language comes from the file extension (.py, .js/.mjs, .ts/.mts) unless you pass --language. Where the code runs is the runtime: Both runtimes enforce the same contract: a 5 second wall clock, source up to 256 KB, arguments up to 5 MB, a result up to 256 KB, labels up to 256 characters, and reasoning truncated to 8 KB. Use the standard library only: a local evaluator that imports a package you have installed would score differently (or fail) anywhere it is missing. A local evaluator’s score must carry the SHA-256 of the evaluator version’s source; the SDK computes it, and a score from different code is refused (409 source_mismatch). Python code evaluators are local-only: Python in the cloud sandbox would not match local results.

Human evaluators

Reviewers score human evaluators in annotation queues. Experiments and the promotion gate take only llm_judge and code evaluators.

Spec files

Keep evaluators in version control as JSON (or YAML, with PyYAML installed) and create them with --file. method_config.source_file is read relative to the spec file:
evaluator.json

Versions

Every edit is a new immutable version; nothing that pinned an earlier version changes meaning. Send only what changes, and method flags are merged into the latest version’s config:
Pin a version as helpful@2 in experiments, gates, and queues; a bare name means the latest.

Scores

A score is one evaluator result on one target. Scores are immutable: a second score for the same evaluator version and target is refused with 409 score_exists. Your own harness can post verdicts on production traffic, and they show up on the trace in the dashboard and in veri datasets import-traces --score-evaluator:
Each score records its source (api, sdk, human, or experiment), its author, passed (the version’s pass rule applied to the value), and the model_version_id that answered the scored request, when known. status is ok, error (with error), or skipped.

Threads

A thread is the requests of one conversation. Send the same X-Veri-Thread-Id header (1-128 characters of A-Z a-z 0-9 . _ : -) on every request of the conversation; the SDK’s session_id argument sends the equivalent X-Veri-Session-Id, which is accepted as an alias. The thread id is stored on each request and trace, and thread evaluators score the whole conversation up to its latest request.

Where to go next

Experiments

Run evaluators over a dataset snapshot against a target.

Human review

Score human evaluators and calibrate judges.