Create a queue
veri annotation-queues list shows every queue with its pending, reserved, completed, and skipped counts.
Add items
Add explicit targets:--passed, --failed, --min-score, and --max-score filter by the score; --experiment limits the selection to one experiment’s items. Targets already in the queue are counted, not added twice.
Review
Reviewers open the queue in the dashboard, take the next item, see the conversation with the queue’s instructions and each evaluator’s allowed values, and then complete, skip, or release it. Taking an item reserves it for that reviewer until the reservation expires, so two people never review the same item at once. Completing an item:- writes one score per pinned evaluator, with
source: humanand the reviewer as author; - records an optional note and a correction (the answer the model should have given);
- with a dataset chosen, appends the correction to that stream as an
evalrow (the item’sinputor threadturns, the correction asexpected, and the queue and reviewer inmetadata), ready for the next experiment or training snapshot. The stream’s format must beeval(or not yet set).
client.annotation_queues.next(queue) reserves and returns the next item with its context (or None when the queue is empty), then complete(queue, item_id, {"tone": "good"}, correction=..., append_to_dataset="golden"), skip, or release. Subscribe to the annotation.completed webhook to react to reviews.
Judge and human agreement
For eachllm_judge evaluator pinned by a lineage’s promotion gate, Veri compares the judge’s verdicts with human scores on the same targets. veri models gate <model> shows the agreement, with a warning when humans and the judge disagree too often (below 80% agreement over at least 10 shared targets). The warning never blocks a promotion; it tells you the judge’s verdicts may not be a trustworthy gate until you fix its prompt.
Where to go next
Evaluators
Write human rubrics and judges.
Dataset streams
Where corrections land, ready to snapshot.

