Skip to main content
An annotation queue holds items for people to review against pinned evaluator versions. Reviewers work through the queue in the dashboard; each completed item writes one human score per evaluator and can turn a reviewer’s corrected answer into a new dataset row. Queues are also how you check an LLM judge: put the judge’s own evaluator in a queue, and the human scores collected on the same targets become the agreement check shown on the promotion gate.

Create a queue

veri annotation-queues list shows every queue with its pending, reserved, completed, and skipped counts.

Add items

Add explicit targets:
Or add the targets of matching scores, for example every trace your judge harness scored as failing on one deployment:
--passed, --failed, --min-score, and --max-score filter by the score; --experiment limits the selection to one experiment’s items. Targets already in the queue are counted, not added twice.

Review

Reviewers open the queue in the dashboard, take the next item, see the conversation with the queue’s instructions and each evaluator’s allowed values, and then complete, skip, or release it. Taking an item reserves it for that reviewer until the reservation expires, so two people never review the same item at once. Completing an item:
  • writes one score per pinned evaluator, with source: human and the reviewer as author;
  • records an optional note and a correction (the answer the model should have given);
  • with a dataset chosen, appends the correction to that stream as an eval row (the item’s input or thread turns, the correction as expected, and the queue and reviewer in metadata), ready for the next experiment or training snapshot. The stream’s format must be eval (or not yet set).
Human scores on experiment items are calibration data: they never change an experiment’s results, verdict, or the gate’s evidence. Building your own review tool? The same flow is in the API and SDK: client.annotation_queues.next(queue) reserves and returns the next item with its context (or None when the queue is empty), then complete(queue, item_id, {"tone": "good"}, correction=..., append_to_dataset="golden"), skip, or release. Subscribe to the annotation.completed webhook to react to reviews.

Judge and human agreement

For each llm_judge evaluator pinned by a lineage’s promotion gate, Veri compares the judge’s verdicts with human scores on the same targets. veri models gate <model> shows the agreement, with a warning when humans and the judge disagree too often (below 80% agreement over at least 10 shared targets). The warning never blocks a promotion; it tells you the judge’s verdicts may not be a trustworthy gate until you fix its prompt.

Where to go next

Evaluators

Write human rubrics and judges.

Dataset streams

Where corrections land, ready to snapshot.