Skip to main content
POST
Create a score

Authorizations

Authorization
string
header
required

API key with the vk_ prefix. Create one from the dashboard.

Body

application/json
evaluator
string
required

Evaluator id (evl_...) or name.

target_type
string
required

trace | thread | experiment_item. A single_turn evaluator scores traces, a thread evaluator scores threads.

evaluator_version
integer<int32> | null

The version to score with; default the latest.

request_id
string | null

Trace targets: the request id (from GET /v1/deployments/{id}/requests or a trace's metadata.veri_request_id).

deployment_id
string | null

Thread targets: the deployment and thread id.

thread_id
string | null
target_revision
string | null

Thread targets: the last request id covered; default the thread's latest request at write time.

value
any

Required when status is ok: a boolean, one of the labels, or a number in [min, max], per the evaluator's output type. Must be absent otherwise.

reasoning
string | null
status
string

ok (default) | error | skipped

error
string | null

Required for status error; optional reason for skipped.

source
string

api (default) | sdk | human. Human scores are unique per author.

source_sha256
string | null

Hex sha256 of the evaluator source that produced this score. Required when the pinned version is a code evaluator with runtime: local (the SDK harness returns it), and must match that version's source; ignored for every other evaluator.

experiment_id
string | null

Experiment-item targets: the experiment and the item (<row_id>.<trial>, or <row_id>.<trial>.t<turn> for one turn of a replayed thread). Outside the runner, only a local code evaluator the experiment pinned takes scores, on an item that is done (see GET /v1/experiments/{id}/items?pending_local=true), plus human labels (source: human, any evaluator, a done item) that never count toward the run.

item_id
string | null

Response

The stored score

One score.

object
string
required

Always "score".

id
string
required

scr_...

evaluator_id
string
required
evaluator_version
integer<int32>
required
name
string
required

The evaluator's name at write time.

target_type
string
required

trace | thread

value
any
required

boolean | label string | number; null for error and skipped scores.

source
string
required

api | sdk | human | experiment | monitor

status
string
required

ok | error | skipped

created_at
string<date-time>
required
request_id
string | null

The scored request (trace targets). Trace id = veri-<request_id>.

deployment_id
string | null
thread_id
string | null

The thread (thread targets; also stamped on trace targets whose request carried a thread id).

target_revision
string | null

Thread targets: the last request id the score covers.

experiment_id
string | null

Experiment-item targets: the experiment.

item_id
string | null

Experiment-item targets: <row_id>.<trial>, or <row_id>.<trial>.t<turn> for one turn of a replayed thread.

passed
boolean | null

The evaluator version's pass rule applied to the value; null when the version has no pass rule or the score carries no value.

reasoning
string | null
error
string | null

Why an error / skipped score has no value.

author
string | null

The user who wrote the score.

model_version_id
string | null

The model version that answered the scored request, when known.

monitor_id
string | null

The monitor that wrote the score (source monitor).

confidence
number<double> | null

The judge's probability for its verdict, 0..1; null unless a decision-mode judge wrote the score.

probabilities
object | null

The Jev judge's distribution behind the value: noul {"true": p, "false": 1 - p}, choice {option: p}, score {"<level>": p}. Null for every other score.