> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veri.studio/llms.txt
> Use this file to discover all available pages before exploring further.

# Create a score

> Records one evaluator result on a trace or a thread of the caller's workspace. The value is checked against the pinned evaluator version's output type and `passed` is derived from its pass rule. Scores are immutable: a second score for the same evaluator version, target and revision (and author, for human scores) is a 409 `score_exists`. A score from a local code evaluator must carry `source_sha256`, the hash of the pinned version's source (409 `source_mismatch` otherwise). An `experiment_item` target takes a local code evaluator the experiment pinned, on an item that is done (the experiment completes when its last owed local score arrives), or a human label (`source: human`) from any evaluator on a done item; human labels feed the judge-vs-human agreement check and never count toward the run.



## OpenAPI

````yaml /api-reference/openapi.json post /v1/scores
openapi: 3.1.0
info:
  title: Veri API
  description: >-
    REST API for the Veri RL post-training platform. All requests require a
    Bearer API key (`vk_` prefix).
  license:
    name: ''
  version: 0.1.0
servers:
  - url: https://api.veri.studio
    description: Production
security: []
tags:
  - name: Training jobs
    description: Create, monitor, and manage training jobs.
  - name: Datasets
    description: Upload and connect training datasets.
  - name: Deployments
    description: Serve trained models and run inference.
  - name: Volumes
    description: Persistent file storage mounted into jobs.
  - name: Models
    description: Custom model registry deployments serve from.
  - name: Regions
    description: Discover available launch regions.
  - name: GPU
    description: Live GPU availability by provider and region.
  - name: Code artifacts
    description: Upload custom training script bundles.
  - name: Billing
    description: Credit balance and transaction history.
  - name: API keys
    description: Create and revoke API keys.
  - name: Account
    description: The authenticated caller's identity.
  - name: Settings
    description: Account-level integrations (Weights & Biases).
  - name: Metrics
    description: Prometheus metrics export for your own observability stack.
  - name: Evaluators
    description: 'Evaluators: versioned scoring rules (LLM judge, code, human).'
  - name: Experiments
    description: >-
      Experiments: offline runs of a pinned dataset snapshot through a target,
      scored by pinned evaluators.
  - name: Annotation queues
    description: >-
      Human review: queue traces, threads and experiment items, reserve one at a
      time, score them and feed corrections back into datasets.
  - name: Monitors
    description: >-
      Monitors: evaluators scoring a sampled share of a deployment's live
      traffic.
  - name: Observability
    description: Agent conversations, agents and their traffic.
  - name: Public runs
    description: Unauthenticated reads of runs their owners published to Explore.
paths:
  /v1/scores:
    post:
      tags:
        - Evaluators
      summary: Create a score
      description: >-
        Records one evaluator result on a trace or a thread of the caller's
        workspace. The value is checked against the pinned evaluator version's
        output type and `passed` is derived from its pass rule. Scores are
        immutable: a second score for the same evaluator version, target and
        revision (and author, for human scores) is a 409 `score_exists`. A score
        from a local code evaluator must carry `source_sha256`, the hash of the
        pinned version's source (409 `source_mismatch` otherwise). An
        `experiment_item` target takes a local code evaluator the experiment
        pinned, on an item that is done (the experiment completes when its last
        owed local score arrives), or a human label (`source: human`) from any
        evaluator on a done item; human labels feed the judge-vs-human agreement
        check and never count toward the run.
      operationId: create
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ScoreCreateRequest'
        required: true
      responses:
        '201':
          description: The stored score
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ScoreResponse'
        '400':
          description: >-
            Value does not match the output type, bad target shape, scope
            mismatch, or evals disabled
        '404':
          description: >-
            No such evaluator, version, trace or thread in the caller's
            workspace
        '409':
          description: >-
            score_exists: the target already has a score from this evaluator
            version (details carry score_id); source_mismatch: a local code
            evaluator's score without the sha256 of the pinned version's source
      security:
        - bearerAuth: []
components:
  schemas:
    ScoreCreateRequest:
      type: object
      required:
        - evaluator
        - target_type
      properties:
        evaluator:
          type: string
          description: Evaluator id (evl_...) or name.
        evaluator_version:
          type:
            - integer
            - 'null'
          format: int32
          description: The version to score with; default the latest.
        target_type:
          type: string
          description: |-
            trace | thread | experiment_item. A single_turn evaluator scores
            traces, a thread evaluator scores threads.
        request_id:
          type:
            - string
            - 'null'
          description: >-
            Trace targets: the request id (from GET
            /v1/deployments/{id}/requests

            or a trace's metadata.veri_request_id).
        deployment_id:
          type:
            - string
            - 'null'
          description: 'Thread targets: the deployment and thread id.'
        thread_id:
          type:
            - string
            - 'null'
        target_revision:
          type:
            - string
            - 'null'
          description: |-
            Thread targets: the last request id covered; default the thread's
            latest request at write time.
        value:
          description: >-
            Required when status is ok: a boolean, one of the labels, or a
            number

            in [min, max], per the evaluator's output type. Must be absent

            otherwise.
        reasoning:
          type:
            - string
            - 'null'
        status:
          type: string
          description: ok (default) | error | skipped
        error:
          type:
            - string
            - 'null'
          description: Required for status error; optional reason for skipped.
        source:
          type: string
          description: api (default) | sdk | human. Human scores are unique per author.
        source_sha256:
          type:
            - string
            - 'null'
          description: >-
            Hex sha256 of the evaluator source that produced this score.
            Required

            when the pinned version is a `code` evaluator with `runtime: local`

            (the SDK harness returns it), and must match that version's source;

            ignored for every other evaluator.
        experiment_id:
          type:
            - string
            - 'null'
          description: >-
            Experiment-item targets: the experiment and the item

            (`<row_id>.<trial>`, or `<row_id>.<trial>.t<turn>` for one turn of a

            replayed thread). Outside the runner, only a local code evaluator
            the

            experiment pinned takes scores, on an item that is done (see

            GET /v1/experiments/{id}/items?pending_local=true), plus human
            labels

            (`source: human`, any evaluator, a done item) that never count

            toward the run.
        item_id:
          type:
            - string
            - 'null'
    ScoreResponse:
      type: object
      description: One score.
      required:
        - object
        - id
        - evaluator_id
        - evaluator_version
        - name
        - target_type
        - value
        - source
        - status
        - created_at
      properties:
        object:
          type: string
          description: Always "score".
        id:
          type: string
          description: '`scr_...`'
        evaluator_id:
          type: string
        evaluator_version:
          type: integer
          format: int32
        name:
          type: string
          description: The evaluator's name at write time.
        target_type:
          type: string
          description: trace | thread
        request_id:
          type:
            - string
            - 'null'
          description: The scored request (trace targets). Trace id = `veri-<request_id>`.
        deployment_id:
          type:
            - string
            - 'null'
        thread_id:
          type:
            - string
            - 'null'
          description: |-
            The thread (thread targets; also stamped on trace targets whose
            request carried a thread id).
        target_revision:
          type:
            - string
            - 'null'
          description: 'Thread targets: the last request id the score covers.'
        experiment_id:
          type:
            - string
            - 'null'
          description: 'Experiment-item targets: the experiment.'
        item_id:
          type:
            - string
            - 'null'
          description: |-
            Experiment-item targets: `<row_id>.<trial>`, or
            `<row_id>.<trial>.t<turn>` for one turn of a replayed thread.
        value:
          description: boolean | label string | number; null for error and skipped scores.
        passed:
          type:
            - boolean
            - 'null'
          description: >-
            The evaluator version's pass rule applied to the value; null when
            the

            version has no pass rule or the score carries no value.
        reasoning:
          type:
            - string
            - 'null'
        source:
          type: string
          description: api | sdk | human | experiment | monitor
        status:
          type: string
          description: ok | error | skipped
        error:
          type:
            - string
            - 'null'
          description: Why an error / skipped score has no value.
        author:
          type:
            - string
            - 'null'
          description: The user who wrote the score.
        model_version_id:
          type:
            - string
            - 'null'
          description: The model version that answered the scored request, when known.
        monitor_id:
          type:
            - string
            - 'null'
          description: The monitor that wrote the score (source monitor).
        confidence:
          type:
            - number
            - 'null'
          format: double
          description: |-
            The judge's probability for its verdict, 0..1; null unless a
            decision-mode judge wrote the score.
        probabilities:
          type:
            - object
            - 'null'
          description: |-
            The Jev judge's distribution behind the value: noul
            `{"true": p, "false": 1 - p}`, choice `{option: p}`, score
            `{"<level>": p}`. Null for every other score.
        created_at:
          type: string
          format: date-time
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: API key
      description: API key with the `vk_` prefix. Create one from the dashboard.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.