> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veri.studio/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluate a version for the promotion gate

> Creates the experiment the lineage's promotion gate requires for this version:                    the gate's dataset snapshot, exactly its pinned evaluator versions and its                    trials, targeting the version (resolved to the fleet adapter or single-model                    deployment serving exactly those weights). When another version is in                    production, its latest matching experiment is the baseline, so the run carries                    its diff. Works whether or not the gate is enabled. Promote once it completes.



## OpenAPI

````yaml /api-reference/openapi.json post /v1/models/{name}/versions/{version}/evaluate
openapi: 3.1.0
info:
  title: Veri API
  description: >-
    REST API for the Veri RL post-training platform. All requests require a
    Bearer API key (`vk_` prefix).
  license:
    name: ''
  version: 0.1.0
servers:
  - url: https://api.veri.studio
    description: Production
security: []
tags:
  - name: Training jobs
    description: Create, monitor, and manage training jobs.
  - name: Datasets
    description: Upload and connect training datasets.
  - name: Deployments
    description: Serve trained models and run inference.
  - name: Volumes
    description: Persistent file storage mounted into jobs.
  - name: Models
    description: Custom model registry deployments serve from.
  - name: Regions
    description: Discover available launch regions.
  - name: GPU
    description: Live GPU availability by provider and region.
  - name: Code artifacts
    description: Upload custom training script bundles.
  - name: Billing
    description: Credit balance and transaction history.
  - name: API keys
    description: Create and revoke API keys.
  - name: Account
    description: The authenticated caller's identity.
  - name: Settings
    description: Account-level integrations (Weights & Biases).
  - name: Metrics
    description: Prometheus metrics export for your own observability stack.
  - name: Evaluators
    description: 'Evaluators: versioned scoring rules (LLM judge, code, human).'
  - name: Experiments
    description: >-
      Experiments: offline runs of a pinned dataset snapshot through a target,
      scored by pinned evaluators.
  - name: Annotation queues
    description: >-
      Human review: queue traces, threads and experiment items, reserve one at a
      time, score them and feed corrections back into datasets.
  - name: Monitors
    description: >-
      Monitors: evaluators scoring a sampled share of a deployment's live
      traffic.
  - name: Observability
    description: Agent conversations, agents and their traffic.
  - name: Public runs
    description: Unauthenticated reads of runs their owners published to Explore.
paths:
  /v1/models/{name}/versions/{version}/evaluate:
    post:
      tags:
        - Models
      summary: Evaluate a version for the promotion gate
      description: >-
        Creates the experiment the lineage's promotion gate requires for this
        version:                    the gate's dataset snapshot, exactly its
        pinned evaluator versions and its                    trials, targeting
        the version (resolved to the fleet adapter or
        single-model                    deployment serving exactly those
        weights). When another version is in                    production, its
        latest matching experiment is the baseline, so the run
        carries                    its diff. Works whether or not the gate is
        enabled. Promote once it completes.
      operationId: evaluate
      parameters:
        - name: name
          in: path
          description: Lineage name
          required: true
          schema:
            type: string
        - name: version
          in: path
          description: 'Version: "v3" or "3"'
          required: true
          schema:
            type: string
      responses:
        '202':
          description: Accepted (status queued); poll GET /v1/experiments/{id}
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ExperimentResponse'
        '400':
          description: Invalid version reference, a purged version, or evals disabled
        '404':
          description: Lineage or version not found in the caller's workspace
        '409':
          description: >-
            gate_not_configured: no gate snapshot and evaluators;
            version_not_servable: no live deployment serves the version
        '429':
          description: >-
            CONCURRENT_EXPERIMENT_LIMIT: the workspace already has the maximum
            number of live experiments
      security:
        - bearerAuth: []
components:
  schemas:
    ExperimentResponse:
      type: object
      description: One experiment.
      required:
        - object
        - id
        - name
        - status
        - dataset_id
        - dataset_snapshot_id
        - target
        - evaluators
        - bindings
        - generation
        - trials
        - items_total
        - items_done
        - created_at
      properties:
        object:
          type: string
          description: Always "experiment".
        id:
          type: string
          description: '`exp_...`'
        name:
          type: string
        status:
          type: string
          description: queued | running | awaiting_local | completed | failed | canceled
        dataset_id:
          type: string
        dataset_snapshot_id:
          type: string
          description: The snapshot the run is pinned to (`<dataset_id>@snap-N`).
        target:
          type: object
          description: The target as requested.
        target_deployment_id:
          type:
            - string
            - 'null'
          description: |-
            Resolved at create: the deployment called, the request `model` sent
            (an adapter's versioned internal name for a model version on a
            fleet), the model version, and the served model's identity.
        target_model:
          type:
            - string
            - 'null'
        target_model_version_id:
          type:
            - string
            - 'null'
        target_served_model:
          type:
            - string
            - 'null'
        evaluators:
          type: array
          items:
            $ref: '#/components/schemas/EvaluatorPin'
          description: The pinned evaluator versions (+ each judge's served model).
        bindings:
          type: object
        generation:
          $ref: '#/components/schemas/Generation'
        trials:
          type: integer
          format: int32
        items_total:
          type: integer
          format: int32
          description: Items = rows x trials; done counts finished items (ok or error).
        items_done:
          type: integer
          format: int32
        aggregates:
          type:
            - object
            - 'null'
          description: |-
            Per evaluator at completion: n, ok, errors, skipped, pass_rate (all
            items in the denominator), mean, label_counts, critical_failures.
        passed:
          type:
            - boolean
            - 'null'
          description: |-
            Every evaluator with a pass rule passed every item with no errors;
            null until completed, or when no evaluator has a pass rule.
        baseline_experiment_id:
          type:
            - string
            - 'null'
        diff:
          type:
            - object
            - 'null'
          description: The diff against the baseline, computed at completion.
        error:
          type:
            - string
            - 'null'
        created_by:
          type:
            - string
            - 'null'
        created_at:
          type: string
          format: date-time
        started_at:
          type:
            - string
            - 'null'
          format: date-time
        finished_at:
          type:
            - string
            - 'null'
          format: date-time
    EvaluatorPin:
      type: object
      description: One pinned evaluator version, stored in `experiments.evaluators`.
      required:
        - evaluator_id
        - version
      properties:
        evaluator_id:
          type: string
        version:
          type: integer
          format: int32
        name:
          type: string
        scope:
          type: string
          description: single_turn | thread
        method:
          type: string
          description: llm_judge | code
        judge_model:
          type:
            - string
            - 'null'
          description: 'llm_judge: the judge deployment''s served model id at create.'
    Generation:
      type: object
      description: Sampling params for the target's calls. Omitted = the server default.
      properties:
        temperature:
          type:
            - number
            - 'null'
          format: double
        max_tokens:
          type:
            - integer
            - 'null'
          format: int32
      additionalProperties: false
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: API key
      description: API key with the `vk_` prefix. Create one from the dashboard.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.