> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veri.studio/llms.txt
> Use this file to discover all available pages before exploring further.

# Create an evaluator

> Creates the evaluator and its version 1. An llm_judge evaluator must name a deployment in the caller's workspace as its judge.



## OpenAPI

````yaml /api-reference/openapi.json post /v1/evaluators
openapi: 3.1.0
info:
  title: Veri API
  description: >-
    REST API for the Veri RL post-training platform. All requests require a
    Bearer API key (`vk_` prefix).
  license:
    name: ''
  version: 0.1.0
servers:
  - url: https://api.veri.studio
    description: Production
security: []
tags:
  - name: Training jobs
    description: Create, monitor, and manage training jobs.
  - name: Datasets
    description: Upload and connect training datasets.
  - name: Deployments
    description: Serve trained models and run inference.
  - name: Volumes
    description: Persistent file storage mounted into jobs.
  - name: Models
    description: Custom model registry deployments serve from.
  - name: Regions
    description: Discover available launch regions.
  - name: GPU
    description: Live GPU availability by provider and region.
  - name: Code artifacts
    description: Upload custom training script bundles.
  - name: Billing
    description: Credit balance and transaction history.
  - name: API keys
    description: Create and revoke API keys.
  - name: Account
    description: The authenticated caller's identity.
  - name: Settings
    description: Account-level integrations (Weights & Biases).
  - name: Metrics
    description: Prometheus metrics export for your own observability stack.
  - name: Evaluators
    description: 'Evaluators: versioned scoring rules (LLM judge, code, human).'
  - name: Experiments
    description: >-
      Experiments: offline runs of a pinned dataset snapshot through a target,
      scored by pinned evaluators.
  - name: Annotation queues
    description: >-
      Human review: queue traces, threads and experiment items, reserve one at a
      time, score them and feed corrections back into datasets.
  - name: Monitors
    description: >-
      Monitors: evaluators scoring a sampled share of a deployment's live
      traffic.
  - name: Observability
    description: Agent conversations, agents and their traffic.
  - name: Public runs
    description: Unauthenticated reads of runs their owners published to Explore.
paths:
  /v1/evaluators:
    post:
      tags:
        - Evaluators
      summary: Create an evaluator
      description: >-
        Creates the evaluator and its version 1. An llm_judge evaluator must
        name a deployment in the caller's workspace as its judge.
      operationId: create
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/EvaluatorCreateRequest'
        required: true
      responses:
        '201':
          description: The evaluator with version 1
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/EvaluatorResponse'
        '400':
          description: >-
            Invalid name, config, bindings, or evals disabled;
            method_config.mode decision while observability is off
        '404':
          description: The judge deployment is not in the caller's workspace
        '409':
          description: An evaluator with this name already exists
      security:
        - bearerAuth: []
components:
  schemas:
    EvaluatorCreateRequest:
      type: object
      required:
        - name
        - method
        - output_type
      properties:
        name:
          type: string
          description: |-
            1-64 lowercase letters, digits, '.', '_' or '-'. Also the Langfuse
            score name.
        description:
          type:
            - string
            - 'null'
        method:
          type: string
          description: llm_judge | code | human | jev
        scope:
          type: string
          description: single_turn (default) | thread
        output_type:
          type: string
          description: boolean | categorical | numeric
        output_config:
          description: |-
            See EvaluatorVersionResponse.output_config. Optional for boolean
            (pass_when defaults to true).
        method_config:
          description: >-
            See EvaluatorVersionResponse.method_config. `runtime: cloud`
            requires

            javascript or typescript.
        bindings:
          description: Variable name -> JSONPath (`$`, `.field`, `['field']`, `[index]`).
        requires_reference:
          type: boolean
    EvaluatorResponse:
      type: object
      required:
        - object
        - id
        - name
        - latest_version
        - created_at
        - updated_at
        - version
      properties:
        object:
          type: string
          description: Always "evaluator".
        id:
          type: string
          description: '`evl_...`'
        name:
          type: string
          description: Unique per workspace; usable in place of the id in every path.
        description:
          type:
            - string
            - 'null'
        latest_version:
          type: integer
          format: int32
        archived_at:
          type:
            - string
            - 'null'
          format: date-time
          description: 'Set when archived: no new versions; hidden from the default list.'
        created_by:
          type:
            - string
            - 'null'
        created_at:
          type: string
          format: date-time
        updated_at:
          type: string
          format: date-time
        version:
          $ref: '#/components/schemas/EvaluatorVersionResponse'
          description: The latest version, or the one asked for with `?version=N`.
        versions:
          type:
            - array
            - 'null'
          items:
            $ref: '#/components/schemas/EvaluatorVersionSummary'
          description: >-
            Version history, oldest first. Only on GET
            /v1/evaluators/{id_or_name}.
    EvaluatorVersionResponse:
      type: object
      description: One immutable evaluator version.
      required:
        - object
        - evaluator_id
        - version
        - method
        - scope
        - output_type
        - output_config
        - method_config
        - bindings
        - requires_reference
        - created_at
      properties:
        object:
          type: string
          description: Always "evaluator_version".
        evaluator_id:
          type: string
        version:
          type: integer
          format: int32
        method:
          type: string
          description: llm_judge | code | human | jev
        scope:
          type: string
          description: single_turn | thread
        output_type:
          type: string
          description: boolean | categorical | numeric
        output_config:
          description: >-
            Normalized output config (defaults filled in): boolean {pass_when},

            categorical {labels, pass_labels}, numeric {min, max,
            pass_threshold,

            higher_is_better}.
        method_config:
          description: |-
            llm_judge {prompt, judge: {deployment_id, model}, mode}, code
            {language, source, runtime}, human {instructions}, jev {model,
            question: {type, instructions, criteria}, state, threshold}.
            llm_judge `mode`: `reasoned` (default, omitted) writes a reasoned
            verdict; `decision` answers with exactly one label and a confidence
            (boolean, categorical, or numeric with integer min/max at most 10
            apart). jev (TypeSafe's Jev): `question.type` noul (boolean, yes at
            P(yes) >= `threshold`, default 0.5), choice (categorical, `criteria`
            = {option: description}, 2-255 options) or score (numeric 0..N-1,
            `criteria` = N ordered level descriptions, 2-10); `state` maps field
            names to JSONPaths into the item (default {input, output}, or
            {transcript} for threads).
        bindings:
          description: Template variable name -> JSONPath.
        requires_reference:
          type: boolean
        created_by:
          type:
            - string
            - 'null'
        created_at:
          type: string
          format: date-time
    EvaluatorVersionSummary:
      type: object
      description: A version in an evaluator's history.
      required:
        - version
        - created_at
      properties:
        version:
          type: integer
          format: int32
        created_by:
          type:
            - string
            - 'null'
        created_at:
          type: string
          format: date-time
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: API key
      description: API key with the `vk_` prefix. Create one from the dashboard.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.