> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veri.studio/llms.txt
> Use this file to discover all available pages before exploring further.

# Update deployment scaling settings



## OpenAPI

````yaml /api-reference/openapi.json patch /v1/deployments/{deployment_id}
openapi: 3.1.0
info:
  title: Veri API
  description: >-
    REST API for the Veri RL post-training platform. All requests require a
    Bearer API key (`vk_` prefix).
  license:
    name: ''
  version: 0.1.0
servers:
  - url: https://api.veri.studio
    description: Production
security: []
tags:
  - name: Training jobs
    description: Create, monitor, and manage training jobs.
  - name: Datasets
    description: Upload and connect training datasets.
  - name: Deployments
    description: Serve trained models and run inference.
  - name: Volumes
    description: Persistent file storage mounted into jobs.
  - name: Models
    description: Custom model registry deployments serve from.
  - name: Regions
    description: Discover available launch regions.
  - name: GPU
    description: Live GPU availability by provider and region.
  - name: Code artifacts
    description: Upload custom training script bundles.
  - name: Billing
    description: Credit balance and transaction history.
  - name: API keys
    description: Create and revoke API keys.
  - name: Account
    description: The authenticated caller's identity.
  - name: Settings
    description: Account-level integrations (Weights & Biases).
  - name: Metrics
    description: Prometheus metrics export for your own observability stack.
  - name: Evaluators
    description: 'Evaluators: versioned scoring rules (LLM judge, code, human).'
  - name: Experiments
    description: >-
      Experiments: offline runs of a pinned dataset snapshot through a target,
      scored by pinned evaluators.
  - name: Annotation queues
    description: >-
      Human review: queue traces, threads and experiment items, reserve one at a
      time, score them and feed corrections back into datasets.
  - name: Monitors
    description: >-
      Monitors: evaluators scoring a sampled share of a deployment's live
      traffic.
  - name: Observability
    description: Agent conversations, agents and their traffic.
  - name: Public runs
    description: Unauthenticated reads of runs their owners published to Explore.
paths:
  /v1/deployments/{deployment_id}:
    patch:
      tags:
        - Deployments
      summary: Update deployment scaling settings
      operationId: update
      parameters:
        - name: deployment_id
          in: path
          description: Deployment ID
          required: true
          schema:
            type: string
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/DeploymentUpdate'
        required: true
      responses:
        '200':
          description: >-
            The updated deployment (desired state; the reconciler converges the
            replica fleet asynchronously). A scaled_to_zero deployment left with
            min_replicas >= 1 is woken by the update (status waking, booting the
            new floor); during a first boot (provisioning) the new scaling
            settings apply once it is serving
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DeploymentResponse'
        '400':
          description: >-
            Invalid bounds, terminal deployment, or a pre-replica deployment
            that cannot scale
        '404':
          description: Not found or not owned by the caller
        '422':
          description: Request body failed validation (immutable or unknown field)
        '429':
          description: >-
            ACTIVE_DEPLOYMENT_LIMIT: raising the replica count, or waking a
            parked deployment at its new floor, would exceed the workspace
            active-replica limit; nothing is saved
      security:
        - bearerAuth: []
components:
  schemas:
    DeploymentUpdate:
      type: object
      description: >-
        PATCH /v1/deployments/{id}: the mutable-field whitelist. Everything else
        on

        a deployment (model, source, gpu, engine, provider) is immutable:
        changing

        those means creating a new deployment. Sending an immutable field in a

        PATCH is a validation error instead of a silent no-op.


        Desired-state semantics: the PATCH validates + persists the new bounds
        and

        returns immediately (desired + observed); the reconciler's scaling pass

        converges the replica fleet asynchronously. A scaled_to_zero deployment

        left with min_replicas >= 1 is woken by the PATCH itself (cap-checked);

        during a first boot the new bounds apply once the deployment serves.

        An omitted field keeps its stored value; only
        scale_to_zero_window_seconds

        and idle_delete_after_days read an explicit null as "clear".
      properties:
        num_replicas:
          type:
            - integer
            - 'null'
          format: int32
        min_replicas:
          type:
            - integer
            - 'null'
          format: int32
        max_replicas:
          type:
            - integer
            - 'null'
          format: int32
        concurrency_target:
          type:
            - number
            - 'null'
          format: double
        scale_to_zero_window_seconds:
          type:
            - integer
            - 'null'
          format: int32
          description: >-
            Idle seconds before a min_replicas = 0 deployment parks (minimum
            300).

            Omit to keep the current value; null clears it back to the default

            (3600). Changing or clearing it while scale-to-zero is on starts a

            fresh idle window.
        idle_delete_after_days:
          type:
            - integer
            - 'null'
          format: int32
          description: >-
            Days a deployment may stay parked with no traffic before it is

            auto-stopped (minimum 1). Omit to keep the current value; null
            clears

            it (never auto-stop).
        self_heal_attempts:
          type:
            - integer
            - 'null'
          format: int32
          description: |-
            Replacements of a lost last GPU box, 0..=5. Any PATCH also
            restarts the replacement budget.
      additionalProperties: false
    DeploymentResponse:
      type: object
      required:
        - object
        - id
        - status
        - name
        - model
        - source
        - created_at
        - updated_at
      properties:
        object:
          type: string
        id:
          type: string
        status:
          $ref: '#/components/schemas/DeploymentStatus'
        name:
          type: string
        model:
          type: string
        source:
          type: string
        source_model_id:
          type:
            - string
            - 'null'
          description: >-
            For a custom_model deployment, the saved model this was created from

            (provenance). Also set on a source=huggingface deployment that took
            the

            transparent cache hit (serving from the caller's cached S3 copy).
            Null

            otherwise.
        cache_model_id:
          type:
            - string
            - 'null'
          description: >-
            HF cache-through: the library model (status importing) this

            deployment's box is uploading its snapshot into. Null unless the

            deployment was created with cache=true and the upload is the one
            this

            boot owns.
        endpoint_url:
          type:
            - string
            - 'null'
          description: >-
            The URL to call once the deployment serves. Managed deployments

            (aws, gcp, digitalocean) get the platform chat URL

            `{api}/v1/deployments/{id}/chat/completions`, authenticated with
            your

            Veri API key; BYOC (huggingface) deployments carry the endpoint URL

            under your own HF account. Null until the deployment first serves,
            and

            null while it is parked (scaled_to_zero) until a wake serves again;
            it

            stays set after stop or failure.
        gpu:
          oneOf:
            - type: 'null'
            - $ref: '#/components/schemas/GpuInfo'
        provider:
          type:
            - string
            - 'null'
        byoc_repo:
          type:
            - string
            - 'null'
          description: >-
            BYOC (provider='huggingface'): the HF hub repository this endpoint

            serves under the owner's account. Null for every non-BYOC
            deployment.
        engine:
          type:
            - string
            - 'null'
        vllm_extra_args:
          type:
            - array
            - 'null'
          items:
            type: string
          description: |-
            The extra vLLM CLI flags this deployment was created with
            (argv list). Null unless supplied at create.
        vllm_image:
          type:
            - string
            - 'null'
          description: |-
            The explicit vLLM serving image this deployment was created
            with. Null unless supplied at create.
        startup_profiling:
          type: boolean
          description: Whether this deployment was created with startup diagnostics on.
        trace_bodies:
          type: boolean
          description: Whether request/response bodies are captured as traces.
        trace_retention:
          type:
            - string
            - 'null'
          description: |-
            Trace retention class ("extended" or null = base). Null unless
            trace_bodies deployments set it.
        startup_metrics:
          description: >-
            Startup phase timings reported by a profiled deployment's first
            healthy

            heartbeat (seconds: model_sync_s, engine_start_to_healthy_s,
            total_s).

            Null unless startup_profiling was set and the engine reached
            healthy.
        num_replicas:
          type: integer
          format: int32
          description: |-
            Desired replica count (GPU boxes) behind this deployment. The
            observed count is ready_replicas; the CLI renders "ready/desired"
            (e.g. 2/3).
        ready_replicas:
          type: integer
          format: int32
          description: 'Observed: replicas currently serving with a routable endpoint.'
        min_replicas:
          type: integer
          format: int32
          description: Replica-count bounds. min == max => fixed N (autoscaling off).
        max_replicas:
          type: integer
          format: int32
        concurrency_target:
          type:
            - number
            - 'null'
          format: double
          description: >-
            Per-replica concurrency setpoint (stored as
            target_ongoing_requests).
        scale_to_zero_window_seconds:
          type:
            - integer
            - 'null'
          format: int32
          description: Idle window before a min=0 deployment parks (None = 3600s).
        idle_delete_after_days:
          type:
            - integer
            - 'null'
          format: int32
          description: |-
            Days a parked deployment is kept before it is stopped (None = kept
            forever).
        self_heal_attempts:
          type: integer
          format: int32
          description: |-
            How many times Veri replaces the GPU box when the deployment's
            last box dies after it has served, 0..=5 (default 3).
        cost_per_hour_usd:
          type:
            - number
            - 'null'
          format: double
        total_cost_usd:
          type:
            - number
            - 'null'
          format: double
        total_requests:
          type: integer
          format: int64
        error:
          type:
            - string
            - 'null'
        failure_code:
          type:
            - string
            - 'null'
          description: >-
            Machine-readable reason for the last failure, e.g.
            `watchdog_timeout`,

            `CAPACITY_UNAVAILABLE`, `replica_lost` (the last GPU box was lost

            after serving: while 'provisioning' Veri is replacing it, with

            'failed' or 'scaled_to_zero' self_heal_attempts was used up or 0),
            or

            the serving agent's `oom` / `model_source_unsupported`. Null when no

            failure is recorded; cleared with `error` on an idle park and on
            wake.

            A failed wake re-parks with its code (e.g. `wake_boot_failed`).
        last_wake_at:
          type:
            - string
            - 'null'
          format: date-time
          description: >-
            When the most recent wake of this scale-to-zero deployment started:
            a

            chat request to the parked deployment, `POST /wake`, or a PATCH that

            left a parked deployment with `min_replicas` >= 1. Kept after the

            wake serves or the deployment parks again. Null if it never woke.
        wake_failure_count:
          type: integer
          format: int32
          description: |-
            Consecutive failed wake boots since the deployment last reached
            serving (0 when the last boot served). A failed wake re-parks the
            deployment (`scaled_to_zero`, still wakeable) and adds 1; reaching
            serving resets it to 0. Starting a new wake does not reset it.
        next_wake_not_before:
          type:
            - string
            - 'null'
          format: date-time
          description: >-
            After a failed wake: the earliest time a chat request triggers the

            next wake (backoff of 60s doubling per consecutive failure, up to
            1h).

            `POST /wake` and a PATCH that raises `min_replicas` to 1 or more

            bypass it. Null until a wake fails; cleared when a wake reaches

            serving. A time in the past means the backoff has elapsed.
        started_at:
          type:
            - string
            - 'null'
          format: date-time
        stopped_at:
          type:
            - string
            - 'null'
          format: date-time
        created_at:
          type: string
          format: date-time
        updated_at:
          type: string
          format: date-time
    DeploymentStatus:
      type: string
      enum:
        - queued
        - provisioning
        - serving
        - stopped
        - failed
        - scaled_to_zero
        - waking
    GpuInfo:
      type: object
      description: |-
        Pydantic `GPUInfo` (with `type` field). Rust uses `gpu_type` internally
        but serializes as "type" for JSON parity.
      required:
        - type
        - count
      properties:
        type:
          type: string
        count:
          type: integer
          format: int32
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: API key
      description: API key with the `vk_` prefix. Create one from the dashboard.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.