> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veri.studio/llms.txt
> Use this file to discover all available pages before exploring further.

# Drain a replica

> Retire one replica of a serving deployment. Default mode shrinks the fleet by one: the replica starts draining at once and is not replaced. replace=true swaps the replica without losing capacity, on any replica including the last one serving: a replacement is launched first, this replica keeps serving (replace_requested_at set) until the replacement serves, then it drains. Both boxes are billed during the overlap. If the replacement fails (crashes, misses its provisioning deadline, or finds no capacity) this replica keeps serving and GET .../replicas shows why in replace_error; it is not retried automatically. A draining replica finishes its in-flight requests before its box is terminated (bounded by the drain timeout); a replica still booting (queued/provisioning) is cancelled outright in either mode.



## OpenAPI

````yaml /api-reference/openapi.json post /v1/deployments/{deployment_id}/replicas/{replica_id}/drain
openapi: 3.1.0
info:
  title: Veri API
  description: >-
    REST API for the Veri RL post-training platform. All requests require a
    Bearer API key (`vk_` prefix).
  license:
    name: ''
  version: 0.1.0
servers:
  - url: https://api.veri.studio
    description: Production
security: []
tags:
  - name: Training jobs
    description: Create, monitor, and manage training jobs.
  - name: Datasets
    description: Upload and connect training datasets.
  - name: Deployments
    description: Serve trained models and run inference.
  - name: Volumes
    description: Persistent file storage mounted into jobs.
  - name: Models
    description: Custom model registry deployments serve from.
  - name: Regions
    description: Discover available launch regions.
  - name: GPU
    description: Live GPU availability by provider and region.
  - name: Code artifacts
    description: Upload custom training script bundles.
  - name: Billing
    description: Credit balance and transaction history.
  - name: API keys
    description: Create and revoke API keys.
  - name: Account
    description: The authenticated caller's identity.
  - name: Settings
    description: Account-level integrations (Weights & Biases).
  - name: Metrics
    description: Prometheus metrics export for your own observability stack.
  - name: Evaluators
    description: 'Evaluators: versioned scoring rules (LLM judge, code, human).'
  - name: Experiments
    description: >-
      Experiments: offline runs of a pinned dataset snapshot through a target,
      scored by pinned evaluators.
  - name: Annotation queues
    description: >-
      Human review: queue traces, threads and experiment items, reserve one at a
      time, score them and feed corrections back into datasets.
  - name: Monitors
    description: >-
      Monitors: evaluators scoring a sampled share of a deployment's live
      traffic.
  - name: Observability
    description: Agent conversations, agents and their traffic.
  - name: Public runs
    description: Unauthenticated reads of runs their owners published to Explore.
paths:
  /v1/deployments/{deployment_id}/replicas/{replica_id}/drain:
    post:
      tags:
        - Deployments
      summary: Drain a replica
      description: >-
        Retire one replica of a serving deployment. Default mode shrinks the
        fleet by one: the replica starts draining at once and is not replaced.
        replace=true swaps the replica without losing capacity, on any replica
        including the last one serving: a replacement is launched first, this
        replica keeps serving (replace_requested_at set) until the replacement
        serves, then it drains. Both boxes are billed during the overlap. If the
        replacement fails (crashes, misses its provisioning deadline, or finds
        no capacity) this replica keeps serving and GET .../replicas shows why
        in replace_error; it is not retried automatically. A draining replica
        finishes its in-flight requests before its box is terminated (bounded by
        the drain timeout); a replica still booting (queued/provisioning) is
        cancelled outright in either mode.
      operationId: drain_replica
      parameters:
        - name: deployment_id
          in: path
          description: Deployment ID
          required: true
          schema:
            type: string
        - name: replica_id
          in: path
          description: Replica ID
          required: true
          schema:
            type: string
        - name: replace
          in: query
          description: >-
            Swap mode: retire this replica WITHOUT shrinking the fleet or losing

            capacity. A replacement is launched first; this replica keeps
            serving

            until the replacement serves, then drains. Both boxes bill during
            the

            overlap.
          required: false
          schema:
            type: boolean
      responses:
        '202':
          description: >-
            Drain accepted. Default mode: the replica is draining (or stopped,
            when it was still booting and was cancelled outright). replace=true
            on a serving replica: it is still serving with replace_requested_at
            set, and drains once its replacement serves
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/DeploymentReplicaResponse'
        '404':
          description: Deployment or replica not found, or not owned by the caller
        '409':
          description: >-
            The deployment is not serving; the replica is loading, draining,
            stopped or failed; in default mode the drain would shrink the fleet
            below max(min_replicas, 1) or remove the last serving replica (use
            replace=true); or a replace is already in progress on this
            deployment
        '429':
          description: >-
            replace=true needs room for one extra replica under the workspace's
            active deployment replica limit (ACTIVE_DEPLOYMENT_LIMIT): drain
            without replace, or free capacity
      security:
        - bearerAuth: []
components:
  schemas:
    DeploymentReplicaResponse:
      type: object
      description: |-
        A frontend-safe view of a serving replica. The endpoint itself remains
        private to the control plane; the topology only needs to know whether an
        endpoint is currently routable.
      required:
        - id
        - status
        - endpoint_available
      properties:
        id:
          type: string
        status:
          type: string
        provider:
          type:
            - string
            - 'null'
        endpoint_available:
          type: boolean
        region:
          type:
            - string
            - 'null'
          description: |-
            Cloud region the box launched into; None when the provider path does
            not expose it (pre-placement rows, non-AWS providers).
        gpu_type:
          type:
            - string
            - 'null'
        gpu_count:
          type:
            - integer
            - 'null'
          format: int32
        cost_per_box_hour_usd:
          type:
            - number
            - 'null'
          format: double
          description: |-
            Per-box hourly rate: the rate snapshotted on the replica at launch,
            else the deployment's per-box fallback.
        started_at:
          type:
            - string
            - 'null'
          format: date-time
        last_heartbeat_at:
          type:
            - string
            - 'null'
          format: date-time
        requests_running:
          type:
            - number
            - 'null'
          format: double
        requests_waiting:
          type:
            - number
            - 'null'
          format: double
        kv_cache_usage:
          type:
            - number
            - 'null'
          format: double
        gpu_utilization:
          type:
            - number
            - 'null'
          format: double
          description: >-
            GPU hardware snapshot (NVML, from the last heartbeat). All

            fields absent when the worker did not report telemetry. Debugging

            metric only: scale on requests_waiting + kv_cache_usage, not this

            (nvidia-smi-style utilization measures kernel residency, not
            pressure).

            Mean utilization percent (0-100) across the replica's GPUs.
        gpu_memory_used_bytes:
          type:
            - number
            - 'null'
          format: double
          description: Summed GPU memory in use across the replica's GPUs, bytes.
        gpu_memory_total_bytes:
          type:
            - number
            - 'null'
          format: double
          description: Summed GPU memory capacity across the replica's GPUs, bytes.
        gpu_power_w:
          type:
            - number
            - 'null'
          format: double
          description: Summed power draw across the replica's GPUs, watts.
        gpu_temperature_c:
          type:
            - number
            - 'null'
          format: double
          description: Hottest GPU on the replica, degrees C.
        gpu_devices:
          type:
            - array
            - 'null'
          items:
            $ref: '#/components/schemas/GpuDeviceMetrics'
          description: Per-GPU readings behind the aggregates above.
        replace_requested_at:
          type:
            - string
            - 'null'
          format: date-time
          description: >-
            Set while a `drain?replace=true` is in flight for this replica: it
            keeps

            serving traffic until its replacement serves, then drains. Null when
            no

            replace is in flight.
        replaces_replica_id:
          type:
            - string
            - 'null'
          description: >-
            On a replacement box launched by `drain?replace=true`: the replica
            it

            replaces. Cleared once it takes over; kept when it failed.
        replace_error:
          type:
            - string
            - 'null'
          description: |-
            Why the last `drain?replace=true` of this replica failed (the
            replacement crashed, did not become ready within its provisioning
            deadline, or could not launch). The replica kept serving; it is not
            retried automatically, so call the drain with replace=true again.
    GpuDeviceMetrics:
      type: object
      description: |-
        One GPU's NVML reading from the replica's last heartbeat. Every
        field except `index` is individually best-effort (some GPUs/VMs withhold
        power or temperature).
      required:
        - index
      properties:
        index:
          type: integer
          format: int32
        utilization:
          type:
            - number
            - 'null'
          format: double
          description: |-
            Utilization percent (0-100) as NVML reports it: the share of time a
            kernel was resident, not how much of the chip it used.
        memory_used_bytes:
          type:
            - number
            - 'null'
          format: double
        memory_total_bytes:
          type:
            - number
            - 'null'
          format: double
        power_w:
          type:
            - number
            - 'null'
          format: double
        temperature_c:
          type:
            - number
            - 'null'
          format: double
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      bearerFormat: API key
      description: API key with the `vk_` prefix. Create one from the dashboard.

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.