Skip to main content
POST
Drain a replica

Authorizations

Authorization
string
header
required

API key with the vk_ prefix. Create one from the dashboard.

Path Parameters

deployment_id
string
required

Deployment ID

replica_id
string
required

Replica ID

Query Parameters

replace
boolean

Swap mode: retire this replica WITHOUT shrinking the fleet or losing capacity. A replacement is launched first; this replica keeps serving until the replacement serves, then drains. Both boxes bill during the overlap.

Response

Drain accepted. Default mode: the replica is draining (or stopped, when it was still booting and was cancelled outright). replace=true on a serving replica: it is still serving with replace_requested_at set, and drains once its replacement serves

A frontend-safe view of a serving replica. The endpoint itself remains private to the control plane; the topology only needs to know whether an endpoint is currently routable.

id
string
required
status
string
required
endpoint_available
boolean
required
provider
string | null
region
string | null

Cloud region the box launched into; None when the provider path does not expose it (pre-placement rows, non-AWS providers).

gpu_type
string | null
gpu_count
integer<int32> | null
cost_per_box_hour_usd
number<double> | null

Per-box hourly rate: the rate snapshotted on the replica at launch, else the deployment's per-box fallback.

started_at
string<date-time> | null
last_heartbeat_at
string<date-time> | null
requests_running
number<double> | null
requests_waiting
number<double> | null
kv_cache_usage
number<double> | null
gpu_utilization
number<double> | null

GPU hardware snapshot (NVML, from the last heartbeat). All fields absent when the worker did not report telemetry. Debugging metric only: scale on requests_waiting + kv_cache_usage, not this (nvidia-smi-style utilization measures kernel residency, not pressure). Mean utilization percent (0-100) across the replica's GPUs.

gpu_memory_used_bytes
number<double> | null

Summed GPU memory in use across the replica's GPUs, bytes.

gpu_memory_total_bytes
number<double> | null

Summed GPU memory capacity across the replica's GPUs, bytes.

gpu_power_w
number<double> | null

Summed power draw across the replica's GPUs, watts.

gpu_temperature_c
number<double> | null

Hottest GPU on the replica, degrees C.

gpu_devices
object[] | null

Per-GPU readings behind the aggregates above.

replace_requested_at
string<date-time> | null

Set while a drain?replace=true is in flight for this replica: it keeps serving traffic until its replacement serves, then drains. Null when no replace is in flight.

replaces_replica_id
string | null

On a replacement box launched by drain?replace=true: the replica it replaces. Cleared once it takes over; kept when it failed.

replace_error
string | null

Why the last drain?replace=true of this replica failed (the replacement crashed, did not become ready within its provisioning deadline, or could not launch). The replica kept serving; it is not retried automatically, so call the drain with replace=true again.