Drain a replica
Retire one replica of a serving deployment. Default mode shrinks the fleet by one: the replica starts draining at once and is not replaced. replace=true swaps the replica without losing capacity, on any replica including the last one serving: a replacement is launched first, this replica keeps serving (replace_requested_at set) until the replacement serves, then it drains. Both boxes are billed during the overlap. If the replacement fails (crashes, misses its provisioning deadline, or finds no capacity) this replica keeps serving and GET …/replicas shows why in replace_error; it is not retried automatically. A draining replica finishes its in-flight requests before its box is terminated (bounded by the drain timeout); a replica still booting (queued/provisioning) is cancelled outright in either mode.
curl --request POST \
--url https://api.veri.studio/v1/deployments/{deployment_id}/replicas/{replica_id}/drain \
--header 'Authorization: Bearer <token>'import requests
url = "https://api.veri.studio/v1/deployments/{deployment_id}/replicas/{replica_id}/drain"
headers = {"Authorization": "Bearer <token>"}
response = requests.post(url, headers=headers)
print(response.text)const options = {method: 'POST', headers: {Authorization: 'Bearer <token>'}};
fetch('https://api.veri.studio/v1/deployments/{deployment_id}/replicas/{replica_id}/drain', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"id": "<string>",
"status": "<string>",
"endpoint_available": true,
"provider": "<string>",
"region": "<string>",
"gpu_type": "<string>",
"gpu_count": 123,
"cost_per_box_hour_usd": 123,
"started_at": "2023-11-07T05:31:56Z",
"last_heartbeat_at": "2023-11-07T05:31:56Z",
"requests_running": 123,
"requests_waiting": 123,
"kv_cache_usage": 123,
"gpu_utilization": 123,
"gpu_memory_used_bytes": 123,
"gpu_memory_total_bytes": 123,
"gpu_power_w": 123,
"gpu_temperature_c": 123,
"gpu_devices": [
{
"index": 123,
"utilization": 123,
"memory_used_bytes": 123,
"memory_total_bytes": 123,
"power_w": 123,
"temperature_c": 123
}
],
"replace_requested_at": "2023-11-07T05:31:56Z",
"replaces_replica_id": "<string>",
"replace_error": "<string>"
}Authorizations
API key with the vk_ prefix. Create one from the dashboard.
Query Parameters
Swap mode: retire this replica WITHOUT shrinking the fleet or losing capacity. A replacement is launched first; this replica keeps serving until the replacement serves, then drains. Both boxes bill during the overlap.
Response
Drain accepted. Default mode: the replica is draining (or stopped, when it was still booting and was cancelled outright). replace=true on a serving replica: it is still serving with replace_requested_at set, and drains once its replacement serves
A frontend-safe view of a serving replica. The endpoint itself remains private to the control plane; the topology only needs to know whether an endpoint is currently routable.
Cloud region the box launched into; None when the provider path does not expose it (pre-placement rows, non-AWS providers).
Per-box hourly rate: the rate snapshotted on the replica at launch, else the deployment's per-box fallback.
GPU hardware snapshot (NVML, from the last heartbeat). All fields absent when the worker did not report telemetry. Debugging metric only: scale on requests_waiting + kv_cache_usage, not this (nvidia-smi-style utilization measures kernel residency, not pressure). Mean utilization percent (0-100) across the replica's GPUs.
Summed GPU memory in use across the replica's GPUs, bytes.
Summed GPU memory capacity across the replica's GPUs, bytes.
Summed power draw across the replica's GPUs, watts.
Hottest GPU on the replica, degrees C.
Per-GPU readings behind the aggregates above.
Show child attributes
Show child attributes
Set while a drain?replace=true is in flight for this replica: it keeps
serving traffic until its replacement serves, then drains. Null when no
replace is in flight.
On a replacement box launched by drain?replace=true: the replica it
replaces. Cleared once it takes over; kept when it failed.
Why the last drain?replace=true of this replica failed (the
replacement crashed, did not become ready within its provisioning
deadline, or could not launch). The replica kept serving; it is not
retried automatically, so call the drain with replace=true again.
curl --request POST \
--url https://api.veri.studio/v1/deployments/{deployment_id}/replicas/{replica_id}/drain \
--header 'Authorization: Bearer <token>'import requests
url = "https://api.veri.studio/v1/deployments/{deployment_id}/replicas/{replica_id}/drain"
headers = {"Authorization": "Bearer <token>"}
response = requests.post(url, headers=headers)
print(response.text)const options = {method: 'POST', headers: {Authorization: 'Bearer <token>'}};
fetch('https://api.veri.studio/v1/deployments/{deployment_id}/replicas/{replica_id}/drain', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"id": "<string>",
"status": "<string>",
"endpoint_available": true,
"provider": "<string>",
"region": "<string>",
"gpu_type": "<string>",
"gpu_count": 123,
"cost_per_box_hour_usd": 123,
"started_at": "2023-11-07T05:31:56Z",
"last_heartbeat_at": "2023-11-07T05:31:56Z",
"requests_running": 123,
"requests_waiting": 123,
"kv_cache_usage": 123,
"gpu_utilization": 123,
"gpu_memory_used_bytes": 123,
"gpu_memory_total_bytes": 123,
"gpu_power_w": 123,
"gpu_temperature_c": 123,
"gpu_devices": [
{
"index": 123,
"utilization": 123,
"memory_used_bytes": 123,
"memory_total_bytes": 123,
"power_w": 123,
"temperature_c": 123
}
],
"replace_requested_at": "2023-11-07T05:31:56Z",
"replaces_replica_id": "<string>",
"replace_error": "<string>"
}
