Update deployment scaling settings
curl --request PATCH \
--url https://api.veri.studio/v1/deployments/{deployment_id} \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"num_replicas": 123,
"min_replicas": 123,
"max_replicas": 123,
"concurrency_target": 123,
"scale_to_zero_window_seconds": 123,
"idle_delete_after_days": 123,
"self_heal_attempts": 123
}
'import requests
url = "https://api.veri.studio/v1/deployments/{deployment_id}"
payload = {
"num_replicas": 123,
"min_replicas": 123,
"max_replicas": 123,
"concurrency_target": 123,
"scale_to_zero_window_seconds": 123,
"idle_delete_after_days": 123,
"self_heal_attempts": 123
}
headers = {
"Authorization": "Bearer <token>",
"Content-Type": "application/json"
}
response = requests.patch(url, json=payload, headers=headers)
print(response.text)const options = {
method: 'PATCH',
headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},
body: JSON.stringify({
num_replicas: 123,
min_replicas: 123,
max_replicas: 123,
concurrency_target: 123,
scale_to_zero_window_seconds: 123,
idle_delete_after_days: 123,
self_heal_attempts: 123
})
};
fetch('https://api.veri.studio/v1/deployments/{deployment_id}', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"object": "<string>",
"id": "<string>",
"status": "queued",
"name": "<string>",
"model": "<string>",
"source": "<string>",
"created_at": "2023-11-07T05:31:56Z",
"updated_at": "2023-11-07T05:31:56Z",
"source_model_id": "<string>",
"cache_model_id": "<string>",
"endpoint_url": "<string>",
"gpu": {
"type": "<string>",
"count": 123
},
"provider": "<string>",
"byoc_repo": "<string>",
"engine": "<string>",
"vllm_extra_args": [
"<string>"
],
"vllm_image": "<string>",
"startup_profiling": true,
"trace_bodies": true,
"trace_retention": "<string>",
"startup_metrics": "<unknown>",
"num_replicas": 123,
"ready_replicas": 123,
"min_replicas": 123,
"max_replicas": 123,
"concurrency_target": 123,
"scale_to_zero_window_seconds": 123,
"idle_delete_after_days": 123,
"self_heal_attempts": 123,
"cost_per_hour_usd": 123,
"total_cost_usd": 123,
"total_requests": 123,
"error": "<string>",
"failure_code": "<string>",
"last_wake_at": "2023-11-07T05:31:56Z",
"wake_failure_count": 123,
"next_wake_not_before": "2023-11-07T05:31:56Z",
"started_at": "2023-11-07T05:31:56Z",
"stopped_at": "2023-11-07T05:31:56Z"
}Authorizations
API key with the vk_ prefix. Create one from the dashboard.
Path Parameters
Deployment ID
Body
PATCH /v1/deployments/{id}: the mutable-field whitelist. Everything else on a deployment (model, source, gpu, engine, provider) is immutable: changing those means creating a new deployment. Sending an immutable field in a PATCH is a validation error instead of a silent no-op.
Desired-state semantics: the PATCH validates + persists the new bounds and returns immediately (desired + observed); the reconciler's scaling pass converges the replica fleet asynchronously. A scaled_to_zero deployment left with min_replicas >= 1 is woken by the PATCH itself (cap-checked); during a first boot the new bounds apply once the deployment serves. An omitted field keeps its stored value; only scale_to_zero_window_seconds and idle_delete_after_days read an explicit null as "clear".
Idle seconds before a min_replicas = 0 deployment parks (minimum 300). Omit to keep the current value; null clears it back to the default (3600). Changing or clearing it while scale-to-zero is on starts a fresh idle window.
Days a deployment may stay parked with no traffic before it is auto-stopped (minimum 1). Omit to keep the current value; null clears it (never auto-stop).
Replacements of a lost last GPU box, 0..=5. Any PATCH also restarts the replacement budget.
Response
The updated deployment (desired state; the reconciler converges the replica fleet asynchronously). A scaled_to_zero deployment left with min_replicas >= 1 is woken by the update (status waking, booting the new floor); during a first boot (provisioning) the new scaling settings apply once it is serving
queued, provisioning, serving, stopped, failed, scaled_to_zero, waking For a custom_model deployment, the saved model this was created from (provenance). Also set on a source=huggingface deployment that took the transparent cache hit (serving from the caller's cached S3 copy). Null otherwise.
HF cache-through: the library model (status importing) this deployment's box is uploading its snapshot into. Null unless the deployment was created with cache=true and the upload is the one this boot owns.
The URL to call once the deployment serves. Managed deployments
(aws, gcp, digitalocean) get the platform chat URL
{api}/v1/deployments/{id}/chat/completions, authenticated with your
Veri API key; BYOC (huggingface) deployments carry the endpoint URL
under your own HF account. Null until the deployment first serves, and
null while it is parked (scaled_to_zero) until a wake serves again; it
stays set after stop or failure.
Pydantic GPUInfo (with type field). Rust uses gpu_type internally
but serializes as "type" for JSON parity.
Show child attributes
Show child attributes
BYOC (provider='huggingface'): the HF hub repository this endpoint serves under the owner's account. Null for every non-BYOC deployment.
The extra vLLM CLI flags this deployment was created with (argv list). Null unless supplied at create.
The explicit vLLM serving image this deployment was created with. Null unless supplied at create.
Whether this deployment was created with startup diagnostics on.
Whether request/response bodies are captured as traces.
Trace retention class ("extended" or null = base). Null unless trace_bodies deployments set it.
Startup phase timings reported by a profiled deployment's first healthy heartbeat (seconds: model_sync_s, engine_start_to_healthy_s, total_s). Null unless startup_profiling was set and the engine reached healthy.
Desired replica count (GPU boxes) behind this deployment. The observed count is ready_replicas; the CLI renders "ready/desired" (e.g. 2/3).
Observed: replicas currently serving with a routable endpoint.
Replica-count bounds. min == max => fixed N (autoscaling off).
Per-replica concurrency setpoint (stored as target_ongoing_requests).
Idle window before a min=0 deployment parks (None = 3600s).
Days a parked deployment is kept before it is stopped (None = kept forever).
How many times Veri replaces the GPU box when the deployment's last box dies after it has served, 0..=5 (default 3).
Machine-readable reason for the last failure, e.g. watchdog_timeout,
CAPACITY_UNAVAILABLE, replica_lost (the last GPU box was lost
after serving: while 'provisioning' Veri is replacing it, with
'failed' or 'scaled_to_zero' self_heal_attempts was used up or 0), or
the serving agent's oom / model_source_unsupported. Null when no
failure is recorded; cleared with error on an idle park and on wake.
A failed wake re-parks with its code (e.g. wake_boot_failed).
When the most recent wake of this scale-to-zero deployment started: a
chat request to the parked deployment, POST /wake, or a PATCH that
left a parked deployment with min_replicas >= 1. Kept after the
wake serves or the deployment parks again. Null if it never woke.
Consecutive failed wake boots since the deployment last reached
serving (0 when the last boot served). A failed wake re-parks the
deployment (scaled_to_zero, still wakeable) and adds 1; reaching
serving resets it to 0. Starting a new wake does not reset it.
After a failed wake: the earliest time a chat request triggers the
next wake (backoff of 60s doubling per consecutive failure, up to 1h).
POST /wake and a PATCH that raises min_replicas to 1 or more
bypass it. Null until a wake fails; cleared when a wake reaches
serving. A time in the past means the backoff has elapsed.
curl --request PATCH \
--url https://api.veri.studio/v1/deployments/{deployment_id} \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: application/json' \
--data '
{
"num_replicas": 123,
"min_replicas": 123,
"max_replicas": 123,
"concurrency_target": 123,
"scale_to_zero_window_seconds": 123,
"idle_delete_after_days": 123,
"self_heal_attempts": 123
}
'import requests
url = "https://api.veri.studio/v1/deployments/{deployment_id}"
payload = {
"num_replicas": 123,
"min_replicas": 123,
"max_replicas": 123,
"concurrency_target": 123,
"scale_to_zero_window_seconds": 123,
"idle_delete_after_days": 123,
"self_heal_attempts": 123
}
headers = {
"Authorization": "Bearer <token>",
"Content-Type": "application/json"
}
response = requests.patch(url, json=payload, headers=headers)
print(response.text)const options = {
method: 'PATCH',
headers: {Authorization: 'Bearer <token>', 'Content-Type': 'application/json'},
body: JSON.stringify({
num_replicas: 123,
min_replicas: 123,
max_replicas: 123,
concurrency_target: 123,
scale_to_zero_window_seconds: 123,
idle_delete_after_days: 123,
self_heal_attempts: 123
})
};
fetch('https://api.veri.studio/v1/deployments/{deployment_id}', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"object": "<string>",
"id": "<string>",
"status": "queued",
"name": "<string>",
"model": "<string>",
"source": "<string>",
"created_at": "2023-11-07T05:31:56Z",
"updated_at": "2023-11-07T05:31:56Z",
"source_model_id": "<string>",
"cache_model_id": "<string>",
"endpoint_url": "<string>",
"gpu": {
"type": "<string>",
"count": 123
},
"provider": "<string>",
"byoc_repo": "<string>",
"engine": "<string>",
"vllm_extra_args": [
"<string>"
],
"vllm_image": "<string>",
"startup_profiling": true,
"trace_bodies": true,
"trace_retention": "<string>",
"startup_metrics": "<unknown>",
"num_replicas": 123,
"ready_replicas": 123,
"min_replicas": 123,
"max_replicas": 123,
"concurrency_target": 123,
"scale_to_zero_window_seconds": 123,
"idle_delete_after_days": 123,
"self_heal_attempts": 123,
"cost_per_hour_usd": 123,
"total_cost_usd": 123,
"total_requests": 123,
"error": "<string>",
"failure_code": "<string>",
"last_wake_at": "2023-11-07T05:31:56Z",
"wake_failure_count": 123,
"next_wake_not_before": "2023-11-07T05:31:56Z",
"started_at": "2023-11-07T05:31:56Z",
"stopped_at": "2023-11-07T05:31:56Z"
}
