Wake a scaled-to-zero deployment
curl --request POST \
--url https://api.veri.studio/v1/deployments/{deployment_id}/wake \
--header 'Authorization: Bearer <token>'import requests
url = "https://api.veri.studio/v1/deployments/{deployment_id}/wake"
headers = {"Authorization": "Bearer <token>"}
response = requests.post(url, headers=headers)
print(response.text)const options = {method: 'POST', headers: {Authorization: 'Bearer <token>'}};
fetch('https://api.veri.studio/v1/deployments/{deployment_id}/wake', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"object": "<string>",
"id": "<string>",
"status": "queued",
"name": "<string>",
"model": "<string>",
"source": "<string>",
"created_at": "2023-11-07T05:31:56Z",
"updated_at": "2023-11-07T05:31:56Z",
"source_model_id": "<string>",
"cache_model_id": "<string>",
"endpoint_url": "<string>",
"gpu": {
"type": "<string>",
"count": 123
},
"provider": "<string>",
"byoc_repo": "<string>",
"engine": "<string>",
"vllm_extra_args": [
"<string>"
],
"vllm_image": "<string>",
"startup_profiling": true,
"trace_bodies": true,
"trace_retention": "<string>",
"startup_metrics": "<unknown>",
"num_replicas": 123,
"ready_replicas": 123,
"min_replicas": 123,
"max_replicas": 123,
"concurrency_target": 123,
"scale_to_zero_window_seconds": 123,
"idle_delete_after_days": 123,
"self_heal_attempts": 123,
"cost_per_hour_usd": 123,
"total_cost_usd": 123,
"total_requests": 123,
"error": "<string>",
"failure_code": "<string>",
"last_wake_at": "2023-11-07T05:31:56Z",
"wake_failure_count": 123,
"next_wake_not_before": "2023-11-07T05:31:56Z",
"started_at": "2023-11-07T05:31:56Z",
"stopped_at": "2023-11-07T05:31:56Z"
}Authorizations
API key with the vk_ prefix. Create one from the dashboard.
Path Parameters
Deployment ID
Response
Wake accepted (no-op if the deployment is already up or booting)
queued, provisioning, serving, stopped, failed, scaled_to_zero, waking For a custom_model deployment, the saved model this was created from (provenance). Also set on a source=huggingface deployment that took the transparent cache hit (serving from the caller's cached S3 copy). Null otherwise.
HF cache-through: the library model (status importing) this deployment's box is uploading its snapshot into. Null unless the deployment was created with cache=true and the upload is the one this boot owns.
The URL to call once the deployment serves. Managed deployments
(aws, gcp, digitalocean) get the platform chat URL
{api}/v1/deployments/{id}/chat/completions, authenticated with your
Veri API key; BYOC (huggingface) deployments carry the endpoint URL
under your own HF account. Null until the deployment first serves, and
null while it is parked (scaled_to_zero) until a wake serves again; it
stays set after stop or failure.
Pydantic GPUInfo (with type field). Rust uses gpu_type internally
but serializes as "type" for JSON parity.
Show child attributes
Show child attributes
BYOC (provider='huggingface'): the HF hub repository this endpoint serves under the owner's account. Null for every non-BYOC deployment.
The extra vLLM CLI flags this deployment was created with (argv list). Null unless supplied at create.
The explicit vLLM serving image this deployment was created with. Null unless supplied at create.
Whether this deployment was created with startup diagnostics on.
Whether request/response bodies are captured as traces.
Trace retention class ("extended" or null = base). Null unless trace_bodies deployments set it.
Startup phase timings reported by a profiled deployment's first healthy heartbeat (seconds: model_sync_s, engine_start_to_healthy_s, total_s). Null unless startup_profiling was set and the engine reached healthy.
Desired replica count (GPU boxes) behind this deployment. The observed count is ready_replicas; the CLI renders "ready/desired" (e.g. 2/3).
Observed: replicas currently serving with a routable endpoint.
Replica-count bounds. min == max => fixed N (autoscaling off).
Per-replica concurrency setpoint (stored as target_ongoing_requests).
Idle window before a min=0 deployment parks (None = 3600s).
Days a parked deployment is kept before it is stopped (None = kept forever).
How many times Veri replaces the GPU box when the deployment's last box dies after it has served, 0..=5 (default 3).
Machine-readable reason for the last failure, e.g. watchdog_timeout,
CAPACITY_UNAVAILABLE, replica_lost (the last GPU box was lost
after serving: while 'provisioning' Veri is replacing it, with
'failed' or 'scaled_to_zero' self_heal_attempts was used up or 0), or
the serving agent's oom / model_source_unsupported. Null when no
failure is recorded; cleared with error on an idle park and on wake.
A failed wake re-parks with its code (e.g. wake_boot_failed).
When the most recent wake of this scale-to-zero deployment started: a
chat request to the parked deployment, POST /wake, or a PATCH that
left a parked deployment with min_replicas >= 1. Kept after the
wake serves or the deployment parks again. Null if it never woke.
Consecutive failed wake boots since the deployment last reached
serving (0 when the last boot served). A failed wake re-parks the
deployment (scaled_to_zero, still wakeable) and adds 1; reaching
serving resets it to 0. Starting a new wake does not reset it.
After a failed wake: the earliest time a chat request triggers the
next wake (backoff of 60s doubling per consecutive failure, up to 1h).
POST /wake and a PATCH that raises min_replicas to 1 or more
bypass it. Null until a wake fails; cleared when a wake reaches
serving. A time in the past means the backoff has elapsed.
curl --request POST \
--url https://api.veri.studio/v1/deployments/{deployment_id}/wake \
--header 'Authorization: Bearer <token>'import requests
url = "https://api.veri.studio/v1/deployments/{deployment_id}/wake"
headers = {"Authorization": "Bearer <token>"}
response = requests.post(url, headers=headers)
print(response.text)const options = {method: 'POST', headers: {Authorization: 'Bearer <token>'}};
fetch('https://api.veri.studio/v1/deployments/{deployment_id}/wake', options)
.then(res => res.json())
.then(res => console.log(res))
.catch(err => console.error(err));{
"object": "<string>",
"id": "<string>",
"status": "queued",
"name": "<string>",
"model": "<string>",
"source": "<string>",
"created_at": "2023-11-07T05:31:56Z",
"updated_at": "2023-11-07T05:31:56Z",
"source_model_id": "<string>",
"cache_model_id": "<string>",
"endpoint_url": "<string>",
"gpu": {
"type": "<string>",
"count": 123
},
"provider": "<string>",
"byoc_repo": "<string>",
"engine": "<string>",
"vllm_extra_args": [
"<string>"
],
"vllm_image": "<string>",
"startup_profiling": true,
"trace_bodies": true,
"trace_retention": "<string>",
"startup_metrics": "<unknown>",
"num_replicas": 123,
"ready_replicas": 123,
"min_replicas": 123,
"max_replicas": 123,
"concurrency_target": 123,
"scale_to_zero_window_seconds": 123,
"idle_delete_after_days": 123,
"self_heal_attempts": 123,
"cost_per_hour_usd": 123,
"total_cost_usd": 123,
"total_requests": 123,
"error": "<string>",
"failure_code": "<string>",
"last_wake_at": "2023-11-07T05:31:56Z",
"wake_failure_count": 123,
"next_wake_not_before": "2023-11-07T05:31:56Z",
"started_at": "2023-11-07T05:31:56Z",
"stopped_at": "2023-11-07T05:31:56Z"
}
