Skip to main content
PATCH
Update deployment scaling settings

Authorizations

Authorization
string
header
required

API key with the vk_ prefix. Create one from the dashboard.

Path Parameters

deployment_id
string
required

Deployment ID

Body

application/json

PATCH /v1/deployments/{id}: the mutable-field whitelist. Everything else on a deployment (model, source, gpu, engine, provider) is immutable: changing those means creating a new deployment. Sending an immutable field in a PATCH is a validation error instead of a silent no-op.

Desired-state semantics: the PATCH validates + persists the new bounds and returns immediately (desired + observed); the reconciler's scaling pass converges the replica fleet asynchronously. A scaled_to_zero deployment left with min_replicas >= 1 is woken by the PATCH itself (cap-checked); during a first boot the new bounds apply once the deployment serves. An omitted field keeps its stored value; only scale_to_zero_window_seconds and idle_delete_after_days read an explicit null as "clear".

num_replicas
integer<int32> | null
min_replicas
integer<int32> | null
max_replicas
integer<int32> | null
concurrency_target
number<double> | null
scale_to_zero_window_seconds
integer<int32> | null

Idle seconds before a min_replicas = 0 deployment parks (minimum 300). Omit to keep the current value; null clears it back to the default (3600). Changing or clearing it while scale-to-zero is on starts a fresh idle window.

idle_delete_after_days
integer<int32> | null

Days a deployment may stay parked with no traffic before it is auto-stopped (minimum 1). Omit to keep the current value; null clears it (never auto-stop).

self_heal_attempts
integer<int32> | null

Replacements of a lost last GPU box, 0..=5. Any PATCH also restarts the replacement budget.

Response

The updated deployment (desired state; the reconciler converges the replica fleet asynchronously). A scaled_to_zero deployment left with min_replicas >= 1 is woken by the update (status waking, booting the new floor); during a first boot (provisioning) the new scaling settings apply once it is serving

object
string
required
id
string
required
status
enum<string>
required
Available options:
queued,
provisioning,
serving,
stopped,
failed,
scaled_to_zero,
waking
name
string
required
model
string
required
source
string
required
created_at
string<date-time>
required
updated_at
string<date-time>
required
source_model_id
string | null

For a custom_model deployment, the saved model this was created from (provenance). Also set on a source=huggingface deployment that took the transparent cache hit (serving from the caller's cached S3 copy). Null otherwise.

cache_model_id
string | null

HF cache-through: the library model (status importing) this deployment's box is uploading its snapshot into. Null unless the deployment was created with cache=true and the upload is the one this boot owns.

endpoint_url
string | null

The URL to call once the deployment serves. Managed deployments (aws, gcp, digitalocean) get the platform chat URL {api}/v1/deployments/{id}/chat/completions, authenticated with your Veri API key; BYOC (huggingface) deployments carry the endpoint URL under your own HF account. Null until the deployment first serves, and null while it is parked (scaled_to_zero) until a wake serves again; it stays set after stop or failure.

gpu
null | object

Pydantic GPUInfo (with type field). Rust uses gpu_type internally but serializes as "type" for JSON parity.

provider
string | null
byoc_repo
string | null

BYOC (provider='huggingface'): the HF hub repository this endpoint serves under the owner's account. Null for every non-BYOC deployment.

engine
string | null
vllm_extra_args
string[] | null

The extra vLLM CLI flags this deployment was created with (argv list). Null unless supplied at create.

vllm_image
string | null

The explicit vLLM serving image this deployment was created with. Null unless supplied at create.

startup_profiling
boolean

Whether this deployment was created with startup diagnostics on.

trace_bodies
boolean

Whether request/response bodies are captured as traces.

trace_retention
string | null

Trace retention class ("extended" or null = base). Null unless trace_bodies deployments set it.

startup_metrics
any

Startup phase timings reported by a profiled deployment's first healthy heartbeat (seconds: model_sync_s, engine_start_to_healthy_s, total_s). Null unless startup_profiling was set and the engine reached healthy.

num_replicas
integer<int32>

Desired replica count (GPU boxes) behind this deployment. The observed count is ready_replicas; the CLI renders "ready/desired" (e.g. 2/3).

ready_replicas
integer<int32>

Observed: replicas currently serving with a routable endpoint.

min_replicas
integer<int32>

Replica-count bounds. min == max => fixed N (autoscaling off).

max_replicas
integer<int32>
concurrency_target
number<double> | null

Per-replica concurrency setpoint (stored as target_ongoing_requests).

scale_to_zero_window_seconds
integer<int32> | null

Idle window before a min=0 deployment parks (None = 3600s).

idle_delete_after_days
integer<int32> | null

Days a parked deployment is kept before it is stopped (None = kept forever).

self_heal_attempts
integer<int32>

How many times Veri replaces the GPU box when the deployment's last box dies after it has served, 0..=5 (default 3).

cost_per_hour_usd
number<double> | null
total_cost_usd
number<double> | null
total_requests
integer<int64>
error
string | null
failure_code
string | null

Machine-readable reason for the last failure, e.g. watchdog_timeout, CAPACITY_UNAVAILABLE, replica_lost (the last GPU box was lost after serving: while 'provisioning' Veri is replacing it, with 'failed' or 'scaled_to_zero' self_heal_attempts was used up or 0), or the serving agent's oom / model_source_unsupported. Null when no failure is recorded; cleared with error on an idle park and on wake. A failed wake re-parks with its code (e.g. wake_boot_failed).

last_wake_at
string<date-time> | null

When the most recent wake of this scale-to-zero deployment started: a chat request to the parked deployment, POST /wake, or a PATCH that left a parked deployment with min_replicas >= 1. Kept after the wake serves or the deployment parks again. Null if it never woke.

wake_failure_count
integer<int32>

Consecutive failed wake boots since the deployment last reached serving (0 when the last boot served). A failed wake re-parks the deployment (scaled_to_zero, still wakeable) and adds 1; reaching serving resets it to 0. Starting a new wake does not reset it.

next_wake_not_before
string<date-time> | null

After a failed wake: the earliest time a chat request triggers the next wake (backoff of 60s doubling per consecutive failure, up to 1h). POST /wake and a PATCH that raises min_replicas to 1 or more bypass it. Null until a wake fails; cleared when a wake reaches serving. A time in the past means the backoff has elapsed.

started_at
string<date-time> | null
stopped_at
string<date-time> | null