Skip to main content
AWS is the default serving provider — omit the provider field to use it.

Serving shapes

AWS has no single-GPU A100/H100 instance — the smallest SKU is the whole 8-GPU node, and you pay for all of it. For a single A100/H100, pick a provider with a genuine 1-GPU shape (GCP, or DigitalOcean when it lands).

Creating a deployment

Both engines (vllm default, max) are supported. Replicas, autoscaling, and scale-to-zero work as documented in Scaling.

Where to go next

All serving providers

Compare shapes, status, and the live CLI catalog.

GPU sizing

Match GPU memory to model size.