Skip to main content
Set gpu_count above 1 on an sft_text or dpo job and Veri runs one training process per GPU: plain data parallelism when the model’s weights, gradients and optimizer states fit one GPU, and ZeRO-3 sharding (PyTorch FSDP2) when they do not. The knobs below use the vocabulary of the Hugging Face Ultra-Scale Playbook, so what you already know transfers one to one. Everything on this page is optional. A job that sets none of these keys gets the automatic plan described under What "auto" picks. The resolved plan is printed at the top of the job log.

Knobs

All keys live under hyperparameters, next to the method’s own parameters. batch_size is still accepted on sft_text and dpo as an alias of micro_batch_size for older configs. Prefer micro_batch_size: on most managed training APIs batch_size means samples per optimizer step, which on 4 GPUs would be four times what this key does. Not available in managed training: ZeRO-1 (use zero_stage: 2, which costs no extra communication), pipeline parallelism, expert parallelism, fp16, FP8.

What "auto" picks

Before launch the worker counts the model’s parameters N from its safetensors files and estimates the model states that must live on the GPUs for the whole run: The budget per GPU is 70% of the GPU’s memory after a 2 GiB reserve for the CUDA context and communication buffers (about 15 GB on an L4-24GB, 54 GB on an H100-80GB). If the states fit the budget the job runs data parallel (zero_stage: 0); otherwise it shards them with zero_stage: 3, which divides the trainable states by the number of GPUs. A DPO reference model is sharded as one unit and is gathered whole on every GPU during its forward pass, so it is counted in full. Worked examples, all with the default "bf16": When even zero_stage: 3 does not fit, the job still launches and the log carries a warning naming the options (lora_rank, cpu_offload, more GPUs, precision). Activations are not part of the estimate: micro_batch_size and the sequence length knobs (max_seq_length, max_length) are what you lower when a job runs out of memory with the states inside the budget.

Throughput

Data parallelism exchanges every gradient between GPUs once per optimizer step, so the gain depends on how much work each GPU does per step. On an AWS L4-24GB x4 (PCIe, no peer-to-peer links) training Qwen3-0.6B: Raise micro_batch_size (or max_seq_length) until a GPU is busy for most of the step, and use gradient_accumulation_steps or global_batch_size to keep the optimizer batch where you want it. GPUs connected with NVLink or NVSwitch (A100-80GB and H100-80GB nodes) scale closer to linearly. Batch semantics change with gpu_count: the default optimizer batch is micro_batch_size x gpu_count, so four GPUs at the default micro_batch_size of 1 train on 4 samples per step, not 1. Set global_batch_size to pin the optimizer batch independent of the GPU count.

Examples

Pin the optimizer batch across GPU counts:
Force sharding and fp32 master weights for a small learning rate:
Fit a model that does not fit the GPUs even when sharded:
The same keys work in a veri.toml under [method]:

Validation at submit

These combinations are refused with a 400 before any GPU is billed:
  • data_parallel_size not equal to gpu_count (TP*CP*DP*PP=... != WORLD_SIZE=...).
  • global_batch_size that is not micro_batch_size x gradient_accumulation_steps x gpu_count (or, without gradient_accumulation_steps, not a multiple of micro_batch_size x gpu_count).
  • zero_stage: 1, pipeline_parallel_size or expert_parallel_size above 1, tensor_parallel_size or context_parallel_size above 1.
  • cpu_offload with zero_stage: 0, or any sharding knob on gpu_count: 1.
  • use_unsloth or load_in_4bit (QLoRA) with gpu_count above 1, and use_liger with gpu_count above 1.
  • The multi-GPU keys on grpo, grpo_harness or sft_video_gen, which still run their own layouts.
Everything else, including a micro_batch_size that does not fit, is the job’s own choice and is reported as a failure below.

What the job reports

  • The resolved plan is the first thing in the job log: plan: dp=4 zero_stage=3 tp=1 cp=1 mbs=1 grad_acc=1 gbs=4 precision=bf16 gradient_checkpointing=on (auto: 8.0 GB states/GPU), preceded by the parameter count, the budget and the decision.
  • The same plan is on the job as resolved_parallelism, under the knob names you submit (data_parallel_size, zero_stage, micro_batch_size, gradient_accumulation_steps, global_batch_size, precision, gradient_checkpointing, cpu_offload), plus zero_stage_requested, the memory estimate and the decision trace. Copy it into hyperparameters to pin an "auto" decision on the next run.
  • A failed job carries error.attribution (customer, platform or pending); platform failures are credited back. See Who the failure belongs to.
  • Progress and metrics come from rank 0 exactly as for a single GPU; step and total_steps count optimizer steps.
  • The final artifact is always a complete model directory (full safetensors for a full fine-tune, an adapter for LoRA), the same files a single-GPU job produces. Intermediate checkpoints are committed only once they are complete.
  • A failure names the GPU it started on: rank 2/4 (node 0, local_rank 2) failed: OutOfMemoryError: CUDA out of memory ... followed by the plan line and the traceback, with error.code set as below.

Where to go next