Intermediate checkpoints are rolling out behind a feature flag. Until it is enabled for your workspace, jobs write only the final artifact,
veri checkpoints reports the feature as unavailable, and resume_from is refused at submit.What a checkpoint is
A managed job (grpo, grpo_harness, sft_text, dpo) saves a checkpoint at a fixed cadence while it trains: the model weights (or the LoRA adapter), plus the optimizer and scheduler state, the RNG state and the trainer state. That is everything needed to continue training exactly where it stopped. Each checkpoint is uploaded in the background while training goes on, and becomes visible only once every file is stored.
Checkpoints are addressed by an id (ckpt_...) or by <job_id>@<step>, where step is the trainer’s global step.
Turning it on and tuning it
Two hyperparameters on every managed method:checkpoint_every yourself to change either.
Each save keeps one checkpoint on the training machine’s local disk and uploads a copy. Uploads run in the background, one at a time with the newest save waiting next; on a slow link an intermediate save can be skipped in favour of a newer one, which the job’s events report as checkpoint_upload_skipped. If three checkpoints would not fit on the local disk (a large full fine-tune on a provider with a small disk), the job turns periodic saves off and records checkpoint_disabled_disk; the final artifact is still saved.
Working with checkpoints
download fetches the model and tokenizer files by default, which load with from_pretrained. Pass --artifact resumable to include the optimizer and trainer state (roughly two thirds of a full fine-tune checkpoint’s bytes), or --artifact adapter for the LoRA adapter files only.
Each checkpoint carries the loss or reward logged at its step (metrics), so picking the best step is a list call.
Resuming a job
A job that failed, was cancelled, or simply stopped atmax_steps can be continued from one of its checkpoints as a new job:
method, base_model, the dataset (and its snapshot), the LoRA settings (lora_rank, lora_alpha, lora_dropout, load_in_4bit) and gpu_count. The source job’s reward functions are reused automatically; do not send new ones. What may change: max_steps or num_epochs (raise them to train further), output_name, the provider, region and GPU type. The source job must have finished, and the checkpoint must still be ready. A request that breaks any rule is refused at submit with the reason.
GRPO jobs regenerate their rollouts after a resume. A step-count checkpoint_every is aligned to a generation boundary automatically so the continued run matches an uninterrupted one.
Retention
Rotation keeps the newestcheckpoint_keep checkpoints of a running job. After a job finishes, its remaining checkpoints are kept for 7 days and then deleted, unless a resumed job that has not finished yet points at one. The final artifact is unaffected and stays with the job. Rolling checkpoints are not billed.
Custom scripts
A custom script has no trainer for Veri to hook, so it uses a directory contract instead: write acheckpoint-<step> directory into $VERI_CHECKPOINT_DIR and Veri uploads it once it stops changing (and once more when the script exits). On a multi-node job only node 0 uploads. Custom-script checkpoints are downloadable like managed ones; resume_from applies to managed jobs only.
