Skip to main content
Two runnable examples, each tuned for a different use case. All snippets work against https://api.veri.studio after veri login.

Math reasoning with GRPO

Train Qwen2.5-1.5B-Instruct on GSM8K to work through grade-school math problems and put the final number in <answer> tags. It runs on a single L4-24GB GPU and trains LoRA adapters, so the whole job costs about $1.40. Save the reward as gsm8k_reward.py. It pays 0.5 for using the <answer> tags and another 0.5 when the number inside them matches GSM8K’s final answer:
gsm8k_reward.py
Then connect the dataset and submit the job:
Time and cost. We ran this exact script on L4-24GB x1 on AWS; the cost column is what that run was billed (about $1.26 per GPU-hour, from instance launch to completion). The quick-trial row is estimated from the same run’s startup and step times. Wall time includes about 6 minutes of startup (instance boot, environment setup, dataset and model download) before the first step; each step then took about 18 seconds. In that run, reward rose from 0.445 (mean of the first 25 steps) to 0.835 (mean of the last 25), with no out-of-memory errors. Why these settings:
  • LoRA on one GPU. lora_rank: 16 trains small adapter matrices instead of all 1.5B weights. Training runs in bf16, where full-finetune updates at the default learning_rate of 1e-6 can be too small to change the weights; LoRA adapters take a normal LoRA learning rate (2e-5 here) and learn visibly within a couple of hundred steps.
  • system_prompt. It turns each GSM8K question into a system + user chat that asks for <answer> tags. Without it the prompt is the bare question, nothing asks the model for the tags the reward checks, and GRPO only learns when completions for the same prompt score differently.
  • max_steps. A GRPO run is max_steps steps long (100 when omitted; num_epochs is not used by GRPO), and each step trains on one prompt and its rollouts_per_prompt completions. One full pass over GSM8K train (7,473 questions) would be 7,473 steps: about 37 hours and $47 on one L4 at the measured step time and rate, which is longer than the 24-hour limit on a job’s running time. Size runs with max_steps.
  • kl_coef left unset. It defaults to 0.0: no KL penalty and no reference model in memory.
What to watch (metric names as client.training_jobs.metrics(job.id) returns them):
  • reward climbs. If it stays flat near 0, the model is not producing the format your reward checks; fix the system_prompt or the reward before training longer.
  • reward_std stays above 0. If it collapses, every completion for a prompt scores the same and GRPO has nothing left to learn from.
  • completions/mean_length settles around the natural answer length instead of growing toward max_response_length.
The trained checkpoint is a LoRA adapter (plus tokenizer files). Download it and load it on top of the base model with PEFT:
For standalone merged weights, push the finished model to your Hugging Face account with the merged artifact (see Hugging Face). Serve it. Because the checkpoint is an adapter, not a standalone model, it can’t be deployed on its own: creating a deployment from this job (or from a library model saved from it) returns a 400. Serve it one of two ways:
  • Merged weights. Add hf_push={"repo": "your-namespace/qwen-1.5b-gsm8k", "artifact": "merged", "private": False} to the job above. When the job finishes, Veri merges the adapter into the base model and pushes standalone weights. Deploy that repo from Hugging Face. Deployments download the repo without your Hugging Face token, so it must be public.
  • Adapter on its base model (beta). Save the job as an adapter (model = client.models.create_from_training_job(job.id, name="qwen-1.5b-gsm8k", artifact="adapter")). Create a deployment of Qwen/Qwen2.5-1.5B-Instruct with "serving_mode": "multi_adapter" in the POST /v1/deployments body (the SDK’s client.deployments.create has no serving_mode argument), then attach the adapter under an alias with POST /v1/deployments/{id}/adapters (SDK: client.deployments.attach_adapter(deployment_id, serve_name="qwen-1.5b-gsm8k", model_id=model.id)). Requests whose model is that alias are answered by the adapter, and one deployment can serve many adapters of the same base. Multi-adapter fleets are in beta: see Multi-adapter fleets for the known limitations.
For larger models, use A100-80GB or H100-80GB; on AWS these come as whole 8-GPU nodes (gpu_count=8). For GRPO, gpu_count above 1 shards one copy of the model across the GPUs: more memory, not more speed. Data-parallel training on several GPUs is available for sft_text and dpo; see Multi-GPU training.

Custom multi-signal reward

A reward that combines correctness, format, and a length penalty. Useful when the model gets the answer right but produces verbose, off-topic preambles.
Total bounded at 1.0. Each weight is tunable. Test locally before submitting:
Then attach it to a job:

Where to go next

GRPO algorithm

Hyperparameter intuition + failure modes.

Deploy your trained model

Serve the result with an OpenAI-compatible endpoint.

Guided demos

Full walkthroughs of these recipes, end to end.