https://api.veri.studio after veri login.
Math reasoning with GRPO
Train Qwen2.5-1.5B-Instruct on GSM8K to work through grade-school math problems and put the final number in<answer> tags. It runs on a single L4-24GB GPU and trains LoRA adapters, so the whole job costs about $1.40.
Save the reward as gsm8k_reward.py. It pays 0.5 for using the <answer> tags and another 0.5 when the number inside them matches GSM8K’s final answer:
gsm8k_reward.py
L4-24GB x1 on AWS; the cost column is what that run was billed (about $1.26 per GPU-hour, from instance launch to completion). The quick-trial row is estimated from the same run’s startup and step times.
Wall time includes about 6 minutes of startup (instance boot, environment setup, dataset and model download) before the first step; each step then took about 18 seconds. In that run,
reward rose from 0.445 (mean of the first 25 steps) to 0.835 (mean of the last 25), with no out-of-memory errors.
Why these settings:
- LoRA on one GPU.
lora_rank: 16trains small adapter matrices instead of all 1.5B weights. Training runs in bf16, where full-finetune updates at the defaultlearning_rateof1e-6can be too small to change the weights; LoRA adapters take a normal LoRA learning rate (2e-5here) and learn visibly within a couple of hundred steps. system_prompt. It turns each GSM8K question into a system + user chat that asks for<answer>tags. Without it the prompt is the bare question, nothing asks the model for the tags the reward checks, and GRPO only learns when completions for the same prompt score differently.max_steps. A GRPO run ismax_stepssteps long (100 when omitted;num_epochsis not used by GRPO), and each step trains on one prompt and itsrollouts_per_promptcompletions. One full pass over GSM8K train (7,473 questions) would be 7,473 steps: about 37 hours and $47 on one L4 at the measured step time and rate, which is longer than the 24-hour limit on a job’s running time. Size runs withmax_steps.kl_coefleft unset. It defaults to0.0: no KL penalty and no reference model in memory.
client.training_jobs.metrics(job.id) returns them):
rewardclimbs. If it stays flat near 0, the model is not producing the format your reward checks; fix thesystem_promptor the reward before training longer.reward_stdstays above 0. If it collapses, every completion for a prompt scores the same and GRPO has nothing left to learn from.completions/mean_lengthsettles around the natural answer length instead of growing towardmax_response_length.
merged artifact (see Hugging Face).
Serve it. Because the checkpoint is an adapter, not a standalone model, it can’t be deployed on its own: creating a deployment from this job (or from a library model saved from it) returns a 400. Serve it one of two ways:
-
Merged weights. Add
hf_push={"repo": "your-namespace/qwen-1.5b-gsm8k", "artifact": "merged", "private": False}to the job above. When the job finishes, Veri merges the adapter into the base model and pushes standalone weights. Deploy that repo from Hugging Face. Deployments download the repo without your Hugging Face token, so it must be public. -
Adapter on its base model (beta). Save the job as an adapter (
model = client.models.create_from_training_job(job.id, name="qwen-1.5b-gsm8k", artifact="adapter")). Create a deployment ofQwen/Qwen2.5-1.5B-Instructwith"serving_mode": "multi_adapter"in thePOST /v1/deploymentsbody (the SDK’sclient.deployments.createhas noserving_modeargument), then attach the adapter under an alias withPOST /v1/deployments/{id}/adapters(SDK:client.deployments.attach_adapter(deployment_id, serve_name="qwen-1.5b-gsm8k", model_id=model.id)). Requests whosemodelis that alias are answered by the adapter, and one deployment can serve many adapters of the same base. Multi-adapter fleets are in beta: see Multi-adapter fleets for the known limitations.
For larger models, use
A100-80GB or H100-80GB; on AWS these come as whole 8-GPU nodes (gpu_count=8). For GRPO, gpu_count above 1 shards one copy of the model across the GPUs: more memory, not more speed. Data-parallel training on several GPUs is available for sft_text and dpo; see Multi-GPU training.Custom multi-signal reward
A reward that combines correctness, format, and a length penalty. Useful when the model gets the answer right but produces verbose, off-topic preambles.1.0. Each weight is tunable. Test locally before submitting:
Where to go next
GRPO algorithm
Hyperparameter intuition + failure modes.
Deploy your trained model
Serve the result with an OpenAI-compatible endpoint.
Guided demos
Full walkthroughs of these recipes, end to end.

