> ## Documentation Index
> Fetch the complete documentation index at: https://docs.veri.studio/llms.txt
> Use this file to discover all available pages before exploring further.

# Examples

> Two end-to-end training recipes: math reasoning with GRPO on a single L4, and a custom multi-signal reward.

Two runnable examples, each tuned for a different use case. All snippets work against `https://api.veri.studio` after `veri login`.

## Math reasoning with GRPO

Train Qwen2.5-1.5B-Instruct on [GSM8K](https://huggingface.co/datasets/openai/gsm8k) to work through grade-school math problems and put the final number in `<answer>` tags. It runs on a single `L4-24GB` GPU and trains LoRA adapters, so the whole job costs about \$1.40.

Save the reward as `gsm8k_reward.py`. It pays `0.5` for using the `<answer>` tags and another `0.5` when the number inside them matches GSM8K's final answer:

```python gsm8k_reward.py theme={null}
import re


def reward(completions, answer, **kwargs):
    scores = []
    for completion, expected in zip(completions, answer):
        text = completion[-1]["content"] if isinstance(completion, list) else str(completion)

        # GSM8K's `answer` is a worked solution ending in "#### <number>".
        gold = expected.split("####")[-1].strip().replace(",", "")

        m = re.search(r"<answer>(.*?)</answer>", text, re.DOTALL)
        pred = m.group(1).strip().replace(",", "").replace("$", "").rstrip(".") if m else None

        format_ok = pred is not None  # used the <answer> tags
        correct = pred == gold        # and the number inside them is right
        scores.append(0.5 * format_ok + 0.5 * correct)
    return scores
```

Then connect the dataset and submit the job:

```python theme={null}
from veri_sdk import Client
client = Client()

dataset = client.datasets.connect(
    name="gsm8k-train",
    source_type="hf",
    huggingface_dataset="openai/gsm8k",
    huggingface_config={
        "split": "train",
        "subset": "main",  # GSM8K requires a config: "main" or "socratic"
        "column_mapping": {"question": "prompt", "answer": "answer"},
    },
)

job = client.training_jobs.create(
    base_model="Qwen/Qwen2.5-1.5B-Instruct",
    dataset_id=dataset.id,
    reward_source=open("gsm8k_reward.py").read(),
    output_name="qwen-1.5b-gsm8k",
    method="grpo",
    hyperparameters={
        "max_steps": 200,             # one prompt (x rollouts_per_prompt) per step
        "rollouts_per_prompt": 8,
        "max_response_length": 512,
        "lora_rank": 16,              # train LoRA adapters, not all 1.5B weights
        "lora_alpha": 32,
        "learning_rate": 2e-5,        # LoRA learning rate
        "system_prompt": (
            "You are a math tutor. Think step by step, then put ONLY the final "
            "numeric answer inside <answer></answer> tags, e.g. <answer>42</answer>."
        ),
    },
    gpu_type="L4-24GB",
    gpu_count=1,
)

job.wait(poll_interval=30)
print(job.status)
```

**Time and cost.** We ran this exact script on `L4-24GB` x1 on AWS; the cost column is what that run was billed (about \$1.26 per GPU-hour, from instance launch to completion). The quick-trial row is estimated from the same run's startup and step times.

| Run | Training steps | Wall time | Cost |
| - | - | - | - |
| Quick trial (`"max_steps": 50`) | 50 | about 22 minutes | about \$0.50 |
| This example (`"max_steps": 200`) | 200 | 66 minutes | \$1.39 |

Wall time includes about 6 minutes of startup (instance boot, environment setup, dataset and model download) before the first step; each step then took about 18 seconds. In that run, `reward` rose from 0.445 (mean of the first 25 steps) to 0.835 (mean of the last 25), with no out-of-memory errors.

**Why these settings:**

* **LoRA on one GPU.** `lora_rank: 16` trains small adapter matrices instead of all 1.5B weights. Training runs in bf16, where full-finetune updates at the default `learning_rate` of `1e-6` can be too small to change the weights; LoRA adapters take a normal LoRA learning rate (`2e-5` here) and learn visibly within a couple of hundred steps.
* **`system_prompt`.** It turns each GSM8K question into a system + user chat that asks for `<answer>` tags. Without it the prompt is the bare question, nothing asks the model for the tags the reward checks, and GRPO only learns when completions for the same prompt score differently.
* **`max_steps`.** A GRPO run is `max_steps` steps long (100 when omitted; `num_epochs` is not used by GRPO), and each step trains on one prompt and its `rollouts_per_prompt` completions. One full pass over GSM8K train (7,473 questions) would be 7,473 steps: about 37 hours and \$47 on one L4 at the measured step time and rate, which is longer than the 24-hour limit on a job's running time. Size runs with `max_steps`.
* **`kl_coef` left unset.** It defaults to `0.0`: no KL penalty and no reference model in memory.

**What to watch** (metric names as `client.training_jobs.metrics(job.id)` returns them):

* `reward` climbs. If it stays flat near 0, the model is not producing the format your reward checks; fix the `system_prompt` or the reward before training longer.
* `reward_std` stays above 0. If it collapses, every completion for a prompt scores the same and GRPO has nothing left to learn from.
* `completions/mean_length` settles around the natural answer length instead of growing toward `max_response_length`.

The trained checkpoint is a LoRA adapter (plus tokenizer files). Download it and load it on top of the base model with PEFT:

```python theme={null}
from peft import AutoPeftModelForCausalLM
from transformers import AutoTokenizer

path = job.download(output_dir="./checkpoints")
model = AutoPeftModelForCausalLM.from_pretrained(path)
tokenizer = AutoTokenizer.from_pretrained(path)
```

For standalone merged weights, push the finished model to your Hugging Face account with the `merged` artifact (see [Hugging Face](/training/huggingface)).

**Serve it.** Because the checkpoint is an adapter, not a standalone model, it can't be deployed on its own: creating a deployment from this job (or from a library model saved from it) returns a `400`. Serve it one of two ways:

* **Merged weights.** Add `hf_push={"repo": "your-namespace/qwen-1.5b-gsm8k", "artifact": "merged", "private": False}` to the job above. When the job finishes, Veri merges the adapter into the base model and pushes standalone weights. Deploy that repo from Hugging Face. Deployments download the repo without your Hugging Face token, so it must be public.

  ```python theme={null}
  dep = client.deployments.create(
      "your-namespace/qwen-1.5b-gsm8k",
      "qwen-1.5b-gsm8k",
      source="huggingface",
      gpu={"gpu_type": "L4-24GB", "gpu_count": 1},
  )
  dep.wait()
  ```

* **Adapter on its base model (beta).** Save the job as an adapter (`model = client.models.create_from_training_job(job.id, name="qwen-1.5b-gsm8k", artifact="adapter")`). Create a deployment of `Qwen/Qwen2.5-1.5B-Instruct` with `"serving_mode": "multi_adapter"` in the `POST /v1/deployments` body (the SDK's `client.deployments.create` has no `serving_mode` argument), then attach the adapter under an alias with `POST /v1/deployments/{id}/adapters` (SDK: `client.deployments.attach_adapter(deployment_id, serve_name="qwen-1.5b-gsm8k", model_id=model.id)`). Requests whose `model` is that alias are answered by the adapter, and one deployment can serve many adapters of the same base. Multi-adapter fleets are in beta: see [Multi-adapter fleets](/deployments/api#multi-adapter-fleets) for the known limitations.

<Note>
  For larger models, use `A100-80GB` or `H100-80GB`; on AWS these come as whole 8-GPU nodes (`gpu_count=8`). For GRPO, `gpu_count` above 1 shards one copy of the model across the GPUs: more memory, not more speed. Data-parallel training on several GPUs is available for `sft_text` and `dpo`; see [Multi-GPU training](/training/multi-gpu).
</Note>

## Custom multi-signal reward

A reward that combines correctness, format, and a length penalty. Useful when the model gets the answer right but produces verbose, off-topic preambles.

```python theme={null}
# multi_signal_reward.py
import re

def reward(completions, answer, **kwargs):
    scores = []
    for completion, expected in zip(completions, answer):
        text = completion[-1]["content"] if isinstance(completion, list) else str(completion)

        # 1. Format: must wrap in <answer></answer>
        format_score = 0.3 if re.search(r"<answer>.*?</answer>", text, re.DOTALL) else 0.0

        # 2. Correctness: the expected answer appears inside the tags.
        #    GSM8K answers end in "#### <number>"; plain answers pass through unchanged.
        m = re.search(r"<answer>(.*?)</answer>", text, re.DOTALL)
        correct_score = 0.0
        if m and expected:
            inner = m.group(1).strip()
            if expected.split("####")[-1].strip() in inner:
                correct_score = 0.5

        # 3. Length: penalize responses > 300 chars
        length_score = 0.2 * max(0.0, 1.0 - max(0, len(text) - 300) / 1000)

        scores.append(format_score + correct_score + length_score)
    return scores
```

Total bounded at `1.0`. Each weight is tunable. Test locally before submitting:

```bash theme={null}
python -c "
from multi_signal_reward import reward
print([round(s, 2) for s in reward(
    ['<answer>4</answer>', 'I think it is 4', '<answer>4</answer> Let me explain at length...' + 'X' * 500],
    ['4', '4', '4'],
)])
# [1.0, 0.2, 0.95]: the first scores full marks, the second has no tags, the third loses part of the length credit
"
```

Then attach it to a job:

```python theme={null}
job = client.training_jobs.create(
    base_model="Qwen/Qwen2.5-0.5B-Instruct",
    dataset_id=dataset.id,
    reward_source=open("multi_signal_reward.py").read(),
    output_name="qwen-0.5b-multi-signal",
    method="grpo",
    hyperparameters={"max_steps": 100},
    gpu_type="L4-24GB",
    gpu_count=1,
)
```

## Where to go next

<CardGroup cols={2}>
  <Card title="GRPO algorithm" icon="brain" href="/training/grpo">
    Hyperparameter intuition + failure modes.
  </Card>

  <Card title="Deploy your trained model" icon="cloud" href="/deployments">
    Serve the result with an OpenAI-compatible endpoint.
  </Card>

  <Card title="Guided demos" icon="route" href="/demos/managed-grpo-quickstart">
    Full walkthroughs of these recipes, end to end.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.