Custom training scripts
Verl GSM8K PPO
A planned PPO math reasoning demo inspired by the Verl GSM8K quickstart.
This demo is not ready yet.
It will cover a GSM8K PPO training workflow inspired by the Verl quickstart: preprocessing GSM8K, training Qwen with a rule-based reward, and reading reward/length metrics.
See the upstream Verl quickstart for the reference workflow: PPO training on GSM8K.

