Skip to main content
This demo is not ready yet. It will cover a GSM8K PPO training workflow inspired by the Verl quickstart: preprocessing GSM8K, training Qwen with a rule-based reward, and reading reward/length metrics. See the upstream Verl quickstart for the reference workflow: PPO training on GSM8K.