GGursimran Singhfeat: unify LoRA support through Megatron Bridge
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat(awex): port Awex weight exchange for Megatron-to-vLLM on GPU/NPU Restore the inline `weight_update_mode=awex` path that was dropped when this branch rebased onto the rl3 infra. Awex pushes Megatron weights to vLLM through a meta server over NCCL/HCCL for weight updates on GPU and NPU. Key changes: - Add AwexConfig and weight_update_mode="awex" choice in cli_args - Add WeightUpdateMeta.from_awex and Awex fields in io_struct - Add Megatron Awex writer adapter and _update_weights_from_awex - Wire Awex runtime bootstrap into PPOTrainer and remote controllers - Add Awex init/update plumbing in vLLM/SGLang remote engines - Add GPU/NPU GSM8K sample configs and Awex tests/benchmark | 3 个月前 | |
feat(distillation): mopd implementation; aime, leetcode, and mopd dataset with examples | 1 个月前 | |
fix: apply_chat_template compatibility with transformers>=5.0 (#1280) * fix: apply_chat_template compatibility with transformers>=5.0 * chore: revert async client in inference controller * chore: fix async client | 4 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat(infra): add HTTP-based Ray scheduler Add a Ray-backed scheduler that allocates accelerator placement groups while keeping worker and inference engine traffic on the existing HTTP RPC path, with batched launcher operations and multi-node rollout support. Key changes: - Add RayScheduler and RayWorkerProcessLauncher for Ray-managed HTTP workers - Wire scheduler.type=ray into infra exports, trainer initialization, logging, docs, and examples - Batch worker startup and status checks by Ray launcher to reduce per-worker actor calls - Split multi-node rollout backend launch and cleanup into a dedicated coordinator - Tighten Ray launcher lifecycle handling for worker shutdown, placement groups, and backend process cleanup | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat: unify LoRA support through Megatron Bridge Route all LoRA variants through Megatron Bridge to avoid duplicate paths and keep registry-specific behavior within the bridge. Key changes: - Consolidate LoRA execution paths - Delegate registry-specific handling to Megatron Bridge - Document FSDP and Megatron LoRA commands in English and Chinese - Update the Megatron LoRA example configuration | 1 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat(megatron): support merged LoRA with Megatron Bridge Qwen models Enable merged LoRA weight export and synchronization through Megatron Bridge. Apply LoRA before DDP wrapping and add compatibility patches for MindSpeed row-parallel layers, Qwen3.5 hybrid attention, and Qwen3-MoE specs. Document the Megatron Bridge LoRA flow and add a merged GSM8K GRPO example. | 1 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
Supporting features for IcePop and KPop (#1405) * docs(cli): sync cli reference for icepop and kpop params * feat(examples): add icepop and kpop configs for gsm8k * style: fix end-of-file newline in icepop/kpop configs * fix(icepop/kpop): detach imp_ratio and KL inputs to prevent gradient leak * fix(kpop): use <= for KL threshold to include boundary * refactor: unify icepop/kpop into rejection_sampling with binary_kl metric - Add binary_kl metric to RejectionSamplingConfig for KPop - Remove enable_icepop/icepop_alpha/icepop_beta/enable_kpop/kpop_phi - Update gsm8k_icepop.yaml to use rejection_sampling with ratio+lower+upper - Update gsm8k_kpop.yaml to use rejection_sampling with binary_kl+upper - Sync CLI docs | 3 个月前 | |
Supporting features for IcePop and KPop (#1405) * docs(cli): sync cli reference for icepop and kpop params * feat(examples): add icepop and kpop configs for gsm8k * style: fix end-of-file newline in icepop/kpop configs * fix(icepop/kpop): detach imp_ratio and KL inputs to prevent gradient leak * fix(kpop): use <= for KL threshold to include boundary * refactor: unify icepop/kpop into rejection_sampling with binary_kl metric - Add binary_kl metric to RejectionSamplingConfig for KPop - Remove enable_icepop/icepop_alpha/icepop_beta/enable_kpop/kpop_phi - Update gsm8k_icepop.yaml to use rejection_sampling with ratio+lower+upper - Update gsm8k_kpop.yaml to use rejection_sampling with binary_kl+upper - Sync CLI docs | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
feat:enable v2 training pipeline with controller parity (#1363) * feat: enable v2 training pipeline with controller parity Bring GatewayTrainController and RolloutControllerV2 to full parity with v1 controllers for RL training paths. Key changes: - Route to RolloutControllerV2 when config._version=="v2" - Add version management, connect_engine, clear_batches to GatewayTrainController - Unify HTTP client session in GatewayTrainController (follows PR #1354) - Switch default workflow to MathAgent in example configs - Add agent config section to all example YAML files - Remove obsolete get_custom_reward_fn from reward module * fix: update wu controller connect method * chore: unblock CI for grpo and grpo_lora with admin key + lora name * chore: unblock CI for v2 parity | 3 个月前 | |
refactor(trainer): move trainer modules from experimental to areal/trainer (#896) * refactor(trainer): move trainer modules from experimental to areal/trainer Move trainer-related modules to establish a cleaner architecture: - Move PPOTrainer and SFTTrainer from areal/experimental/trainer/ to areal/trainer/ - Move PPO actor/critic from areal/engine/ppo/ to areal/trainer/ppo/ - Move SFT lm_engine from areal/engine/sft/ to areal/trainer/sft/ - Move RW engine from areal/engine/rw/ to areal/trainer/rw/ - Export PPOTrainer and SFTTrainer from top-level areal package This refactoring separates training algorithm concerns (trainer/) from backend infrastructure (engine/), making the codebase more modular. The trainers can now be imported directly via `from areal import PPOTrainer`. Updates all imports across examples, tests, docs, and internal modules. * minor fix test * fix * fix missing links | 7 个月前 | |
refactor(api): migrate allocation_mode to per-engine backend fields (#1044) * refactor(api): migrate allocation_mode to per-engine backend fields Replace the centralized `allocation_mode` string with explicit `backend` fields on `TrainEngineConfig` and `InferenceEngineConfig`. Each engine now owns its own backend+parallelism spec (e.g. `fsdp:d4`, `sglang:d4t2`), eliminating implicit auto-backend selection and the shared `AllocationMode` object. Key changes: - Add `backend` field to TrainEngineConfig and InferenceEngineConfig - Add `ModelAllocation.from_str()` for single-component parsing - Remove `AllocationMode` public export (replaced by `ModelAllocation`) - Rename internal `AllocationMode` to `_AllocationMode` for SPMD launcher backward compatibility with FutureWarning - Remove auto-backend selection — explicit backend prefix is now required - Controllers (`TrainController`, `RolloutController`) parse `backend` directly instead of receiving `alloc_mode` from trainers - `WeightUpdateMeta.alloc_mode` replaced by `gen_allocation` (single `ModelAllocation`) - Add `RWTrainer` and `ArchonRWEngine` for reward model training - Remove `get_model_update_meta()` helper (logic moved to trainers) - Update all YAML configs, examples, docs (EN+ZH), and tests BREAKING CHANGE: `AllocationMode` is removed from public API. Users must migrate to per-engine `backend` fields. SPMD launchers emit deprecation warnings. * chore(ci): fix backend specifier for vlm sft test * fix: fix bare dims for actor backends * chore(docs): fix reminder for bare allocation dims | 6 个月前 | |
fix(archon): harden FP8 blockwise training for TP and MoE scenarios (#1118) Improve FP8 robustness: extend shard alignment validation to GroupedExperts, fix DTensor handling in dense FP8 linear forward, add early checkpoint compatibility checks, and clean up config API. Key changes: - Validate GroupedExperts w1/w2/w3 shapes in post-parallelism check - Convert DTensor to local tensor in FP8 linear forward for TP>1 - Restrict FP8 dequant to float8_e4m3fn matching prepare path - Fail fast on Shard(1) FP8 checkpoints before DCP I/O - Add ArchonFP8Config.enabled property to centralize mode checks - Document exclude_modules default list in YAML example | 5 个月前 |
Hyper-parameters for GSM8K Finetuning on Qwen2.5-1.5b-Instruct
The hyperparameters given in gsm8k_grpo.yaml is the set that we found to achieve the
highest max grpo-eval/task_reward/avg during training for Qwen2.5-1.5b-Instruct. You
are free to try out more of the hyperparameters listed below!
| lr | weight decay | group size | max task_reward |
|---|---|---|---|
| 1.70E-05 | 0.017 | 4 | 0.79570 |
| 1.30E-05 | 0.015 | 8 | 0.79355 |
| 1.50E-05 | 0.01 | 4 | 0.79043 |
| 1.50E-05 | 0.02 | 4 | 0.78984 |
| 1.00E-05 | 0.02 | 4 | 0.78311 |
| 1.00E-05 | 0.01 | 8 | 0.78066 |
Other Training Details
- Devices: 8 Nvidia H800 GPUs
- Optimizer: Adam
- LR Scheduler: Constant
- Gradient Clipping: 1.0
- Max_new_tokens: 1024
- Max_head_offpolicyness: 2
- Training Time: ~35 minutes (batchsize 4), ~65 minutes (batchsize 8)
Awex Example Location
Awex-specific GSM8K sample scripts were moved to:
examples/experimental/awex/README.mdexamples/experimental/awex/gsm8k_grpo_awex_sample.yaml
Awex meta server bootstrap is now handled by PPOTrainer when
actor.weight_update_mode=awex; the standard examples/math/gsm8k_rl.py entrypoint can
be used directly with the AWEX sample yaml. Auto-start is only available in
single-controller mode; SPMD runs must provide an explicit awex.meta_server_addr.