已合并
feat(qwen): support Qwen3.5 and Qwen3-VL-MoE parallel training #658
xuxinglei创建于 5月18日
feat(qwen): support Qwen3.5 and Qwen3-VL-MoE parallel training #658
已合并
Pull Request已成功合入, 合并人@MindSpore-Bot
(感谢 xuxinglei 的贡献)5月18日 添加了label:stat/needs-squash
5月18日 添加了label:mindspore-cla/yes
司小南(机器人)
5月18日 评论:
5月18日 评论:
@xuxinglei, 当前/check-pr未通过,原因如下:
以下Pull Request描述检查项未通过:
存在不符合模板的选项: pytest tests/torch/accuracy/test_qwen3_5_accuracy.py::test_qwen3_5_single_card_baseline
存在不符合模板的选项: pytest tests/torch/accuracy/test_qwen3_5_accuracy.py::test_qwen3_5_tp_fully_shard_accuracy
存在不符合模板的选项: pytest tests/torch/accuracy/test_qwen3_5_accuracy.py::test_qwen3_5_tp_cp_fully_shard_accuracy
部分检查项缺失 请重新使用模板
模板中'Test Plan and Test Result' 信息为空,请补充对应信息。
以下issue检查项未通过:
Pull Request未关联issue
请修改好上述检查错误后,重新使用/check-pr触发检查。


5月18日 添加了label:no-pass-all-review
此处折叠了2685条消息 查看更多
7月10日 通过审查
7月10日 通过审查
7月10日 删除了label:no-pass-all-review
7月10日 合入了pull request,合并节点 SHA:fbed2bc1c47a96409962e5141ea14282a0f330e6
What type of PR is this?
/kind feature
What does this PR do / why do we need it:
本 PR 为 Qwen3.5 dense、Qwen3.5-MoE 和 Qwen3-VL-MoE 补齐统一 Trainer 的真实训练链路:模型构造与 checkpoint 转换、数据 registry、meta 参数物化、优化器、以及 TP/CP/EP/FSDP/HSDP/PP/VPP 组合并行。
models/:三个模型族的模型定义、state-dict 转换和模型自有并行方案。trainer/:配置、DeviceMesh、数据加载、PP stage、loss/gradient 汇总和 optimizer step。data/:dummy、VL dummy、HF/JSON、preset tensor 和 Megatron 数据入口。core/:组合并行实际触发的 PP 共享梯度、嵌套 FSDP、Partial Add/Sub 和 output reduction dtype 修复。公共 core 修改的逐文件根因见 Issue #282。
Which issue(s) this PR fixes:
Fixes #282
Test Plan and Test result:What scenarios were tested, and what were the verification results(Function, performance, reliability, etc.):
1. 验证口径
dummy,seq=64dummy,seq=64vl_dummy,grid=2x2x2共同口径:20 个连续 optimizer steps;
param_dtype=bfloat16、reduce_dtype=float32、max_grad_norm=0、global batch=4、micro batch=1、seed=1234。每个候选必须完整输出 steps 1..20 且 loss 全部有限;通过条件为相对同模型单卡基线的最大 loss 绝对误差<= 0.005。2. 支持的数据集
data.typedummytrain.seed + sample index确定性生成 tokenmax_seq_len、train_sizevl_dummymax_seq_len、vl_grid_t/h/w、vl_videohf_datasetstrain_path、subset、text_key、train_sizejson_filetrain_path、template、text_keypreset_pttorch.save(List[Dict[str, Tensor]])train_pathmegatron.bin/.idxprefix 或 weighted blendtrain_path、megatron_seed、pad/eod_token_id.binmmap;支持单 prefix、权重字符串和 pair list当前仅支持
streaming: false;真实数据可配置num_workers、prefetch_factor、pin_memory和shuffle。3. Step-by-step 拉起方式
Step 1:准备环境
export HYPER_PARALLEL_PLATFORM=torch export HCCL_DETERMINISTIC=true export LCCL_DETERMINISTIC=1 export ASCEND_DETERMINISTIC=true export FLASH_ATTENTION_DETERMINISTIC=1 export PYTORCH_NPU_ALLOC_CONF=expandable_segments:TrueStep 2:创建单卡 dense YAML
YAML 可放在仓库外;下面的 checkpoint 路径替换为实际路径。
model: name: qwen3_5 weights_path: /path/to/Qwen3.5-0.8B-Base tokenizer_path: /path/to/Qwen3.5-0.8B-Base config_overrides: num_hidden_layers: 4 data: type: dummy max_seq_len: 64 train: max_steps: 20 num_train_epochs: 1 global_batch_size: 4 micro_batch_size: 1 seed: 1234 backend: torch init_device: meta accelerator: dp_shard: 1 comm_fusion: true optimizer: type: adamw lr: 1.0e-4 lr_min: 1.0e-4 lr_decay_style: constant lr_warmup_ratio: 0.0 max_grad_norm: 0.0 weight_decay: 0.0 loss_aggregation: token_weighted foreach: false mixed_precision: enabled: true param_dtype: bfloat16 reduce_dtype: float32 output_dtype: float32 gradient_checkpointing: activation_checkpoint: none checkpoint: output_dir: outputs/qwen3_5_single save_steps: 0 save_hf_weights: false logging: log_steps: 1 report_throughput: false debug: deterministic: trueStep 3:切换模型
Qwen3.5-MoE 使用
scripts/train_lm.py,把 YAML 中的模型和优化器改为:model: name: qwen3_5_moe weights_path: /path/to/Qwen3.5-35B-A3B tokenizer_path: /path/to/Qwen3.5-35B-A3B config_overrides: num_hidden_layers: 1 train: optimizer: type: adamw lr: 5.0e-6 lr_min: 0.0 lr_decay_style: cosine lr_warmup_ratio: 0.1 max_grad_norm: 0.0 weight_decay: 0.0 loss_aggregation: token_weighted foreach: falseQwen3-VL-MoE 使用
scripts/train_vl.py,替换模型和数据段:model: name: qwen3_vl_moe weights_path: /path/to/Qwen3-VL-30B-A3B-Instruct tokenizer_path: /path/to/Qwen3-VL-30B-A3B-Instruct freeze_modules: - model.visual config_overrides: vl: true text_config: num_hidden_layers: 1 data: type: vl_dummy max_seq_len: 64 vl_grid_t: 2 vl_grid_h: 2 vl_grid_w: 2Step 4:运行单卡 baseline
Step 5:只修改
train.accelerator生成并行 YAMLdense/MoE 表格结果使用
comm_fusion: true,VL-MoE 使用comm_fusion: false;复现时应保持对应模型口径。EP 样例另保留moe_token_dispatcher_type: all_to_all和npu_nums_per_device: 1。train.accelerator修改dp_shard: 1dp_replicate: 2, dp_shard: 1dp_replicate: 2, dp_shard: 2dp_shard: 2tp: 2cp: 2ep: 2, etp: 1tp: 2, dp_shard: 2cp: 2, dp_shard: 2ep: 2, etp: 1, dp_shard: 2tp: 2, cp: 2tp: 2, ep: 2, etp: 1cp: 2, ep: 2, etp: 1pp: 2, dp_shard: 2, pp_micro_batch_num: 2, pp_schedule: 1f1boutput_dtype: float32)pp: 2, pp_vpp: 2, dp_shard: 2, pp_micro_batch_num: 2tp: 2, cp: 2, dp_shard: 2pp: 2, tp: 2, dp_shard: 2, pp_micro_batch_num: 2, pp_schedule: 1f1bStep 6:按卡数运行并行 YAML
其他组合只需同步修改
ASCEND_RT_VISIBLE_DEVICES、--nproc-per-node和 YAML 的train.accelerator;VL 模型将入口换成scripts/train_vl.py。4. 20-step 单卡与并行结果
Qwen3.5 dense
ddp2hsdp2x2fsdp2tp2cp2tp2_fsdp2cp2_fsdp2tp2_cp2pp2_fsdp2vpp2_fsdp2tp2_cp2_fsdp2pp2_tp2_fsdp2Qwen3.5-MoE
ddp2hsdp2x2fsdp2tp2cp2ep2tp2_fsdp2cp2_fsdp2ep2_fsdp2tp2_ep2cp2_ep2Qwen3-VL-MoE
ddp2hsdp2x2fsdp2tp2cp2ep2