| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 年前 | ||
| 1 年前 | ||
[auto settings] fix auto settings feature Co-authored-by: mhlinoer<lvmuheng@h-partners.com> # message auto-generated for no-merge-commit merge: !2926 merge master into master [auto settings] fix auto settings feature Created-by: mhlinoer Commit-by: mhlinoer Merged-by: ascend-robot Description: [auto settings] fix auto settings feature See merge request: Ascend/MindSpeed!2926 | 10 个月前 | |
perf(verl ckpt): ckpt load and save acceleration Co-authored-by: 李鸣沼<lmztju@126.com> # message auto-generated for no-merge-commit merge: !3074 merge verl_load_and_save_ckpt into master perf(verl ckpt): ckpt load and save acceleration Created-by: lmztju Commit-by: 李鸣沼;l30057177 Merged-by: ascend-robot Description: **测试场景1:** verl+megatron后端dapo-qwen3-30b 910A3双机加载和保存ckpt 本地存储  See merge request: Ascend/MindSpeed!3074 | 8 个月前 | |
fix: fix the initialization error of compress-optimizer Co-authored-by: NingGuangyou<ningguangyou@h-partners.com> # message auto-generated for no-merge-commit merge: !3319 merge master into master fix: fix the initialization error of compress-optimizer Created-by: NingGuangyou Commit-by: NingGuangyou Merged-by: ascend-robot Description: fix the initialization error of compress-optimizer See merge request: Ascend/MindSpeed!3319 | 6 个月前 | |
| 1 年前 | ||
fix: te ulysses Co-authored-by: clc2025<chenlucong@huawei.com> # message auto-generated for no-merge-commit merge: !3349 merge fix_te_ulysses into master fix: te ulysses Created-by: clc2025 Commit-by: clc2025 Merged-by: ascend-robot Description: What this PR does / why we need it? Please describe the background and detailed changes of the PR. If it is a bugfix, please attach the related issue. Does this PR introduce any user-facing change? Please describe whether the PR will result in any user-facing usage changes. If there is related documentation, please specify its path. How was this patch tested? Please explain how to verify the correctness and effectiveness of this feature, as well as its usage constraints and limitations. See merge request: Ascend/MindSpeed!3349 | 5 个月前 | |
| 1 年前 | ||
| 1 年前 | ||
Bugfix: 在disable-gloo-groups不开启的情况下不去设置enable_gloo_process_groups Co-authored-by: fishhhqi<moeyfishyq@outlook.com> # message auto-generated for no-merge-commit merge: merge gloo_fix2 into master Bugfix: 在disable-gloo-groups不开启的情况下不去设置enable_gloo_process_groups Created-by: fishhhqi Commit-by: fishhhqi Merged-by: ascend-robot Description: 在disable-gloo-groups不开启的情况下不去设置enable_gloo_process_groups See merge request: Ascend/MindSpeed!2887 | 11 个月前 | |
| 1 年前 | ||
[Bugfix] Fix Megatron checkpoint saving&loading compatibility for torch_dcp format Co-authored-by: 林明哲<linmingzhe3@huawei.com> # message auto-generated for no-merge-commit merge: !3077 merge fix1202 into master [Bugfix] Fix Megatron checkpoint saving&loading compatibility for torch_dcp format Created-by: LinMingZhe Commit-by: 林明哲 Merged-by: ascend-robot Description: Fix Megatron checkpoint saving&loading compatibility for torch_dcp format See merge request: Ascend/MindSpeed!3077 | 9 个月前 | |
fix: NPU datadump level: L0 & mix Co-authored-by: yulelanmei<huangyijie8@huawei.com> # message auto-generated for no-merge-commit merge: !3351 merge master into master fix: NPU datadump level: L0 & mix Created-by: yulelanmei Commit-by: yulelanmei Merged-by: ascend-robot Description: What this PR does / why we need it? 当前--npu-datadump未适配 L0及mix 的dump等级,需要增强功能 Does this PR introduce any user-facing change? N/A How was this patch tested? 开启--npu-datadump,config.json配置level为L0或mix 测试:https://wiki.huawei.com/domains/148330/wiki/296621/WIKI2026032510543405 See merge request: Ascend/MindSpeed!3351 | 5 个月前 | |
feat(triton):sort_chunks_by_idx Co-authored-by: guofanfeng<guofanfeng1@huawei.com> # message auto-generated for no-merge-commit merge: !2997 merge master into master feat(triton):sort_chunks_by_idx Created-by: guofanfeng23 Commit-by: guofanfeng Merged-by: ascend-robot Description: sort_chunks_by_idx triton算子接入 算子验证结果: https://wiki.huawei.com/domains/152732/wiki/307991/WIKI202511219117266 See merge request: Ascend/MindSpeed!2997 | 9 个月前 | |
feat: hccl op mode set Co-authored-by: Jia_Austin<dengjia6@huawei.com> # message auto-generated for no-merge-commit merge: !3376 merge core_r0.12.1_adaptive_hccl_op_v2 into master feat: hccl op mode set Created-by: Jia_Austin Commit-by: Jia_Austin Merged-by: ascend-robot Description: What this PR does / why we need it? feat(torch): hccl op mode set Does this PR introduce any user-facing change? --hccl-op-mode How was this patch tested? Please explain how to verify the correctness and effectiveness of this feature, as well as its usage constraints and limitations. See merge request: Ascend/MindSpeed!3376 | 5 个月前 | |
| 1 年前 | ||
fix: TE LayerNormLinear init weight order align NVTE Co-authored-by: clc2025<chenlucong@huawei.com> # message auto-generated for no-merge-commit merge: !3569 merge fix_lnliner_initweight into master fix: TE LayerNormLinear init weight order align NVTE Created-by: clc2025 Commit-by: clc2025 Merged-by: ascend-robot Description: ## What this PR does / why we need it? 具体查看关联issue ## Does this PR introduce any user-facing change? NA ## How was this patch tested? 基于脚本用例,从GPU上保存权重,NPU加载权重后断点续训,精度能对齐且无功能报错 See merge request: Ascend/MindSpeed!3569 | 2 个月前 | |
Add swap layer input Co-authored-by: JialiZheng<jializheng@huawei.com> # message auto-generated for no-merge-commit merge: !3529 merge master into master Add swap layer input Created-by: JialiZheng1 Commit-by: JialiZheng Merged-by: ascend-robot Description: Add swap input RFC:https://gitcode.com/Ascend/MindSpeed/issues/176 验证wiki:https://wiki.huawei.com/wiki/WIKI2026060411356524 See merge request: Ascend/MindSpeed!3529 | 2 个月前 | |
fix: fix arg of moe_expert_capacity_factor Co-authored-by: yulelanmei<huangyijie8@huawei.com> # message auto-generated for no-merge-commit merge: !3670 merge master into master fix: fix arg of moe_expert_capacity_factor Created-by: yulelanmei Commit-by: yulelanmei Merged-by: ascend-robot Description: ## What this PR does / why we need it? The usage of args.moe_expert_capacity_factor has a problem when running testcase of Verl. ## Does this PR introduce any user-facing change? N/A ## How was this patch tested? Run Verl testcase in MindSpeed/tests_extend/system_tests. See merge request: Ascend/MindSpeed!3670 | 2 个月前 | |
fix: add open component Copyright Co-authored-by: GuoHaifeng1999<guohaifeng12@huawei.com> # message auto-generated for no-merge-commit merge: !3611 merge master into master fix: add open component Copyright Created-by: GuoHaifeng1999 Commit-by: GuoHaifeng1999 Merged-by: ascend-robot Description: ## What this PR does / why we need it? add open component Copyright ## Does this PR introduce any user-facing change? nothing ## How was this patch tested? don't needs See merge request: Ascend/MindSpeed!3611 | 2 个月前 | |
feat: add custom pp layout Co-authored-by: wuweiqiang24<wuweiqiang11@huawei.com> # message auto-generated for no-merge-commit merge: !3496 merge add_pp_layout into master feat: add custom pp layout Created-by: wuweiqiang24 Commit-by: wuweiqiang24 Merged-by: ascend-robot Description: 新增pipeline-model-parallel-layout功能,支持自定义PP每个stage的层排布 验证链接:https://wiki.huawei.com/domains/137239/wiki/268925/WIKI2026052611233549 issue: https://gitcode.com/Ascend/MindSpeed/issues/166 See merge request: Ascend/MindSpeed!3496 | 3 个月前 | |
feat: support W4A8-QAT Co-authored-by: xusiyang<xusiyang2@huawei.com> # message auto-generated for no-merge-commit merge: !3664 merge master into master feat: support W4A8-QAT Created-by: weixin_44492126 Commit-by: xusiyang Merged-by: ascend-robot Description: ## What this PR does / why we need it? 本PR 新增了对 W4A8-QAT(权重MXFP4/激活 MXFP8 量化感知训练,仅限 MoE 场景),并量化粒度分别支持:32(标准MXFP4量化粒度)/128(deepseek-V4采用量化粒度) ## Does this PR introduce any user-facing change? Yes. 用户可以通过在训练命令中添加以下参数来启用W4A8功能,默认量化粒度32: --qat-scheme w4a8-moe-only --qat-quant-block-size 128 ## How was this patch tested? 添加 --qat-scheme w4a8-moe-only 参数启动量化粒度为32的W4A8-QAT训练,添加--qat-quant-block-size 128启动量化粒度为128的W4A8-QAT训练 实现方案和实验结果如下:https://wiki.huawei.com/domains/171785/wiki/358154/WIKI2026071511870516 昇腾 950DT上测试通过:基于减层DeepSeek-V4,分别训练100step对比BF16 loss误差:量化粒度为32时loss误差在0.4%,量化粒度为128时loss误差在0.79% See merge request: Ascend/MindSpeed!3664 | 2 个月前 | |
feat(ut/qos/torch): 补充ut,修复代码遗漏BUG Co-authored-by: Klayyy<wanglei886@h-partners.com> # message auto-generated for no-merge-commit merge: !3309 merge master into master feat(ut/qos/torch): 补充ut,修复代码遗漏BUG Created-by: Klayyy Commit-by: Klayyy Merged-by: ascend-robot Description: 1.补充AI QOS特性feature UT 2.ut补充过程中,自检代码,修复BUG 2.1 torch_npu._C._distributed_c10d.ProcessGroupHCCL.Options()调用名称修改 2.2 qos_feature.py 中 raiseValueError 提示词完善 2.3 qos.py中对于最小冲突度组合中优先级的赋值部分,去掉重复代码,去掉无用库导入,_PARALLEL_TYPES中有逗号未添加 2.4 qos.py中 应是sdma qos 部分的处理,误使用roce 3.补充H2D QOS 对于 PCIE异步通道的使用,对于DCMI接口新建set_h2d_qos接口,提供给python调用 4.修改aiQos Readme中关于DCMI接口的调用,补充DCMI接口SO编译方法 See merge request: Ascend/MindSpeed!3309 | 6 个月前 | |
| 1 年前 | ||
fix: coc feature verification update Co-authored-by: yulelanmei<huangyijie8@huawei.com> # message auto-generated for no-merge-commit merge: !3545 merge master into master fix: coc feature verification update Created-by: yulelanmei Commit-by: yulelanmei Merged-by: ascend-robot Description: ## What this PR does / why we need it? Current coc feature verification isn't correct, what is not supported on A5 is that coc fused kernel instead of coc feature. ## Does this PR introduce any user-facing change? Verificed massege update, coc fused kernel is not supported on A5. ## How was this patch tested? Run MindSpeed/tests_extend/system_tests/feature_tests/coc.sh, turn on coc fused kernel on A5. See merge request: Ascend/MindSpeed!3545 | 3 个月前 | |
| 1 年前 | ||
Feat: adaptor for DeepSeek V4 Co-authored-by: wuweiqiang24<wuweiqiang11@huawei.com> # message auto-generated for no-merge-commit merge: !3427 merge master into master Feat: adaptor for DeepSeek V4 Created-by: wuweiqiang24 Commit-by: wuweiqiang24 Merged-by: ascend-robot Description: What this PR does / why we need it? Adaptor for DeepSeek V4!!! Does this PR introduce any user-facing change? Please describe whether the PR will result in any user-facing usage changes. If there is related documentation, please specify its path. How was this patch tested? Please explain how to verify the correctness and effectiveness of this feature, as well as its usage constraints and limitations. See merge request: Ascend/MindSpeed!3427 | 4 个月前 | |
feat: add last word feature Co-authored-by: Zhaiwenxuan<zhaiwenxuan4@h-partners.com> # message auto-generated for no-merge-commit merge: !3542 merge master into master feat: add last word feature Created-by: Zhaiwenxuan Commit-by: Zhaiwenxuan Merged-by: ascend-robot Description: ## What this PR does / why we need it? 新增适配verl训练场景的临终遗言功能 ## Does this PR introduce any user-facing change? YES ++actor_rollout_ref.actor.megatron.override_transformer_config.ttp_enabled=True 用户可通过增加++actor_rollout_ref.actor.megatron.override_transformer_config.ttp_enabled=True,开启适配verl训练场景的临终遗言功能 ## How was this patch tested? https://wiki.huawei.com/domains/181442/wiki/383165/WIKI2026060911405364 See merge request: Ascend/MindSpeed!3542 | 3 个月前 | |
Add swap layer input Co-authored-by: JialiZheng<jializheng@huawei.com> # message auto-generated for no-merge-commit merge: !3529 merge master into master Add swap layer input Created-by: JialiZheng1 Commit-by: JialiZheng Merged-by: ascend-robot Description: Add swap input RFC:https://gitcode.com/Ascend/MindSpeed/issues/176 验证wiki:https://wiki.huawei.com/wiki/WIKI2026060411356524 See merge request: Ascend/MindSpeed!3529 | 2 个月前 | |
| 1 年前 | ||
| 1 年前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 年前 | ||
| 1 年前 | ||
| 10 个月前 | ||
| 8 个月前 | ||
| 6 个月前 | ||
| 1 年前 | ||
| 5 个月前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 11 个月前 | ||
| 1 年前 | ||
| 9 个月前 | ||
| 5 个月前 | ||
| 9 个月前 | ||
| 5 个月前 | ||
| 1 年前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 2 个月前 | ||
| 3 个月前 | ||
| 2 个月前 | ||
| 6 个月前 | ||
| 1 年前 | ||
| 3 个月前 | ||
| 1 年前 | ||
| 4 个月前 | ||
| 3 个月前 | ||
| 2 个月前 | ||
| 1 年前 | ||
| 1 年前 |