已合并
ShardedGradScale achieves alignment with GPU #36150
我应该是一阵风创建于 5月20日
ShardedGradScale achieves alignment with GPU #36150
已合并
Pull Request已成功合入, 合并人@ascend-robot
(感谢 我应该是一阵风 的贡献)ascend-robot
5月20日 评论:
5月20日 评论:
ascend-robot
5月20日 评论:
5月20日 评论:
Thanks for your pull-request.
The full list of commands accepted by me can be found at here。
You can get sig-info at here
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| torch_npu/npu | ✅ chengpeng25, wangmin0104 (2/2) | ✅ wangmin0104 (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
Windwindzzz, thanks for your pull request. All authors of the commits have signed the CLA. 👍


5月20日 添加了label:ascend-cla/yes
ascend-robot
5月20日 评论:
5月20日 评论:
当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.10.0 | ||
| v2.11.0 | ||
| v2.7.1 | ||
| v2.9.0 | ||
| v2.12.0 |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


此处折叠了53条消息 查看更多
5月22日 添加了label:approved
chengpeng25
5月23日 评论:
5月23日 评论:
/lgtm


5月23日 添加了label:lgtm
ascend-robot
5月23日 评论:
5月23日 评论:
Review Guide
This pull-request passes review.
Committers who wrote a comment of /approve are: wangmin0104.
Reviewers who wrote a comment of /lgtm are: wangmin0104, chengpeng25.


5月23日 合入了pull request
【合入来源】
【修改方案】
原现象:
CPUOffload 场景下,
found_inf_per_device中存在 CPU 上的found_inf。修复前 torch_npu 使用:found_inf_npu = found_inf.to(self._scale.device)但
self._scale.device在 CPUOffload 路径也是 CPU,导致后续对 CPU tensor 执行 HCCLall_reduce,触发:修改方案:
_ShardedGradScaler增加原生torch同构的目标设备语义,默认目标设备为 NPU,将found_inf显式移动到npufound_inf_on_device = found_inf.to(self._device)
【资料变更】
不涉及
【接口变更】
不涉及
【功能验证】
【CheckList】