Pull Request已成功合入, 合并人@ascend-robot
(感谢 luqichao 的贡献)变更摘要
本 PR 旨在优化训练算子的性能:一方面将 aten.embedding_dense_backward.default 加入 NPU native fallback,复用 NPU 原生算子,避免 Inductor 生成低效 kernel;另一方面在 Triton backend 中通过 _matmul_backward_inductor 独立维护完整的 matmul_backward decomposition,覆盖不同输入维度与 batch 场景,降低对 MFusion backend 逻辑的依赖。
主要改动
- 新增
_matmul_backward_inductor分解实现:在torch_npu/_inductor/decomposition.py中新增该函数,按输入张量维度组合(1d-1d、2d-1d、1d-2d、nd 与低维、低维与 nd 及默认torch.matmul分支)分别计算grad_self与grad_other,并通过mask控制是否生成对应梯度。 - 注册
matmul_backward分解:通过register_decomposition([aten.matmul_backward.default])(_matmul_backward_inductor)在 Triton decomposition 注册流程中挂载该实现,使matmul_backward的分解在 Triton backend 内独立完成。 - 扩展 Triton decomposition 列表:在
_register_triton_decompositions()中新增aten.slice_backward、aten.embedding_dense_backward、aten.matmul_backward.default三项分解注册。 - 新增 NPU native fallback 条目:在
torch_npu/_inductor/lowering_fallback_list.py的TORCH_NATIVE_FALLBACK_LIST中新增aten.embedding_dense_backward.default,使该反向算子直接复用 NPU 原生算子实现,避免低效 kernel 生成。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| torch_npu/_inductor | ✅ rain-666, HinPeng (2/2) | ✅ HinPeng (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
luqichao, thanks for your pull request. All authors of the commits have signed the CLA. 👍


当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.12.0 | ||
| v2.11.0 | ||
| v2.7.1 | ||
| v2.10.0 | ||
| v2.9.0 | ||
| v2.7.1-26.1.0 | ||
| v2.9.0-26.1.0 | ||
| v2.10.0-26.1.0 | ||
| v2.11.0-26.1.0 | ||
| v2.12.0-26.1.0 | ||
| ci-test |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | ✅ | >>> |
| Build_ARM | ✅ | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| Build_X86_213 | ✅ | >>> | |
| Build_ARM_213 | ✅ | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | codecheck_pre-commit | ✅ | >>> |
| check_error | ✅ | >>> | |
| lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🛑 | >>> |
| UT_ARM_A3_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_03 | ✅ | >>> | |
| UT_inductor_Part_01 | 🛑 | >>> | |
| UT_inductor_Part_02 | 🛑 | >>> | |
| UT_inductor_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_04 | 🛑 | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Select_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_213 | ✅ | >>> | |
| UT_inductor_Part_213 | 🛑 | >>> | |
| UT_DIST_ARM_Part_213 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_213 | ✅ | >>> | |
| 流水线 | PR-pipeline_pytorch | ✅ | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


/lgtm




aten.embedding_dense_backward.default加入 NPU native fallback,复用 NPU 原生算子,避免 Inductor 生成低效 kernel。matmul_backwarddecomposition,覆盖不同输入维度和 batch 场景,降低对 MFusion backend 逻辑的依赖。