Pull Request已成功合入, 合并人@ascend-robot
(感谢 shi-yufeng99 的贡献)变更摘要
此PR主要针对 torch_npu 的 Inductor 代码生成模块进行适配,使其在 cpp_wrapper 模式下正确支持 CATLASS 与 PyTorch 2.13.0 版本。改动涉及两个文件:一是在 cpp_wrapper_npu.py 中新增 codegen_initialized_kernel_decls 方法,区分 AOT 与 JIT 模式下的 kernel 声明策略,避免 JIT 模式下 extern "C" 函数声明与函数指针变量命名冲突;二是在 npu_combined_scheduling.py 中移除了 benchmark_fused_nodes 中对 node_can_linear() 的预先条件判断,使线性融合 benchmark 的尝试范围更广,失败时统一回退到原始调度逻辑。
主要改动
- 新增
codegen_initialized_kernel_decls方法:在CppWrapperNpu类中重写该方法,JIT 模式下跳过extern "C"声明以规避 C++ 中函数声明与static指针变量同名的编译冲突;AOT 模式下仍调用父类方法生成声明。 benchmark_fused_nodes条件放宽:移除self.node_can_linear()前置检查,始终尝试通过_triton_scheduling.benchmark_fused_nodes进行线性融合 benchmark,失败则回退至原始分组调度逻辑。


代码审查
经过对两个变更文件的逐行审查、上下游调用链分析以及 blast radius 检查,本次 diff 未发现任何需要报告的问题。
审查结论
变更文件 1:torch_npu/_inductor/codegen/cpp_wrapper_npu.py
- 新增方法
codegen_initialized_kernel_decls(第 476–489 行):在 JIT 模式下跳过父类的extern "C"声明(因为 JIT 模式下 CATLASS kernel 通过dlopen/dlsym动态加载,extern "C"函数声明会与static函数指针变量产生 C++ 标识符冲突),在 AOT 模式下委托父类正常生成声明。逻辑正确,注释详尽,与finalize_prefix中已有的 JIT/AOT 分支处理一致。 - 无新增问题。
变更文件 2:torch_npu/_inductor/codegen/npu_combined_scheduling.py
benchmark_fused_nodes方法(第 93–107 行):移除if self.node_can_linear():守卫。node_can_linear在 torch 2.13.0 上游已被移除,旧代码会触发AttributeError。新代码改为始终尝试 linear 路径、异常时回退到 nolinear,与同文件中的codegen_node方法(第 77–91 行)模式完全一致。- 无新增问题。
总结
- P0:0,P1:0,P2:0,P3:0
- 整体风险评估:低风险。两处变更均为适配 torch 2.13.0 上游 API 变更的必要修改,逻辑正确,与现有代码风格一致。
⚠️ 已识别出整体风险,但无法提取行内评论,请参考整体评估。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here。
You can get sig-info at here
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| torch_npu/_inductor | ✅ weizhan4, TonyYA (2/2) | ✅ weizhan4 (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
shi-yufeng99, thanks for your pull request. All authors of the commits have signed the CLA. 👍


Linking Issue Notice
@shi-yufeng99 , the pull request must be linked to at least one issue.
If an issue has already been linked, but the needs-issue label remains, you can remove the label by commenting /check-issue .


当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.11.0 | ||
| v2.12.0-26.1.0 | ||
| v2.10.0-26.1.0 | ||
| v2.10.0 | ||
| v2.9.0-26.1.0 | ||
| v2.7.1-26.1.0 | ||
| v2.9.0 | ||
| v2.7.1 | ||
| v2.11.0-26.1.0 | ||
| v2.12.0 | ||
| ci-test |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


compile


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | ✅ | >>> |
| Build_ARM | ✅ | >>> | |
| Build_LibTorch_x86 | ✅ | >>> | |
| Build_LibTorch_ARM | ✅ | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | CodeCheck | ✅ | >>> |
| check_error | ✅ | >>> | |
| CodeCheck_lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🛑 | >>> |
| UT_ARM_A3_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_01 | 🟨 | >>> | |
| UT_ARM_A2_Part_02 | 🟨 | >>> | |
| UT_ARM_A2_Part_03 | 🟨 | >>> | |
| UT_inductor_Part_01 | 🛑 | >>> | |
| UT_inductor_Part_02 | 🛑 | >>> | |
| UT_inductor_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_04 | 🛑 | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | 🟨 | >>> | |
| UT_ARM_A2_Select_Part_02 | 🟨 | >>> | |
| 流水线 | PR-pipeline_pytorch | 🟨 | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | ✅ | >>> |
| Build_ARM | ✅ | >>> | |
| Build_LibTorch_x86 | ✅ | >>> | |
| Build_LibTorch_ARM | ✅ | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | CodeCheck | ✅ | >>> |
| check_error | ✅ | >>> | |
| CodeCheck_lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🛑 | >>> |
| UT_ARM_A3_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_03 | ✅ | >>> | |
| UT_inductor_Part_01 | 🛑 | >>> | |
| UT_inductor_Part_02 | 🛑 | >>> | |
| UT_inductor_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_04 | 🛑 | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Select_Part_02 | ✅ | >>> | |
| 流水线 | PR-pipeline_pytorch | ✅ | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


/lgtm




【合入来源】
【修改方案】
overwrite codegen_initialized_kernel_decls function, to adapt 2.13.0 torch for catlass cppwrapper, the overwrite reason is stated below:
In AOT mode, CATLASS .o files are linked into model.so, so the parent's
extern "C"declarations are required — same as the community CUTLASS flow.In JIT cpp_wrapper mode, CATLASS kernels are loaded dynamically via dlopen/dlsym through function pointers emitted in finalize_prefix(). An
extern "C"function declaration here would conflict with thestatic <name>_t <name> = nullptr;pointer variable (C++ forbids a function and a variable sharing the same identifier), so we skip it. In JIT mode initialized_kernels only contains CATLASS kernels on NPU.remove node_can_linear check in benchmark_fused_nodes, which is not contained in npu_combined_scheduling in 2.13.0
#2589
【资料变更】
【接口变更】
【功能验证】
【CheckList】