Pull Request已成功合入, 合并人@ascend-robot
(感谢 shi-yufeng99 的贡献)变更摘要
此 PR 适配 PyTorch 社区 v2.13 中两处破坏性变更对 catlass 的影响:一是 SizeVarAllocator 中将 size_hint 重命名为 optimization_hint,将相关调用点统一更新;二是 benchmarker 从 TritonBenchmarker 重构为 InductorBenchmarker,其内部 GPU 判断逻辑无法识别 NPU 设备,通过在 autotune_process.py 中动态替换 ExternKernelGPUBenchmarkRequest 的基类,将 GPUDeviceBenchmarkMixin 替换为 NPUDeviceBenchmarkMixin,确保外部 kernel 基准测试正确使用 torch.npu 而非 torch.cuda。
主要改动
size_hint→optimization_hint重命名适配:在gemm_template.py的CATLASS1xGemmTemplate中,将V.graph.sizevars.size_hint调用改为V.graph.sizevars.optimization_hint,匹配社区 v2.13 的SizeVarAllocator接口变更。size_hint→optimization_hint重命名适配:在kernel_analysis.py的_finalize_stride_collection函数中,同样将V.graph.sizevars.size_hint(stride)改为V.graph.sizevars.optimization_hint(stride)。ExternKernelGPUBenchmarkRequest基类替换:在autotune_process.py的patch_tuning_process中新增代码,将ExternKernelGPUBenchmarkRequest的基类从GPUDeviceBenchmarkMixin动态替换为NPUDeviceBenchmarkMixin,解决InductorBenchmarker无法识别 NPU 设备导致回退到torch.cuda.synchronize()报错的问题。


代码审查
经过彻底审查,本次 PR 中的三个文件变更均为合理且正确的适配修改:
-
torch_npu/_inductor/autotune_process.py:通过__bases__赋值将ExternKernelGPUBenchmarkRequest的基类从GPUDeviceBenchmarkMixin替换为NPUDeviceBenchmarkMixin,确保外部 kernel benchmark(aten::mm,aten::addmm等)使用torch.npu而非torch.cuda。NPUDeviceBenchmarkMixin已存在于文件中,此处仅添加了基类替换逻辑,且该函数在模块初始化时(单线程环境)调用,时机正确。 -
torch_npu/_inductor/codegen/catlass/gemm_template.py:size_hint→optimization_hint重命名,与上游 v2.13 的重构保持一致。代码库中已有 72 处使用optimization_hint,此变更补齐了遗漏的调用点。 -
torch_npu/_inductor/codegen/kernel_analysis.py:同上,size_hint→optimization_hint重命名,补齐最后遗漏的调用点。异常捕获TypeError与原size_hint行为一致。
审查结论
- P0: 0
- P1: 0
- P2: 0
- P3: 0
整体风险评估:低风险。三个变更均为适配 PyTorch v2.13 的必要修改,逻辑正确,与代码库其余部分保持一致。
各文件审查结果:
torch_npu/_inductor/autotune_process.py— 无问题torch_npu/_inductor/codegen/catlass/gemm_template.py— 无问题torch_npu/_inductor/codegen/kernel_analysis.py— 无问题
⚠️ 已识别出整体风险,但无法提取行内评论,请参考整体评估。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here。
You can get sig-info at here
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| torch_npu/_inductor | ✅ weizhan4, TonyYA (2/2) | ✅ weizhan4 (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
shi-yufeng99, thanks for your pull request. All authors of the commits have signed the CLA. 👍


compile


当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.10.0 | ||
| v2.12.0 | ||
| v2.11.0 | ||
| v2.9.0 | ||
| v2.7.1 | ||
| v2.7.1-26.1.0 | ||
| v2.12.0-26.1.0 | ||
| v2.11.0-26.1.0 | ||
| v2.9.0-26.1.0 | ||
| v2.10.0-26.1.0 | ||
| ci-test |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | 🕚 | >>> |
| Build_ARM | 🕚 | >>> | |
| Build_LibTorch_x86 | 🕚 | >>> | |
| Build_LibTorch_ARM | 🕚 | >>> | |
| Build_X86_torchair | 🕚 | >>> | |
| Build_ARM_torchair | 🕚 | >>> | |
| patch_test | 🕚 | >>> | |
| 恶意代码检查 | Antipoison | 🟨 | >>> |
| 编码安全与规范检查 | CodeCheck | 🟨 | >>> |
| check_error | 🟨 | >>> | |
| CodeCheck_lintrunner | 🟨 | >>> | |
| 开源片段检查 | SCA | 🟨 | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🕚 | >>> |
| UT_ARM_A3_Part_02 | 🕚 | >>> | |
| UT_ARM_A2_Part_01 | 🕚 | >>> | |
| UT_ARM_A2_Part_02 | 🕚 | >>> | |
| UT_ARM_A2_Part_03 | 🕚 | >>> | |
| UT_inductor_Part_01 | 🕚 | >>> | |
| UT_inductor_Part_02 | 🕚 | >>> | |
| UT_inductor_Part_03 | 🕚 | >>> | |
| UT_inductor_Part_04 | 🕚 | >>> | |
| UT_DIST_ARM_Part_01 | 🕚 | >>> | |
| UT_DIST_ARM_Part_02 | 🕚 | >>> | |
| UT_DIST_ARM_Part_03 | 🕚 | >>> | |
| UT_DIST_ARM_Part_04 | 🕚 | >>> | |
| UT_ARM_A2_Select_Part_01 | 🕚 | >>> | |
| UT_ARM_A2_Select_Part_02 | 🕚 | >>> | |
| 流水线 | PR-pipeline_pytorch | 🟨 | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


ascend docs pipeline is running...


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


✅ 跳过 docs ci 检查,没有需要检查的文档文件


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | ✅ | >>> |
| Build_ARM | ✅ | >>> | |
| Build_LibTorch_x86 | ✅ | >>> | |
| Build_LibTorch_ARM | ✅ | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | CodeCheck | ✅ | >>> |
| check_error | ✅ | >>> | |
| CodeCheck_lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🛑 | >>> |
| UT_ARM_A3_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_03 | ✅ | >>> | |
| UT_inductor_Part_01 | 🛑 | >>> | |
| UT_inductor_Part_02 | 🛑 | >>> | |
| UT_inductor_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_04 | 🛑 | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Select_Part_02 | ✅ | >>> | |
| 流水线 | PR-pipeline_pytorch | ✅ | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


/lgtm




【合入来源】
【修改方案】
在2.13.0中适配catlass,如下两个点需要适配:
1.
v2.13 社区重构了 SizeVarAllocator,将 size_hint 重命名为 optimization_hint。本次PR将涉及到的size_hint代码都改为了optimization_hint
2.
v2.13 上游将 benchmarker 从 TritonBenchmarker 改为 InductorBenchmarker。InductorBenchmarker.benchmark_gpu() 内部通过 is_gpu() 判断设备类型,而 GPU_TYPES 不含 "npu",导致回退到 "cuda",最终调用 torch.cuda.synchronize() 报 AssertionError: Torch not compiled with CUDA enabled。v2.10 用的是 TritonBenchmarker,委托给 triton 的 do_bench,triton 内部能正确识别 NPU driver,所以不受影响。
所以本次修改对ExternKernelGPUBenchmarkRequest做了基类的替换
#3463
【资料变更】
【接口变更】
【功能验证】
【CheckList】