Pull Request已成功合入, 合并人@ascend-robot
(感谢 Xuan Peng 的贡献)变更摘要
本 PR 修复 PyTorch 2.7.1 下 Inductor NPU cpp_wrapper 不支持 Flex Attention mask_out 动态 backward 的问题。核心改动是将 Triton kernel metadata 中的 workspace_size、lock_num、lock_init_val 通过 NPUCachingAutotuner 传递到 NPU C++ wrapper,由 C++ launcher 按 metadata 在运行时分配 workspace 与同步锁内存并初始化锁,同时为 PyTorch 2.7.1 的 Flex Attention backward template 补充 SymbolicGridFn,使动态 call size 在 cpp_wrapper 下能生成符号化 grid 表达式。
主要改动
- kernel metadata 传递(
torch_npu/_inductor/runtime/triton_heuristics.py):NPUCachingAutotuner在构造params时新增从input_launcher.bin.metadata读取lock_num、lock_init_val、workspace_size三个字段(缺省为 0),供后续 C++ wrapper 使用。 - C++ wrapper 生成 workspace 与同步锁(
torch_npu/_inductor/codegen/cpp_wrapper_npu.py):generate_args_decl新增kernel_params参数,在非force_simt_only且相应 metadata 大于 0 时,通过allocate_workspace按workspace_size * grid_0 * grid_1 * grid_2分配 workspace,并按lock_num * sizeof(int64_t)分配同步锁内存、以lock_init_val初始化后调用aclrtMemcpy(ACL_MEMCPY_HOST_TO_DEVICE)写入设备,且校验其返回值是否为ACL_SUCCESS;同时generate_args_decl的调用处传入kernel_params,并新增包含NPUWorkspaceAllocator.h的头文件引入。 - Flex Attention backward 符号化 grid(
torch_npu/_inductor/kernel/flex_attention.py):新增以SymbolicGridFn装饰的_symbolic_flex_attention_backward_grid,依据BLOCK_M2、BLOCK_N1等 meta 参数计算三维 grid,并将其赋给upstream_flex_attention_backward_template.grid,替代 PyTorch 2.7.1 中无法对符号化 call size 生成 C++ 表达式的普通 grid 函数。


当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.9.0 | ||
| v2.10.0 | ||
| v2.11.0 | ||
| v2.12.0 | ||
| v2.7.1 | ||
| v2.7.1-26.1.0 | ||
| v2.9.0-26.1.0 | ||
| v2.12.0-26.1.0 | ||
| v2.11.0-26.1.0 | ||
| v2.10.0-26.1.0 | ||
| ci-test |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| torch_npu/_inductor | ✅ TonyYA, crazyDannyBoy (2/2) | ✅ crazyDannyBoy (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
HinPeng, thanks for your pull request. All authors of the commits have signed the CLA. 👍


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | ✅ | >>> |
| Build_ARM | ✅ | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| Build_X86_213 | 🛑 | >>> | |
| Build_ARM_213 | 🛑 | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | codecheck_pre-commit | ✅ | >>> |
| check_error | ✅ | >>> | |
| lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | ✅ | >>> |
| UT_ARM_A3_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_03 | ✅ | >>> | |
| UT_inductor_Part_01 | ✅ | >>> | |
| UT_inductor_Part_02 | ✅ | >>> | |
| UT_inductor_Part_03 | ✅ | >>> | |
| UT_inductor_Part_04 | ✅ | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Select_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_213 | 🛑 | >>> | |
| UT_inductor_Part_213 | 🛑 | >>> | |
| UT_DIST_ARM_Part_213 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_213 | 🛑 | >>> | |
| 流水线 | PR-pipeline_pytorch | ✅ | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


The following users do not have permission to comment /lgtm or /approve on any module in this PR:
zhucehw


/lgtm


The following users do not have permission to comment /lgtm or /approve on any module in this PR:
zhucehw


/approve


The following users do not have permission to comment /lgtm or /approve on any module in this PR:
zhucehw




【合入来源】
Fixes https://gitcode.com/Ascend/pytorch/issues/4431
【修改方案】
workspace_size、lock_num和lock_init_val从NPUCachingAutotuner传递到 NPU C++ wrapper。aclrtMemcpy返回值。SymbolicGridFn,使动态 call size 在cpp_wrapper下能够生成符号化 grid 表达式。【资料变更】
不涉及。
【接口变更】
不涉及跨仓或客户可见接口变更;仅调整 Inductor NPU 内部 kernel metadata 传递和 C++ wrapper 代码生成。
【功能验证】
python3 -m py_compile torch_npu/_inductor/codegen/cpp_wrapper_npu.py torch_npu/_inductor/runtime/triton_heuristics.py torch_npu/_inductor/kernel/flex_attention.py:通过。git diff upstream/v2.7.1..HEAD --check:通过。mask_out动态 backward +cpp_wrapper复现路径验证,原始cpp_wrapper requires SymbolicGridFn断言不再出现。【CheckList】