Pull Request已成功合入, 合并人@ascend-robot
(感谢 Xuan Peng 的贡献)变更摘要
本 PR 为 torch_npu Inductor NPU 后端补齐 Flex Attention mask_out 的 C++ wrapper 支持(由 v2.7.1 的 #45507 前向移植到 master):将 Triton kernel metadata 中的运行时资源信息(workspace_size、lock_num、lock_init_val)从 NPU autotuner 传入 C++ wrapper,由生成的 C++ launcher 在运行时分配 workspace 与同步锁并完成锁内存初始化。核心改动位于 torch_npu/_inductor/codegen/cpp_wrapper_npu.py 与 torch_npu/_inductor/runtime/triton_heuristics.py;master 的 Flex Attention 模板已原生使用符号化 grid,因此无需回迁 v2.7.1 的 SymbolicGridFn backport。
主要改动
- kernel metadata 透传: 在
NPUCachingAutotuner与NPUSymbolicGroupedAutotuner的params中新增lock_num、lock_init_val、workspace_size三个字段(通过getattr从metadata读取并转int,缺省为 0),适配新版本 grouped autotune variant 的 metadata 路径。 - workspace 运行时分配:
CppWrapperNpu.generate_args_decl新增kernel_params参数,在非 pure-SIMT 且workspace_size > 0时生成按grid_0 * grid_1 * grid_2缩放大小的allocate_workspace调用,并将workspace_addr指向其 storage 数据。 - 同步锁分配与初始化: 在非 pure-SIMT 且
lock_num > 0时分配同步锁内存,构造lock_init_val初值向量,通过aclrtMemcpy(ACL_MEMCPY_HOST_TO_DEVICE)初始化,并检查返回值,失败时抛出包含错误码的std::runtime_error。 - wrapper 调用链调整:
DeferredNpuTritonCallWrapper改为先调用generate_args_decl(传入is_triton_kernel=True、is_pure_simt、kernel_params=params)再拼装launch_calllambda;同时add_device_include新增torch_npu/csrc/core/npu/NPUWorkspaceAllocator.h头文件引用。 - pure-SIMT 路径隔离: 生成的 workspace 分配与同步锁分配均以
not is_pure_simt为前置条件,pure-SIMT 路径不会生成这些运行时分配代码,保持原有行为不变。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| torch_npu/_inductor | ✅ weizhan4, crazyDannyBoy (2/2) | ✅ weizhan4, crazyDannyBoy (2/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
HinPeng, thanks for your pull request. All authors of the commits have signed the CLA. 👍


当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.7.1-26.1.0 | ||
| v2.10.0 | ||
| v2.9.0 | ||
| v2.7.1 | ||
| v2.11.0 | ||
| v2.12.0 | ||
| v2.9.0-26.1.0 | ||
| v2.12.0-26.1.0 | ||
| v2.11.0-26.1.0 | ||
| v2.10.0-26.1.0 | ||
| ci-test |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | ✅ | >>> |
| Build_ARM | ✅ | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| Build_X86_213 | ✅ | >>> | |
| Build_ARM_213 | ✅ | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | codecheck_pre-commit | ✅ | >>> |
| check_error | ✅ | >>> | |
| lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🛑 | >>> |
| UT_ARM_A3_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_03 | ✅ | >>> | |
| UT_inductor_Part_01 | 🛑 | >>> | |
| UT_inductor_Part_02 | 🛑 | >>> | |
| UT_inductor_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_04 | 🛑 | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Select_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_213 | ✅ | >>> | |
| UT_inductor_Part_213 | 🛑 | >>> | |
| UT_DIST_ARM_Part_213 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_213 | ✅ | >>> | |
| 流水线 | PR-pipeline_pytorch | ✅ | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


The following users do not have permission to comment /lgtm or /approve on any module in this PR:
zhucehw


The following users do not have permission to comment /lgtm or /approve on any module in this PR:
zhucehw


The following users do not have permission to comment /lgtm or /approve on any module in this PR:
zhucehw


/lgtm
/approve


The following users do not have permission to comment /lgtm or /approve on any module in this PR:
zhucehw




【合入来源】
Fixes https://gitcode.com/Ascend/pytorch/issues/4431
Forward-port of https://gitcode.com/Ascend/pytorch/pull/45507 from
v2.7.1tomaster.【修改方案】
workspace_size、lock_num和lock_init_val从 NPU autotuner 传递到 NPU C++ wrapper,并适配新版本的 grouped autotune variant metadata 路径。aclrtMemcpy返回值。master的 Flex Attention 模板已原生使用符号化 grid,因此不再回迁v2.7.1专属的SymbolicGridFnbackport。【资料变更】
不涉及。
【接口变更】
不涉及跨仓或客户可见接口变更;仅调整 Inductor NPU 内部 kernel metadata 传递和 C++ wrapper 代码生成。
【功能验证】
python3 -m py_compile torch_npu/_inductor/codegen/cpp_wrapper_npu.py torch_npu/_inductor/runtime/triton_heuristics.py torch_npu/_inductor/kernel/flex_attention.py torch_npu/_inductor/kernel/flexattention_template.py:通过。git diff --check:通过。torch,且无 NPU 运行环境;PR CI 待执行。【CheckList】