Pull Request已成功合入, 合并人@ascend-robot
(感谢 rain-666 的贡献)变更摘要
本 PR 修复了 torch_npu Inductor 后端中 LayerNorm 使用 welford_reduce 时(尤其是 Ascend 950 上)SIMD 归约代码生成错误的问题。核心改动位于 torch_npu/_inductor/codegen/triton.py:新增 Welford 累加器(_acc_sum/_acc_sum_sq/_acc_count)的 tl.zeros 初始化逻辑,并按"向量化外层轴"与"标量外层轴"区分生成归约循环,避免多行数据在各自 tile 内相互累加;同时将持久化归约判定扩展为支持 A5 上归约长度 ≤ 8192 的 welford_reduce。此外新增默认关闭的 enable_layernorm_v4 开关(环境变量 TORCHINDUCTOR_ENABLE_LAYERNORM_V4),在 W=512 的大行 LayerNorm 场景回退到融合的 CANN 内核,并补充了对应的代码生成测试用例。
主要改动
- Welford 累加器初始化与循环代码生成修复(
triton.py): 新增_is_scalar_welford_outer_axis()判断和initialize_welford_accumulators()辅助函数,在向量化 tile 进入后初始化 Welford 状态;外层归约循环从统一的for loop_r in range(...)改写为按_offset/BLOCK/BLOCK_SUB步进的for <name> in range(...)与for <name>_loop_offset in range(...),防止各行数据相互累加。 - 新增
vectorized_welford_rank()与动态 rank 维度生成: 新增方法计算加入向量化外层轴后的 DSL rank,get_axis_direction()、reduction_resize()及累加器 shape(welford_acc_type/welford_acc_shape)均基于该 rank 生成[block_sub, 1, ...]形式的广播维度。 - 持久化归约判定扩展(
triton.py): 在npu_config.is_ascend950且启用enable_welford、归约类型为welford_reduce且归约长度静态 ≤ 8192 时直接启用持久化归约;同时将find_reduction_node()改为通过getattr(self, "node_schedule", self.features.node_schedule)取值,兼容基类构造函数中node_schedule尚未赋值的情况。 - 新增
enable_layernorm_v4配置与 W=512 回退(config.py、lowering.py): 新增默认关闭的开关enable_layernorm_v4;在lowering.py中通过should_use_layer_norm_v4()判断(A5、启用 Welford 与开关、FP16/BF16、normalized_shape == (512,)且行数可静态证明 ≥ 512)时,回退到融合的native_layer_normCANN 内核,避免 Welford 降级拆分为 SIMD 归约与 SIMT 后处理。 - 新增测试用例(
test/_inductor/test_var_mean.py): 新增test_welford_simd_codegen_above_persistent_threshold,在 Ascend 950 上验证(200, 5036)LayerNorm 编译产物包含npu_kernel_type': 'simd'与vectorized_welford_axis,且不再生成for loop_r循环。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| test | ✅ crazyDannyBoy, TonyYA, weizhan4 (3/2) | ✅ crazyDannyBoy (1/1) |
| torch_npu/_inductor | ✅ weizhan4, crazyDannyBoy, TonyYA (3/2) | ✅ weizhan4, crazyDannyBoy (2/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
rain-666, thanks for your pull request. All authors of the commits have signed the CLA. 👍


当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.9.0 | ||
| v2.12.0 | ||
| v2.11.0 | ||
| v2.10.0-26.1.0 | ||
| v2.12.0-26.1.0 | ||
| v2.10.0 | ||
| v2.7.1-26.1.0 | ||
| v2.11.0-26.1.0 | ||
| v2.9.0-26.1.0 | ||
| v2.7.1 | ||
| ci-test |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


Linking Issue Notice
@rain-666 , the pull request must be linked to at least one issue.
If an issue has already been linked, but the needs-issue label remains, you can remove the label by commenting /check-issue .


compile


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | 🕚 | >>> |
| Build_ARM | 🕚 | >>> | |
| Build_X86_torchair | 🕚 | >>> | |
| Build_ARM_torchair | 🕚 | >>> | |
| patch_test | 🕚 | >>> | |
| Build_X86_213 | 🕚 | >>> | |
| Build_ARM_213 | 🕚 | >>> | |
| 恶意代码检查 | Antipoison | 🟨 | >>> |
| 编码安全与规范检查 | codecheck_pre-commit | 🟨 | >>> |
| check_error | 🟨 | >>> | |
| lintrunner | 🟨 | >>> | |
| 开源片段检查 | SCA | 🟨 | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🕚 | >>> |
| UT_ARM_A3_Part_02 | 🕚 | >>> | |
| UT_ARM_A2_Part_01 | 🕚 | >>> | |
| UT_ARM_A2_Part_02 | 🕚 | >>> | |
| UT_ARM_A2_Part_03 | 🕚 | >>> | |
| UT_inductor_Part_01 | 🕚 | >>> | |
| UT_inductor_Part_02 | 🕚 | >>> | |
| UT_inductor_Part_03 | 🕚 | >>> | |
| UT_inductor_Part_04 | 🕚 | >>> | |
| UT_DIST_ARM_Part_01 | 🕚 | >>> | |
| UT_DIST_ARM_Part_02 | 🕚 | >>> | |
| UT_DIST_ARM_Part_03 | 🕚 | >>> | |
| UT_DIST_ARM_Part_04 | 🕚 | >>> | |
| UT_ARM_A2_Select_Part_01 | 🕚 | >>> | |
| UT_ARM_A2_Select_Part_02 | 🕚 | >>> | |
| UT_ARM_A2_Part_213 | 🕚 | >>> | |
| UT_inductor_Part_213 | 🕚 | >>> | |
| UT_DIST_ARM_Part_213 | 🕚 | >>> | |
| UT_ARM_A2_Select_Part_213 | 🕚 | >>> | |
| 流水线 | PR-pipeline_pytorch | 🟨 | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | ✅ | >>> |
| Build_ARM | ✅ | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| Build_X86_213 | ✅ | >>> | |
| Build_ARM_213 | ✅ | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | codecheck_pre-commit | ✅ | >>> |
| check_error | ✅ | >>> | |
| lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🛑 | >>> |
| UT_ARM_A3_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_03 | ✅ | >>> | |
| UT_inductor_Part_01 | 🛑 | >>> | |
| UT_inductor_Part_02 | 🛑 | >>> | |
| UT_inductor_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_04 | 🛑 | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Select_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_213 | ✅ | >>> | |
| UT_inductor_Part_213 | 🛑 | >>> | |
| UT_DIST_ARM_Part_213 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_213 | ✅ | >>> | |
| 流水线 | PR-pipeline_pytorch | ✅ | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


/lgtm


/approve




本 PR 修复了 torch_npu Inductor 后端中 LayerNorm 使用 welford_reduce 时(尤其是 Ascend 950 上)SIMD 归约代码生成错误的问题。核心改动位于 torch_npu/_inductor/codegen/triton.py:新增 Welford 累加器(_acc_sum/_acc_sum_sq/_acc_count)的 tl.zeros 初始化逻辑,并按"向量化外层轴"与"标量外层轴"区分生成归约循环,避免多行数据在各自 tile 内相互累加;同时将持久化归约判定扩展为支持 A5 上归约长度 ≤ 8192 的 welford_reduce。此外新增默认关闭的 enable_layernorm_v4 开关(环境变量 TORCHINDUCTOR_ENABLE_LAYERNORM_V4),在 W=512 的大行 LayerNorm 场景回退到融合的 CANN 内核,并补充了对应的代码生成测试用例。