已关闭
[Bug]: triton experimental inductor后端,不兼容triton ascend 3.6 #4253
AllenGuan创建于 8月21日关闭于 8月25日
8月21日 添加了label:triage-review
TorchNPU-Bot
8月21日 评论:
8月21日 评论:
issue待分派,添加triage-review标签


8月21日 添加了label:bug
8月25日 关闭了 issue
8月25日 添加了label:resolved
13 天前 关联了里程碑:v26.2.0
13 天前 issue状态由 TODO 改变为 DONE
在提交新问题之前,请确保您已经在社区中搜索过相关问题,并使用了社区中提供的资源/工具后,仍未找到满意的解决方式。
⚠️ 安全信息提醒:请仔细检查提供的文本内容,确保其不包含敏感数据信息,包括但不限于:
在分享配置信息或代码示例时,请将敏感信息脱敏处理,或使用
<TOKEN>等占位符替代原有内容。环境信息
🐛 问题描述
torch_npu 的
triton_experimental后端(torch.compile(model, options={"npu_backend": "triton_experimental"}))在 triton-ascend 3.6.0 上无法完成任何 kernel 的编译/启动,共观察到三类症状(同一环境、不同触发条件)。在 triton-ascend 3.2.x(vendored core 3.2.0)上同样的代码工作正常;差异指向 vendored triton core 3.5.0 的 API 变化:launch hooks 从CompiledKernel类属性迁移到triton.knobs.runtime.*,以及AttrsDescriptor类被整体移除(attrs 改用 plain dict 表达)。症状一:任意 torch.compile 图在 launcher 创建阶段报 AttributeError(最常见的入口症状)
复现步骤:
import torch import torch_npu # noqa: F401 def f(a, b): return a + b compiled = torch.compile(f, options={"npu_backend": "triton_experimental"}) a = torch.randn(1024, device="npu") b = torch.randn(1024, device="npu") print(compiled(a, b))期望行为:编译成功,输出与 eager 一致的
a + b结果。实际行为:抛出异常,任何经过 triton kernel 的图均失败:
决定性栈帧:
torch_npu/_inductor/triton_experimental/npu_triton_heuristics.py的NPUTritonCompileResult.make_launcher(构建 launcher scope 处直接访问binary.__class__.launch_enter_hook/binary.__class__.launch_exit_hook)。triton core 3.5.0 中这两个属性已迁移到triton.knobs.runtime.launch_enter_hook / launch_exit_hook,CompiledKernel类上不再存在。症状二:绕过症状一后,autotuner 报 "All tiling ... are not runnable"(launcher 参数表未剔除 constexpr)
若将 hooks 读取改为兼容写法,同一个最小复现片段继续失败,报错换为:
该异常由 autotuner 在所有候选 tiling 的 PreRun 全部失败后抛出;per-tiling 的底层异常被 debug 日志吞掉,开启 debug 日志后可见根因为:
定位:
npu_triton_heuristics.pymake_launcher中按triton_version_uses_attrs_dict()(core 3.5.0 上为 True)走 attrs_dict 分支构造def_args/call_args时,仅剔除了num_warps/num_stages,未剔除 kernel 的 constexpr 形参(XBLOCK/YBLOCK/R0_BLOCK 等,已由选定的 Config 烘入编译产物,caller 不会传入),导致 launcher 的 Python 签名要求XBLOCK而调用方不提供。症状三:动态 shape 的 OUTER 归约图在 codegen 阶段报 ImportError(rsplit combine kernel 路径)
复现步骤:
import torch import torch_npu # noqa: F401 class Model(torch.nn.Module): def forward(self, mm): view = mm.reshape(mm.shape[0], -1, 128) return view.to(torch.float32).sum(0) model = torch.compile(Model().npu(), dynamic=None) x = torch.randn(16384, 2048, device="npu", dtype=torch.bfloat16) torch._dynamo.mark_dynamic(x, 0) with torch.no_grad(): print(model(x))期望行为:编译成功,输出正确归约结果。
实际行为:整个
torch.compile失败:决定性栈帧:
torch_npu/_inductor/triton_experimental/codegen/triton.py的_npu_emit_rsplit_combine(rsplit-outer 归约的跨核 partial+combine 拆分路径,对满足条件的单 kernel OUTER 归约默认启用),函数内无条件执行from triton.compiler.compiler import AttrsDescriptor并调用AttrsDescriptor.from_dict(...)构造 combine kernel 的triton_meta["configs"]。triton core 3.5.0 已移除AttrsDescriptor(3.2.x 中位于triton.backends.compiler,并由triton.compiler.compilerre-export),因此任何触发该路径的图(动态 shape 的 view+sum、rms_norm backward 等 OUTER 归约)都会失败。除 import 失败外,3.5.0 上 configs 亦要求 plain dict 形态的 attrs 表达,与 3.2.x 的AttrsDescriptor对象形态不同。补充信息
sum(0))。torch_npu._inductor.triton_experimental.npu_triton_heuristicslogger 开启 DEBUG 级别查看底层报错。欢迎加入社区,感谢您对社区的贡献 🎉!