已关闭
[Bug]: 开启 TORCH_LOGS=+all 时 torch.compile(backend="npugraphs") 编译期段错误 #4744
wuyouqi1创建于  23 天前关闭于  11 天前
wuyouqi1
wuyouqi1成员
23 天前 创建

在提交新问题之前,请确保您已经在社区中搜索过相关问题,并使用了社区中提供的资源/工具后,仍未找到满意的解决方式。

已搜索社区,未发现相同问题(get_npu_format + FakeTensor repr 段错误)。

⚠️ 安全信息提醒:请仔细检查提供的文本内容,确保其不包含敏感数据信息,包括但不限于:

  • API 令牌或密钥
  • 密码或身份验证凭证
  • 私有网址或接口地址
  • 个人或机密数据
  • ...

在分享配置信息或代码示例时,请将敏感信息脱敏处理,或使用 <TOKEN> 等占位符替代原有内容。
(以下日志已脱敏:去除主机名/PID/设备指针)

环境信息

- 操作系统:aarch64 Linux(openEuler)
- 昇腾硬件信息:Ascend 950(问题首发现场);Atlas 910B2(本地复现验证,同样触发)
- CANN软件版本:问题与 CANN 版本无关(950 现场与 910B2 + CANN 9.x 均可复现)
- 安装的对应软件版本:python 3.10 / torch 2.7.1 (cpu) / torch_npu 2.7.1.post9

🐛 问题描述

1. 复现脚本

import os
os.environ["TORCH_LOGS"] = "+all"
import torch
import torch.nn.functional as F

def softmax_func(x):
    return F.softmax(x, dim=-1)

compiled_softmax = torch.compile(softmax_func, backend="npugraphs")

x = torch.randn(4, 10, device="npu", dtype=torch.float32)
output = compiled_softmax(x)
print(output.sum(dim=-1))
print("OK")

2. 现象

Segmentation fault (core dumped)。日志止于 Dynamo Step 2(调用 npugraphs 后端)之后的若干条
aten.lift_fresh.default (tensor([]),)(AOT effect token)处,无任何算子下发日志,即崩在编译期:

I ... output_graph.py:1515] [0/0] Step 2: calling compiler function npugraphs
V ... fake_tensor.py:1803] [0/0] aten.lift_fresh.default (tensor([]),) {}
... (共 8 条)
Segmentation fault (core dumped)

3. 触发条件

  • 必要条件:开启 verbose 日志(TORCH_LOGS=+all)。去掉该环境变量后同样脚本运行正常(输出 OK)
  • 与 backend 无关:backend="aot_eager" 同样触发
  • 与硬件无关:950 / 910B2 均可复现

4. 根因分析(faulthandler + gdb 栈定位)

torch_npu 在 import torch_npu 时无条件将 torch.Tensor.__repr__ 替换为
_npu_private_format_repr(torch_npu/utils/tensor_methods.py):对 device.type == "npu"
的张量调用 torch_npu.get_npu_format(self) 以识别私有格式(如 FRACTAL_NZ)。

torch.compile 编译期(Dynamo/AOT trace)会通过 wrap_to_fake 把输入张量转换为 FakeTensor。
FakeTensor 满足 device.type == "npu"(fake_device 继承自真张量),但其底层是普通
c10::StorageImpl,没有 npu_desc_ 成员。

崩溃链:

  1. FakeTensorMode.dispatch(torch/_subclasses/fake_tensor.py:1803)的
    log.debug("%s %s %s", func, tree_flatten(args), kwargs) 格式化参数,需对 FakeTensor 取 repr()
  2. _npu_private_format_repr 对 FakeTensor 调用 get_npu_format
  3. 嵌套调用发生在 _in_kernel_invocation_manager(Python dispatch 被禁用)期间,
    跳过 Python 层拦截、按 PrivateUse1 key 直达 C++ 内核
  4. NPUNativeFunctions::get_npu_format(torch_npu/csrc/aten/common/FormatCastKernelNpu.cpp:383)
    经 NPUBridge::GetNpuStorageImpl 对普通 StorageImpl 做 static_cast<NPUStorageImpl*> 后读
    npu_desc_ → 越界读 → SIGSEGV(try/except 无法捕获,故 repr 补丁中的异常兜底无效)

关键栈(faulthandler):

torch/_ops.py:1158 __call__                       ← torch.ops.npu.get_npu_format(fake)
torch_npu/npu/_format.py:39 patched_get_format
torch_npu/utils/tensor_methods.py:94 _npu_private_format_repr
logging/__init__.py:368 getMessage                ← 日志格式化触发 repr
torch/_subclasses/fake_tensor.py:1803 dispatch    ← fake mode 分发日志

5. 期望行为

  • FakeTensor/MetaTensor 等符号张量的 repr() 正常输出,不触碰 C++ 格式查询
  • get_npu_format 对无 NPU 存储的张量抛出可捕获的 Python 异常,而非段错误
  • 真实 NPU 张量行为完全不变(实测 empty(0)/标量/empty_strided/视图/save-load/拷贝/autograd 等路径均不受影响)

修复 PR:fix/fake-tensor-repr-segv(Ascend/pytorch,基于 v2.7.1,链接见关联 PR)

欢迎加入社区,感谢您对社区的贡献 🎉!

likedislike
wuyouqi1wuyouqi1成员
23 天前 添加了label:bug
TorchNPU-BotTorchNPU-Bot成员
23 天前 添加了label:triage-review
TorchNPU-Bot
TorchNPU-Bot成员
23 天前 评论:

issue待分派,添加triage-review标签

likedislike
wuyouqi1wuyouqi1成员
23 天前 修改标题为 “[Bug]: 开启 TORCH_LOGS=+all 时 torch.compile(backend="npugraphs") 编译期段错误”,原标题为“[Bug]: ”
wuyouqi1wuyouqi1成员
23 天前 修改了issue 的描述
wuyouqi1wuyouqi1成员
23 天前 将 wuyouqi1 设为负责人
wuyouqi1wuyouqi1成员
23 天前 关联了里程碑:v26.2.0
TorchNPU-BotTorchNPU-Bot成员
23 天前 添加了label:bot-triaged;删除了label:triage-review
TorchNPU-Bot
TorchNPU-Bot成员
23 天前 评论:

检测到当前 issue 已关联 PR,自动添加标签:bot-triaged

likedislike
ascend-robotascend-robot成员
16 天前 关联了pull request:[sync] PR-46180: fix(core): guard get_npu_format against symbolic tensors to avoid segfault
此处折叠了17条事件消息 查看更多
ascend-robotascend-robot成员
11 天前 添加了label:resolved