已关闭
[Bug]: torch_npu.npu_grouped_matmul 在W4A8场景下A2开启nz时输出有问题,并且测试用例无法通过 #1516
menogrey创建于  1月22日关闭于  7月4日
menogrey
1月22日 创建

在提交新问题之前,请确保您已经在社区中搜索过相关问题,并使用了社区中提供的资源/工具后,仍未找到满意的解决方式。

⚠️ 安全信息提醒:请仔细检查提供的文本内容,确保其不包含敏感数据信息,包括但不限于:

  • API 令牌或密钥
  • 密码或身份验证凭证
  • 私有网址或接口地址
  • 个人或机密数据
  • ...

在分享配置信息或代码示例时,请将敏感信息脱敏处理,或使用 <TOKEN> 等占位符替代原有内容。

环境信息

例如:
- 操作系统
- 昇腾硬件信息
- CANN软件版本
- 安装的对应软件版本

PyTorch version: 2.8.0+cpu
Is debug build: False

OS: Ubuntu 22.04.5 LTS (aarch64)
GCC version: (Ubuntu 11.4.0-1ubuntu1~22.04.2) 11.4.0
Clang version: Could not collect
CMake version: version 4.2.1
Libc version: glibc-2.35

Python version: 3.11.13 (main, Nov 20 2025, 16:02:27) [GCC 11.4.0] (64-bit runtime)
Python platform: Linux-4.19.90-vhulk2211.3.0.h1543.eulerosv2r10.aarch64-aarch64-with-glibc2.35

CPU:
Architecture: aarch64
CPU op-mode(s): 64-bit
Byte Order: Little Endian
CPU(s): 192
On-line CPU(s) list: 0-191
Vendor ID: HiSilicon
Model name: Kunpeng-920
Model: 0
Thread(s) per core: 1
Core(s) per cluster: 48
Socket(s): -
Cluster(s): 4
Stepping: 0x1
BogoMIPS: 200.00
Flags: fp asimd evtstrm aes pmull sha1 sha2 crc32 atomics fphp asimdhp cpuid asimdrdm jscvt fcma dcpop asimddp asimdfhm ssbs
L1d cache: 12 MiB (192 instances)
L1i cache: 12 MiB (192 instances)
L2 cache: 96 MiB (192 instances)
L3 cache: 192 MiB (8 instances)
NUMA node(s): 8
NUMA node0 CPU(s): 0-23
NUMA node1 CPU(s): 24-47
NUMA node2 CPU(s): 48-71
NUMA node3 CPU(s): 72-95
NUMA node4 CPU(s): 96-119
NUMA node5 CPU(s): 120-143
NUMA node6 CPU(s): 144-167
NUMA node7 CPU(s): 168-191
Vulnerability Itlb multihit: Not affected
Vulnerability L1tf: Not affected
Vulnerability Mds: Not affected
Vulnerability Meltdown: Not affected
Vulnerability Mmio stale data: Not affected
Vulnerability Retbleed: Not affected
Vulnerability Spec store bypass: Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1: Mitigation; __user pointer sanitization
Vulnerability Spectre v2: Not affected
Vulnerability Srbds: Not affected
Vulnerability Tsx async abort: Not affected

Versions of relevant libraries:
[pip3] mypy==1.11.1
[pip3] mypy_extensions==1.1.0
[pip3] numpy==1.26.4
[pip3] pyzmq==27.1.0
[pip3] sentence-transformers==5.2.0
[pip3] torch==2.8.0
[pip3] torch_npu==2.8.0
[pip3] torchvision==0.23.0
[pip3] transformers==4.57.3
[pip3] zmq==0.0.0
[conda] Could not collect
vLLM Version: 0.13.0
vLLM Ascend Version: 0.13.0rc2.dev194+g9337492ed.d20260114 (git sha: 9337492ed, date: 20260114)

ENV Variables:
ATB_OPSRUNNER_KERNEL_CACHE_LOCAL_COUNT=1
ATB_STREAM_SYNC_EVERY_RUNNER_ENABLE=0
ATB_OPSRUNNER_SETUP_CACHE_ENABLE=1
ATB_WORKSPACE_MEM_ALLOC_GLOBAL=1
ATB_DEVICE_TILING_BUFFER_BLOCK_NUM=32
ASCEND_VISIBLE_DEVICES=2,3,4,5
ATB_STREAM_SYNC_EVERY_KERNEL_ENABLE=0
ASCEND_RUNTIME_OPTIONS=
ATB_OPSRUNNER_KERNEL_CACHE_GLOABL_COUNT=5
ATB_HOME_PATH=/usr/local/Ascend/nnal/atb/latest/atb/cxx_abi_1
ASCEND_TOOLKIT_HOME=/usr/local/Ascend/ascend-toolkit/latest
ATB_COMPARE_TILING_EVERY_KERNEL=0
ASCEND_OPP_PATH=/usr/local/Ascend/ascend-toolkit/latest/opp
LD_LIBRARY_PATH=/usr/local/Ascend/nnal/atb/latest/atb/cxx_abi_1/lib:/usr/local/Ascend/nnal/atb/latest/atb/cxx_abi_1/examples:/usr/local/Ascend/nnal/atb/latest/atb/cxx_abi_1/tests/atbopstest:/usr/local/Ascend/ascend-toolkit/latest/tools/aml/lib64:/usr/local/Ascend/ascend-toolkit/latest/tools/aml/lib64/plugin:/usr/local/Ascend/ascend-toolkit/latest/lib64:/usr/local/Ascend/ascend-toolkit/latest/lib64/plugin/opskernel:/usr/local/Ascend/ascend-toolkit/latest/lib64/plugin/nnengine:/usr/local/Ascend/ascend-toolkit/latest/opp/built-in/op_impl/ai_core/tbe/op_tiling/lib/linux/aarch64:/usr/local/Ascend/nnal/atb/latest/atb/cxx_abi_1/lib:/usr/local/Ascend/nnal/atb/latest/atb/cxx_abi_1/examples:/usr/local/Ascend/nnal/atb/latest/atb/cxx_abi_1/tests/atbopstest:/usr/local/Ascend/ascend-toolkit/latest/tools/aml/lib64:/usr/local/Ascend/ascend-toolkit/latest/tools/aml/lib64/plugin:/usr/local/Ascend/ascend-toolkit/latest/lib64:/usr/local/Ascend/ascend-toolkit/latest/lib64/plugin/opskernel:/usr/local/Ascend/ascend-toolkit/latest/lib64/plugin/nnengine:/usr/local/Ascend/ascend-toolkit/latest/opp/built-in/op_impl/ai_core/tbe/op_tiling/lib/linux/aarch64:/usr/local/Ascend/driver/lib64/common/:/usr/local/Ascend/driver/lib64/driver/
ASCEND_AICPU_PATH=/usr/local/Ascend/ascend-toolkit/latest
ATB_STREAM_SYNC_EVERY_OPERATION_ENABLE=0
ASCEND_HOME_PATH=/usr/local/Ascend/ascend-toolkit/latest
ATB_MATMUL_SHUFFLE_K_ENABLE=1
ATB_WORKSPACE_MEM_ALLOC_ALG_TYPE=1
ATB_HOST_TILING_BUFFER_BLOCK_NUM=128
ATB_SHARE_MEMORY_NAME_SUFFIX=
TORCH_DEVICE_BACKEND_AUTOLOAD=1
PYTORCH_NVML_BASED_CUDA_CHECK=1
TORCHINDUCTOR_COMPILE_THREADS=1

NPU:
+------------------------------------------------------------------------------------------------+
| npu-smi 25.3.rc1.2 Version: 25.3.rc1.2 |
+---------------------------+---------------+----------------------------------------------------+
| NPU Name | Health | Power(W) Temp(C) Hugepages-Usage(page)|
| Chip | Bus-Id | AICore(%) Memory-Usage(MB) HBM-Usage(MB) |
+===========================+===============+====================================================+
| 2 910B4 | OK | 87.5 38 0 / 0 |
| 0 | 0000:C2:00.0 | 0 0 / 0 2897 / 32768 |
+===========================+===============+====================================================+
| 3 910B4 | OK | 89.3 40 0 / 0 |
| 0 | 0000:02:00.0 | 0 0 / 0 2893 / 32768 |
+===========================+===============+====================================================+
| 4 910B4 | OK | 87.1 39 0 / 0 |
| 0 | 0000:81:00.0 | 0 0 / 0 2901 / 32768 |
+===========================+===============+====================================================+
| 5 910B4 | OK | 86.0 38 0 / 0 |
| 0 | 0000:41:00.0 | 0 0 / 0 2899 / 32768 |
+===========================+===============+====================================================+
+---------------------------+---------------+----------------------------------------------------+
| NPU Chip | Process id | Process name | Process memory(MB) |
+===========================+===============+====================================================+
| No running processes found in NPU 2 |
+===========================+===============+====================================================+
| No running processes found in NPU 3 |
+===========================+===============+====================================================+
| No running processes found in NPU 4 |
+===========================+===============+====================================================+
| No running processes found in NPU 5 |
+===========================+===============+====================================================+

CANN:
package_name=Ascend-cann-toolkit
version=8.3.RC2
innerversion=V100R001C23SPC002B210
compatible_version=[V100R001C15],[V100R001C18],[V100R001C19],[V100R001C20],[V100R001C21],[V100R001C23]
arch=aarch64
os=linux
path=/usr/local/Ascend/ascend-toolkit/8.3.RC2/aarch64-linux

🐛 问题描述

注释掉用例前的skip并运行命令:
pytest -sv op-plugin/test/test_custom_ops/test_npu_grouped_matmul.py::TestGroupedMatmul::test_npu_grouped_matmul_A8W4

开启nz的时候对不齐,同样的试了A3的,测试用例可以通过

out_nz_shape = out_nz[0].shape
out_nz_dim1 = out_nz_shape[1]
self.assertEqual(out_nz_dim1, golden_dim1)

  self.assertEqual(out_nz[0][:golden_dim0, :], out_golden.npu())

op-plugin/test/test_custom_ops/test_npu_grouped_matmul.py:707:


../.local/lib/python3.11/site-packages/torch_npu/testing/testcase.py:382: in assertEqual
_assertEqual(x, y, prec=prec, message=message, allow_inf=allow_inf, exact_dtype=exact_dtype)
../.local/lib/python3.11/site-packages/torch_npu/testing/testcase.py:354: in _assertEqual
self._assertTensorsEqual(x, y, prec=prec, message=message,
../.local/lib/python3.11/site-packages/torch_npu/testing/testcase.py:334: in _assertTensorsEqual
self._assert_tensor_equal(x, y, message, exact_dtype, allow_inf, prec)
../.local/lib/python3.11/site-packages/torch_npu/testing/testcase.py:270: in _assert_tensor_equal
self.assertLessEqual(max_err, prec, message)
E AssertionError: tensor(161.7500, device='npu:0', dtype=torch.float16) not less than or equal to 1e-05 :
============================================================================================ warnings summary ============================================================================================
../.local/lib/python3.11/site-packages/torch_npu/utils/collect_env.py:58
../.local/lib/python3.11/site-packages/torch_npu/utils/collect_env.py:58
/home/zym/.local/lib/python3.11/site-packages/torch_npu/utils/collect_env.py:58: UserWarning: Warning: The /usr/local/Ascend/ascend-toolkit/latest owner does not match the current owner.
warnings.warn(f"Warning: The {path} owner does not match the current owner.")

../.local/lib/python3.11/site-packages/torch_npu/utils/collect_env.py:58
../.local/lib/python3.11/site-packages/torch_npu/utils/collect_env.py:58
/home/zym/.local/lib/python3.11/site-packages/torch_npu/utils/collect_env.py:58: UserWarning: Warning: The /usr/local/Ascend/ascend-toolkit/8.3.RC2/aarch64-linux/ascend_toolkit_install.info owner does not match the current owner.
warnings.warn(f"Warning: The {path} owner does not match the current owner.")

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
======================================================================================== short test summary info =========================================================================================
FAILED op-plugin/test/test_custom_ops/test_npu_grouped_matmul.py::TestGroupedMatmul::test_npu_grouped_matmul_A8W4 - AssertionError: tensor(161.7500, device='npu:0', dtype=torch.float16) not less than or equal to 1e-05 :
===================================================================================== 1 failed, 4 warnings in 34.63s =====================================================================================

欢迎加入社区,感谢您对社区的贡献 🎉!

likedislike
ascend-robotascend-robot成员
1月22日 添加了label:bug
haiyan8
haiyan8
1月22日 评论:

我们先分析下

likedislike
menogrey
1月26日 评论:

有进展吗 @yuhaiyan

likedislike
neal_xiao
2月3日 评论:

有进展吗,同样的问题@yuhaiyan

likedislike
haiyan8
haiyan8
2月11日 评论:

A3上torch_npu.npu_format_cast不起作用,测试脚本可能有问题,还在确认。

likedislike
neal_xiao
2月14日 评论:

A3上torch_npu.npu_format_cast不起作用,测试脚本可能有问题,还在确认。

@yuhaiyan

不用管A3测试有没有问题把,大家关心的是A2 NZ输出是错的。

likedislike
haiyan8
haiyan8
3月9日 评论:

A3上torch_npu.npu_format_cast不起作用,测试脚本可能有问题,还在确认。

@yuhaiyan

不用管A3测试有没有问题把,大家关心的是A2 NZ输出是错的。

@chuanjiexiao

用例可能有问题,还在定位

likedislike
chenrayray
chenrayray成员
7月4日 评论:

您好,本Issue因长期无有效互动,现暂时关闭。如问题依然存在,请重新打开并提供最新复现信息,我们会继续跟进。感谢理解!

likedislike
chenrayraychenrayray成员
7月4日 issue状态由 TODO 改变为 DONE
chenrayraychenrayray成员
7月4日 关闭了 issue
ascend-robotascend-robot成员
7月4日 添加了label:resolved