已关闭
[Bug]:SubmSparseConv3d算子某些输入报错507015 #368
mengde创建于  19 天前关闭于  11 天前
mengde
mengde
19 天前 创建

环境信息

芯片类型:Ascend910b
CANN版本:8.3.RC1
python版本:3.11
torch版本:2.7.1

🐛 问题描述

实际项目中输入会导致算子崩溃,报错如下

RuntimeError: npuSynchronizeDevice:build/CMakeFiles/torchnpu.dir/compilerdepend.ts:525 NPU function error: AclrtSynchronizeDeviceWithTimeout, error code is 507015
[ERROR] 2026-09-10-15:42:36 (PID:602138, Device:0, RankID:-1) ERR00100 PTA call acl api failed
[Error]: The aicore execution is abnormal.
Rectify the fault based on the error information in the ascend log.
EZ9999: Inner Error!
EZ9999[PID: 602138] 2026-09-10-15:42:36.295.577 (EZ9999): The error from device(chipId:0, dieId:0), serial number is 27, there is an exception of fftsplus aivector error, core id is 37, error code = 0x800000, dump info: pc start: 0x12c09021004c, current: 0x12c09021f0e4, vec error info: 0x8000000a9, mte error info: 0x7d030d7e5c, ifu error info: 0x20000ffffffc0, ccu error info: 0xbc8d3d140100006b, cube error info: 0, biu error info: 0, aic error mask: 0x6500020bd00028c, para base: 0x12c100140080.[FUNC:ProcessStarsCoreErrorInfo][FILE:deviceerrorcoreproc.cc][LINE:333]
TraceBack (most recent call last):
The extend info: errcode:(0x800000, 0, 0) errorStr: The DDR address of the MTE instruction is out of range. fixperror0 info: 0x30d7e5c, fixperror1 info: 0x7d, fsmId:1, tslot:7, thread:0, ctxid:0, blk:23, sublk:1, subErrType:4.[FUNC:ProcessStarsCoreErrorInfo][FILE:deviceerrorcoreproc.cc][LINE:353]
Kernel task happen error, retCode=0x26, [aicore exception].[FUNC:PreCheckTaskErr][FILE:davincikerneltask.cc][LINE:1555]
AICORE Kernel task happen error, retCode=0x26.[FUNC:GetError][FILE:stream.cc][LINE:1191]
[AICINFO] after execute:args print end[FUNC:GetError][FILE:stream.cc][LINE:1191]
[AICINFO] after execute:mixCtx print end[FUNC:GetError][FILE:stream.cc][LINE:1191]
Aicore kernel execute failed, deviceid=0, streamid=47, reportstreamid=47, taskid=1, flipnum=0, fault kernelname=SubmSparseConv3dV3090bc8dcd5655023f20a6bb4bd48cbbd0mixaic, fault kernel info ext=none, program id=1, hash=7976997100038956638.[FUNC:GetError][FILE:stream.cc][LINE:1191]
rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:errormessagemanage.cc][LINE:53]
wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:loginner.cpp][LINE:162]
[W910 15:42:36.754254036 compilerdepend.ts:545] Warning: NPU warning, error code is 507015[Error]:
[Error]: The aicore execution is abnormal.
Rectify the fault based on the error information in the ascend log.
EH9999: Inner Error!
rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:errormessagemanage.cc][LINE:53]
EH9999[PID: 602138] 2026-09-10-15:42:36.302.716 (EH9999): wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:loginner.cpp][LINE:162]
TraceBack (most recent call last):
(function npuSynchronizeUsedDevices)
[W910 15:42:36.757000803 compilerdepend.ts:527] Warning: NPU warning, error code is 507015[Error]:
[Error]: The aicore execution is abnormal.
Rectify the fault based on the error information in the ascend log.
EH9999: Inner Error!
rtDeviceSynchronizeWithTimeout execute failed, reason=[aicore exception][FUNC:FuncErrorReason][FILE:errormessagemanage.cc][LINE:53]
EH9999[PID: 602138] 2026-09-10-15:42:36.306.467 (EH9999): wait for compute device to finish failed, runtime result = 507015.[FUNC:ReportCallError][FILE:loginner.cpp][LINE:162]
TraceBack (most recent call last):
(function npuSynchronizeDevice)

定位到具体代码报错位置为kernels\subm_sparse_conv3d\op_kernel\subm_sparse_conv3d_v3.cpp文件中的ComputeValidPositionOneMap函数,其中的DataCopyPad调用时,map1Offset会越界

    __aicore__ inline void ComputeValidPositionOneMap(const uint32_t &taskOffset, const uint32_t &taskCount) {
        Muls(spatial1Local_, spatial1Local_, spatialShape2_, taskCount);
        Add(spatial1Local_, spatial1Local_, spatial2Local_, taskCount);
        PipeBarrier<PIPE_ALL>();

        for (int16_t i = 0; i < taskCount; i++) {
            int32_t spatial0BaseIdx = spatial0Local_.GetValue(i);
            for (int16_t k0Idx = 0; k0Idx < k0_; k0Idx++) {
                int64_t map1Offset = batchIdxLocal_.GetValue(i) * totalSpatialShape_ + spatial1Local_.GetValue(i) +
                    (k0Idx + spatial0BaseIdx) * spatialShape1_times_2_;

                DataCopyPad(mapValLocal_[i * mapValBufSize_ + k0Idx * k1_ * k2Aligned_], map1GM_[map1Offset],
                    {static_cast<uint16_t>(k1_), static_cast<uint32_t>(k2_ * INT32_BYTE_SIZE),
                        static_cast<uint16_t>((spatialShape2_ - k2_) * INT32_BYTE_SIZE), 0, 0},
                    {true, 0, static_cast<uint8_t>(k2Aligned_ - k2_), -1});
            }
        }

        SetFlag<HardEvent::MTE2_V>(0);
        WaitFlag<HardEvent::MTE2_V>(0);
        WaitFlag<HardEvent::MTE3_V>(0);

        Gather(indicesOffsetLocal_, mapValLocal_, gatherOffsetLocal_, 0u, kernelSize_ * singleLoopTaskAligned_);
    }
likedislike
ascend-robotascend-robot成员
19 天前 添加了label:bug
mengdemengde
19 天前 修改了issue 的描述
xiangyuming
xiangyuming成员
19 天前 评论:

/label add triaged

likedislike
ascend-robotascend-robot成员
19 天前 添加了label:triaged
gitdzy
gitdzy成员
19 天前 评论:

你好,感谢你提出的问题,请问方便提供一下报错case 的 shape信息吗,主要包括 batchsize, spatial_shape,以及num_point。
可以在 <conda路径>/mx_driving/ops/sparse_functional.py的157行后面打印。num_point可以直接print(indices.shape[0])

likedislike
Qqiuqiu成员
19 天前 关联了pull request:docs: 增加sparse和subm系列算子的shape约束
gsoleil成员
11 天前 评论:

你好,因为长时间没有收到回复,缺乏输入数据无法复现当前问题,该issue会先关闭,后续补充输入可以重新提issue

likedislike
Ggsoleil成员
11 天前 issue状态由 TODO 改变为 DONE
Ggsoleil成员
11 天前 关闭了 issue
ascend-robotascend-robot成员
11 天前 添加了label:resolved