已关闭
[Bug-Report|缺陷反馈]: A3 allreduce执行超时 #921
hu-yiliang11创建于  17 天前关闭于  17 天前
hu-yiliang11
17 天前 创建

Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.

Describe the current behavior / 问题描述 (Mandatory / 必填)

A3 大规模集群下,RDMA任务较多的情况下,每个WR生成cqe,打爆CQ队列,QP状态无效后,对端一直重传

Environment / 环境信息 (Mandatory / 必填)

A3 CANN master 跨超

Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)

选一个大的通信域执行allreduce,每个server内的卡数不一致的时候,会选择算法ringComm,rdma任务较多的时候可能会打爆CQ

Describe the expected behavior / 预期结果 (Mandatory / 必填)

allreduce正常执行无报错

[ERROR] KERNEL(8497,sklogd):2026-09-03-20:02:04.131.821 [klogd.c:247][31721263.474438] hns3 0002:71:00.0 hns_0: Local work queue 0xba9 catast error, sub_event type is: 244 SUBSYSTEM=pci DEVICE=+pci:0002:71:00.0

[ERROR] HCCL(8129,aicpu_scheduler):2026-09-03-20:12:19.584.385 [aicpu_hccl_sqcq.cc:313] [13622]Task {devId: 0 streamId: 658 sqId: 678 cqId: 319 type: 2} run failed of exception, idx:[0], info:[streamId :658 taskId :620 errorCode :261 errorType :32 sqeType :3 sqId :678 sqHead :620 matchFlag :0 dropFlag :0 errorBit :1 accError :0]
[ERROR] HCCL(8129,aicpu_scheduler):2026-09-03-20:12:19.598.002 [aicpu_communicator.cc:3519] [13622][TaskException]base information is streamId:1046, sqid:1078, head:537, tail:512, type:NOTIFY WAIT, localRank:397, remoteRank:398, taskId:8729, notifyId:6097, length:0, addr1High:0x0, addr1Low:0x0, addr2High:0x0, addr2Low:0x0.
[ERROR] HCCL(8129,aicpu_scheduler):2026-09-03-20:12:19.604.290 [aicpu_communicator.cc:4217] [13622][TaskException]opData information is tag:AllReduce_group_name_21103ringAllReduceComm_device, group:group_name_21103, isCustom:0, opLaunchIdx:1, opExecIdx:1, count:1, dataType:4, opType:2, rootId:4294967295, dstAddr:0x12c1c0c33c00, srcAddr:0x12c1c0c33c00.
[ERROR] HCCL(8129,aicpu_scheduler):2026-09-03-20:12:19.604.381 [aicpu_communicator.cc:4218] [13622][TaskException]task sequence is OP(1),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099),RS(398,6127),NW(398,6097),RS(396,6125),NW(396,6099)
[ERROR] HCCL(8129,aicpu_scheduler):2026-09-03-20:12:19.875.845 [aicpu_communicator.cc:4908] [13622]Exception happened, group base_emb_inner_15714, sqid 678, cqeStatus 1, sqetype 3, errorCode 5, head 620, tail 672

Special notes for this issue/备注 (Optional / 选填)

likedislike
LLeewis成员
17 天前 关联了看板:HCOMM
Leewis成员
17 天前 评论:

/assign @gcw_NcEfY7mt

likedislike
CANN-robotCANN-robot成员
17 天前 将 gcw_NcEfY7mt 设为负责人
Leewis成员
17 天前 评论:

解决方案:设置集合通信RDMA wr属性,datanotify和dataacknotify才生成cqe, 避免RDMA任务较多的情况下,每个WR生成cqe,QP状态无效后,对端一直重传,打爆CQ队列问题;
对应PR: https://gitcode.com/cann/hcomm/pull/5397

likedislike
LLeewis成员
17 天前 issue类型由 任务 改变为 缺陷
LLeewis成员
17 天前 添加了label:Bug-Report
LLeewis成员
17 天前 issue状态由 待办的 改变为 已完成
LLeewis成员
17 天前 关闭了 issue
CANN-robotCANN-robot成员
17 天前 添加了label:resolved
LLeewis成员
13 天前 移除了看板:HCOMM