已关闭
[Bug]: 分布式训练hccl 通信超时 #470
weixin_46701715创建于  6月29日关闭于  7月6日
weixin_46701715
weixin_46701715
6月29日 创建

在提交新问题之前,请确保您已经在社区中搜索过相关问题,并使用了社区中提供的资源/工具后,仍未找到满意的解决方式。

⚠️ 安全信息提醒:请仔细检查提供的文本内容,确保其不包含敏感数据信息,包括但不限于:

  • API 令牌或密钥
  • 密码或身份验证凭证
  • 私有网址或接口地址
  • 个人或机密数据
  • ...

在分享配置信息或代码示例时,请将敏感信息脱敏处理,或使用 <TOKEN> 等占位符替代原有内容。

环境信息

例如:
- 操作系统 openeuler
- 昇腾硬件信息 atlas A2训练服务器 910b
- CANN软件版本 cann9.0.0
- 安装的对应软件版本 

🐛 问题描述

分布式训练分别启动四个节点的训练任务后,主节点会提示超时,关键的报错信息和命令如下:

开始导入环境信息

[Rank 24 | Local Rank 0] 2026-06-29 15:39:34,528 INFO [torch_npu.env:15] => get env MASTER_ADDR = 172.18.0.31
[Rank 24 | Local Rank 0] 2026-06-29 15:39:34,529 INFO [torch_npu.env:15] => get env MASTER_PORT = 6000
[Rank 24 | Local Rank 0] 2026-06-29 15:39:34,529 INFO [torch_npu.env:15] => get env TORCHELASTIC_USE_AGENT_STORE = True

[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.202 [externalinput.cc:579] [13449][HCCL_ENV] HCCL_CONNECT_TIMEOUT set by environment to [600]s
[INFO] HCCL(13446,python3.10):2026-06-29-15:40:44.118.320 [externalinput.cc:579] [13446][HCCL_ENV] HCCL_CONNECT_TIMEOUT set by environment to [600]s
[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.468 [externalinput.cc:1003] [13449][HCCL_ENV] HCCL_EXEC_TIMEOUT set by environment to [600.00]s
[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.490 [externalinput.cc:646] [13449][HCCL_ENV] HCCL_INTRA_PCIE_ENABLE set by default to [1], HCCL_INTRA_ROCE_ENABLE set by default to [0]
[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.500 [externalinput.cc:725] [13449][HCCL_ENV] environmental variable PROFILING_MODE and GE profiling option is not set, default: false
[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.506 [externalinput.cc:819] [13449][HCCL_ENV] HCCL_WHITELIST_DISABLE set by default to [1]
[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.516 [externalinput.cc:861] [13449][HCCL_ENV] HCCL_IF_IP is set to [172.18.0.37], ip[172.18.0.37].
[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.523 [externalinput.cc:917] [13449][HCCL_ENV] HCCL_SOCKET_IFNAME set by environment to [bond1]
[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.527 [externalinput.cc:886] [13449][HCCL_ENV] HCCL_SOCKET_FAMILY is not set and is used by default [AF_INET]
[INFO] HCCL(13449,python3.10):2026-06-29-15:40:44.118.533 [externalinput.cc:848] [13449][HCCL_ENV] HCCL_IF_BASE_PORT set by default to [60000]
[INFO] HCCL(13445,python3.10):2026-06-29-15:40:44.118.557 [externalinput.cc:579] [13445][HCCL_ENV] HCCL_CONNECT_TIMEOUT set by environment to [600]s

这里开始不知道为什么开始监听这个还有一些没见过的 ip 地址和端口,确定不是所在服务器网卡绑定的 ip

[INFO] HCCL(13445,python3.10):2026-06-29-15:40:45.537.412 [network_manager.cc:687] [13445][Start][Vnic]Listen on ip[192.168.2.198], port[16666] success, devPhyId[1], devLogicId[1], isAutoPort[0]

[INFO] HCCP(13449,python3.10):2026-06-29-15:40:45.857.565 [ra_host.c:704]tid:13449,RaSocketBatchConnect : Input parameters: [0]th, phyId[5], localIp[10.52.211.11], remoteIp[10.52.210.192], port[16666], tag[group_name_0_res_optimize_6_Inter__Inter_MultiSocket_29_0], cnt[0]

这里的 ip 跟我预设的节点间通信的 ip 都不一样,也不是服务器网卡绑定的 ip

子节点这里开始出现报错信息:

[INFO] RUNTIME(13445,python3.10):2026-06-29-15:49:45.344.436 [context.cc:304] 15010 TryToRecycleProgPool: recycle prog pool
[INFO] RUNTIME(13445,python3.10):2026-06-29-15:49:45.344.641 [raw_device.cc:392] 15010 TryToRecycleModulesPool: recycle raw dev pool
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.095 [hccl_socket_manager.cc:797] [15688][Wait][LinkEstablish]wait socket establish timeout, role[1] rank[24] timeout[600 s]
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.119 [hccl_socket_manager.cc:861] [15688][Wait][LinksEstablishCompleted] is failed. ret[9].
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.121 [hccl_socket_manager.cc:797] [15689][Wait][LinkEstablish]wait socket establish timeout, role[1] rank[24] timeout[600 s]
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.133 [hccl_socket_manager.cc:623] [15688]   _________________________LINK_ERROR_INFO___________________________
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.133 [hccl_socket_manager.cc:861] [15689][Wait][LinksEstablishCompleted] is failed. ret[9].
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.138 [hccl_socket_manager.cc:624] [15688]   |  comm error, device[0]
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.140 [hccl_socket_manager.cc:623] [15689]   _________________________LINK_ERROR_INFO___________________________
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.141 [hccl_socket_manager.cc:626] [15688]   |  dest_ip(user_rank)  |   dest_port   |  src_ip(user_rank)   |   src_port   |   MyRole   |   Status   |    TlsStatus   |
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.144 [hccl_socket_manager.cc:624] [15689]   |  comm error, device[0]
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.145 [hccl_socket_manager.cc:628] [15688]   |----------------------|---------------|----------------------|--------------|------------|------------|----------------|
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.147 [hccl_socket_manager.cc:626] [15689]   |  dest_ip(user_rank)  |   dest_port   |  src_ip(user_rank)   |   src_port   |   MyRole   |   Status   |    TlsStatus   |
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.151 [hccl_socket_manager.cc:628] [15689]   |----------------------|---------------|----------------------|--------------|------------|------------|----------------|
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.170 [hccl_socket_manager.cc:583] [15689]   |  10.52.210.99(16)   |  16666  |   10.52.211.6(24)   |  0  |  client  | time out |   UNKNOWN  | LinkInfo
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.181 [hccl_socket_manager.cc:836] [15689][Create][Sockets]Wait links establish completed failed, local role is client. ret[9]
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.186 [transport_manager.cc:1414] [15689][SetMachinePara]call trace: hcclRet -> 9
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.200 [transport_manager.cc:1264] [15689][CreateLink][InitChannelStage][Timeout]SetMachinePara error.
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.171 [hccl_socket_manager.cc:583] [15688]   |  10.52.210.187(8)   |  16666  |   10.52.211.6(24)   |  0  |  client  | time out |   UNKNOWN  | LinkInfo
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.218 [hccl_socket_manager.cc:836] [15688][Create][Sockets]Wait links establish completed failed, local role is client. ret[9]
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.224 [transport_manager.cc:1414] [15688][SetMachinePara]call trace: hcclRet -> 9
[ERROR] HCCL(13444,python3.10):2026-06-29-15:50:45.853.230 [transport_manager.cc:1264] [15688][CreateLink][InitChannelStage][Timeout]SetMachinePara error.
[ERROR] HCCL(13448,python3.10):2026-06-29-15:50:45.857.597 [hccl_socket_manager.cc:797] [15692][Wait][LinkEstablish]wait socket establish timeout, role[1] rank[28] timeout[600 s]
[ERROR] HCCL(13448,python3.10):2026-06-29-15:50:45.857.623 [hccl_socket_manager.cc:797] [15696][Wait][LinkEstablish]wait socket establish timeout, role[1] rank[28] timeout[600 s]
[ERROR] HCCL(13448,python3.10):2026-06-29-15:50:45.857.627 [hccl_socket_manager.cc:861] [15692][Wait][LinksEstablishCompleted] is failed. ret[9].
[ERROR] HCCL(13448,python3.10):2026-06-29-15:50:45.857.634 [hccl_socket_manager.cc:861] [15696][Wait][LinksEstablishCompleted] is failed. ret[9].
[ERROR] HCCL(13448,python3.10):2026-06-29-15:50:45.857.646 [hccl_socket_manager.cc:623] [15692]   _________________________LINK_ERROR_INFO___________________________
[ERROR] HCCL(13448,python3.10):2026-06-29-15:50:45.857.648 [hccl_socket_manager.cc:623] [15696]   _________________________LINK_ERROR_INFO___________________________
[ERROR] HCCL(13448,python3.10):2026-06-29-15:50:45.857.650 [hccl_socket_manager.cc:624] [15692]   |  comm error, device[4]
[ERROR] HCCL(13448,python3.10):2026-06-29-15:50:45.857.652 [hccl_socket_manager.cc:624] [15696]   |  comm error, device[4]
[ERROR] HCCL(13445,python3.10):2026-06-29-15:50:45.857.644 [hccl_socket_manager.cc:797] [15690][Wait][LinkEstablish]wait socket establish timeout, role[1] rank[25] timeout[600 s]
[INFO] RUNTIME(13444,python3.10):2026-06-29-15:51:57.793.855 [api_impl.cc:7480] 13444 PeekLastErr: level=0 err=0.
[ERROR] HCCL(13448,python3.10):2026-06-29-15:51:57.797.726 [detect_connect_anomalies.cc:127] [15692]-------------------CONNECT TIMEOUT DETECT RESULT-----------------------
[ERROR] HCCL(13449,python3.10):2026-06-29-15:51:57.797.744 [detect_connect_anomalies.cc:127] [15697]-------------------CONNECT TIMEOUT DETECT RESULT-----------------------
[ERROR] HCCL(13448,python3.10):2026-06-29-15:51:57.797.757 [detect_connect_anomalies.cc:132] [15692]This node (server 172.18.0.37, device ID 4) detects that srcRank (server 172.18.0.37, device ID 4) fails to connect to dstRank (server 172.18.0.23, device ID 4). Continue to analyze the fault based on the logs of srcRank and dstRank.
[ERROR] HCCL(13445,python3.10):2026-06-29-15:51:57.797.754 [detect_connect_anomalies.cc:127] [15691]-------------------CONNECT TIMEOUT DETECT RESULT-----------------------
[ERROR] HCCL(13448,python3.10):2026-06-29-15:51:57.797.766 [detect_connect_anomalies.cc:132] [15692]This node (server 172.18.0.37, device ID 4) detects that srcRank (server 172.18.0.37, device ID 4) fails to connect to dstRank (server 172.18.0.18, device ID 4). Continue to analyze the fault based on the logs of srcRank and dstRank.
[ERROR] HCCL(13449,python3.10):2026-06-29-15:51:57.797.766 [detect_connect_anomalies.cc:132] [15697]This node (server 172.18.0.37, device ID 5) detects that srcRank (server 172.18.0.37, device ID 5) fails to connect to dstRank (server 172.18.0.23, device ID 5). Continue to analyze the fault based on the logs of srcRank and dstRank.
[ERROR] HCCL(13449,python3.10):2026-06-29-15:51:57.797.775 [detect_connect_anomalies.cc:132] [15697]This node (server 172.18.0.37, device ID 5) detects that srcRank (server 172.18.0.37, device ID 5) fails to connect to dstRank (server 172.18.0.18, device ID 5). Continue to analyze the fault based on the logs of srcRank and dstRank.
[ERROR] HCCL(13445,python3.10):2026-06-29-15:51:57.797.776 [detect_connect_anomalies.cc:132] [15691]This node (server 172.18.0.37, device ID 1) detects that srcRank (server 172.18.0.37, device ID 1) fails to connect to dstRank (server 172.18.0.23, device ID 1). Continue to analyze the fault based on the logs of srcRank and dstRank.
[ERROR] HCCL(13445,python3.10):2026-06-29-15:51:57.797.785 [detect_connect_anomalies.cc:132] [15691]This node (server 172.18.0.37, device ID 1) detects that srcRank (server 172.18.0.37, device ID 1) fails to connect to dstRank (server 172.18.0.18, device ID 1). Continue to analyze the fault based on the logs of srcRank and dstRank.

最后的报错信息如下:

[rank29]: Traceback (most recent call last):
[rank29]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/train/trainer.py", line 432, in <module>
[rank29]:     trainer = Trainer(args=arguments)
[rank29]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/train/trainer.py", line 75, in __init__
[rank29]:     dataloader_result = self.get_dataloader() if dataloader_provider is None else dataloader_provider(args)
[rank29]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/train/trainer.py", line 340, in get_dataloader
[rank29]:     datasets = build_mm_dataset(data_config.dataset_param)
[rank29]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/data/__init__.py", line 27, in build_mm_dataset
[rank29]:     return dataset_cls_or_func(
[rank29]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/data/datasets/huggingface/qwen2vl_dataset.py", line 206, in get_qwen2vl_dataset
[rank29]:     with TrainingArguments(output_dir='./').main_process_first(desc="pre-process dataset"):
[rank29]:   File "/mnt/shangyingda/miniforge3/envs/torch/lib/python3.10/contextlib.py", line 135, in __enter__
[rank29]:     return next(self.gen)
[rank29]:   File "/mnt/shangyingda/miniforge3/envs/torch/lib/python3.10/site-packages/transformers/training_args.py", line 2034, in main_process_first
[rank29]:     dist.barrier()
[rank29]:   File "/mnt/shangyingda/miniforge3/envs/torch/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 81, in wrapper
[rank29]:     return func(*args, **kwargs)
[rank29]:   File "/mnt/shangyingda/miniforge3/envs/torch/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 4635, in barrier
[rank29]:     work = group.barrier(opts=opts)
[rank29]: RuntimeError: InnerRunOpApi:../torch_npu/csrc/framework/OpParamMaker.cpp:286 OPS function error: HcclAllreduce, error code is 9
[rank29]: [ERROR] 2026-06-29-15:51:57 (PID:13449, Device:5, RankID:29) ERR01100 OPS call acl api failed.
[rank29]: [PID: 13449] 2026-06-29-15:51:57.797.841 Communication_Error_Get_Socket(EI0006): Getting socket times out. Reason:
[rank29]: This node (server 172.18.0.37, device ID 5) detects that srcRank (server 172.18.0.37, device ID 5) fails to connect to dstRank (server 172.18.0.23, device ID 5). Continue to analyze the fault based on the logs of srcRank and dstRank.

主节点此时处于等待状态,最后报错:

[INFO] RUNTIME(13991,python3.10):2026-06-29-15:46:07.137.698 [stream.cc:1604] 13991 SynchronizeExecutedTask: report three minutes timeout! stream_id=3, task_id=2, pendingNum=1.
  [WARNING] HCCL(13991,python3.10):2026-06-29-15:50:07.187.715 [heartbeat.cc:356] [16340]establish rank[172.18.0.31/0] to rank[172.18.0.23/0] heartbeat connection failed. Reason: get rasocket timeout,timeout[600s], the HCCL_CONNECT_TIMEOUT may be insufficient. Group[group_name_0].
  [ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.573.111 [stars_engine.cc:1548]15684 ProcLogicCqReport:Task run failed, device_id=2, stream_id=6, task_id=1, sqe_type=0(ffts), errType=0x20(sq sw status error), sqSwStatus=0x111
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.573.224 [device_error_proc.cc:1499]15684 PrintStreamTimeoutSnapshotInfo:stream_id=3, task_id=42, taskType=3 (EVENT_WAIT), EventId=2.
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.574.452 [device_error_proc.cc:1899]15684 ProcessStarsCoreTimeoutDfxInfo:The error from device(chipId:2, dieId:0), serial number is 1, aicore task timeout dfx, falut_stream_id=6, falut_task_id=1, falut_slot_id=0, timeout and own_bitmap=0
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.574.461 [device_error_proc.cc:1919]15684 ProcessStarsCoreTimeoutDfxInfo:task kernel is null, stream_id=6, task_id=1
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.574.469 [device_error_proc.cc:1788]15684 ProcessStarsHcclFftsPlusTimeoutErrorInfo:The error from device(chipId:2, dieId:0), serial number is 2, hccl fftsplus task timeout occurred during task execution, stream_id:6, sq_id:6, task_id:1, stuck notify num:1, timeout=600s.
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.574.556 [device_error_proc.cc:1795]15684 ProcessStarsHcclFftsPlusTimeoutErrorInfo:The 0 stuck notify wait context info:(context_id=21, notify_id=46).
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.574.579 [ffts_task.cc:823]15684 SetStarsResultForFftsPlusTask:FftsPlusTask errorCode=273, logicCq:err=32, errCode=273, stream_id=6, task_id=1
[INFO] RUNTIME(13993,python3.10):2026-06-29-15:50:23.574.584 [runtime.cc:5725] 15684 SetWatchDogDevStatus: There is errInfo of devId=2, tsId=0
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.575.525 [ffts_task.cc:780]15684 DoCompleteSuccForFftsPlusTask:fftsplus report error, retCode=0x111, [fftsplus task timeout].
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.575.557 [ffts_task.cc:754]15684 PrintErrorInfoForFftsPlusTask:fftsplus task execute failed, dev_id=2, stream_id=6, task_id=1, context_id=21, thread_id=0, err_type=13[hccl fftsplus timeout]
[ERROR] RUNTIME(13993,python3.10):2026-06-29-15:50:23.575.574 [ffts_task.cc:733]15684 TaskFailCallBackForFftsPlusTask:fftsplus streamId=6, taskId=1, context_id=21, expandType=1, rtCode=0x715006b,[fftsplus task timeout], psStart=0x0, kernel_name=not found kernel name, binHandle=(nil), binSize=0.
[INFO] IDEDD(13993,python3.10):2026-06-29-15:50:23.575.598 [dump_manager.cpp:45][tid:15684] An exception callback message is received.
[ERROR] RUNTIME(13992,python3.10):2026-06-29-15:50:23.577.322 [stars_engine.cc:1548]15669 ProcLogicCqReport:Task run failed, device_id=1, stream_id=6, task_id=1, sqe_type=0(ffts), errType=0x20(sq sw status error), sqSwStatus=0x111
[ERROR] RUNTIME(13992,python3.10):2026-06-29-15:50:23.577.409 [device_error_proc.cc:1499]15669 PrintStreamTimeoutSnapshotInfo:stream_id=3, task_id=42, taskType=3 (EVENT_WAIT), EventId=2.
[ERROR] RUNTIME(13992,python3.10):2026-06-29-15:50:23.578.718 [device_error_proc.cc:1899]15669 ProcessStarsCoreTimeoutDfxInfo:The error from device(chipId:1, dieId:0), serial number is 1, aicore task timeout dfx, falut_stream_id=6, falut_task_id=1, falut_slot_id=0, timeout and own_bitmap=0
  [rank2]: Traceback (most recent call last):
[rank2]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/train/trainer.py", line 432, in <module>
[rank2]:     trainer = Trainer(args=arguments)
[rank2]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/train/trainer.py", line 75, in __init__
[rank2]:     dataloader_result = self.get_dataloader() if dataloader_provider is None else dataloader_provider(args)
[rank2]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/train/trainer.py", line 340, in get_dataloader
[rank2]:     datasets = build_mm_dataset(data_config.dataset_param)
[rank2]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/data/__init__.py", line 27, in build_mm_dataset
[rank2]:     return dataset_cls_or_func(
[rank2]:   File "/mnt/shangyingda/MindSpeed-MM/mindspeed_mm/fsdp/data/datasets/huggingface/qwen2vl_dataset.py", line 206, in get_qwen2vl_dataset
[rank2]:     with TrainingArguments(output_dir='./').main_process_first(desc="pre-process dataset"):
[rank2]:   File "/mnt/shangyingda/miniforge3/envs/torch/lib/python3.10/contextlib.py", line 135, in __enter__
[rank2]:     return next(self.gen)
[rank2]:   File "/mnt/shangyingda/miniforge3/envs/torch/lib/python3.10/site-packages/transformers/training_args.py", line 2034, in main_process_first
[rank2]:     dist.barrier()
[rank2]:   File "/mnt/shangyingda/miniforge3/envs/torch/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 81, in wrapper
[rank2]:     return func(*args, **kwargs)
[rank2]:   File "/mnt/shangyingda/miniforge3/envs/torch/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 4640, in barrier
[rank2]:     work.wait()
[rank2]: RuntimeError: npuSynchronizeDevice:../torch_npu/csrc/core/npu/NPUStream.cpp:576 NPU function error: AclrtSynchronizeDeviceWithTimeout, error code is 507048
[rank2]: [ERROR] 2026-06-29-15:50:23 (PID:13993, Device:2, RankID:2) ERR00100 PTA call acl api failed
[rank2]: [Error]: The execution of the internal task times out.
  [rank2]:         Rectify the fault based on the error information in the ascend log.
[rank2]: [PID: 13993] 2026-06-29-15:50:23.582.873 Communication_Error_Timeout(EI0002): An timeout occurs when the Notify register waits for execution. Waiting peer end: unknown; task information: streamID:[6], taskID[1], tag[AllReduce_group_name_0], AlgType(level 0-1-2):[fullmesh-NHR-NHR].; communication operator information: ; communicator: group:[group_name_0], user define information[Unspecified], rankSize[32], rankId[2].
[rank2]:         Possible Cause: 1. An exception occurs during the execution on some NPUs in the cluster. As a result, collective communication operation failed.2. The execution speed on some NPU in the cluster is too slow to complete a communication operation within the timeout interval. (The default timeout interval is 1800s, You can set the interval by using HCCL_EXEC_TIMEOUT.)3. The number of training samples of each NPU is inconsistent.4. Packet loss or other connectivity problems occur on the communication link.
[rank2]:         Solution: 1. If this error is reported on part of these ranks, check other ranks to see whether other errors have been reported earlier.2. If this error is reported for all ranks, check whether the error reporting time is consistent (the maximum difference must not exceed 1800s). If not, locate the cause or set the HCCL_EXEC_TIMEOUT environment variable to a larger value. 3. Ensure that the number of training samples of each NPU is consistent. 4. Check whether the completion queue element (CQE) of the error exists in the plog(grep -rn 'error cqe'). If so, check the network connection status. For details about the troubleshooting method, search for the keyword "EI0002" on https://www.hiascend.com/en/document/.
[rank2]:         TraceBack (most recent call last):
[rank2]:         The error from device(chipId:2, dieId:0), serial number is 2, hccl fftsplus task timeout occurred during task execution, stream_id:6, sq_id:6, task_id:1, stuck notify num:1, timeout=600s.[FUNC:ProcessStarsHcclFftsPlusTimeoutErrorInfo][FILE:device_error_proc.cc][LINE:1788]
[rank2]:         The 0 stuck notify wait context info:(context_id=21, notify_id=46).[FUNC:ProcessStarsHcclFftsPlusTimeoutErrorInfo][FILE:device_error_proc.cc][LINE:1795]
[rank2]:         An timeout occurs when the Notify register waits for execution. Waiting peer end: unknown; task information: streamID:[6], taskID[1], tag[AllReduce_group_name_0], AlgType(level 0-1-2):[fullmesh-NHR-NHR].; communication operator information: ; communicator: group:[group_name_0], user define information[Unspecified], rankSize[32], rankId[2].
[rank2]:         rtDeviceSynchronizeWithTimeout execution failed, reason=fftsplus timeout[FUNC:FuncErrorReason][FILE:error_message_manage.cc][LINE:65]
[rank2]:         wait for compute device to finish failed, runtime result = 507048.[FUNC:ReportCallError][FILE:log_inner.cpp][LINE:148]

已经尝试过,服务器间的通信正常,端口监听也显示几个ip 处于使用状态

lsof -i:6000
COMMAND     PID USER   FD   TYPE DEVICE SIZE/OFF NODE NAME
pt_elasti 13223 root   19u  IPv6   1691      0t0  TCP *:x11 (LISTEN)
pt_elasti 13223 root   20u  IPv6   1692      0t0  TCP bms-caiwujihe-910b-001:56440->bms-caiwujihe-910b-001:x11 (ESTABLISHED)
pt_elasti 13223 root   21u  IPv6  29809      0t0  TCP bms-caiwujihe-910b-001:x11->bms-caiwujihe-910b-001:56440 (ESTABLISHED)
pt_elasti 13223 root   22u  IPv6 133563      0t0  TCP bms-caiwujihe-910b-001:x11->172.18.0.18:59910 (ESTABLISHED)
pt_elasti 13223 root   23u  IPv6 127504      0t0  TCP bms-caiwujihe-910b-001:x11->172.18.0.37:48970 (ESTABLISHED)
pt_elasti 13223 root   24u  IPv6 150945      0t0  TCP bms-caiwujihe-910b-001:x11->172.18.0.23:36418 (ESTABLISHED)
pt_elasti 13223 root   25u  IPv6 150994      0t0  TCP bms-caiwujihe-910b-001:x11->172.18.0.37:42190 (ESTABLISHED

子节点访问主节点的端口, telnet 端口连接正常:

 telnet 172.18.0.31 6000
Trying 172.18.0.31...
Connected to 172.18.0.31.
Escape character is '^]'.

请问该如何继续排查问题,还是说我 hccl 中哪里安装或者配置有问题

欢迎加入社区,感谢您对社区的贡献 🎉!

likedislike
weixin_46701715weixin_46701715
6月29日 修改了issue 的描述
weixin_46701715weixin_46701715
6月29日 修改了issue 的描述
weixin_46701715
weixin_46701715
6月29日 评论:

启动命令如下:

source /usr/local/Ascend/cann/set_env.sh
export HCCL_EXEC_TIMEOUT=600
export HCCL_CONNECT_TIMEOUT=600
export NON_MEGATRON=true
export MULTI_STREAM_MEMORY_REUSE=2
export TASK_QUEUE_ENABLE=2
export ASCEND_LAUNCH_BLOCKING=1
export ACLNN_CACHE_LIMIT=100000
export CPU_AFFINITY_CONF=1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export HCCL_IF_IP=172.18.0.31
export ASCEND_SLOG_PRINT_TO_STDOUT=1  # 日志输出到终端 

NPUS_PER_NODE=8
MASTER_ADDR=172.18.0.31
MASTER_PORT=6000
NNODES=4
NODE_RANK=0
WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES))
NODE_ADDR=172.18.0.31

DISTRIBUTED_ARGS="
    --nproc_per_node $NPUS_PER_NODE \
    --nnodes $NNODES \
    --node_rank $NODE_RANK \
    --master_addr $MASTER_ADDR \
    --master_port $MASTER_PORT
"

logfile=$(date +%Y%m%d)_$(date +%H%M%S)
mkdir -p logs
torchrun $DISTRIBUTED_ARGS mindspeed_mm/fsdp/train/trainer.py \
    examples/qwen3_5/qwen3_5_35B_config1.yaml \
    2>&1 | tee logs/train_${logfile}.log
    ```
likedislike
weixin_46701715weixin_46701715
6月29日 修改了issue 的描述
xubin成员
6月29日 评论:

可以试试如下方法
删除export HCCL_IF_IP=172.18.0.31
在ifconfig中查找host ip绑在了哪个网卡上,然后导入如下环境变量
export HCCL_SOCKET_IFNAME=网卡名字
export GLOO_SOCKET_IFNAME=网卡名字

likedislike
weixin_46701715
weixin_46701715
6月29日 评论:

可以试试如下方法
删除export HCCL_IF_IP=172.18.0.31
在ifconfig中查找host ip绑在了哪个网卡上,然后导入如下环境变量
export HCCL_SOCKET_IFNAME=网卡名字
export GLOO_SOCKET_IFNAME=网卡名字

@MoCuishle-M

bond1: flags=5187<UP,BROADCAST,RUNNING,MASTER,MULTICAST>  mtu 1500
        inet 172.18.0.37  netmask 255.255.252.0  broadcast 172.18.3.255
        inet6 fe80::c250:64ff:fead:cb2e  prefixlen 64  scopeid 0x20<link>
        ether c0:50:64:ad:cb:2e  txqueuelen 1000  (Ethernet)
        RX packets 2437963  bytes 2680980065 (2.4 GiB)
        RX errors 0  dropped 0  overruns 0  frame 0
        TX packets 931507  bytes 217369185 (207.2 MiB)
        TX errors 0  dropped 6 overruns 0  carrier 0  collisions 0
source /usr/local/Ascend/cann/set_env.sh
export HCCL_EXEC_TIMEOUT=600
export HCCL_CONNECT_TIMEOUT=600
export NON_MEGATRON=true
export MULTI_STREAM_MEMORY_REUSE=2
export TASK_QUEUE_ENABLE=2
export ASCEND_LAUNCH_BLOCKING=1
export ACLNN_CACHE_LIMIT=100000
export CPU_AFFINITY_CONF=1
export HCCL_SOCKET_IFNAME=bond1
export GLOO_SOCKET_IFNAME=bond1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export ASCEND_SLOG_PRINT_TO_STDOUT=1  # 日志输出到终端 

NPUS_PER_NODE=8
MASTER_ADDR=172.18.0.31
MASTER_PORT=6000
NNODES=4
NODE_RANK=0
WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES))
NODE_ADDR=172.18.0.31

DISTRIBUTED_ARGS="
   --nproc_per_node $NPUS_PER_NODE \
   --nnodes $NNODES \
   --node_rank $NODE_RANK \
   --master_addr $MASTER_ADDR \
   --master_port $MASTER_PORT
"

您好,还是同样的结果,仍然通信超时

likedislike
weixin_46701715
weixin_46701715
6月30日 评论:
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.094 [stream.cc:1472]662555 GetError:Stream Synchronize failed, stream_id=6, retCode=0x111, [fftsplus task timeout].
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.101 [stream.cc:3697]662555 EnterFailureAbort:stream_id=6 enter failure abort.
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.121 [stars_engine.cc:1439]662555 StarsResumeRtsq:stop scheduling in abort failure mode: stream_id=6, sq_id=6, sq_head=1, task_id=1, taskType=52.
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.177 [stream.cc:1613]661185 SynchronizeExecutedTask:context is abort, status=0x715006b.
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.191 [stream.cc:1667]661185 SynchronizeImpl:failed, stream_id=3, error=0x715006b
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.200 [context.cc:569]661185 SyncStreamsWithTimeout:Synchronize stream fail, stream_id=3, errorCode=0x715006b.
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.221 [api_error.cc:2698]661185 DeviceSynchronize:Device synchronize failed.
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.239 [api_c_device.cc:245]661185 rtDeviceSynchronizeWithTimeout:ErrCode=507048, desc=[fftsplus timeout], InnerCode=0x715006b
[ERROR] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.246 [error_message_manage.cc:65]661185 FuncErrorReason:rtDeviceSynchronizeWithTimeout execution failed, reason=fftsplus timeout
[ERROR] ASCENDCL(661185,python3.10):2026-06-30-10:49:34.647.302 [device.cpp:198]661185 aclrtSynchronizeDeviceWithTimeoutImpl:wait for compute device to finish failed, runtime result = 507048.
[INFO] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.362 [api_impl.cc:7480] 661185 PeekLastErr: level=0 err=507048.
[INFO] RUNTIME(661185,python3.10):2026-06-30-10:49:34.647.517 [api_impl.cc:7480] 661185 PeekLastErr: level=0 err=507048.

这里显示错误代码507048,应该是内部任务执行超时,这是为什么呢?

likedislike
xubin成员
6月30日 评论:

可以试试如下方法
删除export HCCL_IF_IP=172.18.0.31
在ifconfig中查找host ip绑在了哪个网卡上,然后导入如下环境变量
export HCCL_SOCKET_IFNAME=网卡名字
export GLOO_SOCKET_IFNAME=网卡名字

@MoCuishle-M

bond1: flags=5187<UP,BROADCAST,RUNNING,MASTER,MULTICAST>  mtu 1500
       inet 172.18.0.37  netmask 255.255.252.0  broadcast 172.18.3.255
       inet6 fe80::c250:64ff:fead:cb2e  prefixlen 64  scopeid 0x20<link>
       ether c0:50:64:ad:cb:2e  txqueuelen 1000  (Ethernet)
       RX packets 2437963  bytes 2680980065 (2.4 GiB)
       RX errors 0  dropped 0  overruns 0  frame 0
       TX packets 931507  bytes 217369185 (207.2 MiB)
       TX errors 0  dropped 6 overruns 0  carrier 0  collisions 0
source /usr/local/Ascend/cann/set_env.sh
export HCCL_EXEC_TIMEOUT=600
export HCCL_CONNECT_TIMEOUT=600
export NON_MEGATRON=true
export MULTI_STREAM_MEMORY_REUSE=2
export TASK_QUEUE_ENABLE=2
export ASCEND_LAUNCH_BLOCKING=1
export ACLNN_CACHE_LIMIT=100000
export CPU_AFFINITY_CONF=1
export HCCL_SOCKET_IFNAME=bond1
export GLOO_SOCKET_IFNAME=bond1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export ASCEND_SLOG_PRINT_TO_STDOUT=1  # 日志输出到终端 

NPUS_PER_NODE=8
MASTER_ADDR=172.18.0.31
MASTER_PORT=6000
NNODES=4
NODE_RANK=0
WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES))
NODE_ADDR=172.18.0.31

DISTRIBUTED_ARGS="
   --nproc_per_node $NPUS_PER_NODE \
   --nnodes $NNODES \
   --node_rank $NODE_RANK \
   --master_addr $MASTER_ADDR \
   --master_port $MASTER_PORT
"

您好,还是同样的结果,仍然通信超时

@weixin_46701715
四台机器都是这样写的吗?还是只有一台服务器是这样写的,这只是一个示例?四台服务器对应IP的网卡是不是都叫bond1。

likedislike
weixin_46701715
weixin_46701715
6月30日 评论:

可以试试如下方法
删除export HCCL_IF_IP=172.18.0.31
在ifconfig中查找host ip绑在了哪个网卡上,然后导入如下环境变量
export HCCL_SOCKET_IFNAME=网卡名字
export GLOO_SOCKET_IFNAME=网卡名字

@MoCuishle-M

bond1: flags=5187<UP,BROADCAST,RUNNING,MASTER,MULTICAST>  mtu 1500
       inet 172.18.0.37  netmask 255.255.252.0  broadcast 172.18.3.255
       inet6 fe80::c250:64ff:fead:cb2e  prefixlen 64  scopeid 0x20<link>
       ether c0:50:64:ad:cb:2e  txqueuelen 1000  (Ethernet)
       RX packets 2437963  bytes 2680980065 (2.4 GiB)
       RX errors 0  dropped 0  overruns 0  frame 0
       TX packets 931507  bytes 217369185 (207.2 MiB)
       TX errors 0  dropped 6 overruns 0  carrier 0  collisions 0
source /usr/local/Ascend/cann/set_env.sh
export HCCL_EXEC_TIMEOUT=600
export HCCL_CONNECT_TIMEOUT=600
export NON_MEGATRON=true
export MULTI_STREAM_MEMORY_REUSE=2
export TASK_QUEUE_ENABLE=2
export ASCEND_LAUNCH_BLOCKING=1
export ACLNN_CACHE_LIMIT=100000
export CPU_AFFINITY_CONF=1
export HCCL_SOCKET_IFNAME=bond1
export GLOO_SOCKET_IFNAME=bond1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export ASCEND_SLOG_PRINT_TO_STDOUT=1  # 日志输出到终端 

NPUS_PER_NODE=8
MASTER_ADDR=172.18.0.31
MASTER_PORT=6000
NNODES=4
NODE_RANK=0
WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES))
NODE_ADDR=172.18.0.31

DISTRIBUTED_ARGS="
   --nproc_per_node $NPUS_PER_NODE \
   --nnodes $NNODES \
   --node_rank $NODE_RANK \
   --master_addr $MASTER_ADDR \
   --master_port $MASTER_PORT
"

您好,还是同样的结果,仍然通信超时

@weixin_46701715
四台机器都是这样写的吗?还是只有一台服务器是这样写的,这只是一个示例?四台服务器对应IP的网卡是不是都叫bond1。

@MoCuishle-M
四台机器的网卡都是 bond1,只有 bond1对应的网卡在集群通信网段

likedislike
xubin成员
7月2日 评论:

可以试试如下方法
删除export HCCL_IF_IP=172.18.0.31
在ifconfig中查找host ip绑在了哪个网卡上,然后导入如下环境变量
export HCCL_SOCKET_IFNAME=网卡名字
export GLOO_SOCKET_IFNAME=网卡名字

@MoCuishle-M

bond1: flags=5187<UP,BROADCAST,RUNNING,MASTER,MULTICAST>  mtu 1500
       inet 172.18.0.37  netmask 255.255.252.0  broadcast 172.18.3.255
       inet6 fe80::c250:64ff:fead:cb2e  prefixlen 64  scopeid 0x20<link>
       ether c0:50:64:ad:cb:2e  txqueuelen 1000  (Ethernet)
       RX packets 2437963  bytes 2680980065 (2.4 GiB)
       RX errors 0  dropped 0  overruns 0  frame 0
       TX packets 931507  bytes 217369185 (207.2 MiB)
       TX errors 0  dropped 6 overruns 0  carrier 0  collisions 0
source /usr/local/Ascend/cann/set_env.sh
export HCCL_EXEC_TIMEOUT=600
export HCCL_CONNECT_TIMEOUT=600
export NON_MEGATRON=true
export MULTI_STREAM_MEMORY_REUSE=2
export TASK_QUEUE_ENABLE=2
export ASCEND_LAUNCH_BLOCKING=1
export ACLNN_CACHE_LIMIT=100000
export CPU_AFFINITY_CONF=1
export HCCL_SOCKET_IFNAME=bond1
export GLOO_SOCKET_IFNAME=bond1
export PYTORCH_NPU_ALLOC_CONF=expandable_segments:True
export ASCEND_SLOG_PRINT_TO_STDOUT=1  # 日志输出到终端 

NPUS_PER_NODE=8
MASTER_ADDR=172.18.0.31
MASTER_PORT=6000
NNODES=4
NODE_RANK=0
WORLD_SIZE=$(($NPUS_PER_NODE*$NNODES))
NODE_ADDR=172.18.0.31

DISTRIBUTED_ARGS="
   --nproc_per_node $NPUS_PER_NODE \
   --nnodes $NNODES \
   --node_rank $NODE_RANK \
   --master_addr $MASTER_ADDR \
   --master_port $MASTER_PORT
"

您好,还是同样的结果,仍然通信超时

@weixin_46701715
四台机器都是这样写的吗?还是只有一台服务器是这样写的,这只是一个示例?四台服务器对应IP的网卡是不是都叫bond1。

@MoCuishle-M
四台机器的网卡都是 bond1,只有 bond1对应的网卡在集群通信网段

@weixin_46701715

那不同机器上NODE_RANK应该都改了吧

likedislike
yaoyaoxuyaoyaoxu成员
7月2日 将 MoCuishle-M 设为负责人
ascend-robotascend-robot成员
7月3日 关联了看板:MindStudio ISSUE管理
xubin成员
7月6日 评论:

因为长时间未回复,先关闭此issue。如果后续有什么问题,欢迎提新issue或者再次开启此issue。

likedislike
Xxubin成员
7月6日 issue状态由 TODO 改变为 DONE
Xxubin成员
7月6日 关闭了 issue
ascend-robotascend-robot成员
7月6日 添加了label:resolved
weixin_46701715
weixin_46701715
7月9日 评论:

底层驱动问题,已自行解决

likedislike