在提交新问题之前,请确保您已经在社区中搜索过相关问题,并使用了社区中提供的资源/工具后,仍未找到满意的解决方式。
⚠️ 安全信息提醒:请仔细检查提供的文本内容,确保其不包含敏感数据信息,包括但不限于:
在分享配置信息或代码示例时,请将敏感信息脱敏处理,或使用 <TOKEN> 等占位符替代原有内容。
<TOKEN>
A5机器 FrameworkPTAdapter FrameworkPTAdapter 26.1.0.B080 CANN 9.1.0.B050
背景:跑mindspeed训练调用到torch.gather报错
[rank2]: Traceback (most recent call last): [rank2]: File "/home/tqy/develop/RecSDK/benchmark/models/recsys_ranking/recsys-examples/examples/hstu/pretrain_gr_ranking.py", line 234, in <module> [rank2]: main() [rank2]: File "/home/tqy/develop/RecSDK/benchmark/models/recsys_ranking/recsys-examples/examples/hstu/pretrain_gr_ranking.py", line 222, in main [rank2]: train_with_pipeline( [rank2]: File "/home/tqy/develop/RecSDK/benchmark/models/recsys_ranking/recsys-examples/examples/hstu/training/training.py", line 286, in train_with_pipeline [rank2]: flops = cal_flops( [rank2]: ^^^^^^^^^^ [rank2]: File "/home/tqy/develop/RecSDK/benchmark/models/recsys_ranking/recsys-examples/examples/hstu/training/utils.py", line 140, in cal_flops [rank2]: torch.distributed.gather(seqlens_tensor, gathered_seqlens, dst=0) [rank2]: File "/usr/local/python3.11.0/lib/python3.11/site-packages/torch_npu/distributed/distributed_c10d.py", line 211, in _gather [rank2]: work = _group.gather(output_tensors, input_tensors, opts) [rank2]: ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ [rank2]: RuntimeError: ProcessGroupHCCL does not support gather [rank2]: [ERROR] 2026-06-10-12:10:24 (PID:716, Device:2, RankID:2) ERR02007 DIST feature not supported
该版本pta在A5机器上torch_npu.npu.use_compatible_impl()默认为True,走入cpp端的gather侧,在该处遇到A5机型判断报不支持报错 解决:gather的原实现是在python侧调用,所以在python增加机型判断,A5机器时无效化该接口,直接走原实现 复现用例:pytorch\test\distributed\test_gather.py
欢迎加入社区,感谢您对社区的贡献 🎉!
通过日志:8086828a13cb47a3819720b3a57fe169.log 原报错日志:988315632d3f44c48e0644e292d16a4b.log
在提交新问题之前,请确保您已经在社区中搜索过相关问题,并使用了社区中提供的资源/工具后,仍未找到满意的解决方式。
⚠️ 安全信息提醒:请仔细检查提供的文本内容,确保其不包含敏感数据信息,包括但不限于:
在分享配置信息或代码示例时,请将敏感信息脱敏处理,或使用
<TOKEN>等占位符替代原有内容。环境信息
A5机器
FrameworkPTAdapter FrameworkPTAdapter 26.1.0.B080
CANN 9.1.0.B050
🐛 问题描述
背景:跑mindspeed训练调用到torch.gather报错
该版本pta在A5机器上torch_npu.npu.use_compatible_impl()默认为True,走入cpp端的gather侧,在该处遇到A5机型判断报不支持报错
解决:gather的原实现是在python侧调用,所以在python增加机型判断,A5机器时无效化该接口,直接走原实现
复现用例:pytorch\test\distributed\test_gather.py
欢迎加入社区,感谢您对社区的贡献 🎉!