已关闭
[Bug-Report|缺陷反馈]: UBX部分算子爬坡性能不达预期 #663
zhaojiayu创建于 8月27日关闭于 29 天前
8月27日 关联了pull request:ubx perf bugfix
8月27日 将 yjzz1007 设为负责人
8月27日 添加了label:bug
Leewis
29 天前 评论:
29 天前 评论:
https://gitcode.com/cann/hccl/pull/2727
代码已修改上库,本问题闭环;
修改点:
- AllReduceAutoSelector::SelectMeshAlgoAicpuUBX 小数据判断调整:将 isClosNumMultipleOfMeshNum && !IsSmallData(dataSize) 改为 isClosNumMultipleOfMeshNum && dataSize > SMALL_COUNT_512KB,统一小数据量判定口径,影响是否选择 AicpuAllReducePipeLineUBX 等算法。
- ReduceScatterAutoSelector::SelectMeshAlgoAicpuForMesh1DClos 小数据判断调整:同样的条件由 !IsSmallData(dataSize) 改为 dataSize > SMALL_COUNT_512KB,影响 AicpuReduceScatterPipeLineUBX 等算法的选择。
- ReduceAutoSelector::SelectMeshAlgoAicpu 增加数据量分支:在非 64 位数据类型且非 HCCL_REDUCE_PROD 的分支中,按 dataSize 区分算法——大于 SMALL_COUNT_512KB 时选择 ReduceParallelMesh1DNHRUBX,否则选择 AicpuReduceSoleNHR。
- ReduceNHR::CalcRes 通道计算优化:当 topoInfo->level0Topo == Level0Shape::MESH_1D_CLOS 且 !topoInfo->level0PcieMix 时,改用 CalcChannelRequestNhrMultiJetty 计算通道并仅保留协议为 COMM_PROTOCOL_UBC_CTP 的通道加入 level1Channels,其余场景仍走原 CalcChannelRequestNhr 路径。


29 天前 添加了label:resolved
Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.
Describe the current behavior / 问题描述 (Mandatory / 必填)
部分UBX算子在单跑1G场景和爬坡场景性能不一致
Environment / 环境信息 (Mandatory / 必填)
950机型
Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)
爬坡必现
Describe the expected behavior / 预期结果 (Mandatory / 必填)
预期一致
Related log / screenshot / 日志 / 截图 (Mandatory / 必填)
爬坡性能低于单跑
Special notes for this issue/备注 (Optional / 选填)