已合并
perf(sort): 优化归并、非末轴排序及索引输出 #5672
黄晓彬创建于 22 天前
perf(sort): 优化归并、非末轴排序及索引输出 #5672
已合并
黄晓彬创建于 22 天前
黄晓彬成员
22 天前

描述

优化 Sort 在多行归并、非末轴小轴及 UB 可容纳整行的大轴场景中的执行开销,保持稳定排序、升降序及 values/indices 语义。

  • 引入公共 ping-pong 归并,减少阶段间回拷;在适用路径批量归并多行,并优化多核末轮归并。
  • 按实际 UB 布局统一选择整行归并路径,优化非末轴批次规划、NDDMA 搬运与向量化索引输出。
  • 对 RegBase 平台排序轴后全为单例维的输入,通过 Reshape 消除冗余转置;其他输入沿用原路径。
  • 同步 TopKV2 依赖模板,补充 NaN、正负零、尾块及稳定索引看护。

关联的Issue

https://gitcode.com/cann/ops-math/issues/3230

测试

  • Ascend950PR_9579 :Sort kernel 395 条、正负零逐位比较 18 条、Sort ACLNN 40 条、TopKV2 受影响用例 177 条,合计 630 条精度全部通过,Sort 同时比较 values 和 indices。
  • 相对前一版已验证实现,设备耗时异常项复测后未确认稳定超过 5% 且超过 0.5 μs 的劣化
  • 二级冒烟通过

文档更新

不涉及

类型标签

likedislike
Pull Request已成功合入, 合并人@CANN-robot
(感谢 黄晓彬 的贡献)
黄黄晓彬成员
22 天前 创建了 pull request,commit 5f7ff43c
atomgit-bot
atomgit-bot
22 天前 评论:

变更摘要

本 PR 将累积的 ping-pong 归并优化适配到 master 分支,核心是新增统一的 UB 内 ping-pong 归并工具 PingPongMergeSortCommon::SortToProposal,用它替换 sort 算子各路径原先的 Concat + Sort 流程,并对 merge more-core 调度做重构:支持 direct(单轮直接并行)、sync-merge(keyParams1 逻辑块)与 multi-round(多轮)三种模式,同时新增 SORT_SCHID_12 直接调度内核实例。改动覆盖 tiling 侧(sort_tiling_common.cpp/.h、sort_tiling_arch35.cpp)与内核侧(merge_more_core_base.h、merge_intra_core_base.h、sort_merge_sort.h、non_last_small_axis_base.h、merge_sort_big_size.h),并配套更新了单测 test_sort_tiling.cpp。

主要改动

  • 新增 ping-pong 归并工具 ping_pong_merge_sort.h:新增 SortToProposal 与 MergeStage,先用 Sort32 生成 DEALING_SORT_NUM_ONCE 长度的有序 run,再在 ping/pong 两个 proposal buffer 间迭代做 4 路归并,返回最终结果所在 buffer,替代原先依赖 Concat 临时缓冲的排序流程。
  • 重构 merge more-core 调度:tiling 侧以 IsMergeMoreCoreProfitable、SelectMergeMoreCoreDataSize、SelectMergeSyncMergeBlockSize 取代 IsMergeMoreCoreSupported,ComputeMergeMoreCorePlan 按 direct → sync-block → multi-round 顺序选择,新增常量 MERGE_MORE_CORE_DATA_SIZE_BASE/LARGE 与 MERGE_MORE_CORE_MIN_CORES_PER_ROW_FOR_MULTI_ROUND;内核侧 merge_more_core_base.h 重写 Process() 支持多轮循环,新增 ProcessDirect()、SortLogicalBlock/MergeLogicalBlock、MergeAcrossCores,keyParams1 承载 sync-merge 逻辑块大小(0 表示普通 more-core),并新增 SORT_SCHID_12 直接调度模板参数。
  • 重写 sort_merge_sort.h 主归并内核:移除 Concat 相关缓冲,改用 SortToProposal 生成 proposal 后 Extract 输出 value/index;新增 ProcessDoubleBuffered 双缓冲流水(MTE2/MTE3 与 Vector 重叠),并为 bf16 增加 SortRowsBf16 在 cast buffer 内就地输出结果。
  • 增加单块快速路径与紧凑输出:merge_intra_core_base.h 在 blocksPerRow_ == 1 时跳过 merge 阶段(省去 concatTmpBuf_ 及 Phase2 UB);tiling 侧新增 ComputeMergeIntraCoreSingleBlockSortSize、ComputeMergeSortBatchCapacity,并为 FillMergeSortInfo/ComputeMergeSortTiling 增加 compactOutput 选项;non_last_small_axis_base.h 的 merge 路径改用 SortToProposal + Extract,并将 sortInput_ 复用为 value 输出缓冲。
  • 调整调度选择顺序与数据类型的阈值:sort_tiling_arch35.cpp 中 TryMerge 引入 fp16/bf16 的 merge axis 阈值(2048/1536),新增 PreferSortMergeMoreCore 决策函数,整体选择顺序改为 TryMerge → TryMergeIntraCore → TryRadixOneCore,并更新了 tiling 单测用例与期望 tiling key。
likedislike
不准确?
atomgit-bot
atomgit-bot
22 天前 评论:

🤖 AI 代码检视正在进行中,请稍候…

likedislike
不准确?
CANN-robotCANN-robot成员
22 天前 添加了label:cann-cla/yes
CANN-robot
CANN-robot成员
22 天前 评论:

Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.


PR Approval Progress

✅ Congratulations! All modules have met the lgtm and approve requirements.

Module Approval Details

module lgtm status approve status
math/sort ✅ yue-ma, 王瑞 (2/2) ✅ 王瑞, yue-ma (2/1)
repo-cann/ops-math ✅ 王瑞, yue-ma (2/2) ✅ 王瑞, yue-ma (2/1)

💡 Tip:

  • Committer can comment /approve or /lgtm
  • Commenting /approve implies both code review (lgtm) and intent to merge (approve)

CLA Signature Pass

ConanHuang, thanks for your pull request. All authors of the commits have signed the CLA. 👍

likedislike
此处折叠了172条消息 查看更多
RuiWang_成员
18 天前 评论:

/lgtm
/approve

likedislike
CANN-robotCANN-robot成员
18 天前 添加了label:approved
yue-ma成员
18 天前 评论:

/lgtm
/approve

likedislike
CANN-robotCANN-robot成员
18 天前 添加了label:lgtm
CANN-robotCANN-robot成员
18 天前 合入了pull request