已关闭
[Requirement|需求建议]: 统一 QBMM Tensor API 单 batch 内核并优化 Ascend 950 tiling #5624
ddssz创建于  22 天前关闭于  6 天前
ddssz
ddssz成员
22 天前 创建

Thanks for sending an requirement! Please fill in the following template to help quickly solve your problem.

Backgroud(背景信息)

Ascend 950 的 QuantBatchMatmulV3 与 QuantBatchMatmulInplaceAdd 在 Cube、MIX 和 MX 场景中存在分散的有 batch/无 batch Tensor API 内核入口及 tiling 数据路径。

需要统一相关实现,并完善非 MX Cube 单 batch Tensor API 分发及 Adaptive Sliding Window tiling 策略。

Origin(信息来源)

代码提交:f9ffee7bd46711d6a9016b5a566f06d1081a5a7e
提交者:smdbha
具体提出部门/团队:待补充。

Benefit / Necessity (价值/作用)

  • 减少 CMCT、Blaze 及独立无 batch 内核入口之间的重复实现。
  • 统一 QuantBatchMatmulV3 与 QuantBatchMatmulInplaceAdd 的单 batch tiling 数据路径。
  • 根据 L1 容量、MTE2/计算耗时估算及负载均衡结果选择 tiling 策略。
  • 具体性能收益待补充性能测试数据。

Design(设计方案)

  1. 将 Cube、MIX、MX 的无 batch 实现并入对应 Tensor API 内核头文件。
  2. 为 Ascend 950 非 MX Cube 场景增加紧凑的无 batch tiling 数据和 tiling key 分发。
  3. 为 Cube tiling 增加 L1 2/3/4 buffer 选择、stepK 调整和 A 全载判定。
  4. 扩展基础块负载均衡并归一化单轮尾块切分。
  5. 更新相关 Host/Kernel UT 数据及预期 tiling key。

关联 PR:https://gitcode.com/cann/ops-nn/pull/9970

likedislike
ddsszddssz成员
22 天前 添加了label:requirement
ddsszddssz成员
22 天前 将 smdbha 设为负责人
CANN-robotCANN-robot成员
6 天前 关闭了 issue
CANN-robotCANN-robot成员
5 天前 添加了label:resolved