已合并
refactor(matmul): unify Tensor API kernels and tune QBMM tiling #9970
ddssz创建于 9月7日
refactor(matmul): unify Tensor API kernels and tune QBMM tiling #9970
已合并
ddssz创建于 9月7日
ddssz
ddssz成员
9月7日

描述

统一 Ascend 950 上 QuantBatchMatmulV3 与 QuantBatchMatmulInplaceAdd 的 Tensor API 单 batch 内核及 tiling 数据路径:

  • 将 Cube、MIX 和 MX 的有 batch/无 batch Tensor API 入口合并到对应 Blaze 内核头文件,清理重复的 CMCT 与独立无 batch实现。
  • 为非 MX Cube 路径增加无 batch Tensor API tiling 数据、tiling key 和内核分发。
  • 优化 Adaptive Sliding Window Cube tiling,包括基础块负载均衡、单轮尾块归一化、L1 2/3/4 buffer 选择、stepK 调整以及 A 全载策略。
  • 同步更新 QuantBatchMatmulV3 和 QuantBatchMatmulInplaceAdd 的 Host/Kernel UT 数据及预期 tiling key。

关联的Issue

关联 Issue:#5624

测试

  • git diff --check upstream/master f9ffee7bd46711d6a9016b5a566f06d1081a5a7e 通过。
  • 提交中已更新相关 Host/Kernel UT 用例及预期数据。

文档更新

无。

类型标签

AI/Agent生成声明

likedislike
Pull Request已成功合入, 合并人@CANN-robot
(感谢 ddssz 的贡献)
ddsszddssz成员
9月7日 创建了 pull request,commit d38a7025
atomgit-bot
atomgit-bot
9月7日 评论:

变更摘要

本 PR 统一 Ascend 950(DAV_3510)上 QuantBatchMatmulV3 与 QuantBatchMatmulInplaceAdd 的 Tensor API 单 batch(无 batch)内核入口与 tiling 数据路径:将有 batch/无 batch 的 Cube、MIX、MX Tensor API 入口合并进对应的 Blaze 内核头文件,删除独立的无 batch 文件与 CMCT 内核;QuantBatchMatmulInplaceAdd 侧将双 tiling 数据结构收敛为统一的 QbmmiaWithoutBatchTilingData;同时为非 MX(HiFloat8)Cube 路径补齐无 batch Tensor API 的 tiling 数据、tiling key 与内核分发,并对 Adaptive Sliding Window Cube tiling 的 L1 buffer 选择、stepK、A 全载策略及基础块负载均衡进行调优。

主要改动

  • Tensor API 内核入口合并与 CMCT 清理:将 QbmmMixWithoutBatchTensorApiKernel、QbmmMxWithoutBatchTensorApiKernel 分别并入 qbmm_mix_tensor_api_blaze.h、qbmm_mx_tensor_api_blaze.h 并删除对应独立头文件;quant_batch_matmul_inplace_add.cpp 改用 Blaze Tensor API 分发(QbmmiaCubeWithoutBatchTensorApiKernel),删除 qbmmia_cube_basic_api_cmct.h、qbmmia_mx_basic_api_cmct.h 等 CMCT 实现及旧 IS_BLAZE 条件宏,统一走 TPL_*_WITH_MMAPI_WITHOUT_BATCH 两条 kernel type。
  • Qbmmia tiling 数据结构与生成路径统一:QuantBatchMatmulInplaceAddTilingData/QuantBatchMatmulInplaceAddTensorAPIWithoutBatchTilingData 更名为 QbmmiaTilingData/QbmmiaWithoutBatchTilingData;Cube/MX Basic API tiling 类移除 tilingData_ 引用与 SetWithoutBatchTilingData 分支,改用 IsTensorApiEnabled() 返回 true,并新增 CopyV3WithoutBatchTilingData 从 V3 基类的 withoutBatchTilingData_ 拷贝数据,使 Tensor API 路径始终产出无 batch tiling 数据。
  • 非 MX Cube 无 batch Tensor API 路径补齐:AdaptiveSlidingWindowCubeBasicAPITiling 新增 withoutBatchTilingData_/useWithoutBatchTilingData_/SetWithoutBatchTilingData/GetBatchMode/CalcBasicBlock,并在 quant_batch_matmul_v3.cpp 增加 QUANT_BMMV3_CUBE_WITHOUT_BATCH_TENSOR_API_IMPL_CLASS 分发;tiling key 头文件新增 SUPPORT_CUBE_WITHOUT_BATCH_TILING_KEY(HiFloat8 Cube ND/WeightNz)并将实例化条件扩展为 SUPPORT_NO_VEC_WITHOUT_BATCH_TILING_KEY。
  • Adaptive Sliding Window Cube tiling 调优:以 CalculateNBufferNum() 取代 CalculateNBufferNum4Cube(),按 L1 占用估算选择 2/3/4 buffer,并引入 CanReduceStepKToTwo(stepK 归约到 2)、IsCubeMte2Bound/EstimateCubeMte2TimeUs/EstimateCubeTimeUs(带宽利用率 0.9)等启发式;UpdateAFullLoadStatus 改为基于 CanOpenMultiBufferByL1Estimate 与重复 A 读取占比(REPEAT_A_LOAD_RATIO_THRESHOLD = 0.20)决策;新增 NormalizeSingleRoundTailSplitBasicBlock 归一化单轮尾块,OptimizeBaseBlockForLoadBalance 扩展支持新增的 BaseBlockMode::CUBE_BASIC 模式,常量 MXFP8_CUBE_MACS_PER_CYCLE 更名为通用的 B8_CUBE_MACS_PER_CYCLE。
  • UT 与预期数据同步更新:Host/Kernel UT 移除仅 MX 场景的跳过逻辑,统一断言 QbmmiaWithoutBatchTilingData 大小及 m/n/k 字段,test_quant_batch_matmul_inplace_add_apt_hif8.cpp 改用 TestOneParamCase950 复用公共用例,tiling key 预期值同步调整。
likedislike
不准确?
atomgit-bot
atomgit-bot
9月7日 评论:

代码审查

✅ 未发现问题

likedislike
不准确?
CANN-robotCANN-robot成员
9月7日 添加了label:cann-cla/yes
CANN-robot
CANN-robot成员
9月7日 评论:

Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.


PR Approval Progress

✅ Congratulations! All modules have met the lgtm and approve requirements.

Module Approval Details

module lgtm status approve status
matmul ✅ 王子韬, 杨阳 (2/2) ✅ 王子韬, 杨阳 (2/1)

💡 Tip:

  • Committer can comment /approve or /lgtm
  • Commenting /approve implies both code review (lgtm) and intent to merge (approve)

CLA Signature Pass

smdbha, thanks for your pull request. All authors of the commits have signed the CLA. 👍

likedislike
此处折叠了212条消息 查看更多
wangzitao
wangzitao成员
16 天前 评论:

/lgtm
/approve

likedislike
CANN-robotCANN-robot成员
16 天前 添加了label:lgtm
CANN-robotCANN-robot成员
16 天前 关闭了关联的issue
CANN-robotCANN-robot成员
16 天前 合入了pull request
CANN-robot
CANN-robot成员
16 天前 评论:

Pull Request 已合并或已关闭。

If you want to solve this problem, you can click here to do it in the FAQs.

likedislike