Pull Request已成功合入, 合并人@CANN-robot
(感谢 ddssz 的贡献)变更摘要
本 PR 统一 Ascend 950(DAV_3510)上 QuantBatchMatmulV3 与 QuantBatchMatmulInplaceAdd 的 Tensor API 单 batch(无 batch)内核入口与 tiling 数据路径:将有 batch/无 batch 的 Cube、MIX、MX Tensor API 入口合并进对应的 Blaze 内核头文件,删除独立的无 batch 文件与 CMCT 内核;QuantBatchMatmulInplaceAdd 侧将双 tiling 数据结构收敛为统一的 QbmmiaWithoutBatchTilingData;同时为非 MX(HiFloat8)Cube 路径补齐无 batch Tensor API 的 tiling 数据、tiling key 与内核分发,并对 Adaptive Sliding Window Cube tiling 的 L1 buffer 选择、stepK、A 全载策略及基础块负载均衡进行调优。
主要改动
- Tensor API 内核入口合并与 CMCT 清理:将
QbmmMixWithoutBatchTensorApiKernel、QbmmMxWithoutBatchTensorApiKernel分别并入qbmm_mix_tensor_api_blaze.h、qbmm_mx_tensor_api_blaze.h并删除对应独立头文件;quant_batch_matmul_inplace_add.cpp改用 Blaze Tensor API 分发(QbmmiaCubeWithoutBatchTensorApiKernel),删除qbmmia_cube_basic_api_cmct.h、qbmmia_mx_basic_api_cmct.h等 CMCT 实现及旧IS_BLAZE条件宏,统一走TPL_*_WITH_MMAPI_WITHOUT_BATCH两条 kernel type。 Qbmmiatiling 数据结构与生成路径统一:QuantBatchMatmulInplaceAddTilingData/QuantBatchMatmulInplaceAddTensorAPIWithoutBatchTilingData更名为QbmmiaTilingData/QbmmiaWithoutBatchTilingData;Cube/MX Basic API tiling 类移除tilingData_引用与SetWithoutBatchTilingData分支,改用IsTensorApiEnabled()返回 true,并新增CopyV3WithoutBatchTilingData从 V3 基类的withoutBatchTilingData_拷贝数据,使 Tensor API 路径始终产出无 batch tiling 数据。- 非 MX Cube 无 batch Tensor API 路径补齐:
AdaptiveSlidingWindowCubeBasicAPITiling新增withoutBatchTilingData_/useWithoutBatchTilingData_/SetWithoutBatchTilingData/GetBatchMode/CalcBasicBlock,并在quant_batch_matmul_v3.cpp增加QUANT_BMMV3_CUBE_WITHOUT_BATCH_TENSOR_API_IMPL_CLASS分发;tiling key 头文件新增SUPPORT_CUBE_WITHOUT_BATCH_TILING_KEY(HiFloat8 Cube ND/WeightNz)并将实例化条件扩展为SUPPORT_NO_VEC_WITHOUT_BATCH_TILING_KEY。 - Adaptive Sliding Window Cube tiling 调优:以
CalculateNBufferNum()取代CalculateNBufferNum4Cube(),按 L1 占用估算选择 2/3/4 buffer,并引入CanReduceStepKToTwo(stepK 归约到 2)、IsCubeMte2Bound/EstimateCubeMte2TimeUs/EstimateCubeTimeUs(带宽利用率 0.9)等启发式;UpdateAFullLoadStatus改为基于CanOpenMultiBufferByL1Estimate与重复 A 读取占比(REPEAT_A_LOAD_RATIO_THRESHOLD = 0.20)决策;新增NormalizeSingleRoundTailSplitBasicBlock归一化单轮尾块,OptimizeBaseBlockForLoadBalance扩展支持新增的BaseBlockMode::CUBE_BASIC模式,常量MXFP8_CUBE_MACS_PER_CYCLE更名为通用的B8_CUBE_MACS_PER_CYCLE。 - UT 与预期数据同步更新:Host/Kernel UT 移除仅 MX 场景的跳过逻辑,统一断言
QbmmiaWithoutBatchTilingData大小及 m/n/k 字段,test_quant_batch_matmul_inplace_add_apt_hif8.cpp改用TestOneParamCase950复用公共用例,tiling key 预期值同步调整。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| matmul | ✅ 王子韬, 杨阳 (2/2) | ✅ 王子韬, 杨阳 (2/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
smdbha, thanks for your pull request. All authors of the commits have signed the CLA. 👍


/lgtm
/approve


Pull Request 已合并或已关闭。
If you want to solve this problem, you can click here to do it in the FAQs.


描述
统一 Ascend 950 上 QuantBatchMatmulV3 与 QuantBatchMatmulInplaceAdd 的 Tensor API 单 batch 内核及 tiling 数据路径:
关联的Issue
关联 Issue:#5624
测试
git diff --check upstream/master f9ffee7bd46711d6a9016b5a566f06d1081a5a7e通过。文档更新
无。
类型标签
AI/Agent生成声明