合并受阻
变更摘要
本 PR 完成 Challenge05-MulCustom 逐元素乘法算子开发,实现两个 float32 Tensor 的 z[i] = x[i] * y[i] 运算。核心代码位于新增的 mul_custom.asc 文件:Kernel 侧定义 KernelMul 类,按 CopyIn → Compute → CopyOut 流程使用 GlobalTensor 管理 GM 数据、LocalTensor 与 TQue 管理 UB 数据,通过 AscendC::DataCopy 完成数据搬运、AscendC::Mul 完成逐元素乘法;Host 侧通过 kernel_mul 封装基于 ACL Runtime 完成 Device 初始化、内存申请、H2D/D2H 拷贝、Kernel 直调、Stream 同步与资源释放,并提供 VerifyResult 与 main 做精度验证。
主要改动
- Kernel 实现(
mul_custom.asc中的KernelMul类): 新增Init/Process/CopyIn/Compute/CopyOut方法,按 8 个 AI Vector Core 划分blockLength,设置BUFFER_NUM = 2双缓冲(tileLength = blockLength / tileNum / BUFFER_NUM),循环tileNum * BUFFER_NUM次执行搬运、计算、搬出;Compute中使用AscendC::Mul完成逐元素乘法。 - Kernel 入口与 Tiling 数据结构: 新增
__global__ __aicore__入口函数mul_custom(声明KERNEL_TASK_TYPE_DEFAULT(KERNEL_TYPE_AIV_ONLY)仅使用 AI Vector Core),并定义MulCustomTilingData结构体携带totalLength与tileNum两个 tiling 参数。 - Host 侧 Kernel 直调封装(
kernel_mul函数): 新增基于 ACL Runtime 的完整调用流程,依次执行aclInit、aclrtSetDevice、aclrtCreateStream、Host/Device 内存申请、aclrtMemcpyH2D 拷贝、mul_custom<<<blockDim, nullptr, stream>>>直调、aclrtSynchronizeStream同步、D2H 拷贝及内存与运行时资源释放,返回计算结果 vector。 - 精度验证与主函数(
VerifyResult与main): 新增打印 Output/Golden 前 20 个元素并逐元素比较的验证逻辑;main中以totalLength = 8 * 2048、常量valueX = 1.2f、valueY = 2.3f构造输入并计算期望结果valueX * valueY进行校验。 - 构建运行脚本(
run.sh): 新增编译运行脚本,激活 CANN 环境后执行cmake、make并运行./mul_test;同时提交了.ipynb_checkpoints/下对应的mul_custom-checkpoint.asc与run-checkpoint.sh副本(内容与正式文件一致)。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.
PR Approval Progress
⚠️ This PR does not yet meet the following requirements:lgtm (requires ≥ 2 person(s) per module)、approve (requires ≥ 1 person(s) per module)
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| repo-cann/cann-outreach | ❌ (0/2)(You can also ask: 田晓亮, li-shengxian3, gcw_tgPdOLZT, shaoyf, 傅涛) | ❌ (0/1)(You can also ask: jxlang, luzx66, Carolina_yuan, yanhf, li-shengxian3) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
FineL1, thanks for your pull request. All authors of the commits have signed the CLA. 👍


本 PR 完成 Challenge05-MulCustom 算子开发,实现两个 float32 Tensor 的逐元素乘法:
z=x×y
实现内容
Kernel 实现
在
mul_custom.asc中完成 Ascend C Kernel 开发:GlobalTensor管理 GM 中的输入输出;LocalTensor和TQue管理 UB 数据;BUFFER_NUM = 2实现双缓冲;CopyIn → Compute → CopyOut流程;AscendC::DataCopy完成 GM 与 UB 之间的数据搬运;AscendC::Mul完成逐元素乘法计算。Host 侧调用
Host 侧通过 ACL Runtime 完成:
改动类型 / Change Type
测试信息 / Testing
实际运行结果:
验证结果:
PASS