已开启
【代码侦探Challenge05】实现 MulCustom 逐元素乘法算子 #3563
【代码侦探Challenge05】实现 MulCustom 逐元素乘法算子 #3563
已开启
FineL1创建于 29 天前
FineL1
FineL1
29 天前

本 PR 完成 Challenge05-MulCustom 算子开发,实现两个 float32 Tensor 的逐元素乘法:

z=x×yz = x \times y

实现内容

Kernel 实现

在 mul_custom.asc 中完成 Ascend C Kernel 开发:

  • 使用 8 个 AI Vector Core 并行处理输入数据;
  • 使用 GlobalTensor 管理 GM 中的输入输出;
  • 使用 LocalTensor 和 TQue 管理 UB 数据;
  • 设置 BUFFER_NUM = 2 实现双缓冲;
  • Kernel 采用 CopyIn → Compute → CopyOut 流程;
  • 使用 AscendC::DataCopy 完成 GM 与 UB 之间的数据搬运;
  • 使用 AscendC::Mul 完成逐元素乘法计算。

Host 侧调用

Host 侧通过 ACL Runtime 完成:

  • Device 初始化;
  • Host / Device 内存申请;
  • H2D 数据拷贝;
  • Ascend C Kernel 直调;
  • Stream 同步;
  • D2H 结果拷贝;
  • 资源释放。

改动类型 / Change Type

  • 新功能 / New Feature
  • Ascend C Kernel
  • Kernel 直调验证

测试信息 / Testing

实际运行结果:

Output: 2.76 2.76 2.76 2.76 ...
Golden: 2.76 2.76 2.76 2.76 ...
[Success] Case accuracy is verification passed.

3e47f051-fe66-4e44-92cc-457a9d75c343.png
验证结果:PASS

likedislike
合并受阻
FineL1FineL1
29 天前 创建了 pull request,commit 305240b5
atomgit-bot
atomgit-bot
29 天前 评论:

变更摘要

本 PR 完成 Challenge05-MulCustom 逐元素乘法算子开发,实现两个 float32 Tensor 的 z[i] = x[i] * y[i] 运算。核心代码位于新增的 mul_custom.asc 文件:Kernel 侧定义 KernelMul 类,按 CopyIn → Compute → CopyOut 流程使用 GlobalTensor 管理 GM 数据、LocalTensor 与 TQue 管理 UB 数据,通过 AscendC::DataCopy 完成数据搬运、AscendC::Mul 完成逐元素乘法;Host 侧通过 kernel_mul 封装基于 ACL Runtime 完成 Device 初始化、内存申请、H2D/D2H 拷贝、Kernel 直调、Stream 同步与资源释放,并提供 VerifyResult 与 main 做精度验证。

主要改动

  • Kernel 实现(mul_custom.asc 中的 KernelMul 类): 新增 Init/Process/CopyIn/Compute/CopyOut 方法,按 8 个 AI Vector Core 划分 blockLength,设置 BUFFER_NUM = 2 双缓冲(tileLength = blockLength / tileNum / BUFFER_NUM),循环 tileNum * BUFFER_NUM 次执行搬运、计算、搬出;Compute 中使用 AscendC::Mul 完成逐元素乘法。
  • Kernel 入口与 Tiling 数据结构: 新增 __global__ __aicore__ 入口函数 mul_custom(声明 KERNEL_TASK_TYPE_DEFAULT(KERNEL_TYPE_AIV_ONLY) 仅使用 AI Vector Core),并定义 MulCustomTilingData 结构体携带 totalLength 与 tileNum 两个 tiling 参数。
  • Host 侧 Kernel 直调封装(kernel_mul 函数): 新增基于 ACL Runtime 的完整调用流程,依次执行 aclInit、aclrtSetDevice、aclrtCreateStream、Host/Device 内存申请、aclrtMemcpy H2D 拷贝、mul_custom<<<blockDim, nullptr, stream>>> 直调、aclrtSynchronizeStream 同步、D2H 拷贝及内存与运行时资源释放,返回计算结果 vector。
  • 精度验证与主函数(VerifyResult 与 main): 新增打印 Output/Golden 前 20 个元素并逐元素比较的验证逻辑;main 中以 totalLength = 8 * 2048、常量 valueX = 1.2f、valueY = 2.3f 构造输入并计算期望结果 valueX * valueY 进行校验。
  • 构建运行脚本(run.sh): 新增编译运行脚本,激活 CANN 环境后执行 cmake、make 并运行 ./mul_test;同时提交了 .ipynb_checkpoints/ 下对应的 mul_custom-checkpoint.asc 与 run-checkpoint.sh 副本(内容与正式文件一致)。
likedislike
不准确?
atomgit-bot
atomgit-bot
29 天前 评论:

代码审查

✅ 未发现问题

likedislike
不准确?
CANN-robotCANN-robot成员
29 天前 添加了label:cann-cla/yes
CANN-robot
CANN-robot成员
29 天前 评论:

Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.


PR Approval Progress

⚠️ This PR does not yet meet the following requirements:lgtm (requires ≥ 2 person(s) per module)、approve (requires ≥ 1 person(s) per module)

Module Approval Details

module lgtm status approve status
repo-cann/cann-outreach ❌ (0/2)(You can also ask: 田晓亮, li-shengxian3, gcw_tgPdOLZT, shaoyf, 傅涛) ❌ (0/1)(You can also ask: jxlang, luzx66, Carolina_yuan, yanhf, li-shengxian3)

💡 Tip:

  • Committer can comment /approve or /lgtm
  • Commenting /approve implies both code review (lgtm) and intent to merge (approve)

CLA Signature Pass

FineL1, thanks for your pull request. All authors of the commits have signed the CLA. 👍

likedislike
FineL1FineL1
24 天前 关联了issue:【代码侦探Challenge05】实现 MulCustom 逐元素乘法算子