| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
docs目录基本概念md文件名改为英文 Co-authored-by: chenjiao<chenjiao31@huawei.com> # message auto-generated for no-merge-commit merge: !4182 merge master into master docs目录基本概念md文件名改为英文 Created-by: gitcode-chenjiao Commit-by: chenjiao Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> docs目录基本概念md文件名改为英文: 避免link中的中文字符引发的跳转异常,例如两段式接口.md变成%E4%B8%A4%E6%AE%B5%E5%BC%8F%E6%8E%A5%E5%8F%A3.md,不易于维护,可能导致其他平台跳转有问题。 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。--> <!-- 如果这个PR是为了解决特定的问题单,请在这里描述问题单单号。--> [#2283](https://gitcode.com/cann/ops-math/issues/2283) ## 测试 <!--描述进行了哪些测试来验证你的改动。包括但不限于二级冒烟、算子泛化等。--> ok ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> docs/zh/context所有md和对应的link ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x ] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-math!4182 | 2 个月前 | |
refactor: 统一 experimental 目录代码风格(全仓 clang-format) Co-authored-by: songkai111<songkai16@huawei.com> # message auto-generated for no-merge-commit merge: !3382 merge master into master refactor: 统一 experimental 目录代码风格(全仓 clang-format) Created-by: songkai111 Commit-by: songkai111 Merged-by: cann-robot Description: ## 描述 对 experimental/ 目录下的算子源码统一执行 clang-format 格式化,使全仓代码风格与根目录 .clang-format 配置保持一致。 本次改动范围: - 共 1279 个文件,涵盖 experimental/math(1151)与 experimental/conversion(128)两大目录。 - 涉及代码层级:op_kernel(490)、op_host(375)、op_api(200)、examples(171)、tests(156)、op_graph(37)。 - 文件类型:.cpp(834)、.h(445)。 - 涉及硬件分支目录:arch35(48)、arch32(2)。 改动内容为纯格式化(缩进、空格、对齐、单行语句拆分为多行、指针 * 贴合类型等),不涉及任何逻辑修改、新增文件或删除文件,代码语义保持不变。 ## 关联的Issue #1986 ## 测试 clang-format 格式化不改变代码语义,无需新增功能测试;格式化后的算子代码在编译与既有单测/二级冒烟用例下保持原有行为。 ## 文档更新 无。本次仅对源码做格式化,未修改任何文档文件。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:代码格式化(clang-format 全仓风格统一) See merge request: cann/ops-math!3382 | 3 个月前 | |
rsqrt add support int datatype in A2/A3 Co-authored-by: Nice_try<nicetryzzw@163.com> # message auto-generated for no-merge-commit merge: !3863 merge submit-InplaceSqrt into master 【CANN开源开放社区任务】【社区任务】InplaceRsqrt 新增整型数据类型支持 Created-by: Nice_try Commit-by: Nice_try Merged-by: cann-robot Description: ## 描述 ### 改动背景 Rsqrt/InplaceRsqrt 算子原仅支持浮点类型(float16、float、bfloat16),本次 PR 将其数据类型支持扩展至全部整型和布尔类型(int8、int16、int32、uint8、bool),对齐 AscendC 算子的通用数据覆盖范围。 ### 改动内容 #### 1. 算子定义层(op_def / op_proto) - **op_host/rsqrt_def.cpp**:Input/Output 的 DataType 列表新增 DT_INT8、DT_INT16、DT_INT32、DT_UINT8、DT_BOOL - **op_graph/rsqrt_proto.h**:INPUT/OUTPUT 的 TensorType 列表新增上述整型和布尔类型 #### 2. 算子 Tiling 层(op_host/rsqrt_tiling.cpp) - 新增 DATA_NUM_INT32 = 2、DATA_NUM_INT16 = 2、DATA_NUM_INT8 = 6、DATA_NUM_UINT8 = 4、DATA_NUM_BOOL = 2 五个常量,用于各数据类型在 UB 空间中的 Buffer 分配计算 - 在 switch 分支中新增对应数据类型到 ubDataNumber 的映射 - tiling 函数已扩展覆盖全部新增类型,无需额外修改注册逻辑 #### 3. 算子 ACLNN API 层(op_api/aclnn_rsqrt.cpp) - DTYPE_SUPPORT_LIST 移除旧版的独立 DTYPE_OUT_LIST(仅浮点),统一为一个支持全部 8 种类型的列表 - 移除 CheckDtypeValid 中不必要的 OUTPUT 类型检查重复逻辑 - 移除多余的 Cast 中间转换序列(由 rsqrt.cpp 内部统一处理) - **新增关键校验**:CheckDtypeValid 增加 self->GetDataType() != out->GetDataType() 检查,确保 self 和 out 的数据类型必须一致,避免 ViewCopy 在 dtype 不匹配时出现未定义行为 #### 4. 算子 Kernel 层(op_kernel/rsqrt.h / rsqrt.cpp) - rsqrt_tiling_data.h:RsqrtTilingData 结构体将 tile/tail 相关字段从 uint32_t 升级为 uint64_t,避免大 shape 下溢出 - rsqrt.h:KernelRsqrt 模板类新增 ComputeImpl 分派函数,按 TYPE_X 类型分发至以下计算函数: - **int32**:Cast 到 float → Rsqrt → Cast(CAST_RINT) 写回 int32 - **int16**:Cast 到 half → Rsqrt → ComputeIntCorrection(Mins/Adds/Maxs/Muls/Sub 校正)→ Cast(CAST_RINT) 写回 int16 - **int8**:Cast 到 half → Rsqrt → ComputeIntCorrection → Cast(CAST_RINT) 写回 int8 - **uint8**:Cast 到 half → Rsqrt → Cast(CAST_RINT) 写回 uint8 - **bool**:直接 Duplicate(0x0101) 填充 int16_t 模式,每个输入元素输出 true - ComputeIntCorrection 为 int16/int8 类型提供分段线性校正:将 [0, 2] 范围的 Rsqrt 结果通过 min(2)→add(-1)→max(0)→mul(3)→sub 的校正公式压缩到正确范围 - rsqrt.cpp:扩展 REGISTER_TILING_DEFAULT 和 DTYPE_X 宏路径,单入口 rsqrt<schMode> 模板支持 schMode=0(双缓冲)和 schMode=1(单缓冲) #### 5. 测试层(tests/ut) **op_api 测试**(test_aclnn_rsqrt.cpp): - 16 个测试用例,覆盖 int16、float16、bf16、float、int32、uint8 等类型的正常入参 - 覆盖不同 shape、空 tensor、非连续、低维/高维、inplace 等场景 - 覆盖 dtype 不匹配(ACLNN_ERR_PARAM_INVALID)、shape 不一致、nullptr 等错误路径 **op_host 测试**: - test_rsqrt_infershape.cpp:3 个测试用例覆盖 float、float16、bf16 的 shape 推导 - test_rsqrt_tiling.cpp:5 个测试用例覆盖不同 shape(8x8、8x2048、1023x2047)和数据类型(float、float16、bf16)的 tiling 参数计算 **op_kernel 测试**: - test_rsqrt.cpp:float32 单缓冲(单流水)测试 + float32 双缓冲(双流水)测试,均检查 system() 返回值 - test_rsqrt_int8.cpp:int8 类型 kernel 端到端测试 - test_rsqrt_int16.cpp:int16 类型 kernel 端到端测试 - test_rsqrt_int32.cpp:int32 类型 kernel 端到端测试 - test_rsqrt_uint8.cpp:uint8 类型 kernel 端到端测试 - test_rsqrt_bool.cpp:bool 类型 kernel 端到端测试 - 使用 rsqrt_test_entry.h 共享入口模板(宏拼接生成唯一函数名 rsqrt_int8、rsqrt_int16 等),避免多 DTYPE_X 文件的 ODR 冲突 - gen_data.py / compare_data.py:同时支持 8 种数据类型的 golden 数据生成与结果比对,整型使用 float32 精度计算对齐 AscendC Rsqrt 结果 ## 关联的Issue [#2156](https://gitcode.com/cann/ops-math/issues/2156) ## 测试 - **UT 测试**(全部通过): - --opapi:16/16 ✅ - --ophost:8/8 ✅ - --opkernel:7/7 ✅ - **AscendOpTest 测试**:已通过算子泛化测试 - **测试覆盖的数据类型**:float16、float、bfloat16、int8、int16、int32、uint8、bool ## 文档更新 - 更新了 README.md 文件,补充新增的整型/布尔类型支持说明 - 更新了 docs/aclnnRsqrt&aclnnInplaceRsqrt.md 文件 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-math!3863 | 2 个月前 | |
rsqrt add support int datatype in A2/A3 Co-authored-by: Nice_try<nicetryzzw@163.com> # message auto-generated for no-merge-commit merge: !3863 merge submit-InplaceSqrt into master 【CANN开源开放社区任务】【社区任务】InplaceRsqrt 新增整型数据类型支持 Created-by: Nice_try Commit-by: Nice_try Merged-by: cann-robot Description: ## 描述 ### 改动背景 Rsqrt/InplaceRsqrt 算子原仅支持浮点类型(float16、float、bfloat16),本次 PR 将其数据类型支持扩展至全部整型和布尔类型(int8、int16、int32、uint8、bool),对齐 AscendC 算子的通用数据覆盖范围。 ### 改动内容 #### 1. 算子定义层(op_def / op_proto) - **op_host/rsqrt_def.cpp**:Input/Output 的 DataType 列表新增 DT_INT8、DT_INT16、DT_INT32、DT_UINT8、DT_BOOL - **op_graph/rsqrt_proto.h**:INPUT/OUTPUT 的 TensorType 列表新增上述整型和布尔类型 #### 2. 算子 Tiling 层(op_host/rsqrt_tiling.cpp) - 新增 DATA_NUM_INT32 = 2、DATA_NUM_INT16 = 2、DATA_NUM_INT8 = 6、DATA_NUM_UINT8 = 4、DATA_NUM_BOOL = 2 五个常量,用于各数据类型在 UB 空间中的 Buffer 分配计算 - 在 switch 分支中新增对应数据类型到 ubDataNumber 的映射 - tiling 函数已扩展覆盖全部新增类型,无需额外修改注册逻辑 #### 3. 算子 ACLNN API 层(op_api/aclnn_rsqrt.cpp) - DTYPE_SUPPORT_LIST 移除旧版的独立 DTYPE_OUT_LIST(仅浮点),统一为一个支持全部 8 种类型的列表 - 移除 CheckDtypeValid 中不必要的 OUTPUT 类型检查重复逻辑 - 移除多余的 Cast 中间转换序列(由 rsqrt.cpp 内部统一处理) - **新增关键校验**:CheckDtypeValid 增加 self->GetDataType() != out->GetDataType() 检查,确保 self 和 out 的数据类型必须一致,避免 ViewCopy 在 dtype 不匹配时出现未定义行为 #### 4. 算子 Kernel 层(op_kernel/rsqrt.h / rsqrt.cpp) - rsqrt_tiling_data.h:RsqrtTilingData 结构体将 tile/tail 相关字段从 uint32_t 升级为 uint64_t,避免大 shape 下溢出 - rsqrt.h:KernelRsqrt 模板类新增 ComputeImpl 分派函数,按 TYPE_X 类型分发至以下计算函数: - **int32**:Cast 到 float → Rsqrt → Cast(CAST_RINT) 写回 int32 - **int16**:Cast 到 half → Rsqrt → ComputeIntCorrection(Mins/Adds/Maxs/Muls/Sub 校正)→ Cast(CAST_RINT) 写回 int16 - **int8**:Cast 到 half → Rsqrt → ComputeIntCorrection → Cast(CAST_RINT) 写回 int8 - **uint8**:Cast 到 half → Rsqrt → Cast(CAST_RINT) 写回 uint8 - **bool**:直接 Duplicate(0x0101) 填充 int16_t 模式,每个输入元素输出 true - ComputeIntCorrection 为 int16/int8 类型提供分段线性校正:将 [0, 2] 范围的 Rsqrt 结果通过 min(2)→add(-1)→max(0)→mul(3)→sub 的校正公式压缩到正确范围 - rsqrt.cpp:扩展 REGISTER_TILING_DEFAULT 和 DTYPE_X 宏路径,单入口 rsqrt<schMode> 模板支持 schMode=0(双缓冲)和 schMode=1(单缓冲) #### 5. 测试层(tests/ut) **op_api 测试**(test_aclnn_rsqrt.cpp): - 16 个测试用例,覆盖 int16、float16、bf16、float、int32、uint8 等类型的正常入参 - 覆盖不同 shape、空 tensor、非连续、低维/高维、inplace 等场景 - 覆盖 dtype 不匹配(ACLNN_ERR_PARAM_INVALID)、shape 不一致、nullptr 等错误路径 **op_host 测试**: - test_rsqrt_infershape.cpp:3 个测试用例覆盖 float、float16、bf16 的 shape 推导 - test_rsqrt_tiling.cpp:5 个测试用例覆盖不同 shape(8x8、8x2048、1023x2047)和数据类型(float、float16、bf16)的 tiling 参数计算 **op_kernel 测试**: - test_rsqrt.cpp:float32 单缓冲(单流水)测试 + float32 双缓冲(双流水)测试,均检查 system() 返回值 - test_rsqrt_int8.cpp:int8 类型 kernel 端到端测试 - test_rsqrt_int16.cpp:int16 类型 kernel 端到端测试 - test_rsqrt_int32.cpp:int32 类型 kernel 端到端测试 - test_rsqrt_uint8.cpp:uint8 类型 kernel 端到端测试 - test_rsqrt_bool.cpp:bool 类型 kernel 端到端测试 - 使用 rsqrt_test_entry.h 共享入口模板(宏拼接生成唯一函数名 rsqrt_int8、rsqrt_int16 等),避免多 DTYPE_X 文件的 ODR 冲突 - gen_data.py / compare_data.py:同时支持 8 种数据类型的 golden 数据生成与结果比对,整型使用 float32 精度计算对齐 AscendC Rsqrt 结果 ## 关联的Issue [#2156](https://gitcode.com/cann/ops-math/issues/2156) ## 测试 - **UT 测试**(全部通过): - --opapi:16/16 ✅ - --ophost:8/8 ✅ - --opkernel:7/7 ✅ - **AscendOpTest 测试**:已通过算子泛化测试 - **测试覆盖的数据类型**:float16、float、bfloat16、int8、int16、int32、uint8、bool ## 文档更新 - 更新了 README.md 文件,补充新增的整型/布尔类型支持说明 - 更新了 docs/aclnnRsqrt&aclnnInplaceRsqrt.md 文件 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-math!3863 | 2 个月前 | |
rsqrt add support int datatype in A2/A3 Co-authored-by: Nice_try<nicetryzzw@163.com> # message auto-generated for no-merge-commit merge: !3863 merge submit-InplaceSqrt into master 【CANN开源开放社区任务】【社区任务】InplaceRsqrt 新增整型数据类型支持 Created-by: Nice_try Commit-by: Nice_try Merged-by: cann-robot Description: ## 描述 ### 改动背景 Rsqrt/InplaceRsqrt 算子原仅支持浮点类型(float16、float、bfloat16),本次 PR 将其数据类型支持扩展至全部整型和布尔类型(int8、int16、int32、uint8、bool),对齐 AscendC 算子的通用数据覆盖范围。 ### 改动内容 #### 1. 算子定义层(op_def / op_proto) - **op_host/rsqrt_def.cpp**:Input/Output 的 DataType 列表新增 DT_INT8、DT_INT16、DT_INT32、DT_UINT8、DT_BOOL - **op_graph/rsqrt_proto.h**:INPUT/OUTPUT 的 TensorType 列表新增上述整型和布尔类型 #### 2. 算子 Tiling 层(op_host/rsqrt_tiling.cpp) - 新增 DATA_NUM_INT32 = 2、DATA_NUM_INT16 = 2、DATA_NUM_INT8 = 6、DATA_NUM_UINT8 = 4、DATA_NUM_BOOL = 2 五个常量,用于各数据类型在 UB 空间中的 Buffer 分配计算 - 在 switch 分支中新增对应数据类型到 ubDataNumber 的映射 - tiling 函数已扩展覆盖全部新增类型,无需额外修改注册逻辑 #### 3. 算子 ACLNN API 层(op_api/aclnn_rsqrt.cpp) - DTYPE_SUPPORT_LIST 移除旧版的独立 DTYPE_OUT_LIST(仅浮点),统一为一个支持全部 8 种类型的列表 - 移除 CheckDtypeValid 中不必要的 OUTPUT 类型检查重复逻辑 - 移除多余的 Cast 中间转换序列(由 rsqrt.cpp 内部统一处理) - **新增关键校验**:CheckDtypeValid 增加 self->GetDataType() != out->GetDataType() 检查,确保 self 和 out 的数据类型必须一致,避免 ViewCopy 在 dtype 不匹配时出现未定义行为 #### 4. 算子 Kernel 层(op_kernel/rsqrt.h / rsqrt.cpp) - rsqrt_tiling_data.h:RsqrtTilingData 结构体将 tile/tail 相关字段从 uint32_t 升级为 uint64_t,避免大 shape 下溢出 - rsqrt.h:KernelRsqrt 模板类新增 ComputeImpl 分派函数,按 TYPE_X 类型分发至以下计算函数: - **int32**:Cast 到 float → Rsqrt → Cast(CAST_RINT) 写回 int32 - **int16**:Cast 到 half → Rsqrt → ComputeIntCorrection(Mins/Adds/Maxs/Muls/Sub 校正)→ Cast(CAST_RINT) 写回 int16 - **int8**:Cast 到 half → Rsqrt → ComputeIntCorrection → Cast(CAST_RINT) 写回 int8 - **uint8**:Cast 到 half → Rsqrt → Cast(CAST_RINT) 写回 uint8 - **bool**:直接 Duplicate(0x0101) 填充 int16_t 模式,每个输入元素输出 true - ComputeIntCorrection 为 int16/int8 类型提供分段线性校正:将 [0, 2] 范围的 Rsqrt 结果通过 min(2)→add(-1)→max(0)→mul(3)→sub 的校正公式压缩到正确范围 - rsqrt.cpp:扩展 REGISTER_TILING_DEFAULT 和 DTYPE_X 宏路径,单入口 rsqrt<schMode> 模板支持 schMode=0(双缓冲)和 schMode=1(单缓冲) #### 5. 测试层(tests/ut) **op_api 测试**(test_aclnn_rsqrt.cpp): - 16 个测试用例,覆盖 int16、float16、bf16、float、int32、uint8 等类型的正常入参 - 覆盖不同 shape、空 tensor、非连续、低维/高维、inplace 等场景 - 覆盖 dtype 不匹配(ACLNN_ERR_PARAM_INVALID)、shape 不一致、nullptr 等错误路径 **op_host 测试**: - test_rsqrt_infershape.cpp:3 个测试用例覆盖 float、float16、bf16 的 shape 推导 - test_rsqrt_tiling.cpp:5 个测试用例覆盖不同 shape(8x8、8x2048、1023x2047)和数据类型(float、float16、bf16)的 tiling 参数计算 **op_kernel 测试**: - test_rsqrt.cpp:float32 单缓冲(单流水)测试 + float32 双缓冲(双流水)测试,均检查 system() 返回值 - test_rsqrt_int8.cpp:int8 类型 kernel 端到端测试 - test_rsqrt_int16.cpp:int16 类型 kernel 端到端测试 - test_rsqrt_int32.cpp:int32 类型 kernel 端到端测试 - test_rsqrt_uint8.cpp:uint8 类型 kernel 端到端测试 - test_rsqrt_bool.cpp:bool 类型 kernel 端到端测试 - 使用 rsqrt_test_entry.h 共享入口模板(宏拼接生成唯一函数名 rsqrt_int8、rsqrt_int16 等),避免多 DTYPE_X 文件的 ODR 冲突 - gen_data.py / compare_data.py:同时支持 8 种数据类型的 golden 数据生成与结果比对,整型使用 float32 精度计算对齐 AscendC Rsqrt 结果 ## 关联的Issue [#2156](https://gitcode.com/cann/ops-math/issues/2156) ## 测试 - **UT 测试**(全部通过): - --opapi:16/16 ✅ - --ophost:8/8 ✅ - --opkernel:7/7 ✅ - **AscendOpTest 测试**:已通过算子泛化测试 - **测试覆盖的数据类型**:float16、float、bfloat16、int8、int16、int32、uint8、bool ## 文档更新 - 更新了 README.md 文件,补充新增的整型/布尔类型支持说明 - 更新了 docs/aclnnRsqrt&aclnnInplaceRsqrt.md 文件 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-math!3863 | 2 个月前 | |
rsqrt add support int datatype in A2/A3 Co-authored-by: Nice_try<nicetryzzw@163.com> # message auto-generated for no-merge-commit merge: !3863 merge submit-InplaceSqrt into master 【CANN开源开放社区任务】【社区任务】InplaceRsqrt 新增整型数据类型支持 Created-by: Nice_try Commit-by: Nice_try Merged-by: cann-robot Description: ## 描述 ### 改动背景 Rsqrt/InplaceRsqrt 算子原仅支持浮点类型(float16、float、bfloat16),本次 PR 将其数据类型支持扩展至全部整型和布尔类型(int8、int16、int32、uint8、bool),对齐 AscendC 算子的通用数据覆盖范围。 ### 改动内容 #### 1. 算子定义层(op_def / op_proto) - **op_host/rsqrt_def.cpp**:Input/Output 的 DataType 列表新增 DT_INT8、DT_INT16、DT_INT32、DT_UINT8、DT_BOOL - **op_graph/rsqrt_proto.h**:INPUT/OUTPUT 的 TensorType 列表新增上述整型和布尔类型 #### 2. 算子 Tiling 层(op_host/rsqrt_tiling.cpp) - 新增 DATA_NUM_INT32 = 2、DATA_NUM_INT16 = 2、DATA_NUM_INT8 = 6、DATA_NUM_UINT8 = 4、DATA_NUM_BOOL = 2 五个常量,用于各数据类型在 UB 空间中的 Buffer 分配计算 - 在 switch 分支中新增对应数据类型到 ubDataNumber 的映射 - tiling 函数已扩展覆盖全部新增类型,无需额外修改注册逻辑 #### 3. 算子 ACLNN API 层(op_api/aclnn_rsqrt.cpp) - DTYPE_SUPPORT_LIST 移除旧版的独立 DTYPE_OUT_LIST(仅浮点),统一为一个支持全部 8 种类型的列表 - 移除 CheckDtypeValid 中不必要的 OUTPUT 类型检查重复逻辑 - 移除多余的 Cast 中间转换序列(由 rsqrt.cpp 内部统一处理) - **新增关键校验**:CheckDtypeValid 增加 self->GetDataType() != out->GetDataType() 检查,确保 self 和 out 的数据类型必须一致,避免 ViewCopy 在 dtype 不匹配时出现未定义行为 #### 4. 算子 Kernel 层(op_kernel/rsqrt.h / rsqrt.cpp) - rsqrt_tiling_data.h:RsqrtTilingData 结构体将 tile/tail 相关字段从 uint32_t 升级为 uint64_t,避免大 shape 下溢出 - rsqrt.h:KernelRsqrt 模板类新增 ComputeImpl 分派函数,按 TYPE_X 类型分发至以下计算函数: - **int32**:Cast 到 float → Rsqrt → Cast(CAST_RINT) 写回 int32 - **int16**:Cast 到 half → Rsqrt → ComputeIntCorrection(Mins/Adds/Maxs/Muls/Sub 校正)→ Cast(CAST_RINT) 写回 int16 - **int8**:Cast 到 half → Rsqrt → ComputeIntCorrection → Cast(CAST_RINT) 写回 int8 - **uint8**:Cast 到 half → Rsqrt → Cast(CAST_RINT) 写回 uint8 - **bool**:直接 Duplicate(0x0101) 填充 int16_t 模式,每个输入元素输出 true - ComputeIntCorrection 为 int16/int8 类型提供分段线性校正:将 [0, 2] 范围的 Rsqrt 结果通过 min(2)→add(-1)→max(0)→mul(3)→sub 的校正公式压缩到正确范围 - rsqrt.cpp:扩展 REGISTER_TILING_DEFAULT 和 DTYPE_X 宏路径,单入口 rsqrt<schMode> 模板支持 schMode=0(双缓冲)和 schMode=1(单缓冲) #### 5. 测试层(tests/ut) **op_api 测试**(test_aclnn_rsqrt.cpp): - 16 个测试用例,覆盖 int16、float16、bf16、float、int32、uint8 等类型的正常入参 - 覆盖不同 shape、空 tensor、非连续、低维/高维、inplace 等场景 - 覆盖 dtype 不匹配(ACLNN_ERR_PARAM_INVALID)、shape 不一致、nullptr 等错误路径 **op_host 测试**: - test_rsqrt_infershape.cpp:3 个测试用例覆盖 float、float16、bf16 的 shape 推导 - test_rsqrt_tiling.cpp:5 个测试用例覆盖不同 shape(8x8、8x2048、1023x2047)和数据类型(float、float16、bf16)的 tiling 参数计算 **op_kernel 测试**: - test_rsqrt.cpp:float32 单缓冲(单流水)测试 + float32 双缓冲(双流水)测试,均检查 system() 返回值 - test_rsqrt_int8.cpp:int8 类型 kernel 端到端测试 - test_rsqrt_int16.cpp:int16 类型 kernel 端到端测试 - test_rsqrt_int32.cpp:int32 类型 kernel 端到端测试 - test_rsqrt_uint8.cpp:uint8 类型 kernel 端到端测试 - test_rsqrt_bool.cpp:bool 类型 kernel 端到端测试 - 使用 rsqrt_test_entry.h 共享入口模板(宏拼接生成唯一函数名 rsqrt_int8、rsqrt_int16 等),避免多 DTYPE_X 文件的 ODR 冲突 - gen_data.py / compare_data.py:同时支持 8 种数据类型的 golden 数据生成与结果比对,整型使用 float32 精度计算对齐 AscendC Rsqrt 结果 ## 关联的Issue [#2156](https://gitcode.com/cann/ops-math/issues/2156) ## 测试 - **UT 测试**(全部通过): - --opapi:16/16 ✅ - --ophost:8/8 ✅ - --opkernel:7/7 ✅ - **AscendOpTest 测试**:已通过算子泛化测试 - **测试覆盖的数据类型**:float16、float、bfloat16、int8、int16、int32、uint8、bool ## 文档更新 - 更新了 README.md 文件,补充新增的整型/布尔类型支持说明 - 更新了 docs/aclnnRsqrt&aclnnInplaceRsqrt.md 文件 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-math!3863 | 2 个月前 | |
rsqrt add support int datatype in A2/A3 Co-authored-by: Nice_try<nicetryzzw@163.com> # message auto-generated for no-merge-commit merge: !3863 merge submit-InplaceSqrt into master 【CANN开源开放社区任务】【社区任务】InplaceRsqrt 新增整型数据类型支持 Created-by: Nice_try Commit-by: Nice_try Merged-by: cann-robot Description: ## 描述 ### 改动背景 Rsqrt/InplaceRsqrt 算子原仅支持浮点类型(float16、float、bfloat16),本次 PR 将其数据类型支持扩展至全部整型和布尔类型(int8、int16、int32、uint8、bool),对齐 AscendC 算子的通用数据覆盖范围。 ### 改动内容 #### 1. 算子定义层(op_def / op_proto) - **op_host/rsqrt_def.cpp**:Input/Output 的 DataType 列表新增 DT_INT8、DT_INT16、DT_INT32、DT_UINT8、DT_BOOL - **op_graph/rsqrt_proto.h**:INPUT/OUTPUT 的 TensorType 列表新增上述整型和布尔类型 #### 2. 算子 Tiling 层(op_host/rsqrt_tiling.cpp) - 新增 DATA_NUM_INT32 = 2、DATA_NUM_INT16 = 2、DATA_NUM_INT8 = 6、DATA_NUM_UINT8 = 4、DATA_NUM_BOOL = 2 五个常量,用于各数据类型在 UB 空间中的 Buffer 分配计算 - 在 switch 分支中新增对应数据类型到 ubDataNumber 的映射 - tiling 函数已扩展覆盖全部新增类型,无需额外修改注册逻辑 #### 3. 算子 ACLNN API 层(op_api/aclnn_rsqrt.cpp) - DTYPE_SUPPORT_LIST 移除旧版的独立 DTYPE_OUT_LIST(仅浮点),统一为一个支持全部 8 种类型的列表 - 移除 CheckDtypeValid 中不必要的 OUTPUT 类型检查重复逻辑 - 移除多余的 Cast 中间转换序列(由 rsqrt.cpp 内部统一处理) - **新增关键校验**:CheckDtypeValid 增加 self->GetDataType() != out->GetDataType() 检查,确保 self 和 out 的数据类型必须一致,避免 ViewCopy 在 dtype 不匹配时出现未定义行为 #### 4. 算子 Kernel 层(op_kernel/rsqrt.h / rsqrt.cpp) - rsqrt_tiling_data.h:RsqrtTilingData 结构体将 tile/tail 相关字段从 uint32_t 升级为 uint64_t,避免大 shape 下溢出 - rsqrt.h:KernelRsqrt 模板类新增 ComputeImpl 分派函数,按 TYPE_X 类型分发至以下计算函数: - **int32**:Cast 到 float → Rsqrt → Cast(CAST_RINT) 写回 int32 - **int16**:Cast 到 half → Rsqrt → ComputeIntCorrection(Mins/Adds/Maxs/Muls/Sub 校正)→ Cast(CAST_RINT) 写回 int16 - **int8**:Cast 到 half → Rsqrt → ComputeIntCorrection → Cast(CAST_RINT) 写回 int8 - **uint8**:Cast 到 half → Rsqrt → Cast(CAST_RINT) 写回 uint8 - **bool**:直接 Duplicate(0x0101) 填充 int16_t 模式,每个输入元素输出 true - ComputeIntCorrection 为 int16/int8 类型提供分段线性校正:将 [0, 2] 范围的 Rsqrt 结果通过 min(2)→add(-1)→max(0)→mul(3)→sub 的校正公式压缩到正确范围 - rsqrt.cpp:扩展 REGISTER_TILING_DEFAULT 和 DTYPE_X 宏路径,单入口 rsqrt<schMode> 模板支持 schMode=0(双缓冲)和 schMode=1(单缓冲) #### 5. 测试层(tests/ut) **op_api 测试**(test_aclnn_rsqrt.cpp): - 16 个测试用例,覆盖 int16、float16、bf16、float、int32、uint8 等类型的正常入参 - 覆盖不同 shape、空 tensor、非连续、低维/高维、inplace 等场景 - 覆盖 dtype 不匹配(ACLNN_ERR_PARAM_INVALID)、shape 不一致、nullptr 等错误路径 **op_host 测试**: - test_rsqrt_infershape.cpp:3 个测试用例覆盖 float、float16、bf16 的 shape 推导 - test_rsqrt_tiling.cpp:5 个测试用例覆盖不同 shape(8x8、8x2048、1023x2047)和数据类型(float、float16、bf16)的 tiling 参数计算 **op_kernel 测试**: - test_rsqrt.cpp:float32 单缓冲(单流水)测试 + float32 双缓冲(双流水)测试,均检查 system() 返回值 - test_rsqrt_int8.cpp:int8 类型 kernel 端到端测试 - test_rsqrt_int16.cpp:int16 类型 kernel 端到端测试 - test_rsqrt_int32.cpp:int32 类型 kernel 端到端测试 - test_rsqrt_uint8.cpp:uint8 类型 kernel 端到端测试 - test_rsqrt_bool.cpp:bool 类型 kernel 端到端测试 - 使用 rsqrt_test_entry.h 共享入口模板(宏拼接生成唯一函数名 rsqrt_int8、rsqrt_int16 等),避免多 DTYPE_X 文件的 ODR 冲突 - gen_data.py / compare_data.py:同时支持 8 种数据类型的 golden 数据生成与结果比对,整型使用 float32 精度计算对齐 AscendC Rsqrt 结果 ## 关联的Issue [#2156](https://gitcode.com/cann/ops-math/issues/2156) ## 测试 - **UT 测试**(全部通过): - --opapi:16/16 ✅ - --ophost:8/8 ✅ - --opkernel:7/7 ✅ - **AscendOpTest 测试**:已通过算子泛化测试 - **测试覆盖的数据类型**:float16、float、bfloat16、int8、int16、int32、uint8、bool ## 文档更新 - 更新了 README.md 文件,补充新增的整型/布尔类型支持说明 - 更新了 docs/aclnnRsqrt&aclnnInplaceRsqrt.md 文件 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-math!3863 | 2 个月前 | |
提交Ascend C实现的Rsqrt算子 Co-authored-by: skywang2<727854256@qq.com> # message auto-generated for no-merge-commit merge: !643 merge math-Rsqrt into master 提交Ascend C实现的Rsqrt算子 Created-by: skywang2 Commit-by: skywang2 Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 使用Ascend C实现rsqrt算子。 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> <!-- 如果这个PR是为了解决特定的问题单,请在这里描述问题单单号。--> [https://gitcode.com/cann/ops-math/issues/369](https://gitcode.com/cann/ops-math/issues/369) ## 测试 <!--描述进行了哪些测试来验证你的改动。包括但不限于二级冒烟、算子泛化等。--> 测试命令: bash build.sh --run_example rsqrt eager cust --vendor_name=custom_wang ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 无 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-math!643 | 5 个月前 | |
random类和conver类算子aclnn API md增加产品标签 Co-authored-by: chenjiao<chenjiao31@huawei.com> # message auto-generated for no-merge-commit merge: !4164 merge master into master random类和conver类算子aclnn API md增加产品标签 Created-by: gitcode-chenjiao Commit-by: chenjiao Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 所有aclnn前缀的md文件需要增加产品标签注释对,方便后续按产品型号筛选 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。--> <!-- 如果这个PR是为了解决特定的问题单,请在这里描述问题单单号。--> [#2272](https://gitcode.com/cann/ops-math/issues/2272) ## 测试 <!--描述进行了哪些测试来验证你的改动。包括但不限于二级冒烟、算子泛化等。--> ok ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> math类算子对应的所有aclnn API MD ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x ] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-math!4164 | 2 个月前 |
Rsqrt
支持的产品型号
- Atlas A2 训练系列产品
产品形态详细说明请参见昇腾产品形态说明
算子描述
-
功能描述
Rsqrt算子将数据进行开方并取倒数运算。 -
原型信息
算子类型(OpType) Rsqrt name Type data type format 算子输入 x tensor float32,float16,bfloat16,int8,int16,int32,uint8,bool ND 算子输出 y tensor float32,float16,bfloat16,int8,int16,int32,uint8,bool ND 核函数名 rsqrt
约束与限制
- x,y,out的数据类型只支持 float32,float16,bfloat16,int8,int16,int32,uint8,bool,数据格式只支持ND
运行验证
测试命令调用方式:build.sh
| 目录 | 描述 |
|---|---|
| test_aclnn_rsqrt.cpp | 通过aclnn调用的方式调用Rsqrt算子。 |
贡献说明
| 贡献者 | 贡献方 | 贡献算子 | 贡献时间 | 贡献内容 |
|---|---|---|---|---|
| skywang2 | 个人开发者 | Rsqrt | 2026/03/30 | 新增Rsqrt算子 |
| Nice_try | 个人开发者 | Rsqrt | 2026/07/06 | Rsqrt新增支持整型数据支持(A2/A3) |