| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
feat: 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 Co-authored-by: xuejinghui<xuejinghui@huawei.com> # message auto-generated for no-merge-commit merge: !8821 merge InferShape into master feat: 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 Created-by: xuejinghui Commit-by: xuejinghui Merged-by: cann-robot Description: ## 描述 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5250 <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:Unknown Shape/ Unknown Rank迁移适配 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!8821 | 1 个月前 | |
fix(apply_top_k_top_p): 拦截 p 和 k 同时为空 Co-authored-by: yangziqi_007<yangziqi4@huawei.com> # message auto-generated for no-merge-commit merge: !11213 merge ApplyTopKTopP into master fix(apply_top_k_top_p): 拦截 p 和 k 同时为空 Created-by: yangziqi_007 Commit-by: yangziqi_007 Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> fix(apply_top_k_top_p): 拦截 p 和 k 同时为空 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> 关联Issue [#5999](https://gitcode.com/cann/ops-nn/issues/5999) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> st、ut测试通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 无 ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!11213 | 11 天前 | |
fix: arg_max_grad算子cleancode修复&&清理部分敏感词问题 Co-authored-by: LuckySun<sunwenlong8@huawei.com> # message auto-generated for no-merge-commit merge: !10803 merge cleancode_fix into master fix: arg_max_grad算子cleancode修复 Created-by: LuckySun Commit-by: LuckySun Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 修复arg_max_grad算子的cleancode问题 包含: 不要使用难以理解的字面常量 对于不会修改成员变量的成员函数,应使用const修饰 清理部分敏感词问题 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/6016 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> 编译通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述: cleancode修复 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10803 | 19 天前 | |
fix(proto): 修复 norm/softmax 算子 proto 注释拼写/语法/描述错误 Co-authored-by: sun-yibo1<sunyibo1@hisilicon.com> # message auto-generated for no-merge-commit merge: !9032 merge fix-proto-comment-typos into master fix(proto): 修复 norm/softmax 算子 proto 注释拼写/语法/描述错误 Created-by: sun-yibo1 Commit-by: sun-yibo1 Merged-by: cann-robot Description: ## 描述 修复 21 个算子 proto 头文件中 Doxygen 注释的拼写、语法、描述错误与公式括号不匹配问题。所有改动仅涉及注释文本,不改动任何 REG_OP 定义。 ### 改动原因 proto 头文件的 Doxygen 注释存在拼写错误、语法错误(冠词/单复数/主谓一致)、描述误导(输入描述与实际语义不符)以及 rstd 公式多余的右括号等问题,会误导接口使用者。 ### 改动方法 仅修改注释文本,不改动 REG_OP/INPUT/OUTPUT/ATTR/DATATYPE 等代码定义。同时按 pre-commit clang-format 要求对部分文件做格式化。 ### 涉及修复 **拼写错误:** - normlization→normalization、while→will(AddLayerNormQuant) - smooth_scales1/2→smooth_scale1/2(AddRmsNormDynamicQuant,4 处) - listBool→ListBool(AddRmsNormDynamicQuant/V2) - secend→second、opertaion→operation(AddRmsNormQuant) - reseluts→results(BatchNormGrad) - support→supported(BatchNormV3) - dype→type(GroupNormV2,4 处) - Layernorm→LayerNorm、中文或→or、全角,→ASCII(LayerNorm) - \file 文件名错误 nn_norm_ops.h→layer_norm_grad_proto.h(LayerNormGrad) **语法错误(~20 处):** - A optional→An optional(AddLayerNorm/InplaceAddRmsNorm/InstanceNorm) - A ND tensor→An ND tensor(SoftmaxV2/LogSoftmaxGrad/ConfusionSoftmaxGrad) - An required→A required(GroupNormV2) - support→supports(SoftmaxV2/BroadcastGradientArgs/AddRmsNormDynamicMxQuant) - 缺复数 s:input→inputs、output→outputs、group→groups(GroupNormGrad/V2) **描述错误:** - GroupNormGrad:x 输入描述从 the offset 改为 input tensor **公式括号(7 处):** - rstd 公式多余的右括号 ))→)(AddRmsNorm/AddRmsNormCast/AddRmsNormDynamicMxQuant/AddRmsNormDynamicQuant/AddRmsNormQuant/AddRmsNormQuantV2/InplaceAddRmsNorm) ## 关联的Issue - #4989 ## 测试 - 仅注释修改,无代码变更,不影响编译与功能。 - pre-commit 全部 hook 通过(trailing-whitespace/end-of-file-fixer/clang-format/codespell/check-*/detect-private-key);OAT 合规检查 0 issues。 ## 文档更新 无。 ## 类型标签 - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!9032 | 21 天前 | |
fix(aicpu): 整改代码质量告警 Co-authored-by: zhaowenrui666<zhaowenrui7@huawei.com> # message auto-generated for no-merge-commit merge: !10899 merge fix-aicpu-codecheck-warnings into master fix(aicpu): 整改代码质量告警 Created-by: zhaowenrui666 Commit-by: zhaowenrui666 Merged-by: cann-robot Description: ## 描述 整改 SoftmaxV2、Bucketize 和 ReverseSequence AICPU 实现中的代码质量告警: - 为并行 lambda 显式列出捕获变量,移除默认引用捕获。 - 将 void* 到目标类型指针的转换由 reinterpret_cast 调整为 static_cast,并移除不必要的 void* 转换。 - 为派生类析构函数补充正确的 override 声明。 - 不改变算子功能语义。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6076 ## 测试 - cmake --build build --target aicpu_kernels -j8:通过,成功生成 libnn_aicpu_kernels.so。 - 对本次修改文件执行 pre-commit run --files ...:全部通过,包括 clang-format、codespell 和 OAT 检查。 - git diff --check:通过。 ## 文档更新 无文档变更。 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:AICPU 代码质量告警整改 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10899 | 18 天前 | |
modify pic name Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !10839 merge 0920pic into master modify pic name Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 修改图片名称为小写+下划线格式 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> 关联Issue [#6007](https://gitcode.com/cann/ops-nn/issues/6007) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 相关文档已更新 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10839 | 19 天前 | |
feat(unique): add Regbase values-only AI Core support Co-authored-by: ConanHuang<huangxiaobin1@huawei.com> # message auto-generated for no-merge-commit merge: !10993 merge feature_op_unique into master feat(unique): add Regbase values-only AI Core support Created-by: ConanHuang Commit-by: ConanHuang Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 为 aclnnUnique / aclnnUnique2 增加 Regbase values-only AI Core 路径,在单次 Kernel 调用中完成排序与去重,减少不必要的索引搬运和中间存储。 - 新增 arch35 Tiling、Kernel 和动态输出 shape 回传,接入 ACLNN 分流。 - 整理算子 L0、原型及 Ascend 950/350 构建注册;调整共享 radix 和 UniqueConsecutive 的大输入计算。 - 新增 ACLNN/Kernel 测试用例、golden 和 GEIR 示例。 新路径仅处理不返回 inverse/counts 且满足准入条件的调用,其余场景沿用既有路径;GE/ONNX 仍使用 AICPU。详细 Tiling 与 Kernel 设计见关联 Issue。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6187 ## 测试 - 二级冒烟通过 - UniqueWithCountsAndSorting 和 KthValue 各自 300+ kernel ST通过 - aclnnUnique aclnnUnique2 各自 170+ aclnn用例通过 - onnx geir 通路兼容性验证通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 不涉及 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10993 | 15 天前 | |
erfinv等3个算子新增ascend350平台支持 Co-authored-by: c00887544<chenshengze@huawei.com> # message auto-generated for no-merge-commit merge: !9912 merge 350 into master erfinv等3个算子新增ascend350平台支持 Created-by: chengzhi1120 Commit-by: c00887544 Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> erfinv等3个算子新增ascend350平台支持 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/5664 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!9912 | 1 个月前 | |
legacy下线:整改NN算子图模式注册与日志规范 Co-authored-by: zhouxuan78<zhouxuan78@huawei.com> # message auto-generated for no-merge-commit merge: !8388 merge master into master legacy下线:整改NN算子图模式注册与日志规范 Created-by: zhouxuan78 Commit-by: zhouxuan78 Merged-by: cann-robot Description: ## 描述 本 PR 用于 ops-nn 仓 Legacy 下线整改,补齐部分算子的直调/图模式注册、InferShape / InferDataType 拆分、日志规范整改及必要的编译/预提交问题修复。 1. 补齐多个算子的 op_graph 侧 InferDataType / graph infer 注册文件,并同步补充对应 CMakeLists.txt: - 如 FastGeluV2、HardShrinkGrad、HardSigmoidGrad、HardSwish、HardSwishGrad、Selu、SeluGrad、Shrink、SoftplusGrad、SoftplusV2、SoftplusV2Grad、Softsign、SoftsignGrad、SmoothL1Loss、SoftMarginLoss、LpNormUpdateV2、ApplyFtrlV2、ApplyMomentum、ApplyRMSProp 等。 2. 调整部分算子的 op_host infer 文件: - 将 InferDataType 逻辑从 infershape 文件中拆分/迁移到 op_graph 侧; - 保留 InferShape 逻辑,避免 shape/type 混在同一处注册; - 收敛 IndexCheck infershape 日志,仅保留入口日志。 3. 补充 SoftMarginLoss 图模式原型: - 新增 loss/soft_margin_loss/op_graph/soft_margin_loss_proto.h; - 新增 OPS_PROTO_DEF_SOFTMARGINLOSS 去重宏,避免重复注册; - 补充 SoftMarginLoss 的 op_graph/CMakeLists.txt 和 graph infer 注册。 4. 日志整改: - 为 Arch35 tiling 入口补充统一入口日志; - 对本 PR 修改范围内新增/调整的外部输入错误日志,按规范使用 OP_LOGE_FOR_INVALID_* 系列上报接口; - 收敛部分过多的 tiling / infershape 日志打印,避免循环或冗余打印。 5. 回退/收敛不属于本次整改范围的 Softshrink 重命名相关改动: - Softshrink -> SoftShrink 的命名调整不放在本 PR 中,后续单独 PR 处理。 6. 修复 pre-commit 格式问题: - 按仓库 .clang-format 修复格式; - 不修改 .pre-commit-config.yaml。 影响范围 主要影响 activation、index、loss、norm、optim、quant、vfusion 等目录下 Legacy 下线相关算子的 op_graph 注册、InferDataType 拆分、tiling 入口日志和错误日志规范。 ## 变更内容 ## 关联的Issue ## 测试 ## 文档更新 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他:Legacy包下线整改 See merge request: cann/ops-nn!8388 | 30 天前 | |
feat: Embedding算子新增SIMD模板支持pice through场景(连续二维+非连续场景) Co-authored-by: Apricityh<wangxu359@huawei.com> # message auto-generated for no-merge-commit merge: !8200 merge feat/embedding-simd-template-clean into master feat: Embedding算子新增SIMD模板支持pice through场景(连续二维+非连续场景) Created-by: Apricityh Commit-by: Apricityh Merged-by: cann-robot Description: ## 描述 Embedding算子新增SIMD模板支持pice through场景(连续二维+非连续场景) ## 关联的Issue 关联Issue #5027 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 白盒用例、端对端通路验证:包括geir通路、aclnn、aclgraph通路 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!8200 | 1 个月前 | |
本 PR 为 6 个算子新增 ascend350 平台的适配支持(EmbeddingBag、EmbeddingDenseGrad、EmbeddingDenseGradV2、AvgPool3D、AvgPool3DGrad、MaxPoolWithArgmax),和 1 个 算子 fusion pass 通路(max_pool_fusion_pass.cpp)的 ascend350 适配。 Co-authored-by: duxinlei<duxinlei@h-partners.com> Co-authored-by: u010470851<shangguanqinnan@huawei.com> # message auto-generated for no-merge-commit merge: !10135 merge ascend3500909 into master ascend350 算子适配350平台 Created-by: duxinlei Commit-by: duxinlei;u010470851 Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 给 6 个算子新增 ascend350 平台的适配支持(EmbeddingBag、EmbeddingDenseGrad、EmbeddingDenseGradV2、AvgPool3D、AvgPool3DGrad、MaxPoolWithArgmax),给 1 个 算子 fusion pass 通路(max_pool_fusion_pass.cpp)新增 ascend350 适配。 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> 关联Issue [#5820](https://gitcode.com/cann/ops-nn/issues/5820) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> david冒烟 david通路测试 dvlite通路测试 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10135 | 28 天前 | |
modify file name Co-authored-by: liuguoyue<liuguoyue@huawei.com> # message auto-generated for no-merge-commit merge: !10835 merge modify_head into master modify file name Created-by: liuguoyue Commit-by: liuguoyue Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 修改EmbeddingDenseGrad重复头文件 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> 关联Issue [#6044](https://gitcode.com/cann/ops-nn/issues/6044) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> 完成算子编译和ST用例验证 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10835 | 19 天前 | |
docs: 统一原型注释中的产品名称(master) Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !11186 merge 923soc into master docs: 统一原型注释中的产品名称(master) Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 原型注释里的产品名称仍沿用 Series Product、Ascend 950 AI Processor、Ascend 950PR/Ascend 950DT 等旧表述,与当前产品命名不一致。本次只改注释,不改算子接口和计算逻辑。 合入目标:cann/ops-nn 的 master。 覆盖各算子 op_graph/*_proto.h,以及 common/inc/op_graph/op_nn_proto_extend.h、common/stub/op_graph/math_proto_stub.cpp。 ## 关联的Issue 关联 Issue #6213(Close #6213) https://gitcode.com/cann/ops-nn/issues/6213 ## 测试 仅修改原型注释中的产品名称,不涉及算子 host/kernel 计算逻辑,无需算子泛化。已在本地用仓库 pre-commit 的 end-of-file-fixer 与 clang-format 核对格式,且没有改动算子输入、输出和属性定义。 ## 文档更新 更新了 op_graph 原型注释中的产品名称,无独立 markdown 文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11186 | 11 天前 | |
[CANNBot]ExpandIntoJaggedPermute算子支持Ascend950 Co-authored-by: yourealize<chenzhanxi1@huawei.com> # message auto-generated for no-merge-commit merge: !8583 merge codex/add_expand_into_jagged_permute_ascend950 into master [CANNBot]ExpandIntoJaggedPermute算子支持Ascend950 Created-by: yourealize Commit-by: yourealize Merged-by: cann-robot Description: ## 描述 ExpandIntoJaggedPermute算子新增Ascend950(arch35)SIMT实现,同时保留既有Ascend910B(arch22)实现和现有ACLNN/ST/示例交付件。 - 新增arch35 Kernel、Host tiling、binary配置与双架构CMake路由。 - InferShape支持UnknownRank、UnknownDim及混合已知/未知shape,输出固定推导为[output_size];InferShapeRange同步输出确定的min/max range。 - Tiling侧按仓库报错规范做根因级拦截,覆盖dtype、format、rank、shape关系、offset/P关系、output_size、view边界、物理字节溢出、平台资源等异常,避免重复包装错误信息。 - 保持40B TilingData ABI、TilingKey与SIMT算法等价,既有910B实现及交付件不回归。 ## 关联的issue Close #6134 ## 测试 - pre-commit clean PASS:19个变更文件,clang-format、codespell、OAT等适用hook全部通过。 - Host UT Ascend910B:13/13 PASS。 - Host UT Ascend950:23/23 PASS,覆盖known/unknown shape、unknown range及tiling异常边界。 - Kernel UT Ascend910B:1/1 PASS,48个模拟核均正常退出。 - 本地增量编译:single package与Ascend950 package PASS;两个.run包独立安装成功,4次共享库ldd均无not found。 - TTK使用Bash定向执行:v12 6/6 PASS;Host修复后v13调整用例3/3 PASS,int32 binary_equal均为100%,OOB均为OK/OK,未发现精度问题。v13 driver日志仍显示vendors/custom_math,因此built-in Host实际加载归因记为NOT_PROVEN;本PR通过正式仓Ascend950 Host UT 23/23闭环Host逻辑。 - ST设计生成与QA:PASS 21、FAIL 0、SKIP 7;按要求未执行ST全量。 - 按要求未执行ACLNN测试;远端CI的ARM/独立镜像/硬件smoke等仍以PR门禁为准。 ## 文档更新 - README与接口文档支持矩阵更新为Ascend950支持。 - 补充UnknownShape/UnknownRange能力与本次平台适配说明。 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [x] 性能优化 - [x] 文档更新 - [ ] 其他,请描述 ## AI/Agent 生成声明 代码来源: - [ ] 纯人工手写 - [ ] AI 辅助编写 - [x] AI 完全生成 人工审查:已完成 测试验证:已通过 合规检查:已完成 See merge request: cann/ops-nn!8583 | 11 天前 | |
gather elements文档更新,代码增加判空 Co-authored-by: yuhao_<yuhao93@huawei.com> # message auto-generated for no-merge-commit merge: !11250 merge 0928-gather into master gather elements文档更新,代码增加判空 Created-by: huang-jz Commit-by: yuhao_ Merged-by: cann-robot Description: ## 描述 gather elements 文档更新,代码增加判空 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000-->关联Issue #6000 ## 测试 <!--描述进行了哪些测试来验证你的改动。-->本地精度通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。-->更新了gather elements 文档 ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!11250 | 11 天前 | |
补充aclnnGather资料中缺少的dtype Co-authored-by: sikaiwei<sikaiwei1@h-partners.com> # message auto-generated for no-merge-commit merge: !10244 merge aclnnGather_md into master 补充aclnnGather资料中缺少的dtype Created-by: sikaiwei Commit-by: sikaiwei Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 资料补充 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/5789 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> NA ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 更新了index/gather_elements_v2/docs/aclnnGather.md ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10244 | 28 天前 | |
【mc62】tiling侧workspaceSize硬编码16M大小,替换为通过接口获取,ascendcPlatform.GetLibApiWorkSpaceSize(); Co-authored-by: xufeng12121<1074805447@qq.com> # message auto-generated for no-merge-commit merge: !10878 merge nn_wo into master 【mc62】tiling侧workspaceSize硬编码16M大小,替换为通过接口获取,ascendcPlatform.GetLibApiWorkSpaceSize(); Created-by: xufeng12121 Commit-by: xufeng12121 Merged-by: cann-robot Description: ## 描述 1. 背景:mc62算子workspace占用对比1951增大 2. 由于mc62代码和A5是onetrack, 所有workspacesize硬编码为16M, 在mc62上也同样是16M,asc 提供接口,对mc62进行区分,从原来的16M 变成1K,A5和A2大小不变。 3. 修改:算子侧tiling中硬编码为16M大小的workspacesize统一改为用接口获取。 GetLibApiWorkSpaceSize实现如下: const static uint32_t WORK_SPACE_SIZE_910B = 16 * 1024 * 1024; const static uint32_t WORK_SPACE_SIZE_950 = 16 * 1024 * 1024; const static uint32_t WORK_SPACE_SIZE = 1024; uint32_t PlatformAscendC::GetLibApiWorkSpaceSize(void) const { auto npuArch = GetCurNpuArch(); if (npuArch == NpuArch::DAV_RESV) { PF_LOGE("get platform failed, CurNpuArch is NpuArch::DAV_RESV"); return -1; } else if (npuArch == NpuArch::DAV_2201) { return WORK_SPACE_SIZE_910B; } else if (npuArch == NpuArch::DAV_3510) { return WORK_SPACE_SIZE_950; } return WORK_SPACE_SIZE; } ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6135 ## 测试 1.单算子框架测试通过 2.二级冒烟通过 文档更新 类型标签 - [x] Bug修复 See merge request: cann/ops-nn!10878 | 17 天前 | |
modify pic name Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !10839 merge 0920pic into master modify pic name Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 修改图片名称为小写+下划线格式 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> 关联Issue [#6007](https://gitcode.com/cann/ops-nn/issues/6007) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 相关文档已更新 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10839 | 19 天前 | |
format cpp Co-authored-by: yang-di52<yangdi52@huawei.com> # message auto-generated for no-merge-commit merge: !6784 merge issue_fix into master format cpp Created-by: yang-di52 Commit-by: yang-di52 Merged-by: cann-robot Description: ## 描述 批量刷新cpp代码格式 ## 关联的Issue [#3791](https://gitcode.com/cann/ops-nn/issues/3791) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:代码格式化 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!6784 | 3 个月前 | |
feat(index): 新增模板支持index算子连续和非连续的广播场景 Co-authored-by: sikaiwei<sikaiwei1@h-partners.com> # message auto-generated for no-merge-commit merge: !9927 merge index_broadcast into master feat(index): 新增模板支持index算子连续和非连续的广播场景 Created-by: sikaiwei Commit-by: sikaiwei Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 新增index_nocon_broadcast.h和index_broadcast.h模板分别支持index算子非连续广播场景和连续广播场景 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/6183 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> 白盒用例、冒烟等均已通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> aclnnIndex.md和算子目录下的README.md ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!9927 | 15 天前 | |
fix(index_check): enforce Trap violation semantics and complete tiling input validation Co-authored-by: Bright0313<liwenbo84@huawei.com> # message auto-generated for no-merge-commit merge: !10445 merge codex/indexcheck-hardening into master fix(index_check): enforce Trap violation semantics and complete tiling input validation Created-by: Bright0313 Commit-by: Bright0313 Merged-by: cann-robot Description: ## 描述 修复转测代码检视(P1 红线 4 FAIL)与准入扫描发现的共性问题,并按全量代码复检(四轮,共 560+ 条例次)补齐新发现问题。IndexCheck 为索引类 aclnn L2 接口内部使用的 L0 校验算子,本 PR 不改变其接口与分发结构,仅做缺陷修复。 **1. kernel 校验语义修复(P1 红线:assert 释放态退化 no-op——算子唯一功能失效)** - 基线 4 处 ascendc_assert 替换为 Trap() 硬件异常终止 + 新增 bound<=0 防御 Trap(共 5 处);int32 批量路径辅以 ReduceMin/ReduceMax 向量化聚合 **2. tiling 输入校验补全**:判空/格式符/LOG/死存储清理;CheckFormat 运行时 ND 校验;bounds 1-D/rank≤8 校验 **3. op_api 加固**:L0 入口判空(含 executor) **4. 文档对齐(15.4)**:README 约束补齐;产品表对齐 **5. UT 存活修复(顺带)**:scatter_elements UT 守卫 **6-7. 复检第一/二轮修复**:API-1(GetValue→DataCopyPad/裸指针)、RL-6 判空前移、GEN-4.2 魔鬼数字×2、GEN-1.3 死分支、GEN-4.3/SEC-8.4/GEN-2.3×2/style **8. 第三轮修复**:**TIL-3** int64 死队列条件化分配(5 队列 if constexpr 跳过 + tiling 预算 int64=8B)——maxBatchSize 8336→29184(3.5×);SEC-8.4 日志句式;GEN-2.3 op_def.h **9. 第四轮修复(终扫确认轮)**: - **TOPK-1(FAIL 85%)**:SetBlockDim 返回值未校验(tiling.cpp:182,SDK 头文件实证 SetSimdNumBlocks→GetOutputPointer 存在 nullptr→GRAPH_FAILED 失败路径)→ 补 OP_CHECK_IF 校验(与 BesselI1e/Expint 加固同款) - **GEN-1.1(SUSP 70%)**:L0 入口 executor 形参未判空(bounds/indices 已判,防御不对称)→ 纳入既有判空守卫 - **SEC-1.3/8.2(SUSP LOW)**:op_api %zu 传 int64_t(varargs 形式 UB,LP64 下无害)→ static_cast<size_t> 对齐 - **S7(design-check ⚠️)**:补齐 config/ascend950/index_check_binary.json(仓内 26 个注册 950 的 index/ 算子中唯一缺失者,内容同 910b) **不修复项(设计权衡)**:PERF-1 int64 标量主链(四轮判定 PASS/FAIL/SUSP/PASS——向量归约仅支持 half/float 属正确性约束);PERF-3 单缓冲(SUSP LOW 35%,guard 算子延迟优先);PERF-4 PIPE_ALL 屏障(MTE2_S 细粒度留后续) **规则冲突豁免(.clang-format 强制)**:style 2.3 */& 位置(PointerAlignment: Left);style 2.7 单行短函数体(AllowShortFunctionsOnASingleLine: true) **误报项说明**:15.1 产品(旧芯片只查 canndev TBE ini,本算子经 AddConfig 注册);15.3 aclnn 文档(L0-only 无公开 aclnn API,仓内 11 个同类算子同惯例) ## 关联的Issue #5888 ## 测试 - TTK 两方(stat_rel_err,ascend950):all_cases.csv **201/201 PASS**——共五轮复测零退化(含 TIL-3 优化后与本轮修复后各一次终验) - 异常套件 **6/6 有效拦截**(release 真机实测):rank9×2→OPTILING_FAILURE;bool→BINARY_MATCH_FAILURE;oob/bounds_zero×3→VEC_ERROR(kernel Trap 实测触发) - int64 优化效果:maxBatchSize 8336→29184(bin_tiling_data 可核) - build.sh --soc=ascend950/910b 双编译通过(含新增 950 config) - pre-commit 全钩子绿 - 四轮全量代码复检:第四轮红线 10/10 PASS(TOPK-1 修复后);API 条例 0 FAIL;TIL-3 修复经两轮独立复核确认完备 See merge request: cann/ops-nn!10445 | 11 天前 | |
docs: 统一原型注释中的产品名称(master) Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !11186 merge 923soc into master docs: 统一原型注释中的产品名称(master) Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 原型注释里的产品名称仍沿用 Series Product、Ascend 950 AI Processor、Ascend 950PR/Ascend 950DT 等旧表述,与当前产品命名不一致。本次只改注释,不改算子接口和计算逻辑。 合入目标:cann/ops-nn 的 master。 覆盖各算子 op_graph/*_proto.h,以及 common/inc/op_graph/op_nn_proto_extend.h、common/stub/op_graph/math_proto_stub.cpp。 ## 关联的Issue 关联 Issue #6213(Close #6213) https://gitcode.com/cann/ops-nn/issues/6213 ## 测试 仅修改原型注释中的产品名称,不涉及算子 host/kernel 计算逻辑,无需算子泛化。已在本地用仓库 pre-commit 的 end-of-file-fixer 与 clang-format 核对格式,且没有改动算子输入、输出和属性定义。 ## 文档更新 更新了 op_graph 原型注释中的产品名称,无独立 markdown 文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11186 | 11 天前 | |
fix: 修正文档中 aclnnInplaceIndexFill 参数名 Co-authored-by: fazhenyao<fazhenyao@h-partners.com> # message auto-generated for no-merge-commit merge: !11287 merge fix-inplace-index-fill-doc-param into master fix: 修正文档中 aclnnInplaceIndexFill 参数名 Created-by: fazhenyao123 Commit-by: fazhenyao Merged-by: cann-robot Description: ## 描述 修正 aclnnInplaceIndexFillGetWorkspaceSize 文档原型的首个参数名,将 self 改为 selfRef,使其与头文件和实现保持一致。 ## 关联的Issue 无。 ## 测试 - git diff --check - 扫描仓库内该接口的头文件、实现和文档原型,确认均使用 selfRef ## 文档更新 更新 index/index_fill_d/docs/aclnnIndexFill&aclnnInplaceIndexFill.md 中的接口原型。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!11287 | 11 天前 | |
format cpp Co-authored-by: yang-di52<yangdi52@huawei.com> # message auto-generated for no-merge-commit merge: !6784 merge issue_fix into master format cpp Created-by: yang-di52 Commit-by: yang-di52 Merged-by: cann-robot Description: ## 描述 批量刷新cpp代码格式 ## 关联的Issue [#3791](https://gitcode.com/cann/ops-nn/issues/3791) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:代码格式化 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!6784 | 3 个月前 | |
add InferDataType for GatherElements/IndexPutV2/UnsortedSegmentSum Co-authored-by: 贾剑勇<jiajianyong123@hisilicon.com> # message auto-generated for no-merge-commit merge: !10531 merge master into master add InferDataType for GatherElements/IndexPutV2/UnsortedSegmentSum Created-by: jia-jianyong Commit-by: 贾剑勇 Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> add InferDataType for GatherElements/IndexPutV2/UnsortedSegmentSum ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/6172 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> GEIR验证ok,冒烟ok ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 不涉及 ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10531 | 16 天前 | |
算子日志可读性与代码规范性的优化 Co-authored-by: xingtaowu<wuxingtao3@h-partners.com> # message auto-generated for no-merge-commit merge: !9376 merge fix_nn_log into master 算子日志可读性与代码规范性的优化 Created-by: xingtaowu Commit-by: xingtaowu Merged-by: cann-robot Description: ## 描述 批量修正日志/报错信息中的英文拼写错误与语法问题,并补充更明确的参数越界诊断信息,同时修正示例测试程序中的日志级别与代码格式问题。改动不涉及业务逻辑与算法流程变化,均为日志可读性与代码规范性的优化。 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5291 <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 冒烟通过 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ x] 其他,请描述:算子日志可读性与代码规范性的优化 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!9376 | 1 个月前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
fix(cleancode): AICPU算子告警整改(第二批) Co-authored-by: liu-wei<lovline.liuwei@huawei.com> # message auto-generated for no-merge-commit merge: !10662 merge fix/warning-cleanup-v2 into master fix(cleancode): AICPU算子告警整改(第二批) Created-by: liu-wei Commit-by: liu-wei Merged-by: cann-robot Description: ## 描述 ops-nn AICPU算子告警整改第二批,修复3类告警。 ## 改动内容 - G.EXP.15: reinterpret_cast → PtrToPtr (scatter_nd_min/scatter_nd_max/sparse_segment_sum/sparse_to_dense/sparse_fill_empty_rows) - G.CLS.12: 析构函数加 override (index_to_addr) - G.RES.06: lambda默认捕获改为显式捕获 (scatter_nd_max/log_softmax_v2) ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5951 ## 测试 - pre-commit 全部 Passed - aicpu kernel UT: 80/80 PASSED - graph example: 7/7 success - eager example: log_softmax_v2 success ## 类型标签 - [x] ♻️ 重构 - [x] 🧹 代码清理 See merge request: cann/ops-nn!10662 | 21 天前 | |
test(inplace_add): register TensorFlow golden spec Co-authored-by: raoliang_sac<raoliang4@huawei.com> # message auto-generated for no-merge-commit merge: !10750 merge test/inplace-add-tf-e2e-spec into master test(inplace_add): register TensorFlow golden spec Created-by: raoliang_sac Commit-by: raoliang_sac Merged-by: cann-robot Description: ## 描述 为 InplaceAdd 的 TTK golden 补充 TensorFlow E2E TestSpec 注册: - 注册 tf.raw_ops.InplaceAdd 与 tensorflow.raw_ops.InplaceAdd 两种 CSV api_name; - 增加 TensorFlow E2E CPU 高精度 golden 及原生 TensorFlow third-party 映射; - 保留现有 kernel/GEIR、legacy golden 和 Torch third-party 行为不变。 该文件与已验证的 operator-artifacts InplaceAdd golden 内容完全一致。 ## 关联的Issue 无。 ## 测试 - ast.parse 语法检查通过; - Spec 注册键、third-party 映射及最小 FP32 数值用例通过; - ruff check、ruff format --check、git diff --check 通过; - 提交阶段 pre-commit(含 OAT 增量检查)通过; - 仅修改 golden 测试资产,未重新编译或部署算子。 ## 文档更新 未修改用户文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:测试 golden/TTK 注册完善 ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!10750 | 21 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
算子DVLite芯片适配 Co-authored-by: gcw_DS4cmz2b<1312925094@qq.com> # message auto-generated for no-merge-commit merge: !10012 merge DVLite into master 算子DVLite芯片适配 Created-by: gcw_DS4cmz2b Commit-by: gcw_DS4cmz2b Merged-by: cann-robot Description: ## 描述 repeat_interleave_grad和where等4个算子新增DVLite适配,where修改了fusion_pass融合规则,repeat_interleave_grad和cross_entropy_sum_exp_and_index_logit改成了新版的Cmakelists同时适配了DVLite芯片版本。 ## 关联的Issue [算子缺少350芯片适配](https://gitcode.com/cann/ops-nn/issues/5748) ## 测试 算子二级冒烟通过,aclnn通路验证通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10012 | 28 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
docs: 补齐repeat_interleave与scatter_nd_update文档数据类型,统一index_add日志格式 Co-authored-by: AdogDboy<liuchenghao14@huawei.com> # message auto-generated for no-merge-commit merge: !10237 merge feat9_10 into master docs: 补齐repeat_interleave与scatter_nd_update文档数据类型,统一index_add日志格式 Created-by: AdogDboy Commit-by: AdogDboy Merged-by: cann-robot Description: 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 本 PR 修复文档与代码实际支持的数据类型不一致问题,并统一 aclnnIndexAdd 报错日志格式。用户按文档使用文档中缺失的数据类型时,实际代码可以正常执行,文档与实现不一致。 改动内容: 1. aclnnRepeatInterleave / aclnnRepeatInterleaveInt / aclnnRepeatInterleaveIntWithDim / aclnnRepeatInterleaveWithDim 4 个接口文档:self/out 参数支持的数据类型从 9 种补齐为 12 种,新增 UINT16、UINT32、UINT64。代码依据:index/repeat_interleave/op_api/aclnn_repeat_interleave.cpp 中 ASCEND950_DTYPE_SUPPORT_LIST_SELF(Ascend 950 平台共支持 12 种类型),4 个变体接口共用同一 dtype 支持列表,且 out 强制与 self 类型一致。 2. aclnnScatterNdUpdate 接口文档:varRef/updates 参数补齐 UINT8。代码依据:index/scatter_nd_update/op_api/aclnn_scatter_nd_update.cpp 中 ASCEND950_DTYPE_DTYPE_SUPPORT_LIST,且代码要求 updates 与 varRef 数据类型一致。 3. aclnnIndexAdd:CheckDtypeValid 中 4 条 dtype 校验报错日志去掉支持列表的方括号,与仓库主流日志风格统一,仅日志文案调整,无逻辑改动。 平台说明:新增的 UINT16/UINT32/UINT64/UINT8 仅 Ascend 950PR/Ascend 950DT 系列产品支持(A2/A3 平台走各自的 dtype 支持列表,不含上述类型)。 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> [Documentation|文档反馈]: aclnnRepeatInterleave系列及aclnnScatterNdUpdate文档支持的数据类型与代码不一致 #5769 [Bug-Report|缺陷反馈]: aclnnIndexAdd dtype校验失败时日志打印变量与文案和实际校验对象不符 #5770 测试 <!--描述进行了哪些测试来验证你的改动。--> - 逐行对照 op_api 层代码的 dtype 支持列表核验文档描述(ASCEND950_DTYPE_SUPPORT_LIST_SELF / ASCEND950_DTYPE_DTYPE_SUPPORT_LIST); - git diff 确认改动范围:6 个文件、14 行(4 个 repeat_interleave 文档 self/out 行、scatter_nd_update 文档 varRef/updates 行、index_add 4 条日志字符串); - 文档改动不涉及功能逻辑;aclnn_index_add.cpp 仅修改日志字符串字面量,不影响编译产物行为; - 保留各文档所在分支基线的官方排版规范(中英文空格),无冲突标记残留。 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> - index/repeat_interleave/docs/aclnnRepeatInterleave.md(self/out 数据类型) - index/repeat_interleave/docs/aclnnRepeatInterleaveInt.md(self/out 数据类型) - index/repeat_interleave/docs/aclnnRepeatInterleaveIntWithDim.md(self/out 数据类型) - index/repeat_interleave/docs/aclnnRepeatInterleaveWithDim.md(self/out 数据类型) - index/scatter_nd_update/docs/aclnnScatterNdUpdate.md(varRef/updates 数据类型) 类型标签 <!-- [x] 表示选中 --> - Bug修复 - 新特性 - 性能优化 - 文档更新 - 其他,请描述: AI/Agent生成声明 <!-- [x] 表示选中 --> - AI辅助编写 See merge request: cann/ops-nn!10237 | 28 天前 | |
fix: ApplyAdagrad,HardSwishGradV2,HardSigmoid,InplaceSub检视意见闭环 Co-authored-by: tianyu52<tianyu52@huawei.com> # message auto-generated for no-merge-commit merge: !10724 merge master into master fix: ApplyAdagrad,HardSwishGradV2,HardSigmoid,InplaceSub检视意见闭环 Created-by: tianyu52 Commit-by: tianyu52 Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ApplyAdagrad,HardSwishGradV2,HardSigmoid,InplaceSub检视意见闭环 1.补充通路golden及测试用例 2.增加校验保护 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> UT/ST 通路验证 PASS ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10724 | 21 天前 | |
feat: 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 Co-authored-by: xuejinghui<xuejinghui@huawei.com> # message auto-generated for no-merge-commit merge: !8821 merge InferShape into master feat: 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 Created-by: xuejinghui Commit-by: xuejinghui Merged-by: cann-robot Description: ## 描述 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5250 <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:Unknown Shape/ Unknown Rank迁移适配 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!8821 | 1 个月前 | |
feat(unique): add Regbase values-only AI Core support Co-authored-by: ConanHuang<huangxiaobin1@huawei.com> # message auto-generated for no-merge-commit merge: !10993 merge feature_op_unique into master feat(unique): add Regbase values-only AI Core support Created-by: ConanHuang Commit-by: ConanHuang Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 为 aclnnUnique / aclnnUnique2 增加 Regbase values-only AI Core 路径,在单次 Kernel 调用中完成排序与去重,减少不必要的索引搬运和中间存储。 - 新增 arch35 Tiling、Kernel 和动态输出 shape 回传,接入 ACLNN 分流。 - 整理算子 L0、原型及 Ascend 950/350 构建注册;调整共享 radix 和 UniqueConsecutive 的大输入计算。 - 新增 ACLNN/Kernel 测试用例、golden 和 GEIR 示例。 新路径仅处理不返回 inverse/counts 且满足准入条件的调用,其余场景沿用既有路径;GE/ONNX 仍使用 AICPU。详细 Tiling 与 Kernel 设计见关联 Issue。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6187 ## 测试 - 二级冒烟通过 - UniqueWithCountsAndSorting 和 KthValue 各自 300+ kernel ST通过 - aclnnUnique aclnnUnique2 各自 170+ aclnn用例通过 - onnx geir 通路兼容性验证通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 不涉及 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10993 | 15 天前 | |
算子日志可读性与代码规范性的优化 Co-authored-by: xingtaowu<wuxingtao3@h-partners.com> # message auto-generated for no-merge-commit merge: !9376 merge fix_nn_log into master 算子日志可读性与代码规范性的优化 Created-by: xingtaowu Commit-by: xingtaowu Merged-by: cann-robot Description: ## 描述 批量修正日志/报错信息中的英文拼写错误与语法问题,并补充更明确的参数越界诊断信息,同时修正示例测试程序中的日志级别与代码格式问题。改动不涉及业务逻辑与算法流程变化,均为日志可读性与代码规范性的优化。 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5291 <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 冒烟通过 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ x] 其他,请描述:算子日志可读性与代码规范性的优化 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!9376 | 1 个月前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
docs: 统一原型注释中的产品名称(master) Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !11186 merge 923soc into master docs: 统一原型注释中的产品名称(master) Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 原型注释里的产品名称仍沿用 Series Product、Ascend 950 AI Processor、Ascend 950PR/Ascend 950DT 等旧表述,与当前产品命名不一致。本次只改注释,不改算子接口和计算逻辑。 合入目标:cann/ops-nn 的 master。 覆盖各算子 op_graph/*_proto.h,以及 common/inc/op_graph/op_nn_proto_extend.h、common/stub/op_graph/math_proto_stub.cpp。 ## 关联的Issue 关联 Issue #6213(Close #6213) https://gitcode.com/cann/ops-nn/issues/6213 ## 测试 仅修改原型注释中的产品名称,不涉及算子 host/kernel 计算逻辑,无需算子泛化。已在本地用仓库 pre-commit 的 end-of-file-fixer 与 clang-format 核对格式,且没有改动算子输入、输出和属性定义。 ## 文档更新 更新了 op_graph 原型注释中的产品名称,无独立 markdown 文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11186 | 11 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
fix: 修复文件重名问题 Co-authored-by: duxinlei<duxinlei@h-partners.com> # message auto-generated for no-merge-commit merge: !9675 merge checkRepeat2 into master fix: 修复文件重名问题 Created-by: duxinlei Commit-by: duxinlei Merged-by: cann-robot Description: ## 描述 解决单仓内和多仓间文件名重复问题, 1、对于将公共文件引入单算子目录的重命名问题,改引用公共目录下的文件 2、对于单纯重名问题,重命名 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> 关联issue #5460 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!9675 | 1 个月前 | |
md大模型检测低错修复 Co-authored-by: gitcode-chenjiao<chenjiao31@huawei.com> # message auto-generated for no-merge-commit merge: !9420 merge master into master md大模型检测低错修复 Created-by: gitcode-chenjiao Commit-by: gitcode-chenjiao Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> md aidd大模型检测低错修复 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/5254 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ok ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> acl*.md ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!9420 | 1 个月前 | |
perf(index): 优化 KthValue、Median 与 NanMedian 选择及归并路径 Co-authored-by: ConanHuang<huangxiaobin1@huawei.com> # message auto-generated for no-merge-commit merge: !10411 merge master into master perf(index): 优化 KthValue、Median 与 NanMedian 选择及归并路径 Created-by: ConanHuang Commit-by: ConanHuang Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 优化 KthValue、Median、NanMedian 的选择与共享归并路径,减少完整排序、重复扫描和逐行控制开销,保持各算子的目标秩、NaN 和索引语义。 - 改进 radix select:累积字节直方图定位 bucket、UB 驻留 key、相同高位跳过及候选位置压缩,减少无效扫描。 - 补充窄轴 SIMT select、短秩选择和适用场景的驻留直方图路径,按 dtype、轴长、行数及 UB 容量选择路由。 - 同步公共 ping-pong 归并和批处理优化,完善 INT16 merge 路径与 host/kernel UB 布局对应关系。 - 精简相关 tiling 结构与分发,使用编译期 Median 模式分支;补充浮点 key、正负零、尾块及同步约束说明。 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/5977 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> - Ascend950PR_9579:KthValue 338 条、Median 260 条、NanMedian 265 条全量 ST,加正负零逐位比较 32 条,合计 **895 条精度全部通过**。 - 相对前一版已验证实现,性能异常项复测及旧包/当前包同机对照后,未确认稳定超过 5% 且超过 0.5 μs 的设备耗时劣化。 - 二级冒烟通过。 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 不涉及 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [x] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10411 | 21 天前 | |
format cpp Co-authored-by: yang-di52<yangdi52@huawei.com> # message auto-generated for no-merge-commit merge: !6784 merge issue_fix into master format cpp Created-by: yang-di52 Commit-by: yang-di52 Merged-by: cann-robot Description: ## 描述 批量刷新cpp代码格式 ## 关联的Issue [#3791](https://gitcode.com/cann/ops-nn/issues/3791) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:代码格式化 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!6784 | 3 个月前 | |
fix(index): KthValue公共头文件命名及代码质量整改 Co-authored-by: ConanHuang<huangxiaobin1@huawei.com> # message auto-generated for no-merge-commit merge: !10851 merge master into master fix(index): KthValue公共头文件命名及代码质量整改 Created-by: ConanHuang Commit-by: ConanHuang Merged-by: cann-robot Description: ## 描述 为 KthValue 的 12 个 common 头文件增加 kth_value_ 前缀,并同步更新引用和头文件保护宏;修正 aclnnKthvalue 数据类型说明;为只读 tiling 参数增加 const;区分 NanMedian 三处 ConvertToTensor 失败日志。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5998 https://gitcode.com/cann/ops-nn/issues/6002 ## 测试 - 二级冒烟通过 - 使用仓内 CSV 和 golden 进行 TTK release 验证:KthValue 30/30、Median 26/26、NanMedian 26/26,合计 82/82 通过。 ## 文档更新 更新 aclnn_kthvalue.h 中的参数数据类型说明。 ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10851 | 19 天前 | |
NonZero算子ascend350平台适配修改 Co-authored-by: zhangxiyan7<zhangxiyan7@huawei.com> # message auto-generated for no-merge-commit merge: !11124 merge DVLite into master NonZero算子ascend350平台适配修改 Created-by: zhangxiyan7 Commit-by: zhangxiyan7 Merged-by: cann-robot Description: ## 描述 aclnnNonzeroV2 在 Ascend350 上报错: AclNN_Parameter_Error(EZ1001):Tensor self not inplemented for DT_BFLOAT16, should be in dtype support list [DT_FLOAT,DT_FLOAT16,DT_INT8,DT_UINT8,DT_INT16,DT_UINT16,DT_INT32,DT_UINT32,DT_INT64,DT_UINT64,DT_DOUBLE,DT_BOOL,]. 原因是适配350平台时未将socversion更新为npuarch格式,导致350平台无法正确识别版本导致走到错误分支。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6259 ## 测试 950 350平台泛化测试对比一致 冒烟测试通过 ## 文档更新 不涉及 ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!11124 | 11 天前 | |
docs: 统一原型注释中的产品名称(master) Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !11186 merge 923soc into master docs: 统一原型注释中的产品名称(master) Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 原型注释里的产品名称仍沿用 Series Product、Ascend 950 AI Processor、Ascend 950PR/Ascend 950DT 等旧表述,与当前产品命名不一致。本次只改注释,不改算子接口和计算逻辑。 合入目标:cann/ops-nn 的 master。 覆盖各算子 op_graph/*_proto.h,以及 common/inc/op_graph/op_nn_proto_extend.h、common/stub/op_graph/math_proto_stub.cpp。 ## 关联的Issue 关联 Issue #6213(Close #6213) https://gitcode.com/cann/ops-nn/issues/6213 ## 测试 仅修改原型注释中的产品名称,不涉及算子 host/kernel 计算逻辑,无需算子泛化。已在本地用仓库 pre-commit 的 end-of-file-fixer 与 clang-format 核对格式,且没有改动算子输入、输出和属性定义。 ## 文档更新 更新了 op_graph 原型注释中的产品名称,无独立 markdown 文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11186 | 11 天前 | |
feat(aicpu): enable constant folding for migrated ops Co-authored-by: hid29204727<wangboyu20@huawei.com> # message auto-generated for no-merge-commit merge: !9280 merge codex/aicpu-const-folding into master feat(aicpu): enable constant folding for migrated ops Created-by: hid29204727 Commit-by: hid29204727 Merged-by: cann-robot Description: ## 描述 补充 ops-nn 仓 AICPU host 常量折叠注册通路,并为当前已标记 OPEN_OPS_FLAG 的 18 个迁移算子统一开启常量折叠支持。 主要改动: - 新增仓级 OPS_NN_REGISTER_CPU_KERNELV2 注册包装宏,host 构建优先注册 V2,device/custom/UT 保持 V1 回退。 - 为 add_modules_sources 和 add_all_modules_sources 增加 HOSTCPU 参数及 host OBJECT 收集逻辑。 - 为 18 个 OPEN_OPS_FLAG 算子补充 HOSTCPU TRUE、注册头文件和 V2 注册。 - host OBJECT 汇入 libopconstant_folding_nn.so。 ## 关联的Issue 无。 ## 测试 已完成以下静态验证: - git diff --check 通过。 - 18 个 OPEN_OPS_FLAG 算子均具备 HOSTCPU TRUE。 - 18 个算子均使用 OPS_NN_REGISTER_CPU_KERNELV2,无直接 V1 注册残留。 - 注册头文件、host 编译宏和仓级 so 收集通路检查通过。 当前 Windows 环境缺少 CANN、NPU、CMake 和远端构建桥,未执行 bash build.sh --pkg -j16;完整打包构建及常量折叠功能验证待 CANN Linux 环境/CI 完成。 ## 文档更新 无仓内文档变更。 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!9280 | 21 天前 | |
feat(aicpu): enable constant folding for migrated ops Co-authored-by: hid29204727<wangboyu20@huawei.com> # message auto-generated for no-merge-commit merge: !9280 merge codex/aicpu-const-folding into master feat(aicpu): enable constant folding for migrated ops Created-by: hid29204727 Commit-by: hid29204727 Merged-by: cann-robot Description: ## 描述 补充 ops-nn 仓 AICPU host 常量折叠注册通路,并为当前已标记 OPEN_OPS_FLAG 的 18 个迁移算子统一开启常量折叠支持。 主要改动: - 新增仓级 OPS_NN_REGISTER_CPU_KERNELV2 注册包装宏,host 构建优先注册 V2,device/custom/UT 保持 V1 回退。 - 为 add_modules_sources 和 add_all_modules_sources 增加 HOSTCPU 参数及 host OBJECT 收集逻辑。 - 为 18 个 OPEN_OPS_FLAG 算子补充 HOSTCPU TRUE、注册头文件和 V2 注册。 - host OBJECT 汇入 libopconstant_folding_nn.so。 ## 关联的Issue 无。 ## 测试 已完成以下静态验证: - git diff --check 通过。 - 18 个 OPEN_OPS_FLAG 算子均具备 HOSTCPU TRUE。 - 18 个算子均使用 OPS_NN_REGISTER_CPU_KERNELV2,无直接 V1 注册残留。 - 注册头文件、host 编译宏和仓级 so 收集通路检查通过。 当前 Windows 环境缺少 CANN、NPU、CMake 和远端构建桥,未执行 bash build.sh --pkg -j16;完整打包构建及常量折叠功能验证待 CANN Linux 环境/CI 完成。 ## 文档更新 无仓内文档变更。 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!9280 | 21 天前 | |
docs: 统一原型注释中的产品名称(master) Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !11186 merge 923soc into master docs: 统一原型注释中的产品名称(master) Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 原型注释里的产品名称仍沿用 Series Product、Ascend 950 AI Processor、Ascend 950PR/Ascend 950DT 等旧表述,与当前产品命名不一致。本次只改注释,不改算子接口和计算逻辑。 合入目标:cann/ops-nn 的 master。 覆盖各算子 op_graph/*_proto.h,以及 common/inc/op_graph/op_nn_proto_extend.h、common/stub/op_graph/math_proto_stub.cpp。 ## 关联的Issue 关联 Issue #6213(Close #6213) https://gitcode.com/cann/ops-nn/issues/6213 ## 测试 仅修改原型注释中的产品名称,不涉及算子 host/kernel 计算逻辑,无需算子泛化。已在本地用仓库 pre-commit 的 end-of-file-fixer 与 clang-format 核对格式,且没有改动算子输入、输出和属性定义。 ## 文档更新 更新了 op_graph 原型注释中的产品名称,无独立 markdown 文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11186 | 11 天前 | |
repeat_interleave性能优化 Co-authored-by: 张鑫<zhangxin660@h-partners.com> # message auto-generated for no-merge-commit merge: !10057 merge REG_OP_Gather into master repeat_interleave性能优化 Created-by: z30075199 Commit-by: 张鑫 Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> repeat_interleave性能优化 优化四个场景的性能(新增2个模板,修改2个模板) 对于大batch小cp场景,使用前缀和+simt实现 对于切repeat场景,使用输出分核策略+前缀和获取输出偏移 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/6243 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> 相关场景st用例100条,平均性能提升44倍 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [x] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10057 | 11 天前 | |
重命名RIG算子中的重复文件名 platform.h Co-authored-by: wow5522<wuao1@huawei.com> # message auto-generated for no-merge-commit merge: !10798 merge master into master 重命名RIG算子中的重复文件名 platform.h Created-by: wow5523 Commit-by: wow5522 Merged-by: cann-robot Description: ## 描述 把repeat_interleave_grad算子中的文件名 platform.h改成repeat_interleave_grad_platform.h <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6030 <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10798 | 19 天前 | |
fix(cleancode): AICPU算子告警整改(第三批) Co-authored-by: liu-wei<lovline.liuwei@huawei.com> # message auto-generated for no-merge-commit merge: !10997 merge fix/warning-cleanup-v3 into master fix(cleancode): AICPU算子告警整改(第三批) Created-by: liu-wei Commit-by: liu-wei Merged-by: cann-robot Description: ## 描述 ops-nn AICPU 算子告警整改第三批,修复 18 条告警(G.CNS.03 ×12 / G.EXP.27+36 ×5 / G.CNS.04 ×1),共 14 个文件 +29/-29: | 告警 | 算子 | 修复 | |---|---|---| | G.CNS.03 ×12 | log_softmax_v2、scatter_elements、sparse_segment_mean、sparse_segment_sum、sparse_to_dense、tensor_scatter_update | 12 处成员函数补充 const(声明与定义同步,被调成员均为 const 或自由函数) | | G.EXP.27/36 ×4 | softmax_v2 | NormalCheck(...)/detail::SoftmaxV2Check(...) 两处 ?: 条件改为显式 != KERNEL_STATUS_OK | | G.EXP.27 ×1 | reverse_sequence | reverseNum % kEven 显式比较 != 0 | | G.CNS.04 ×1 | tensor_scatter_update | GetBatchStrides 的 outer_shape 参数加 const(仅读取) | ## 不处理项(附判定依据) - 浮点 == static_cast<T>(0) ×2(log_softmax_v2/softmax_v2 的除零防护):保持语义,建议工具豁免; - sparse_to_dense G.STD.02 ×4:char* 为字节缓冲非字符串,误报; - EigenSparseToDense 的 SparseTensor& st:st.ToDense<>() 受 Eigen API 限制,疑似误报。 ## 验证 - 8 个改动 TU 语法级编译通过(g++ -std=c++14); - 二进制级对比(父提交 467b22d6c vs 本提交 7f12ebcbc,同编译参数 g++ -O2 -std=c++14):8 个 TU 全部机器码/数据节**内容多重集逐字节一致**(合计 1298 节),仅 const 成员函数 mangled 名(符号表及 scatter_elements 15 节、sparse_to_dense 1 节的内嵌签名串 ZN→ZNK)变化——**零功能影响**; - CI 流水线待跑。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6114 ## 类型标签 - [x] ♻️ 重构 - [x] 🧹 代码清理 See merge request: cann/ops-nn!10997 | 17 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
docs(index): scatter Mul/Div/Max/Min 资料与 UT 注释同步 int64 直排后的支持面 Co-authored-by: tangpingchuan<tangpingchuan@huawei.com> # message auto-generated for no-merge-commit merge: !10794 merge fix/scatter-doc-ut-sync into master docs(index): scatter Mul/Div/Max/Min 资料与 UT 注释同步 int64 直排后的支持面 Created-by: zl_hw Commit-by: tangpingchuan Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10794 | 19 天前 | |
docs: 统一原型注释中的产品名称(master) Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !11186 merge 923soc into master docs: 统一原型注释中的产品名称(master) Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 原型注释里的产品名称仍沿用 Series Product、Ascend 950 AI Processor、Ascend 950PR/Ascend 950DT 等旧表述,与当前产品命名不一致。本次只改注释,不改算子接口和计算逻辑。 合入目标:cann/ops-nn 的 master。 覆盖各算子 op_graph/*_proto.h,以及 common/inc/op_graph/op_nn_proto_extend.h、common/stub/op_graph/math_proto_stub.cpp。 ## 关联的Issue 关联 Issue #6213(Close #6213) https://gitcode.com/cann/ops-nn/issues/6213 ## 测试 仅修改原型注释中的产品名称,不涉及算子 host/kernel 计算逻辑,无需算子泛化。已在本地用仓库 pre-commit 的 end-of-file-fixer 与 clang-format 核对格式,且没有改动算子输入、输出和属性定义。 ## 文档更新 更新了 op_graph 原型注释中的产品名称,无独立 markdown 文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11186 | 11 天前 | |
docs: 统一原型注释中的产品名称(master) Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !11186 merge 923soc into master docs: 统一原型注释中的产品名称(master) Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 原型注释里的产品名称仍沿用 Series Product、Ascend 950 AI Processor、Ascend 950PR/Ascend 950DT 等旧表述,与当前产品命名不一致。本次只改注释,不改算子接口和计算逻辑。 合入目标:cann/ops-nn 的 master。 覆盖各算子 op_graph/*_proto.h,以及 common/inc/op_graph/op_nn_proto_extend.h、common/stub/op_graph/math_proto_stub.cpp。 ## 关联的Issue 关联 Issue #6213(Close #6213) https://gitcode.com/cann/ops-nn/issues/6213 ## 测试 仅修改原型注释中的产品名称,不涉及算子 host/kernel 计算逻辑,无需算子泛化。已在本地用仓库 pre-commit 的 end-of-file-fixer 与 clang-format 核对格式,且没有改动算子输入、输出和属性定义。 ## 文档更新 更新了 op_graph 原型注释中的产品名称,无独立 markdown 文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11186 | 11 天前 | |
[Test] 补充 ScatterListTorchSpec 的 torch 三方标杆注册 Co-authored-by: chendunyang<chendunyang1@huawei.com> # message auto-generated for no-merge-commit merge: !10238 merge fix/scatter-list-torch-third-party into master [Test] 补充 ScatterListTorchSpec 的 torch 三方标杆注册 Created-by: qq_37913898 Commit-by: chendunyang Merged-by: cann-robot Description: ## 问题 ScatterListTorchSpec 缺少 third_party 注册,E2E 三方比对无法从该 Spec 获取 torch 组合实现。 ## 修改 - 补充 third_party = {"torch": _Compose}。 - 新增最小参数名适配,将 Torch API 的 input/indices/updates 转交现有 _ScatterListCompose,mask、axis 等参数原样透传。 - 不修改现有计算逻辑、CPU golden、Kernel/ACLNN 注册或精度标准。仅修改 index/scatter_list/tests/assets/golden.py,新增 5 行。 ## 验证 - 使用真实 TTK get_spec_attr 成功加载 torch_npu.npu_scatter_list 的 third_party。 See merge request: cann/ops-nn!10238 | 29 天前 | |
docs(index): scatter Mul/Div/Max/Min 资料与 UT 注释同步 int64 直排后的支持面 Co-authored-by: tangpingchuan<tangpingchuan@huawei.com> # message auto-generated for no-merge-commit merge: !10794 merge fix/scatter-doc-ut-sync into master docs(index): scatter Mul/Div/Max/Min 资料与 UT 注释同步 int64 直排后的支持面 Created-by: zl_hw Commit-by: tangpingchuan Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10794 | 19 天前 | |
feat: 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 Co-authored-by: xuejinghui<xuejinghui@huawei.com> # message auto-generated for no-merge-commit merge: !8821 merge InferShape into master feat: 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 Created-by: xuejinghui Commit-by: xuejinghui Merged-by: cann-robot Description: ## 描述 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5250 <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:Unknown Shape/ Unknown Rank迁移适配 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!8821 | 1 个月前 | |
docs(index): scatter Mul/Div/Max/Min 资料与 UT 注释同步 int64 直排后的支持面 Co-authored-by: tangpingchuan<tangpingchuan@huawei.com> # message auto-generated for no-merge-commit merge: !10794 merge fix/scatter-doc-ut-sync into master docs(index): scatter Mul/Div/Max/Min 资料与 UT 注释同步 int64 直排后的支持面 Created-by: zl_hw Commit-by: tangpingchuan Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10794 | 19 天前 | |
docs(index): scatter Mul/Div/Max/Min 资料与 UT 注释同步 int64 直排后的支持面 Co-authored-by: tangpingchuan<tangpingchuan@huawei.com> # message auto-generated for no-merge-commit merge: !10794 merge fix/scatter-doc-ut-sync into master docs(index): scatter Mul/Div/Max/Min 资料与 UT 注释同步 int64 直排后的支持面 Created-by: zl_hw Commit-by: tangpingchuan Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10794 | 19 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
fix(cleancode): AICPU算子告警整改(第二批) Co-authored-by: liu-wei<lovline.liuwei@huawei.com> # message auto-generated for no-merge-commit merge: !10662 merge fix/warning-cleanup-v2 into master fix(cleancode): AICPU算子告警整改(第二批) Created-by: liu-wei Commit-by: liu-wei Merged-by: cann-robot Description: ## 描述 ops-nn AICPU算子告警整改第二批,修复3类告警。 ## 改动内容 - G.EXP.15: reinterpret_cast → PtrToPtr (scatter_nd_min/scatter_nd_max/sparse_segment_sum/sparse_to_dense/sparse_fill_empty_rows) - G.CLS.12: 析构函数加 override (index_to_addr) - G.RES.06: lambda默认捕获改为显式捕获 (scatter_nd_max/log_softmax_v2) ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5951 ## 测试 - pre-commit 全部 Passed - aicpu kernel UT: 80/80 PASSED - graph example: 7/7 success - eager example: log_softmax_v2 success ## 类型标签 - [x] ♻️ 重构 - [x] 🧹 代码清理 See merge request: cann/ops-nn!10662 | 21 天前 | |
告警清理 Co-authored-by: zcy123456x<zhouchenyang5@huawei.com> # message auto-generated for no-merge-commit merge: !10737 merge 0918-codewarning into master 告警清理 Created-by: zcy123456x Commit-by: zcy123456x Merged-by: cann-robot Description: ## 描述 本次变更用于清理 AICPU 算子代码告警,主要包括: 1. 避免直接使用 reinterpret_cast - 将部分 GetData() 后的指针转换改为仓内通用的 PtrToPtr<void, T> 写法。 - 涉及 ScatterElements、SparseSegmentMean、AvgPool1DAvgMatrix 等算子。 2. 避免 lambda 使用默认捕获模式 - 将 [&] 默认捕获改为显式捕获,明确 lambda 依赖的局部变量或成员。 - 涉及 ScatterElements、ScatterNdMin、SparseToDense、TensorScatterUpdate 等算子。 3. 避免宏定义依赖宏外部局部变量 - SparseFillEmptyRows 的数据类型分发宏改为显式传入 ret、ctx、indices、values、denseShape、defaultValue。 4. 补充虚函数重写标识 - 为 AvgPool1DAvgMatrixCpuKernel 析构函数补充 override。 本次修改不涉及算子功能逻辑变更,主要为静态检查/代码规范告警清理。 ## 关联的Issue 本次不关联ISSUE ## 测试 黄区验证门禁,LLT,build构建 ## 文档更新 不涉及文档更新 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10737 | 19 天前 | |
feat(scatter_nd_sub): indices支持1D(单条深度K索引),与A2语义对齐 Co-authored-by: aiteer<lujiale4@huawei.com> # message auto-generated for no-merge-commit merge: !10563 merge scatterndsub_916 into master feat(scatter_nd_sub): indices支持1D(单条深度K索引),与A2语义对齐 Created-by: aiteer Commit-by: aiteer Merged-by: cann-robot Description: ## 描述 ScatterNdSub 算子 A5(ascend950/arch35)host 侧 tiling 强制要求 indices rank ≥ 2,与 A2 不兼容,本修改支持indices 1D(单条深度K索引),与A2语义对齐。 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> [#5980](https://gitcode.com/cann/ops-nn/issues/5980) ## 测试 ut测试,根据问题单用例进行ttk测试 ## 文档更新 更新scatter_nd_sub算子README.md,支持一维indices。 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10563 | 21 天前 | |
fix:scatter_update算子修复update为scalar标量时pcie through通路aicore问题&修复scatter_nd_update算子重复索引情况下的精度问题 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !10220 merge dev_zl_2026_0806 into master fix:scatter_update算子修复update为scalar标量时pcie through通路aicore问题&修复scatter_nd_update算子重复索引情况下的精度问题 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 scatter_update算子修复update为scalar标量时pcie through通路aicore问题 scatter_nd_update算子修复pcie through场景,同一个核处理重复索引情况下的精度问题 本PR包含以下改动: ### 修复scatter_update算子标量更新pcie through通路问题 - 在tiling侧将 isPcieThrough标记存入tilingData结构体,使kernel侧可感知pcie through模式 - kernel侧在标量更新场景下,pcie through通路改用DataCopyPad搬运数据,代替GetValue+SyncStoV方式 ### 修复scatter_nd_update算子重复索引情况下的精度问题 - 确定性语义要求重复索引同地址"后写覆盖先写",需显式等待MTE3写落盘. ## 关联的Issue #5819 ## 测试 问题用例已回归通过,白盒用例,冒烟测试已通过 ## 文档更新 NA ## 类型标签 - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!10220 | 17 天前 | |
feat: 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 Co-authored-by: xuejinghui<xuejinghui@huawei.com> # message auto-generated for no-merge-commit merge: !8821 merge InferShape into master feat: 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 Created-by: xuejinghui Commit-by: xuejinghui Merged-by: cann-robot Description: ## 描述 18个迁移算子适配unknown shape/rank InferShape并补UT及入口日志 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5250 <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:Unknown Shape/ Unknown Rank迁移适配 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!8821 | 1 个月前 | |
ScatterMul/Div/Max/Min 排序路径改直排 int64 key,消除 var 首维能力上限 Co-authored-by: tangpingchuan<tangpingchuan@huawei.com> # message auto-generated for no-merge-commit merge: !10655 merge feat/scatter-int64-index into master ScatterMul/Div/Max/Min 排序路径改直排 int64 key,消除 var 首维能力上限 Created-by: zl_hw Commit-by: tangpingchuan Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10655 | 21 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
fix:scatter_update算子修复update为scalar标量时pcie through通路aicore问题&修复scatter_nd_update算子重复索引情况下的精度问题 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !10220 merge dev_zl_2026_0806 into master fix:scatter_update算子修复update为scalar标量时pcie through通路aicore问题&修复scatter_nd_update算子重复索引情况下的精度问题 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 scatter_update算子修复update为scalar标量时pcie through通路aicore问题 scatter_nd_update算子修复pcie through场景,同一个核处理重复索引情况下的精度问题 本PR包含以下改动: ### 修复scatter_update算子标量更新pcie through通路问题 - 在tiling侧将 isPcieThrough标记存入tilingData结构体,使kernel侧可感知pcie through模式 - kernel侧在标量更新场景下,pcie through通路改用DataCopyPad搬运数据,代替GetValue+SyncStoV方式 ### 修复scatter_nd_update算子重复索引情况下的精度问题 - 确定性语义要求重复索引同地址"后写覆盖先写",需显式等待MTE3写落盘. ## 关联的Issue #5819 ## 测试 问题用例已回归通过,白盒用例,冒烟测试已通过 ## 文档更新 NA ## 类型标签 - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!10220 | 17 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
告警清理 Co-authored-by: zcy123456x<zhouchenyang5@huawei.com> # message auto-generated for no-merge-commit merge: !10737 merge 0918-codewarning into master 告警清理 Created-by: zcy123456x Commit-by: zcy123456x Merged-by: cann-robot Description: ## 描述 本次变更用于清理 AICPU 算子代码告警,主要包括: 1. 避免直接使用 reinterpret_cast - 将部分 GetData() 后的指针转换改为仓内通用的 PtrToPtr<void, T> 写法。 - 涉及 ScatterElements、SparseSegmentMean、AvgPool1DAvgMatrix 等算子。 2. 避免 lambda 使用默认捕获模式 - 将 [&] 默认捕获改为显式捕获,明确 lambda 依赖的局部变量或成员。 - 涉及 ScatterElements、ScatterNdMin、SparseToDense、TensorScatterUpdate 等算子。 3. 避免宏定义依赖宏外部局部变量 - SparseFillEmptyRows 的数据类型分发宏改为显式传入 ret、ctx、indices、values、denseShape、defaultValue。 4. 补充虚函数重写标识 - 为 AvgPool1DAvgMatrixCpuKernel 析构函数补充 override。 本次修改不涉及算子功能逻辑变更,主要为静态检查/代码规范告警清理。 ## 关联的Issue 本次不关联ISSUE ## 测试 黄区验证门禁,LLT,build构建 ## 文档更新 不涉及文档更新 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10737 | 19 天前 | |
perf: 优化 SparseSegmentMean 和 SparseSegmentSum AICPU 性能 Co-authored-by: Ding_Jing<dingjing19@huawei.com> # message auto-generated for no-merge-commit merge: !11057 merge perf/sparse-segment-eigen5 into master perf: 优化 SparseSegmentMean 和 SparseSegmentSum AICPU 性能 Created-by: Ding_Jing Commit-by: Ding_Jing Merged-by: cann-robot Description: ## 描述 Eigen 5.0 下 SparseSegmentMean 和 SparseSegmentSum 的 TensorMap/chip 热路径存在明显性能回退。 本 PR 将两个算子改为连续行 packet 循环与标量尾部,只对 segment 间隙清零,并保持逐元素累加顺序; 同步补充向量尾部、segment gap、负索引和有符号溢出等真实 kernel UT。 首轮流水线暴露 SparseSegmentSum 头文件依赖未自包含及 C++17 写法与生产 C++14 不兼容,已修复。 两个热函数的 CodeCheck NBNC 均已收敛:Mean 为 50,Sum 为 48。Mean 的最终修复保留原 while (true) 热循环;普通 helper 候选曾稳定回退约 5.6%,已否决,未进入本 PR。 本 PR 变更文件: - index/sparse_segment_mean/op_kernel_aicpu/sparse_segment_mean_aicpu.cpp - index/sparse_segment_mean/tests/ut/op_kernel_aicpu/test_sparse_segment_mean.cpp - index/sparse_segment_sum/op_kernel_aicpu/sparse_segment_sum_aicpu.cpp - index/sparse_segment_sum/op_kernel_aicpu/sparse_segment_sum_aicpu.h - index/sparse_segment_sum/tests/ut/op_kernel_aicpu/test_sparse_segment_sum.cpp ## 关联的Issue 关联 Issue #6138 ## 测试 - Mean 指定场景 FLOAT[104,105] + INT32[10922] + INT32[10922] -> FLOAT[3,105]: 最终 5x101 最慢均值 335.563 us,比 Eigen 3.4 的 722.35 us 门槛低 53.55%; 21x101 最终版/调整前复测比为 1.00093,95% CI [0.99960,1.00227]。 - Sum 指定场景 FLOAT[32,512] + INT32[8192] + INT32[8192] -> FLOAT[3,512]: CodeCheck 修复后 9x101 p50 平均 692.812 us,比 Eigen 3.4 的 924.43 us 门槛低约 24.8%; 修复后/修复前为 1.00130,95% CI [0.99905,1.00356]。 - 完整矩阵:Mean/Sum 共 56 个 dtype x shape 组合无确认性能回退;56/56 普通值和 6/6 特殊值 输出哈希一致。Mean 最终行数修复另做 12 组同机 A/B,输出哈希逐位一致且无确认回退。 - 独立 ABI=0 真实 kernel UT:Mean 12/12、Sum 9/9;ASan+UBSan 均通过,无报告。 - 当前源码 repo-native AICPU UT:Mean 12/12、Sum 9/9 均编译、链接并运行通过。 - 两个 graph example 均成功且输出与各自 host 参考一致。Mean 的 host/device plog 记录 aicpu_ascend_kernel::CheckSupported 和从 libcpu_kernels.so 取得 RunCpuKernel;Sum 的同一 会话记录 aicpu_custom_scheduler 与 CUSTOM_COMPUTE,确认实际执行 AICPU 路径。 - 本轮检视修复:x 第 0 维非法日志已补充算子名前缀和实际 dim0,并同步 canndev。 修复前后两仓各 68 个 ComputeKernelWithType 模板符号大小逐项一致,指定 FLOAT 实例 规范化指令一致;ops-nn Sum 真实 kernel UT 9/9、canndev 非文件驱动 UT 15/15 通过。 - 两仓均通过 -O2 -ftrapv -Wall -Wextra -Werror 严格编译;5 个变更文件连续两次通过 pre-commit(clang-format、OAT、codespell),git diff --check 通过。 - 当前单提交已通过服务端完整流水线,标签为 ci-pipeline-passed。 ## 文档更新 无产品文档变更;本地完整验证报告为 /home/ding-jing/ops_nn_sparse_segment_eigen5_perf_verification.md。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [x] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11057 | 15 天前 | |
feat(aicpu): enable constant folding for migrated ops Co-authored-by: hid29204727<wangboyu20@huawei.com> # message auto-generated for no-merge-commit merge: !9280 merge codex/aicpu-const-folding into master feat(aicpu): enable constant folding for migrated ops Created-by: hid29204727 Commit-by: hid29204727 Merged-by: cann-robot Description: ## 描述 补充 ops-nn 仓 AICPU host 常量折叠注册通路,并为当前已标记 OPEN_OPS_FLAG 的 18 个迁移算子统一开启常量折叠支持。 主要改动: - 新增仓级 OPS_NN_REGISTER_CPU_KERNELV2 注册包装宏,host 构建优先注册 V2,device/custom/UT 保持 V1 回退。 - 为 add_modules_sources 和 add_all_modules_sources 增加 HOSTCPU 参数及 host OBJECT 收集逻辑。 - 为 18 个 OPEN_OPS_FLAG 算子补充 HOSTCPU TRUE、注册头文件和 V2 注册。 - host OBJECT 汇入 libopconstant_folding_nn.so。 ## 关联的Issue 无。 ## 测试 已完成以下静态验证: - git diff --check 通过。 - 18 个 OPEN_OPS_FLAG 算子均具备 HOSTCPU TRUE。 - 18 个算子均使用 OPS_NN_REGISTER_CPU_KERNELV2,无直接 V1 注册残留。 - 注册头文件、host 编译宏和仓级 so 收集通路检查通过。 当前 Windows 环境缺少 CANN、NPU、CMake 和远端构建桥,未执行 bash build.sh --pkg -j16;完整打包构建及常量折叠功能验证待 CANN Linux 环境/CI 完成。 ## 文档更新 无仓内文档变更。 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!9280 | 21 天前 | |
perf: 优化 SparseSegmentMean 和 SparseSegmentSum AICPU 性能 Co-authored-by: Ding_Jing<dingjing19@huawei.com> # message auto-generated for no-merge-commit merge: !11057 merge perf/sparse-segment-eigen5 into master perf: 优化 SparseSegmentMean 和 SparseSegmentSum AICPU 性能 Created-by: Ding_Jing Commit-by: Ding_Jing Merged-by: cann-robot Description: ## 描述 Eigen 5.0 下 SparseSegmentMean 和 SparseSegmentSum 的 TensorMap/chip 热路径存在明显性能回退。 本 PR 将两个算子改为连续行 packet 循环与标量尾部,只对 segment 间隙清零,并保持逐元素累加顺序; 同步补充向量尾部、segment gap、负索引和有符号溢出等真实 kernel UT。 首轮流水线暴露 SparseSegmentSum 头文件依赖未自包含及 C++17 写法与生产 C++14 不兼容,已修复。 两个热函数的 CodeCheck NBNC 均已收敛:Mean 为 50,Sum 为 48。Mean 的最终修复保留原 while (true) 热循环;普通 helper 候选曾稳定回退约 5.6%,已否决,未进入本 PR。 本 PR 变更文件: - index/sparse_segment_mean/op_kernel_aicpu/sparse_segment_mean_aicpu.cpp - index/sparse_segment_mean/tests/ut/op_kernel_aicpu/test_sparse_segment_mean.cpp - index/sparse_segment_sum/op_kernel_aicpu/sparse_segment_sum_aicpu.cpp - index/sparse_segment_sum/op_kernel_aicpu/sparse_segment_sum_aicpu.h - index/sparse_segment_sum/tests/ut/op_kernel_aicpu/test_sparse_segment_sum.cpp ## 关联的Issue 关联 Issue #6138 ## 测试 - Mean 指定场景 FLOAT[104,105] + INT32[10922] + INT32[10922] -> FLOAT[3,105]: 最终 5x101 最慢均值 335.563 us,比 Eigen 3.4 的 722.35 us 门槛低 53.55%; 21x101 最终版/调整前复测比为 1.00093,95% CI [0.99960,1.00227]。 - Sum 指定场景 FLOAT[32,512] + INT32[8192] + INT32[8192] -> FLOAT[3,512]: CodeCheck 修复后 9x101 p50 平均 692.812 us,比 Eigen 3.4 的 924.43 us 门槛低约 24.8%; 修复后/修复前为 1.00130,95% CI [0.99905,1.00356]。 - 完整矩阵:Mean/Sum 共 56 个 dtype x shape 组合无确认性能回退;56/56 普通值和 6/6 特殊值 输出哈希一致。Mean 最终行数修复另做 12 组同机 A/B,输出哈希逐位一致且无确认回退。 - 独立 ABI=0 真实 kernel UT:Mean 12/12、Sum 9/9;ASan+UBSan 均通过,无报告。 - 当前源码 repo-native AICPU UT:Mean 12/12、Sum 9/9 均编译、链接并运行通过。 - 两个 graph example 均成功且输出与各自 host 参考一致。Mean 的 host/device plog 记录 aicpu_ascend_kernel::CheckSupported 和从 libcpu_kernels.so 取得 RunCpuKernel;Sum 的同一 会话记录 aicpu_custom_scheduler 与 CUSTOM_COMPUTE,确认实际执行 AICPU 路径。 - 本轮检视修复:x 第 0 维非法日志已补充算子名前缀和实际 dim0,并同步 canndev。 修复前后两仓各 68 个 ComputeKernelWithType 模板符号大小逐项一致,指定 FLOAT 实例 规范化指令一致;ops-nn Sum 真实 kernel UT 9/9、canndev 非文件驱动 UT 15/15 通过。 - 两仓均通过 -O2 -ftrapv -Wall -Wextra -Werror 严格编译;5 个变更文件连续两次通过 pre-commit(clang-format、OAT、codespell),git diff --check 通过。 - 当前单提交已通过服务端完整流水线,标签为 ci-pipeline-passed。 ## 文档更新 无产品文档变更;本地完整验证报告为 /home/ding-jing/ops_nn_sparse_segment_eigen5_perf_verification.md。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [x] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11057 | 15 天前 | |
修改SparseSegmentSumGrad、Dilation2D infershape并增加对应UT Co-authored-by: wang-shilong32<wangshilong23@huawei.com> # message auto-generated for no-merge-commit merge: !8931 merge fix_dilation_sparse_infer into master 修改SparseSegmentSumGrad、Dilation2D infershape并增加对应UT Created-by: wang-shilong32 Commit-by: wang-shilong32 Merged-by: cann-robot Description: ## 描述 修改SparseSegmentSumGrad、Dilation2D infershape -1/-2场景 并增加对应UT <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/4950 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> 本地验证对应UT用例,出包验证,冒烟测试。 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!8931 | 1 个月前 | |
nn仓vector算子适配350平台 Co-authored-by: zhouliang<zhouliang88@huawei.com> # message auto-generated for no-merge-commit merge: !10004 merge master into master nn仓vector算子适配350平台 Created-by: zhou-liang1027 Commit-by: zhouliang Merged-by: cann-robot Description: ## 描述 p_relu等算子适配350平台 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5755 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10004 | 29 天前 | |
docs: 统一原型注释中的产品名称(master) Co-authored-by: yolic<chenyuning1@huawei.com> # message auto-generated for no-merge-commit merge: !11186 merge 923soc into master docs: 统一原型注释中的产品名称(master) Created-by: yolic Commit-by: yolic Merged-by: cann-robot Description: ## 描述 原型注释里的产品名称仍沿用 Series Product、Ascend 950 AI Processor、Ascend 950PR/Ascend 950DT 等旧表述,与当前产品命名不一致。本次只改注释,不改算子接口和计算逻辑。 合入目标:cann/ops-nn 的 master。 覆盖各算子 op_graph/*_proto.h,以及 common/inc/op_graph/op_nn_proto_extend.h、common/stub/op_graph/math_proto_stub.cpp。 ## 关联的Issue 关联 Issue #6213(Close #6213) https://gitcode.com/cann/ops-nn/issues/6213 ## 测试 仅修改原型注释中的产品名称,不涉及算子 host/kernel 计算逻辑,无需算子泛化。已在本地用仓库 pre-commit 的 end-of-file-fixer 与 clang-format 核对格式,且没有改动算子输入、输出和属性定义。 ## 文档更新 更新了 op_graph 原型注释中的产品名称,无独立 markdown 文档。 ## 类型标签 - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!11186 | 11 天前 | |
fix(cleancode): AICPU算子告警整改(第三批) Co-authored-by: liu-wei<lovline.liuwei@huawei.com> # message auto-generated for no-merge-commit merge: !10997 merge fix/warning-cleanup-v3 into master fix(cleancode): AICPU算子告警整改(第三批) Created-by: liu-wei Commit-by: liu-wei Merged-by: cann-robot Description: ## 描述 ops-nn AICPU 算子告警整改第三批,修复 18 条告警(G.CNS.03 ×12 / G.EXP.27+36 ×5 / G.CNS.04 ×1),共 14 个文件 +29/-29: | 告警 | 算子 | 修复 | |---|---|---| | G.CNS.03 ×12 | log_softmax_v2、scatter_elements、sparse_segment_mean、sparse_segment_sum、sparse_to_dense、tensor_scatter_update | 12 处成员函数补充 const(声明与定义同步,被调成员均为 const 或自由函数) | | G.EXP.27/36 ×4 | softmax_v2 | NormalCheck(...)/detail::SoftmaxV2Check(...) 两处 ?: 条件改为显式 != KERNEL_STATUS_OK | | G.EXP.27 ×1 | reverse_sequence | reverseNum % kEven 显式比较 != 0 | | G.CNS.04 ×1 | tensor_scatter_update | GetBatchStrides 的 outer_shape 参数加 const(仅读取) | ## 不处理项(附判定依据) - 浮点 == static_cast<T>(0) ×2(log_softmax_v2/softmax_v2 的除零防护):保持语义,建议工具豁免; - sparse_to_dense G.STD.02 ×4:char* 为字节缓冲非字符串,误报; - EigenSparseToDense 的 SparseTensor& st:st.ToDense<>() 受 Eigen API 限制,疑似误报。 ## 验证 - 8 个改动 TU 语法级编译通过(g++ -std=c++14); - 二进制级对比(父提交 467b22d6c vs 本提交 7f12ebcbc,同编译参数 g++ -O2 -std=c++14):8 个 TU 全部机器码/数据节**内容多重集逐字节一致**(合计 1298 节),仅 const 成员函数 mangled 名(符号表及 scatter_elements 15 节、sparse_to_dense 1 节的内嵌签名串 ZN→ZNK)变化——**零功能影响**; - CI 流水线待跑。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6114 ## 类型标签 - [x] ♻️ 重构 - [x] 🧹 代码清理 See merge request: cann/ops-nn!10997 | 17 天前 | |
修改aclnn文档示例代码问题 Co-authored-by: lihang_13123<lihang106@h-partners.com> # message auto-generated for no-merge-commit merge: !9489 merge scatter_add into master 修改aclnn文档示例代码问题 Created-by: lihang_13123 Commit-by: lihang_13123 Merged-by: cann-robot Description: ## 描述 修改aclnnTfScatterAdd.md文档示例代码问题 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5326 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 aclnnScatterAdd.md和aclnnTfScatterAdd.md ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!9489 | 1 个月前 | |
Fix top_k_top_p_sample docs and example Co-authored-by: licheng261<licheng261@huawei.com> # message auto-generated for no-merge-commit merge: !9906 merge top_k_top_p_sample_doc_0904 into master Fix top_k_top_p_sample docs and example Created-by: licheng261 Commit-by: licheng261 Merged-by: cann-robot Description: ## 描述 修复top_k_top_p_sample文档和示例代码,保持一致 ## 关联的Issue [#5591](https://gitcode.com/cann/ops-nn/issues/5591) ## 测试 ## 文档更新 ops-nn/index/top_k_top_p_sample/docs/aclnnTopKTopPSample.md ops-nn/index/top_k_top_p_sample/examples/test_aclnn_top_k_top_p_sample.cpp ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!9906 | 1 个月前 | |
修复aclnnTopKTopPSampleV2文档示例代码与参数表不一致问题 Co-authored-by: 陈赵旻熠<chenzhaominyi1@huawei.com> # message auto-generated for no-merge-commit merge: !10065 merge master into master 修复aclnnTopKTopPSampleV2文档示例代码与参数表不一致问题 Created-by: zerosaki_admin Commit-by: 陈赵旻熠 Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 修复aclnnTopKTopPSampleV2文档示例代码与参数表不一致问题 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> [#5681](https://gitcode.com/cann/ops-nn/issues/5681) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 更新了aclnnTopKTopPSampleV2文档 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!10065 | 1 个月前 | |
feat(unique): add Regbase values-only AI Core support Co-authored-by: ConanHuang<huangxiaobin1@huawei.com> # message auto-generated for no-merge-commit merge: !10993 merge feature_op_unique into master feat(unique): add Regbase values-only AI Core support Created-by: ConanHuang Commit-by: ConanHuang Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 为 aclnnUnique / aclnnUnique2 增加 Regbase values-only AI Core 路径,在单次 Kernel 调用中完成排序与去重,减少不必要的索引搬运和中间存储。 - 新增 arch35 Tiling、Kernel 和动态输出 shape 回传,接入 ACLNN 分流。 - 整理算子 L0、原型及 Ascend 950/350 构建注册;调整共享 radix 和 UniqueConsecutive 的大输入计算。 - 新增 ACLNN/Kernel 测试用例、golden 和 GEIR 示例。 新路径仅处理不返回 inverse/counts 且满足准入条件的调用,其余场景沿用既有路径;GE/ONNX 仍使用 AICPU。详细 Tiling 与 Kernel 设计见关联 Issue。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6187 ## 测试 - 二级冒烟通过 - UniqueWithCountsAndSorting 和 KthValue 各自 300+ kernel ST通过 - aclnnUnique aclnnUnique2 各自 170+ aclnn用例通过 - onnx geir 通路兼容性验证通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 不涉及 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10993 | 15 天前 | |
refactor: 为公共op_api头文件增加NN后缀 Co-authored-by: clinglai0517<laichangling@huawei.com> # message auto-generated for no-merge-commit merge: !10791 merge refactor/nn-suffix-common-headers into master refactor: 为公共op_api头文件增加NN后缀 Created-by: clinglai0517 Commit-by: clinglai0517 Merged-by: cann-robot Description: ## 描述 ops-nn 的 common/inc/op_api 目录中仍有 level2_base.h 和 level2_base_caculation.h 两个通用名称。它们与其他算子仓的同名文件一起进入软件包时存在覆盖风险。 本 MR 是 ops-nn 分阶段迁移的第一阶段: - 新增完整实现头 level2_base_nn.h 和 level2_base_caculation_nn.h; - 保留原文件为轻量转发头,兼容尚未迁移的存量 ACLNN 源文件和开发中的 MR; - 新计算头只依赖新基础头,避免新命名链路回退到旧文件名; - op_api_def_nn.h 已具有 NN 标识,本次不修改; - classify_rule.yaml 中没有这两个头文件的现有条目,也没有证据表明该文件负责软件包头文件筛选,因此本次不修改; - 先迁移 Mish、UniqueConsecutive、Silu、LogSigmoid 四个简单 ACLNN 接口作为样例。 本阶段不删除兼容头,也不声明已完成最终软件包去重。当前仍有 32 处 level2_base.h 和 5 处 level2_base_caculation.h 的旧引用,后续可按模块分批迁移;旧引用归零并验证实际制品清单后,再移除兼容头。 ## 关联的Issue 关联并在合入后关闭 #6008: https://gitcode.com/cann/ops-nn/issues/6008 ## 测试 已完成: - git diff --check; - 新旧完整实现内容等价性检查(仅头文件保护宏和内部 include 路径变化); - 两个旧文件均为 16 行轻量转发头; - 新命名头文件之间仅通过新文件名依赖; - 四个样例 ACLNN 文件均为单行 include 替换; - 使用 clang++ -M -MG 完成头文件依赖解析冒烟检查; - 统计剩余旧引用:level2_base.h 32 处、level2_base_caculation.h 5 处; - 8 个变更文件通过 OAT 合规检查; - origin 同名分支与本地提交一致。 ## 文档更新 无。 ## 类型标签 - [ ] 新功能(不影响现有功能的新特性) - [ ] Bug 修复(不影响现有功能的问题修复) - [ ] 重大变更(会导致现有功能无法正常工作的修复或功能) - [ ] 文档更新 - [ ] 性能优化 - [ ] 代码重构 - [x] 其他,请描述:头文件命名与兼容迁移 ## AI/Agent生成声明 - [ ] 完全由开发者手工编写 - [x] AI辅助编写 - [ ] AI生成 See merge request: cann/ops-nn!10791 | 19 天前 | |
将卷积和matmul的ut根据版本隔离开, 整改950 ophost ut Co-authored-by: 18811725231<yangdi52@huawei.com> # message auto-generated for no-merge-commit merge: !9549 merge master into master 将卷积和matmul的ut根据版本隔离开, 整改950 ophost ut Created-by: yang-di52 Commit-by: 18811725231 Merged-by: cann-robot Description: ## 描述 主要修改内容: 1. nn仓完成和legacy common 动态库解耦后,卷积和matmul的不同版本 ophost ut 已经不能混合一起跑了。需要版本隔离 2. ascend950的ut全量编译 失败,需要修改 ## 关联的Issue [https://gitcode.com/cann/ops-nn/issues/5425](https://gitcode.com/cann/ops-nn/issues/5425) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!9549 | 1 个月前 | |
feat(unique): add Regbase values-only AI Core support Co-authored-by: ConanHuang<huangxiaobin1@huawei.com> # message auto-generated for no-merge-commit merge: !10993 merge feature_op_unique into master feat(unique): add Regbase values-only AI Core support Created-by: ConanHuang Commit-by: ConanHuang Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> 为 aclnnUnique / aclnnUnique2 增加 Regbase values-only AI Core 路径,在单次 Kernel 调用中完成排序与去重,减少不必要的索引搬运和中间存储。 - 新增 arch35 Tiling、Kernel 和动态输出 shape 回传,接入 ACLNN 分流。 - 整理算子 L0、原型及 Ascend 950/350 构建注册;调整共享 radix 和 UniqueConsecutive 的大输入计算。 - 新增 ACLNN/Kernel 测试用例、golden 和 GEIR 示例。 新路径仅处理不返回 inverse/counts 且满足准入条件的调用,其余场景沿用既有路径;GE/ONNX 仍使用 AICPU。详细 Tiling 与 Kernel 设计见关联 Issue。 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/6187 ## 测试 - 二级冒烟通过 - UniqueWithCountsAndSorting 和 KthValue 各自 300+ kernel ST通过 - aclnnUnique aclnnUnique2 各自 170+ aclnn用例通过 - onnx geir 通路兼容性验证通过 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> 不涉及 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!10993 | 15 天前 | |
docs:修复Unique系列 ACLNN文档中的格式与示例错误 Co-authored-by: jerry_gd<jerry.luo@huawei.com> # message auto-generated for no-merge-commit merge: !9697 merge master into master docs:修复Unique系列 ACLNN文档中的格式与示例错误 Created-by: jerry_gd Commit-by: jerry_gd Merged-by: cann-robot Description: ## 描述 修复Unique系列 ACLNN文档中的格式与示例错误。 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> https://gitcode.com/cann/ops-nn/issues/5491 ## 测试 <!--描述进行了哪些测试来验证你的改动。--> 本次仅修改 Markdown 文档和文档内示例代码,未执行编译或运行测试 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> - index/unique_consecutive/docs/aclnnUnique.md - index/unique_consecutive/docs/aclnnUnique2.md - index/unique_consecutive/docs/aclnnUniqueConsecutive.md - index/unique_with_counts_ext2/docs/aclnnUniqueDim.md ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!9697 | 1 个月前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Co-authored-by: klein8793<zhangli24@huawei.com> # message auto-generated for no-merge-commit merge: !9754 merge dev_zl_2026_0902 into master Feature: 为 index 和 pooling 类算子新增 Ascend350 平台支持 Created-by: klein8793 Commit-by: klein8793 Merged-by: cann-robot Description: ## 描述 为 index 和 pooling 类算子新增 Ascend350 平台支持。具体变更包括: 1. **新增 Ascend350 平台配置文件**(42 个 JSON):为每个算子新增 op_host/config/ascend350/ 下的二进制配置 JSON 2. **注册 Ascend350 平台**:在各算子 _def.cpp 中添加 his->AICore().AddConfig(ascend350, aicoreConfig) 3. **更新编译配置**:修改各算子 CMakeLists.txt,添加 Ascend350 编译选项;更新 scendc_config.json 中对应算子的 compute_units 列表,将 [ascend950] 扩展为 [ascend950, ascend350] 4. **更新 Fusion Pass 平台判断**:在 bucketize_v2_fusion_pass、inplace_add_fusion_pass、gather_fusion_pass、tensor_scatter_add_fusion_pass、sorted_sparse_segment_mean_grad_fusion_pass 中将 Ascend350 加入平台白名单 涉及算子(35+): - index: bucketize_v2, concat_offset, gather_elements, gather_nd, gather_v2, index_fill, index_fill_d, index_put_v2, inplace_index_add, inplace_index_fill, linear_index_v2, masked_scatter, repeat_interleave, reverse_sequence, reverse_v2, scatter, scatter_add, scatter_add_with_sorted, scatter_elements, scatter_elements_v2, scatter_nd, scatter_nd_add, scatter_nd_max, scatter_nd_min, scatter_nd_update, scatter_sub, scatter_update, segment_sum, sorted_sparse_segment_mean_grad, sparse_segment_mean, sparse_to_dense, unsorted_segment_max, unsorted_segment_min - pooling: avg_pool, avg_pool_v2 - 公共库: scatter_nd_common, unsorted_segment_common, sort_lib ## 关联的Issue - #5651 ## 测试 obp冒烟、david冒烟已通过 ## 文档更新 NA ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [ ] AI辅助编写 See merge request: cann/ops-nn!9754 | 22 天前 | |
UnsortedSegmentProd性能优化 Co-authored-by: zhangxiyan7<zhangxiyan7@huawei.com> # message auto-generated for no-merge-commit merge: !9677 merge UnsortedSegmentProd into master UnsortedSegmentProd性能优化 Created-by: zhangxiyan7 Commit-by: zhangxiyan7 Merged-by: cann-robot Description: ## 描述 针对 UnsortedSegmentProd 算子性能对标竞品部分用例不达标的问题,从单一 Simt 路径扩展为多模板路径体系,并增加全无效 id 向量化快路径。分模板优化说明: 1. OutFl / OutFlWsMerge(输出全载)模板 - 输入行数远超输出规模( inputOuterDim > outputSize × 总核数 × 2)时启用:输出整块驻留 UB,输入行按批流入直接累乘,消除逐行原子写 - WsMerge 变体:各核在 workspace 私有段累乘局部积,最后按输出行跨核 merge,突破单核 UB 容量限制 2. InputPart(输入分区)模板 - 输入行数 ≥ 1024 且输出可驻留 workspace 时启用:按输入行多核分区并行,各核产出局部积后按输出行 merge,解决输入长、输出短的负载并行度问题 - strided flag 跨核聚合:各核将"是否存在有效 id"写入带 stride 的 GM flag,清缓存 + SyncAll 后汇聚,全局无有效 id 时直接跳过 merge 写初值 1 3. SegmentSort(段排序)模板 - float32 且 128 ≤ innerDim < 512、行数 ≥ 1024 时启用:RadixSort 对 segmentIds 排序(多核本地排序 + 归并树合并),归约阶段按连续 segment 顺次连乘,消除随机访存 4. SplitCol / SortSimt 模板:复用公共路径 + Prod 特化 - SplitCol 的 baseA ;sLoopNum == 1 时走单块快路径——一次性读入全部 ids 并先做全无效检测,再决定是否搬运数据 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/5636 ## 测试 目标用例性能达标(换算带宽后实现 竞品耗时/NPU耗时 > 0.2),冒烟通过。部分用例性能对比结果如下: | # | 用例 / dtype | 输入形状(seg数) | GPU耗时(µs) | 初始NPU耗时(µs) | 初始耗时比 | 最终NPU耗时(µs) | 最终耗时比 | |---|---|---|---|---|---|---|---| | 1 | 000132 fp32/int32 | [5410,77,4] (554) | 19.35 | 381.08 | 0.051 | 30.09 | 0.643 | | 2 | 000040 bf16/int32 | [20,1,8,9,12] (906) | 13.32 | 235.62 | 0.057 | 10.40 | 1.280 | | 3 | 000047 bf16/int64 | [2,2019,33,2,2,2,2,2] (231) | 451.06 | 12101.55 | 0.037 | 110.87 | 4.068 | | 4 | 000116 fp16/int32 | [63,5,10044] (834) | 28.21 | 518.22 | 0.054 | 27.30 | 1.033 | | 5 | 000062 int32/int32 | [3,2,1110,2,2,11,2,2] (224) | 14.59 | 1248.42 | 0.012 | 29.71 | 0.491 | | 6 | 000140 fp32/int32 | [2,4,4,1,8,5,8] (308) | 4.98 | 76.65 | 0.065 | 4.97 | 1.003 | | 7 | 000131 fp32/int64 | [60,5717] (262) | 6.69 | 144.32 | 0.046 | 7.02 | 0.953 | | 8 | 000163 int32/int64 | [101,32383] (853) | 70.59 | 1705.76 | 0.041 | 148.06 | 0.477 | | 9 | 000130 fp32/int32 | [1605,1054] (724) | 620.00 | 34430.41 | 0.018 | 141.21 | 4.391 | | 10 | 000171 int32/int64 | [10090,3,3,3,5,3] (822) | 490.38 | 26759.76 | 0.018 | 270.01 | 1.816 | | 11 | 000017 fp32/int64 | [770] (909) | 8.04 | 1084.10 | 0.007 | 20.81 | 0.386 | | 12 | 000007 fp32/int32 | [52,9739,4,4] (97) | 434.26 | 19934.99 | 0.022 | 66.02 | 6.577 | | 13 | 000134 fp32/int32 | [5,28,34,19] (66) | 160.63 | 34359.26 | 0.005 | 638.77 | 0.251 | | 14 | 000024 fp32/int32 | [20,4,5,21,9] (769) | 5.45 | 198.86 | 0.027 | 7.66 | 0.711 | | 15 | 000006 fp16/int32 | [68,4,4,13603] (208) | 208.56 | 4464.44 | 0.047 | 113.08 | 1.844 | | 16 | 000126 fp16/int32 | [136,2,2,2,645,2,2,2] (965) | 58.30 | 1237.38 | 0.047 | 46.24 | 1.261 | | 17 | 000046 bf16/int32 | [2,2,2,272,2,2,2,63] (406) | 143.06 | 3622.09 | 0.039 | 63.18 | 2.264 | | 18 | 000021 fp32/int64 | [5,12,3171] (383) | 23.28 | 152.96 | 0.152 | 26.94 | 0.864 | | 19 | 000011 fp32/int64 | [2,7081,2,2,2,6,3,2] (632) | 57.57 | 4150.62 | 0.014 | 140.37 | 0.410 | | 20 | 000037 bf16/int64 | [145,4,8254] (1000) | 78.42 | 1037.15 | 0.076 | 96.01 | 0.817 | | 21 | 000135 fp32/int64 | [5,5,4,41] (839) | 45.12 | 3645.09 | 0.012 | 15.85 | 2.847 | | 22 | 000146 bf16/int32 | [23,47] (269) | 14.94 | 480.06 | 0.031 | 15.58 | 0.959 | | 23 | 000017 fp32/int64 | [9718] (954) | 44.76 | 2030.42 | 0.022 | 19.37 | 2.310 | | 24 | 000056 int32/int32 | [4,4,17,4,6535] (396) | 3092.56 | 209483.03 | 0.015 | 552.16 | 5.601 | | 25 | 000227 int32/int64 | [116,317,10,446] (788) | 1254.56 | 83651.63 | 0.015 | 2683.99 | 0.467 | ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [x] 其他,请描述:UnsortedSegmentProd算子性能优化 ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!9677 | 1 个月前 | |
docs: 更新 UnsortedSegmentSum 产品命名及确定性说明 Co-authored-by: aiteer<lujiale4@huawei.com> # message auto-generated for no-merge-commit merge: !11260 merge aclnn_uss_doc into master docs: 更新 UnsortedSegmentSum 产品命名及确定性说明 Created-by: aiteer Commit-by: aiteer Merged-by: cann-robot Description: ## 描述 <!--在这里详细描述你的改动,包括改动的原因和所采取的方法。--> * 根据最新仓文档模板修改产品名称 * 删除算子层文档中的确定性约束说明 * 修正aclnn接口文档中参数取值支持的dtype,int64_t -> INT64 * 修正aclnn接口文档中确定性说明的语句 ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> [#6251](https://gitcode.com/cann/ops-nn/issues/6251) ## 测试 <!--描述进行了哪些测试来验证你的改动。--> ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> * uss 的算子层文档 * uss 的aclnn接口文档 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [x] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [ ] AI辅助编写 See merge request: cann/ops-nn!11260 | 11 天前 | |
feat(aicpu): enable constant folding for migrated ops Co-authored-by: hid29204727<wangboyu20@huawei.com> # message auto-generated for no-merge-commit merge: !9280 merge codex/aicpu-const-folding into master feat(aicpu): enable constant folding for migrated ops Created-by: hid29204727 Commit-by: hid29204727 Merged-by: cann-robot Description: ## 描述 补充 ops-nn 仓 AICPU host 常量折叠注册通路,并为当前已标记 OPEN_OPS_FLAG 的 18 个迁移算子统一开启常量折叠支持。 主要改动: - 新增仓级 OPS_NN_REGISTER_CPU_KERNELV2 注册包装宏,host 构建优先注册 V2,device/custom/UT 保持 V1 回退。 - 为 add_modules_sources 和 add_all_modules_sources 增加 HOSTCPU 参数及 host OBJECT 收集逻辑。 - 为 18 个 OPEN_OPS_FLAG 算子补充 HOSTCPU TRUE、注册头文件和 V2 注册。 - host OBJECT 汇入 libopconstant_folding_nn.so。 ## 关联的Issue 无。 ## 测试 已完成以下静态验证: - git diff --check 通过。 - 18 个 OPEN_OPS_FLAG 算子均具备 HOSTCPU TRUE。 - 18 个算子均使用 OPS_NN_REGISTER_CPU_KERNELV2,无直接 V1 注册残留。 - 注册头文件、host 编译宏和仓级 so 收集通路检查通过。 当前 Windows 环境缺少 CANN、NPU、CMake 和远端构建桥,未执行 bash build.sh --pkg -j16;完整打包构建及常量折叠功能验证待 CANN Linux 环境/CI 完成。 ## 文档更新 无仓内文档变更。 ## 类型标签 - [ ] Bug修复 - [x] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 - [x] AI辅助编写 See merge request: cann/ops-nn!9280 | 21 天前 | |
License Change Split2 Co-authored-by: huohuo_wy<wangyan389@huawei.com> # message auto-generated for no-merge-commit merge: !371 merge spilit2License3 into master License Change Split2 Created-by: huohuo_wangyan Commit-by: huohuo_wy Merged-by: cann-robot Description: ## 描述 批量整改文件copyright注释 ## 关联的Issue https://gitcode.com/cann/ops-nn/issues/177 ## 测试 COPYRIGHT注释头按新要求刷新 ## 文档更新 <!--如果这个PR包含文档的更新,请在这里指出。例如:更新了README.md文件。--> ## 类型标签 <!-- [x] 表示选中 --> - [x] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: See merge request: cann/ops-nn!371 | 9 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 个月前 | ||
| 11 天前 | ||
| 19 天前 | ||
| 21 天前 | ||
| 18 天前 | ||
| 19 天前 | ||
| 15 天前 | ||
| 1 个月前 | ||
| 30 天前 | ||
| 1 个月前 | ||
| 28 天前 | ||
| 19 天前 | ||
| 11 天前 | ||
| 11 天前 | ||
| 11 天前 | ||
| 28 天前 | ||
| 17 天前 | ||
| 19 天前 | ||
| 3 个月前 | ||
| 15 天前 | ||
| 11 天前 | ||
| 11 天前 | ||
| 11 天前 | ||
| 3 个月前 | ||
| 16 天前 | ||
| 1 个月前 | ||
| 22 天前 | ||
| 21 天前 | ||
| 21 天前 | ||
| 22 天前 | ||
| 28 天前 | ||
| 22 天前 | ||
| 28 天前 | ||
| 21 天前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 1 个月前 | ||
| 22 天前 | ||
| 11 天前 | ||
| 22 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 21 天前 | ||
| 3 个月前 | ||
| 19 天前 | ||
| 11 天前 | ||
| 11 天前 | ||
| 21 天前 | ||
| 21 天前 | ||
| 11 天前 | ||
| 11 天前 | ||
| 19 天前 | ||
| 17 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 19 天前 | ||
| 11 天前 | ||
| 11 天前 | ||
| 29 天前 | ||
| 19 天前 | ||
| 1 个月前 | ||
| 19 天前 | ||
| 19 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 21 天前 | ||
| 19 天前 | ||
| 21 天前 | ||
| 17 天前 | ||
| 1 个月前 | ||
| 21 天前 | ||
| 22 天前 | ||
| 17 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 19 天前 | ||
| 15 天前 | ||
| 21 天前 | ||
| 15 天前 | ||
| 1 个月前 | ||
| 29 天前 | ||
| 11 天前 | ||
| 17 天前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 19 天前 | ||
| 1 个月前 | ||
| 15 天前 | ||
| 1 个月前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 22 天前 | ||
| 1 个月前 | ||
| 11 天前 | ||
| 21 天前 | ||
| 9 个月前 |