| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
grouped_dynamic_mx_quant算子支持group_index int64类型 Co-authored-by: TH<taohai2@huawei.com> # message auto-generated for no-merge-commit merge: !9023 merge grouped_dynamic_mx_quantAdaptINT64 into master grouped_dynamic_mx_quant算子支持group_index int64类型 Created-by: TH_HW Commit-by: TH Merged-by: cann-robot Description: ## 描述 使 GroupedDynamicMxQuant 算子(含 aclnnGroupedDynamicMxQuant 与 aclnnGroupedDynamicMxQuantV2 接口)的 groupIndex 输入不再仅支持 INT32,而是同时支持 INT32、INT64 两种数据类型,满足上层框架以 int64 索引调用算子的需求。同时保持对既有 INT32 场景的完整兼容。 主要改动如下: 1. **算子描述与接口校验**:op_graph/grouped_dynamic_mx_quant_proto.h 中 group_index 的 TensorType 由 {DT_Int32} 扩展为 {DT_Int32, DT_Int64};op_api 与 op_host/op_api 头文件中补充 groupIndex 的 int64 支持说明;aclnn_grouped_dynamic_mx_quant.cpp / aclnn_grouped_dynamic_mx_quant_v2.cpp 的GROUP_INDEX_DTYPE_SUPPORT_LIST 增加 DT_INT64。 2. **Host 侧(Tiling/定义)**:tiling 校验集合 GROUPIDX_SUPPORT_DTYPE_SET 增加 DT_INT64,GroupedDynamicMxQuantTilingParam 新增 groupIndexType 字段,并在 SetTilingKeyParam 中将其编入 tiling key;grouped_dynamic_mx_quant_def.cpp 的数据类型组合表扩展为 int32/int64 与所有输入、输出类型的组合(共16组)。 3. **Binary 配置**:ascend950/grouped_dynamic_mx_quant_binary.json 为各 x 输入类型(fp16/bf16)与输出类型(fp8_e4m3fn、fp8_e5m2、fp4_e2m1、fp4_e1m2)组合补充了 group_index 为 int64 的 _int64 二进制算子配置。 4. **Kernel 侧**:GroupedDynamicMxQuantCombine 模板新增 GroupIndexT 类型参数,GlobalTensor<GroupIndexT> 读取分组索引;公共切分基类 GroupedSplitBase::GroupedSplit 同步模板化 GroupIndexT;grouped_dynamic_mx_quant_struct.h 新增 TPL_GROUP_INDEX_INT32/TPL_GROUP_INDEX_INT64 模板实参选择;grouped_dynamic_mx_quant.cpp 入口通过 if constexpr 依据 tiling key 中的 groupIndexType 分发实例化。 5. **文档更新**:README 与 aclnnGroupedDynamicMxQuant / aclnnGroupedDynamicMxQuantV2 接口文档中 groupIndex 支持类型由 INT32 更新为 INT32、INT64。 6. **单测适配**:test_grouped_dynamic_mx_quant.cpp 中 kernel 调用补充新增的 groupIndexType 模板实参。 ## 关联的Issue 无。 ## 测试 - 对 group_index 均为 int32、int64 时逐用例执行了 tests/ut/op_kernel/test_grouped_dynamic_mx_quant.cpp 中现有用例验证(bf16→fp8/fp4、fp16→fp8/fp4 等多种输入输出组合),结果均通过。 - 新增/回归验证了 tiling key 中 groupIndexType 的取整与算子 bin 选择在 asc 引擎上的绑定正确性。 ## 文档更新 更新了以下文档: - quant/grouped_dynamic_mx_quant/README.md - quant/grouped_dynamic_mx_quant/docs/aclnnGroupedDynamicMxQuant.md - quant/grouped_dynamic_mx_quant/docs/aclnnGroupedDynamicMxQuantV2.md ## 类型标签 <!-- [x] 表示选中 --> - [x] 新特性 - [ ] Bug修复 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: ## AI/Agent生成声明 <!-- [x] 表示选中 --> - [x] AI辅助编写 See merge request: cann/ops-nn!9023 | 1 个月前 | |
将norm,quant目录下 MicroAPI 命名空间统一替换为 Reg Co-authored-by: m0_71149169<1907825139@qq.com> # message auto-generated for no-merge-commit merge: !9291 merge norm-quant-microapi-to-reg into master 将norm,quant目录下 MicroAPI 命名空间统一替换为 Reg Created-by: m0_71149169 Commit-by: m0_71149169 Merged-by: cann-robot Description: ## 描述 MicroAPI → Reg 命名空间迁移: norm,quant CANN 头文件 kernel_macros.h 中存在命名空间别名 namespace MicroAPI = Reg;,MicroAPI 只是 Reg 的别名。为统一命名规范,将 nn 仓所有算子 kernel 代码中的 MicroAPI 替换为 Reg,并通过对比替换前后 binary .o 文件 MD5 验证替换无影响。 覆盖范围 - nn 主仓目录:norm,quant ## 关联的Issue <!-- 如果这个PR是为了解决特定的Issue,请在这里提供Issue链接。例如:关联Issue #000--> 关联Issue [#5098](https://gitcode.com/cann/ops-nn/issues/5098) ## 测试 验证方法 严格按四步流程,逐算子验证: 1. Step 1:编译原始代码,保存 baseline binary .o 文件 MD5 2. Step 2:执行 MicroAPI → Reg 替换 3. Step 3:编译替换后代码,保存 after binary .o 文件 MD5 4. Step 4:对比 A_before.md5 与 B_after.md5 平台:ascend950,编译参数:--soc=ascend950 --ops=<op> -j16 验证替换前后 binary 产物 MD5 完全一致。 ## 类型标签 <!-- [x] 表示选中 --> - [ ] Bug修复 - [ ] 新特性 - [ ] 性能优化 - [ ] 文档更新 - [ ] 其他,请描述: norm,quant目录下MicroAPI → Reg 命名空间迁移 See merge request: cann/ops-nn!9291 | 1 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 1 个月前 | ||
| 1 个月前 |