Pull Request已成功合入, 合并人@CANN-robot
(感谢 gaoxin 的贡献)变更摘要
本 PR 为 Transpose、IsNan、IsFinite 三个算子新增 V2 ATT 性能建模公式,替换原有的 GetPerfFunc(kUnitVector) 通用评估:新增 UnaryBitWidthChangeNodeParams、TransposeNodeParams 参数结构,打通从 codegen 参数填充、parser 解析到 V2 性能注册的完整链路;在 ascendc_regbase_perf.cpp 中实现 IsNanPerf/IsFinitePerf(共享 AddUnaryBitWidthChangeCommonPerf 公式)与 TransposePerf(按 total_dim/inner_dim 分 7 个场景的索引生成 + 扩展搬运建模),并将注册入口切换为专用 V2 函数,以贴近实际 AscendC codegen 指令序列,提升性能评估精度。
主要改动
- 新增参数结构及注册/填充链路: 在
ascir_node_param.h中新增UnaryBitWidthChangeNodeParams(valid、cal_count、outer_repeats、input_strides、output_strides)与TransposeNodeParams(valid、inner_dim、total_dim、outer_loop_axes、output_dims及 strides),扩展AnySpecificParamsvariant,并在NodeDetail/NodeInfo中新增对应字段;ascir_param_builder.cpp新增RegisterUnaryBitWidthChangeAscirNodeParams/RegisterTransposeAscirNodeParams,specific_params_builder.cpp新增FillUnaryBitWidthChangeParams/FillTransposeParams,codegen 侧reg_transpose_api_call.cpp与unary_bitwidth_change_api_call_v2.cpp分别新增FillTransposeNodeParams/FillUnaryBitWidthChangeNodeParams完成参数填充。 - 新增
IsNanPerf/IsFinitePerf公式: 共享AddUnaryBitWidthChangeCommonPerf(kDuplicate→kInt8、kUpdateMask/kLoad→输入 dtype、kNe→kUInt8、kSelect/kStore→kFloat32),IsFinitePerf额外增加kCompareScalarEQ与kMaskOr;call_count置于最外层乘法,最终公式为(GetVFHeadCost() + max_latency + all_vf_instruct_cost) * call_count,GetUnaryBitWidthChangeCallCount在参数无效时 fallback 使用input_dims全量乘积。 - 新增
TransposePerf公式: 按total_dim/inner_dim组合分 Dim2Inner1、Dim3Inner1/2、Dim4Inner1/2/3、Fallback 共 7 个场景,通过AddGenOne/Two/ThreeInnerDimTransposeIndexPerf建模 1-3 个内层维度的索引序列生成(Arange/kPlaceholder、kMuls、kAdds、kDiv、kMul、kCompareScalarGE、kSelect等指令),通过AddTransposeExtendPerf及 One/Two/ThreeOuterDim 变体建模多层外层循环的kLoad+ Gather(kPlaceholder) +kStore搬运,最终结果乘以outer_count。 - 注册切换与常数/性能表补充:
ascir_api_perf_v2.cpp新增IsNanApi/IsFiniteApi/TransposeApi,将kIsnan/kIsFinite/kTranspose的 V2 注册由GetPerfFunc(kUnitVector)切换为GetPerfFunc(kIsnan + "V2")/kIsFinite + "V2"/kTranspose + "V2";att_const_values.h新增kSymFive常量,perf_param_v2.cpp新增kCompareScalarGE性能表项。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| repo-cann/graph-autofusion | ✅ yangyongqiang0606, zhang_shengjie, xchu42 (3/2) | ✅ yangyongqiang0606 (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
gcw_V3YyYBt1, thanks for your pull request. All authors of the commits have signed the CLA. 👍


/lgtm


/approve


Pull Request
描述
新增 Transpose、IsNan、IsFinite 三个算子的 V2 ATT 性能建模公式,替换原有的
GetPerfFunc(kUnitVector)通用评估,提升性能评估精度。Transpose:按
total_dim/inner_dim分场景(Dim2Inner1、Dim3Inner1/2、Dim4Inner1/2/3、Fallback)构建索引生成与扩展搬运公式,支持 1-3 个内层维度的索引序列生成建模及多层外层循环的 Gather/Store 搬运建模。IsNan/IsFinite:提取
UnaryBitWidthChange通用公式,call_count放在最外层乘法;各算子类型统一 dtype(Duplicate→kInt8、Ne→kUInt8、Select/Store→kFloat32、MaskPack→kUInt8)。背景
当前 IsNan、IsFinite、Transpose 在 V2 性能注册中使用
GetPerfFunc(kUnitVector)通用评估,无法反映实际 AscendC codegen 的指令序列,导致性能公式精度不足。修改方案
UnaryBitWidthChangeNodeParams、TransposeNodeParams及对应的构建和填充逻辑AddUnaryBitWidthChangeCommonPerf,公式为(vfhead + max_latency + all_cost) * call_counttail_repeat参数,补充kSymFive/kSymFour常量,GetUnaryBitWidthChangeCallCountfallback 使用input_dims全量乘积CompareScalarGE性能表项测试
合入前:
合入后:
回正44条
变更类型
核对清单
其他信息