Pull Request已成功合入, 合并人@CANN-robot
(感谢 zhang_shengjie 的贡献)Thanks for your pull-request.
The full list of commands accepted by me can be found at here。
You can get sig-info at here
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| repo-cann/graph-autofusion | ✅ wangxiaotian995, xchu42, xuyafei (3/2) | ✅ wangxiaotian995 (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
zhang_shengjie, thanks for your pull request. All authors of the commits have signed the CLA. 👍


/compile


流水线任务触发成功
任务链接 [d284fc7dfdcc4ee69ea81085068693fd][流水线指导]
| 任务名称 | 状态 | 日志 | 下载链接 |
|---|---|---|---|
| codecheck | ✅ SUCCESS | >>>>> | |
| SCA | ✅ SUCCESS | >>>>> | |
| antipoison | ✅ SUCCESS | >>>>> | |
| codecheck_checkpr | ✅ SUCCESS | ||
| pre_comment | ✅ SUCCESS | >>>>> | |
| codecheck_codestyle | ⚠️ WARNING | >>>>> | |
| codecheck_precommit | ⚠️ WARNING | >>>>> | >>>>> |
[2026-06-23 18:16:36] CI执行结束


流水线任务触发成功
任务链接 [1176b81511b1474ead0f7dd48c5536ef][流水线指导]
| 任务名称 | 状态 | 日志 | 下载链接 |
|---|---|---|---|
| UT_Test_Python_superkernel | ✅ SUCCESS | >>>>> | |
| ST_Test_Python_superkernel | ✅ SUCCESS | >>>>> | |
| UT_Test_superkernel | ✅ SUCCESS | >>>>> | |
| UT_Test_autofuse_framework | ✅ SUCCESS | >>>>> | |
| ST_Test_autofuse_framework | ✅ SUCCESS | >>>>> | |
| UT_Test_autofuse_ascendc_api | ✅ SUCCESS | >>>>> | |
| ST_Test_autofuse_ascendc_api | ✅ SUCCESS | >>>>> | |
| ST_Test_autofuse_e2e | ✅ SUCCESS | >>>>> | |
| pre_comment | ✅ SUCCESS | >>>>> |
[2026-06-23 18:26:02] CI执行结束


/lgtm


变更摘要
本 PR 主要解决 Reduce 算子性能模型中的三类问题:(1) RA 非对齐 tail 场景下 VF Head Cost 被按 tail 迭代次数重复计算,导致性能评估偏大——通过新增 BuildRepeatedVfGroupCost、BuildVfGroupBodyCost、BuildB64VfGroupBodyCost 等函数将 VF Head 与 body cost 分离,并引入 TailInplaceCostBuilder 回调使 tree reduce 路径可为 unaligned tail 指定专用 cost builder;(2) Reduce tile reorder 使用硬编码阈值判断轴交换,无法适配不同架构——改为参数化阈值并支持运行时动态 shape 的 reorder 规则下发;(3) 架构硬件参数(cache_line_size、vector_len_size)分散在多个文件中——新增 arch_param.h 集中管理。
主要改动
-
拆分 Reduce Tail VF Head Cost:在
reduce_api_perf_v2.cpp中新增BuildRepeatedVfGroupCost、BuildVfGroupBodyCost、BuildB64VfGroupBodyCost函数,将 VF Head Cost(常量 64 cycles)从 body cost 中分离,按min(repeats, 1) * kReduceVfHeadCost + repeats * body_cost计算,同时新增TailInplaceCostBuilder回调类型使 tree reduce 路径可指定专用 inplace cost builder。 -
Tile Reorder 参数化与运行时动态选择:在
arg_list_reorder.cpp/.h 和axes_reorder_solver_gen.cpp/.h 中重构 tile reorder 逻辑,用可配置阈值替代硬编码常量,并为动态 shape 场景增加运行时 reorder 规则下发机制,使其能根据实际 shape 动态选择最优轴序。 -
集中化架构硬件参数:新增
autofuse/att/base/arch_param.h,将 V1/V2 架构的cache_line_size和vector_len_size参数统一管理,消除原先分散在perf_param.h、perf_param_v1.h、perf_param_v2.h等多处硬编码。 -
模型信息接口扩展:
model_info.h和gen_model_info.cpp配合上述改动,增加对架构参数和 tail cost 新模型的适配字段,确保性能模型信息生成链路完整。


代码审查
审查总结
我对全部 16 个变更文件逐一进行了审查,以下按优先级汇总:
P2(1 个)
reduce_api_perf_v2.cpp:27— VF head cost 不一致:tail 路径使用硬编码kReduceVfHeadCost = 64,而 main-fold 路径通过VfPerfUtils::GetVFHeadCost()使用 20。两者代表不同层级的开销,但 body builder 完全去掉了 per-group 的 20 cycles,可能影响性能评估准确性。
P3(2 个)
arg_list_reorder.cpp:96-101—CeilDiv除零时静默返回 0,缺少错误日志,可能在未来调用方遗漏校验时掩盖 bug。axes_reorder_solver_gen.cpp:1079-1099—GenRuntimeCompoundExprValue使用字符串find+replace做占位名替换,在极端情况下可能错误匹配子串。
各文件审查结论
| 文件 | 结论 |
|---|---|
autofuse/att/base/arch_param.h |
无问题 |
autofuse/att/base/model_info.h |
无问题 |
autofuse/att/gen_model_info/api_perf_register/perf_param.h |
无问题 |
autofuse/att/gen_model_info/api_perf_register/v1/perf_param_v1.h |
无问题 |
autofuse/att/gen_model_info/expr_gen/arg_list_reorder.cpp |
P3: CeilDiv 除零静默 |
autofuse/att/gen_model_info/expr_gen/arg_list_reorder.h |
无问题 |
autofuse/att/gen_model_info/gen_model_info.cpp |
无问题 |
autofuse/att/generator/solver_pass_gen/axes_reorder_solver/axes_reorder_solver_gen.cpp |
P3: 字符串替换健壮性 |
autofuse/att/generator/solver_pass_gen/axes_reorder_solver/axes_reorder_solver_gen.h |
无问题 |
autofuse/att/generator/solver_pass_gen/solver_pass_manager.cpp |
无问题 |
autofuse/tests/common/stub/stub_solver_model_info.h |
无问题 |
autofuse/tests/ut/att/testcase/gen_model_info/expr_gen/test_arg_list_reorder.cpp |
无问题 |
autofuse/tests/ut/att/testcase/solver_pass_gen/axes_reorder_gen/test_axes_reorder_gen.cpp |
无问题 |
autofuse/tests/v35/ut/att/gen_model_info/api_perf_register/test_reduce_min_max_api_perf_v2.cpp |
无问题 |
autofuse/v35/att/api_perf_register/ascendc_api_perf/reduce_api_perf_v2.cpp |
P2: VF head cost 不一致 |
autofuse/v35/att/api_perf_register/perf_param_v2.h |
无问题 |
总体风险评估
中低风险。核心逻辑变更(VF head cost 分离、runtime reorder 机制)设计合理,无崩溃、数据损坏或安全漏洞。P2 的 head cost 不一致问题需要领域专家确认 64 vs 20 的语义差异是否 intentional,若不一致则可能影响性能模型精度。其余 P3 项为防御性改进建议。
| 类型 | 数量 |
|---|---|
| 🔴 阻塞 | 0 |
| 🟡 建议 | 3 |
💬 仅评论


🔵 Low Priority
CeilDiv(arg_list_reorder.cpp 第 96–101 行)在 divisor == 0U 时直接返回 0U,不记录任何错误。这会在调用方意外传入 0 除数时静默产生 0 结果,可能掩盖数据问题。当前所有调用点在传入 divisor 前均已校验(例如 HasSmallTailLargeReduceTile 第 296 行检查 tensor->data_type_size == 0U),但该函数作为工具函数,静默吞错的设计不够安全,未来新增调用点可能遗漏校验。
建议:在 divisor==0 分支增加错误日志,并返回一个标示异常的上界值(例如 UINT32_MAX),或者使用 assert/GE_ASSERT_TRUE 确保调用方不会传入零除数。


🔵 Low Priority
GenRuntimeCompoundExprValue(axes_reorder_solver_gen.cpp 第 1079–1099 行)使用字符串 find + replace 的方式将临时占位名(__runtime_expr_N__)替换为运行时求值表达式。该方案依赖占位名不会出现在任何替换值的字符串表示中。如果某个 GenRuntimeExprValue(arg) 的返回值恰好包含形如 __runtime_expr_M__ 的子串,则可能发生错误替换,导致生成的 C++ 代码不正确。
虽然当前占位名格式(带序号 + 双下划线)与正常 Expr 名称冲突概率极低,但随着系统演进,存在潜在风险。建议使用更安全的替换方式,例如按 token 边界替换,或在替换前对占位名做更唯一的编码(如加入随机后缀或使用不合法 C++ 标识符字符)。
建议:考虑在占位名中加入不会出现在正常 Expr 名中的字符(如 $),或用一次 regex_replace 整体替换,避免先后替换时的子串匹配问题。


🟡 Medium Priority
BuildRepeatedVfGroupCost(第 420 行)使用硬编码常量 kReduceVfHeadCost = 64(第 27 行)作为 VF 头开销。但系统中已有的 VF 头开销获取途径 VfPerfUtils::GetVFHeadCost() 返回 PerfParamTableV2::GetVectorFunctionHeadCost() 的值 20(perf_param_v2.cpp 第 418–420 行)。旧代码中,tail 路径通过 VfOpCost / GetVfGroupCost 间接使用了 20 这个值(每次调用加 20),新代码的 body builder(BuildVfGroupBodyCost、BuildB64VfGroupBodyCost)完全去掉了 per-group 的 20,改为在 BuildRepeatedVfGroupCost 中一次性加 64。
具体影响:对于 tail repeats=N 的场景,旧模型为 N * (20 * group_count + body_instructions),新模型为 1 * 64 + N * body_instructions。这意味着 per-instruction-group 的 20 cycles 头开销完全消失,取而代之的是一次性的 64 cycles。如果 20 和 64 对应的是不同层级的开销(per-instruct-group vs per-VF-invocation),需要确认 body cost 是否应当保留 per-group 的 20。目前 main-fold 路径仍使用 GetVfGroupCost(含 20),而所有 tail 路径均使用新模型(含 64),两者不一致。
触发条件:任何 Reduce tail 场景(AR tail inplace、RA tree reduce tail、RA B64 const tree cost、RA normal aligned/unaligned tail)。
建议:确认 kReduceVfHeadCost = 64 的语义:如果 64 是 Reduce VF 调用级头开销,应将其统一定义到 arch_param 或通过 PerfParamTable 查询,避免硬编码;同时确认 body builder 中是否需要恢复 per-instruct-group 的 GetVectorFunctionHeadCost()(20)。如果 20 本身就不该在 body 中出现,那 main-fold 路径也需同步调整,消除两者不一致。


/approve


Pull Request
描述
一、主要解决的问题
1.1 Reduce RA Unaligned Tail 性能模型重复计算 VF Head Cost
Reduce 算子在 RA(Reduce-Axis)非对齐 tail 场景下,性能模型将 VF Head Cost(向量函数头部开销)按 tail 迭代次数重复计算,导致性能评估偏大。实际硬件执行时 VF Head 仅执行一次,不随 tail 迭代重复。此问题同时影响 AR(Axis-Reduce)tail inplace 路径和 RA tree reduce 路径。
1.2 Reduce Tile Reorder 缺少参数化阈值和运行时动态选择
原有 Reduce tile reorder 逻辑使用硬编码阈值判断是否交换 tail/reduce 轴顺序,无法适配不同架构的 cache_line_size 和 vector_len_size。同时,对于动态 shape 场景(编译期无法确定轴大小),缺少运行时 reorder 规则下发机制,导致无法在运行时根据实际 shape 动态选择最优轴序。
1.3 架构硬件参数硬编码分散
V1/V2 架构的 cache_line_size 和 vector_len_size 分散在多个文件中以硬编码常量形式使用,维护不便且容易不一致。
二、修改方案
2.1 拆分 Reduce Tail VF Head Cost
BuildRepeatedVfGroupCost函数,将 VF Head Cost(常量 64 cycles)与 body cost 分离:cost = min(repeats, 1) * kReduceVfHeadCost + repeats * body_costBuildVfGroupBodyCost/BuildB64VfGroupBodyCost仅计算 body 部分(load + binary + store)的指令开销,不含 headTailInplaceCostBuilder回调类型,使 tree reduce 路径可为 unaligned tail 指定专用的 inplace cost builderBuildTreeBodyCost支持 tail_builder 回调,对 unaligned 场景使用专用 tail cost(如 4 load + 1 binary + 2 store),避免统一使用通用路径2.2 Reduce Tile Reorder 参数化与运行时规则
arch_param.h,定义ArchHardwareConfig结构体和 V1/V2 架构常量(kV1ArchHardwareConfig{512, 256}、kV2ArchHardwareConfig{128, 256})TilingScheduleConfig新增vector_len_size字段,TilingScheduleConfigTable新增GetVectorLenSize()虚方法PerfParamTable/PerfParamTableV1/PerfParamTableV2的GetMicroApiLen()统一从arch_param获取RuntimeReorderRule结构体,描述运行时 reorder 条件(preferred_axis、fallback_axis、condition_axis、compare_axis 及阈值)arg_list_reorder.cpp新增 reduce tile reorder 逻辑:RuntimeReorderRule下发至 solver codegen,使用原始轴(orig_axis)大小作为判断条件axes_reorder_solver_gen.cpp新增GenRuntimeReorderRules()方法,生成运行时 reorder 的 C++ 代码(条件判断 + local_buffer_vars 交换)solver_pass_manager.cpp将runtime_reorder_rules传递到 solver gen2.3 架构参数统一收归
TilingScheduleConfigTable子类使用arch_param::kV1ArchHardwareConfig/kV2ArchHardwareConfig替代硬编码TilingScheduleConfig默认值使用arch_param::kDefaultCacheLineSize/kDefaultVectorLenSize三、代码修改流程图
3.1 多文件调用关系图
graph TD A[arch_param.h<br/>新增: 架构硬件参数] B[model_info.h<br/>新增: RuntimeReorderRule<br/>新增: vector_len_size] C[perf_param.h / v1 / v2<br/>GetMicroApiLen 参数化] D[arg_list_reorder.cpp<br/>Reduce tile reorder 逻辑<br/>生成 RuntimeReorderRule] E[axes_reorder_solver_gen.cpp<br/>GenRuntimeReorderRules<br/>生成运行时 reorder 代码] F[solver_pass_manager.cpp<br/>传递 runtime_reorder_rules] G[reduce_api_perf_v2.cpp<br/>拆分 VF Head/Body Cost] A --> B A --> C B --> D B --> F D -->|生成 RuntimeReorderRule| F F -->|设置到 solver_gen| E E -->|codegen 输出| H[运行时 reorder C++ 代码] style A fill:#90EE90,stroke:#006400,stroke-width:2px style B fill:#FFD700,stroke:#B8860B,stroke-width:2px style D fill:#FFD700,stroke:#B8860B,stroke-width:2px style E fill:#FFD700,stroke:#B8860B,stroke-width:2px style G fill:#FFD700,stroke:#B8860B,stroke-width:2px3.2 Reduce Tile Reorder 决策流程
flowchart TD Start(["FindSpecialArgs 遍历节点"]) --> CheckBlock{"Reduce 轴<br/>Block Split?"} CheckBlock -->|"是"| SetTilingR["tiling_R_ = true<br/>保持 Tiling_R 排序"] CheckBlock -->|"否"| CheckTile{"Reduce 轴<br/>Tile Split?"} CheckTile -->|"否"| Default["保持默认轴顺序"] CheckTile -->|"是"| CheckThreshold{"tail < cache_line<br/>且 reduce > vector_len?"} CheckThreshold -->|"不满足"| Default CheckThreshold -->|"满足"| CheckStatic{"shape 静态可知?"} CheckStatic -->|"是"| StaticReorder["prefer_reduce_tile_ = true<br/>编译期提升 Tail 轴优先级"] CheckStatic -->|"否"| DynamicRule["记录 RuntimeReorderRule<br/>传递给 codegen"] DynamicRule --> Codegen["solver 函数中生成<br/>运行时条件交换代码"] style Start fill:#90EE90,stroke:#006400,stroke-width:2px style StaticReorder fill:#e1f5ff,stroke:#0288d1,stroke-width:2px style DynamicRule fill:#e1f5ff,stroke:#0288d1,stroke-width:2px style Codegen fill:#e1f5ff,stroke:#0288d1,stroke-width:2px style SetTilingR fill:#FFD700,stroke:#B8860B,stroke-width:2px3.3 Reduce Tail VF Cost 拆分流程
flowchart TD Start([BuildTreeBodyCost]) --> HasTail{tail_builder != nullptr?} HasTail -->|Yes| TailBuild[调用 tail_builder<br/>BuildNormalUnalignedInplaceCost] HasTail -->|No| DefaultBuild[BuildRepeatedVfGroupCost<br/>head + repeats * body] TailBuild --> RepeatedCost[BuildRepeatedVfGroupCost] DefaultBuild --> RepeatedCost RepeatedCost --> HeadCost[min repeats 1 * kReduceVfHeadCost 64] RepeatedCost --> BodyCost[repeats * body_cost] HeadCost --> Total[cost = head + body + main_fold + tail_fold] BodyCost --> Total Total --> Return([返回总 cost]) style Start fill:#90EE90,stroke:#006400,stroke-width:2px style RepeatedCost fill:#e1f5ff,stroke:#0066cc,stroke-width:2px style HeadCost fill:#FFD700,stroke:#B8860B,stroke-width:2px变更类型
关联的Issue
如何测试
一、测试用例说明
1.1 单元测试
test_arg_list_reorder.cpp:新增 7 个测试用例keep_tiling_r_arg_list_when_reduce_block_split:验证 reduce 分核场景保持 tiling_R_arg_list 顺序keep_tiling_r_arg_list_when_reduce_block_split_without_threshold_match:验证阈值不匹配时不触发 reorderreorder_single_template_for_reduce_tile_small_tail_large_reduce:验证小 tail 大 reduce 单模板 reorderkeep_default_single_template_for_reduce_tile_without_threshold_match:验证多种阈值不匹配场景record_runtime_reorder_for_dynamic_reduce_tile:验证动态 shape 生成 RuntimeReorderRulev2_micro_api_len_equals_schedule_vector_len:验证 V2 MicroApi 长度与 vector_len_size 一致v1_micro_api_len_equals_schedule_vector_len:验证 V1 MicroApi 长度与 vector_len_size 一致test_axes_reorder_gen.cpp:新增 3 个测试用例GenRuntimeReorderRuleForDynamicReduceTile:验证运行时 reorder 代码生成GenRuntimeReorderRuleUsesOriginalAxesForDynamicReduceTile:验证使用原始轴生成条件GenRuntimeReorderRuleUsesOriginalAxisProductForDynamicReduceTile:验证多轴乘积场景test_reduce_min_max_api_perf_v2.cpp:新增 5 个测试用例ReduceSumRaUnalignedTailRDoesNotRepeatVfHead:验证 RA 非对齐 tail 不重复 VF headReduceSumRaAlignedTailRDoesNotRepeatVfHead:验证 RA 对齐 tail 不重复 VF headReduceSumRaSymbolicAlignedTailDoesNotRepeatVfHead:验证符号表达式对齐场景ReduceRaB64TailRDoesNotRepeatVfHead:验证 B64 特化路径ReduceArTailRDoesNotRepeatVfHead:验证 AR tail 路径核对清单
其他信息
验证方法
运行相关 UT:
注意事项
min(repeats, 1) * 64 + repeats * body,对 repeats=0 的边界场景无影响(cost=0)arch_param.h中的硬件参数为inline constexpr,不引入额外链接依赖提交记录
修改文件清单