已合并
fix: 优化 Reduce RA 未对齐尾部开销和 tile reorder 轴优先级 #1006
fix: 优化 Reduce RA 未对齐尾部开销和 tile reorder 轴优先级 #1006
已合并
zhang_shengjie创建于 6月23日
zhang_shengjie成员
6月23日

Pull Request

描述

一、主要解决的问题

1.1 Reduce RA Unaligned Tail 性能模型重复计算 VF Head Cost

Reduce 算子在 RA(Reduce-Axis)非对齐 tail 场景下,性能模型将 VF Head Cost(向量函数头部开销)按 tail 迭代次数重复计算,导致性能评估偏大。实际硬件执行时 VF Head 仅执行一次,不随 tail 迭代重复。此问题同时影响 AR(Axis-Reduce)tail inplace 路径和 RA tree reduce 路径。

1.2 Reduce Tile Reorder 缺少参数化阈值和运行时动态选择

原有 Reduce tile reorder 逻辑使用硬编码阈值判断是否交换 tail/reduce 轴顺序,无法适配不同架构的 cache_line_size 和 vector_len_size。同时,对于动态 shape 场景(编译期无法确定轴大小),缺少运行时 reorder 规则下发机制,导致无法在运行时根据实际 shape 动态选择最优轴序。

1.3 架构硬件参数硬编码分散

V1/V2 架构的 cache_line_size 和 vector_len_size 分散在多个文件中以硬编码常量形式使用,维护不便且容易不一致。

二、修改方案

2.1 拆分 Reduce Tail VF Head Cost

  • 新增 BuildRepeatedVfGroupCost 函数,将 VF Head Cost(常量 64 cycles)与 body cost 分离:cost = min(repeats, 1) * kReduceVfHeadCost + repeats * body_cost
  • 新增 BuildVfGroupBodyCost / BuildB64VfGroupBodyCost 仅计算 body 部分(load + binary + store)的指令开销,不含 head
  • 引入 TailInplaceCostBuilder 回调类型,使 tree reduce 路径可为 unaligned tail 指定专用的 inplace cost builder
  • 修改 BuildTreeBodyCost 支持 tail_builder 回调,对 unaligned 场景使用专用 tail cost(如 4 load + 1 binary + 2 store),避免统一使用通用路径

2.2 Reduce Tile Reorder 参数化与运行时规则

  • 新增 arch_param.h,定义 ArchHardwareConfig 结构体和 V1/V2 架构常量(kV1ArchHardwareConfig{512, 256}、kV2ArchHardwareConfig{128, 256})
  • TilingScheduleConfig 新增 vector_len_size 字段,TilingScheduleConfigTable 新增 GetVectorLenSize() 虚方法
  • PerfParamTable / PerfParamTableV1 / PerfParamTableV2 的 GetMicroApiLen() 统一从 arch_param 获取
  • 新增 RuntimeReorderRule 结构体,描述运行时 reorder 条件(preferred_axis、fallback_axis、condition_axis、compare_axis 及阈值)
  • arg_list_reorder.cpp 新增 reduce tile reorder 逻辑:
    • 编译期常量 shape:基于阈值(condition_threshold=64, compare_threshold=128)直接交换单模板内轴序
    • 动态 shape:生成 RuntimeReorderRule 下发至 solver codegen,使用原始轴(orig_axis)大小作为判断条件
  • axes_reorder_solver_gen.cpp 新增 GenRuntimeReorderRules() 方法,生成运行时 reorder 的 C++ 代码(条件判断 + local_buffer_vars 交换)
  • solver_pass_manager.cpp 将 runtime_reorder_rules 传递到 solver gen

2.3 架构参数统一收归

  • V1/V2 的 TilingScheduleConfigTable 子类使用 arch_param::kV1ArchHardwareConfig / kV2ArchHardwareConfig 替代硬编码
  • TilingScheduleConfig 默认值使用 arch_param::kDefaultCacheLineSize / kDefaultVectorLenSize

三、代码修改流程图

3.1 多文件调用关系图

graph TD
    A[arch_param.h<br/>新增: 架构硬件参数]
    B[model_info.h<br/>新增: RuntimeReorderRule<br/>新增: vector_len_size]
    C[perf_param.h / v1 / v2<br/>GetMicroApiLen 参数化]
    D[arg_list_reorder.cpp<br/>Reduce tile reorder 逻辑<br/>生成 RuntimeReorderRule]
    E[axes_reorder_solver_gen.cpp<br/>GenRuntimeReorderRules<br/>生成运行时 reorder 代码]
    F[solver_pass_manager.cpp<br/>传递 runtime_reorder_rules]
    G[reduce_api_perf_v2.cpp<br/>拆分 VF Head/Body Cost]

    A --> B
    A --> C
    B --> D
    B --> F
    D -->|生成 RuntimeReorderRule| F
    F -->|设置到 solver_gen| E
    E -->|codegen 输出| H[运行时 reorder C++ 代码]

    style A fill:#90EE90,stroke:#006400,stroke-width:2px
    style B fill:#FFD700,stroke:#B8860B,stroke-width:2px
    style D fill:#FFD700,stroke:#B8860B,stroke-width:2px
    style E fill:#FFD700,stroke:#B8860B,stroke-width:2px
    style G fill:#FFD700,stroke:#B8860B,stroke-width:2px

3.2 Reduce Tile Reorder 决策流程

flowchart TD
    Start(["FindSpecialArgs 遍历节点"]) --> CheckBlock{"Reduce 轴<br/>Block Split?"}
    CheckBlock -->|"是"| SetTilingR["tiling_R_ = true<br/>保持 Tiling_R 排序"]
    CheckBlock -->|"否"| CheckTile{"Reduce 轴<br/>Tile Split?"}
    CheckTile -->|"否"| Default["保持默认轴顺序"]
    CheckTile -->|"是"| CheckThreshold{"tail < cache_line<br/>且 reduce > vector_len?"}
    CheckThreshold -->|"不满足"| Default
    CheckThreshold -->|"满足"| CheckStatic{"shape 静态可知?"}
    CheckStatic -->|"是"| StaticReorder["prefer_reduce_tile_ = true<br/>编译期提升 Tail 轴优先级"]
    CheckStatic -->|"否"| DynamicRule["记录 RuntimeReorderRule<br/>传递给 codegen"]
    DynamicRule --> Codegen["solver 函数中生成<br/>运行时条件交换代码"]

    style Start fill:#90EE90,stroke:#006400,stroke-width:2px
    style StaticReorder fill:#e1f5ff,stroke:#0288d1,stroke-width:2px
    style DynamicRule fill:#e1f5ff,stroke:#0288d1,stroke-width:2px
    style Codegen fill:#e1f5ff,stroke:#0288d1,stroke-width:2px
    style SetTilingR fill:#FFD700,stroke:#B8860B,stroke-width:2px

3.3 Reduce Tail VF Cost 拆分流程

flowchart TD
    Start([BuildTreeBodyCost]) --> HasTail{tail_builder != nullptr?}
    HasTail -->|Yes| TailBuild[调用 tail_builder<br/>BuildNormalUnalignedInplaceCost]
    HasTail -->|No| DefaultBuild[BuildRepeatedVfGroupCost<br/>head + repeats * body]
    TailBuild --> RepeatedCost[BuildRepeatedVfGroupCost]
    DefaultBuild --> RepeatedCost
    RepeatedCost --> HeadCost[min repeats 1 * kReduceVfHeadCost 64]
    RepeatedCost --> BodyCost[repeats * body_cost]
    HeadCost --> Total[cost = head + body + main_fold + tail_fold]
    BodyCost --> Total
    Total --> Return([返回总 cost])

    style Start fill:#90EE90,stroke:#006400,stroke-width:2px
    style RepeatedCost fill:#e1f5ff,stroke:#0066cc,stroke-width:2px
    style HeadCost fill:#FFD700,stroke:#B8860B,stroke-width:2px

变更类型

关联的Issue

如何测试

一、测试用例说明

1.1 单元测试

  • test_arg_list_reorder.cpp:新增 7 个测试用例

    • keep_tiling_r_arg_list_when_reduce_block_split:验证 reduce 分核场景保持 tiling_R_arg_list 顺序
    • keep_tiling_r_arg_list_when_reduce_block_split_without_threshold_match:验证阈值不匹配时不触发 reorder
    • reorder_single_template_for_reduce_tile_small_tail_large_reduce:验证小 tail 大 reduce 单模板 reorder
    • keep_default_single_template_for_reduce_tile_without_threshold_match:验证多种阈值不匹配场景
    • record_runtime_reorder_for_dynamic_reduce_tile:验证动态 shape 生成 RuntimeReorderRule
    • v2_micro_api_len_equals_schedule_vector_len:验证 V2 MicroApi 长度与 vector_len_size 一致
    • v1_micro_api_len_equals_schedule_vector_len:验证 V1 MicroApi 长度与 vector_len_size 一致
  • test_axes_reorder_gen.cpp:新增 3 个测试用例

    • GenRuntimeReorderRuleForDynamicReduceTile:验证运行时 reorder 代码生成
    • GenRuntimeReorderRuleUsesOriginalAxesForDynamicReduceTile:验证使用原始轴生成条件
    • GenRuntimeReorderRuleUsesOriginalAxisProductForDynamicReduceTile:验证多轴乘积场景
  • test_reduce_min_max_api_perf_v2.cpp:新增 5 个测试用例

    • ReduceSumRaUnalignedTailRDoesNotRepeatVfHead:验证 RA 非对齐 tail 不重复 VF head
    • ReduceSumRaAlignedTailRDoesNotRepeatVfHead:验证 RA 对齐 tail 不重复 VF head
    • ReduceSumRaSymbolicAlignedTailDoesNotRepeatVfHead:验证符号表达式对齐场景
    • ReduceRaB64TailRDoesNotRepeatVfHead:验证 B64 特化路径
    • ReduceArTailRDoesNotRepeatVfHead:验证 AR tail 路径

核对清单

其他信息

验证方法

运行相关 UT:

sh build.sh -u --module=autofuse_framework -j 8

注意事项

  • VF Head Cost 拆分为 min(repeats, 1) * 64 + repeats * body,对 repeats=0 的边界场景无影响(cost=0)
  • Runtime reorder 仅在动态 shape 且满足 orig_axis 阈值条件时生成,不影响编译期可确定的场景
  • arch_param.h 中的硬件参数为 inline constexpr,不引入额外链接依赖

提交记录

Commit 描述 修改文件数 修改行数
89aa3f2 perf: parameterize att tiling reorder thresholds 9 +303/-79
ade509a perf: adjust att reduce tile reorder 2 +63/-18
a1f1458 fix: resolve att warning issues 4 +284/-280
36a00ca fix: add reduce tile reorder selection logs 4 +46/-12
f8651e5 fix: use original axes for reduce tile reorder 5 +148/-9
2fa5eb0 fix(att): adjust reduce ra unaligned tail cost 2 +63/-18
050018e fix(att): split reduce tail vf head cost 2 +139/-40

修改文件清单

文件路径 修改类型 说明
autofuse/att/base/arch_param.h 新增 架构硬件参数定义(V1/V2 cache_line_size, vector_len_size)
autofuse/att/base/model_info.h 修改 新增 RuntimeReorderRule、vector_len_size、GetVectorLenSize()
autofuse/att/gen_model_info/api_perf_register/perf_param.h 修改 GetMicroApiLen() 使用 arch_param 常量
autofuse/att/gen_model_info/api_perf_register/v1/perf_param_v1.h 修改 V1 GetMicroApiLen/CacheLineSize/VectorLenSize 参数化
autofuse/att/gen_model_info/expr_gen/arg_list_reorder.cpp 修改 新增 reduce tile reorder 逻辑、RuntimeReorderRule 生成
autofuse/att/gen_model_info/expr_gen/arg_list_reorder.h 修改 SortArgList 新增 runtime_reorder_rules 参数
autofuse/att/gen_model_info/gen_model_info.cpp 修改 传递 runtime_reorder_rules 到 SortArgList
autofuse/att/generator/solver_pass_gen/axes_reorder_solver/axes_reorder_solver_gen.cpp 修改 新增 GenRuntimeReorderRules() 代码生成
autofuse/att/generator/solver_pass_gen/axes_reorder_solver/axes_reorder_solver_gen.h 修改 新增 RuntimeReorderRule 相关成员和方法
autofuse/att/generator/solver_pass_gen/solver_pass_manager.cpp 修改 传递 runtime_reorder_rules 到 solver_gen
autofuse/tests/common/stub/stub_solver_model_info.h 修改 stub 补充 vector_len_size
autofuse/tests/ut/att/testcase/gen_model_info/expr_gen/test_arg_list_reorder.cpp 修改 新增 7 个 reduce tile reorder 和 vector_len 测试
autofuse/tests/ut/att/testcase/solver_pass_gen/axes_reorder_gen/test_axes_reorder_gen.cpp 修改 新增 3 个 runtime reorder codegen 测试
autofuse/tests/v35/ut/att/gen_model_info/api_perf_register/test_reduce_min_max_api_perf_v2.cpp 修改 新增 5 个 VF head cost 拆分验证测试
autofuse/v35/att/api_perf_register/ascendc_api_perf/reduce_api_perf_v2.cpp 修改 拆分 VF Head/Body Cost、新增 BuildRepeatedVfGroupCost
autofuse/v35/att/api_perf_register/perf_param_v2.h 修改 V2 GetMicroApiLen/CacheLineSize/VectorLenSize 参数化
likedislike
Pull Request已成功合入, 合并人@CANN-robot
(感谢 zhang_shengjie 的贡献)
Zzhang_shengjie成员
6月23日 创建了 pull request,commit 1e14e70e
CANN-robotCANN-robot成员
6月23日 添加了label:stat/needs-squash
CANN-robotCANN-robot成员
6月23日 添加了label:cann-cla/yes
CANN-robot
CANN-robot成员
6月23日 评论:

Thanks for your pull-request.
The full list of commands accepted by me can be found at here。
You can get sig-info at here


PR Approval Progress

✅ Congratulations! All modules have met the lgtm and approve requirements.

Module Approval Details

module lgtm status approve status
repo-cann/graph-autofusion ✅ wangxiaotian995, xchu42, xuyafei (3/2) ✅ wangxiaotian995 (1/1)

💡 Tip:

  • Committer can comment /approve or /lgtm
  • Commenting /approve implies both code review (lgtm) and intent to merge (approve)

CLA Signature Pass

zhang_shengjie, thanks for your pull request. All authors of the commits have signed the CLA. 👍

likedislike
zhang_shengjie成员
6月23日 评论:

/compile

likedislike
Zzhang_shengjie成员
6月23日 预合并成功(commit_id: e54e5139d62a72815f66e4558d677b5a8809c36e)
CANN-robotCANN-robot成员
6月23日 添加了label:ci-pipeline-running
CANN-robot
CANN-robot成员
6月23日 评论:

流水线任务触发成功
任务链接 [d284fc7dfdcc4ee69ea81085068693fd][流水线指导]

任务名称状态日志下载链接
codecheck ✅ SUCCESS >>>>>
SCA ✅ SUCCESS >>>>>
antipoison ✅ SUCCESS >>>>>
codecheck_checkpr ✅ SUCCESS
pre_comment ✅ SUCCESS >>>>>
codecheck_codestyle ⚠️ WARNING >>>>>
codecheck_precommit ⚠️ WARNING >>>>> >>>>>

[2026-06-23 18:16:36]    CI执行结束

likedislike
CANN-robot
CANN-robot成员
6月23日 评论:

流水线任务触发成功
任务链接 [1176b81511b1474ead0f7dd48c5536ef][流水线指导]

任务名称状态日志下载链接
UT_Test_Python_superkernel ✅ SUCCESS >>>>>
ST_Test_Python_superkernel ✅ SUCCESS >>>>>
UT_Test_superkernel ✅ SUCCESS >>>>>
UT_Test_autofuse_framework ✅ SUCCESS >>>>>
ST_Test_autofuse_framework ✅ SUCCESS >>>>>
UT_Test_autofuse_ascendc_api ✅ SUCCESS >>>>>
ST_Test_autofuse_ascendc_api ✅ SUCCESS >>>>>
ST_Test_autofuse_e2e ✅ SUCCESS >>>>>
pre_comment ✅ SUCCESS >>>>>

[2026-06-23 18:26:02]    CI执行结束

likedislike
CANN-robot
CANN-robot成员
6月23日 评论:

流水线任务触发成功
任务链接 [5666f4e7aeda4fccb244528c11f60b19][流水线指导]

任务名称状态日志下载链接
Compile_Ascend_X86 ✅ SUCCESS >>>>> >>>>>
Compile_Ascend_ARM ✅ SUCCESS >>>>> >>>>>
pre_comment ✅ SUCCESS >>>>>
PreSmoke_A900_npupool ✅ SUCCESS >>>>>

[2026-06-23 18:24:02]    CI执行结束

likedislike
CANN-robotCANN-robot成员
6月23日 删除了label:ci-pipeline-running
CANN-robotCANN-robot成员
6月23日 添加了label:ci-pipeline-passed
Zzhang_shengjie成员
6月23日 修改了pull request 的描述
xchu42
xchu42成员
6月23日 评论:

/lgtm

likedislike
xuyafei成员
6月23日 评论:

/lgtm

likedislike
CANN-robotCANN-robot成员
6月23日 添加了label:lgtm
atomgit-bot
atomgit-bot
6月23日 评论:

变更摘要

本 PR 主要解决 Reduce 算子性能模型中的三类问题:(1) RA 非对齐 tail 场景下 VF Head Cost 被按 tail 迭代次数重复计算,导致性能评估偏大——通过新增 BuildRepeatedVfGroupCost、BuildVfGroupBodyCost、BuildB64VfGroupBodyCost 等函数将 VF Head 与 body cost 分离,并引入 TailInplaceCostBuilder 回调使 tree reduce 路径可为 unaligned tail 指定专用 cost builder;(2) Reduce tile reorder 使用硬编码阈值判断轴交换,无法适配不同架构——改为参数化阈值并支持运行时动态 shape 的 reorder 规则下发;(3) 架构硬件参数(cache_line_size、vector_len_size)分散在多个文件中——新增 arch_param.h 集中管理。

主要改动

  • 拆分 Reduce Tail VF Head Cost:在 reduce_api_perf_v2.cpp 中新增 BuildRepeatedVfGroupCost、BuildVfGroupBodyCost、BuildB64VfGroupBodyCost 函数,将 VF Head Cost(常量 64 cycles)从 body cost 中分离,按 min(repeats, 1) * kReduceVfHeadCost + repeats * body_cost 计算,同时新增 TailInplaceCostBuilder 回调类型使 tree reduce 路径可指定专用 inplace cost builder。

  • Tile Reorder 参数化与运行时动态选择:在 arg_list_reorder.cpp/.h 和 axes_reorder_solver_gen.cpp/.h 中重构 tile reorder 逻辑,用可配置阈值替代硬编码常量,并为动态 shape 场景增加运行时 reorder 规则下发机制,使其能根据实际 shape 动态选择最优轴序。

  • 集中化架构硬件参数:新增 autofuse/att/base/arch_param.h,将 V1/V2 架构的 cache_line_size 和 vector_len_size 参数统一管理,消除原先分散在 perf_param.h、perf_param_v1.h、perf_param_v2.h 等多处硬编码。

  • 模型信息接口扩展:model_info.h 和 gen_model_info.cpp 配合上述改动,增加对架构参数和 tail cost 新模型的适配字段,确保性能模型信息生成链路完整。

likedislike
不准确?
atomgit-bot
atomgit-bot
6月23日 评论:

代码审查

审查总结

我对全部 16 个变更文件逐一进行了审查,以下按优先级汇总:

P2(1 个)

  • reduce_api_perf_v2.cpp:27 — VF head cost 不一致:tail 路径使用硬编码 kReduceVfHeadCost = 64,而 main-fold 路径通过 VfPerfUtils::GetVFHeadCost() 使用 20。两者代表不同层级的开销,但 body builder 完全去掉了 per-group 的 20 cycles,可能影响性能评估准确性。

P3(2 个)

  • arg_list_reorder.cpp:96-101 — CeilDiv 除零时静默返回 0,缺少错误日志,可能在未来调用方遗漏校验时掩盖 bug。
  • axes_reorder_solver_gen.cpp:1079-1099 — GenRuntimeCompoundExprValue 使用字符串 find+replace 做占位名替换,在极端情况下可能错误匹配子串。

各文件审查结论

文件 结论
autofuse/att/base/arch_param.h 无问题
autofuse/att/base/model_info.h 无问题
autofuse/att/gen_model_info/api_perf_register/perf_param.h 无问题
autofuse/att/gen_model_info/api_perf_register/v1/perf_param_v1.h 无问题
autofuse/att/gen_model_info/expr_gen/arg_list_reorder.cpp P3: CeilDiv 除零静默
autofuse/att/gen_model_info/expr_gen/arg_list_reorder.h 无问题
autofuse/att/gen_model_info/gen_model_info.cpp 无问题
autofuse/att/generator/solver_pass_gen/axes_reorder_solver/axes_reorder_solver_gen.cpp P3: 字符串替换健壮性
autofuse/att/generator/solver_pass_gen/axes_reorder_solver/axes_reorder_solver_gen.h 无问题
autofuse/att/generator/solver_pass_gen/solver_pass_manager.cpp 无问题
autofuse/tests/common/stub/stub_solver_model_info.h 无问题
autofuse/tests/ut/att/testcase/gen_model_info/expr_gen/test_arg_list_reorder.cpp 无问题
autofuse/tests/ut/att/testcase/solver_pass_gen/axes_reorder_gen/test_axes_reorder_gen.cpp 无问题
autofuse/tests/v35/ut/att/gen_model_info/api_perf_register/test_reduce_min_max_api_perf_v2.cpp 无问题
autofuse/v35/att/api_perf_register/ascendc_api_perf/reduce_api_perf_v2.cpp P2: VF head cost 不一致
autofuse/v35/att/api_perf_register/perf_param_v2.h 无问题

总体风险评估

中低风险。核心逻辑变更(VF head cost 分离、runtime reorder 机制)设计合理,无崩溃、数据损坏或安全漏洞。P2 的 head cost 不一致问题需要领域专家确认 64 vs 20 的语义差异是否 intentional,若不一致则可能影响性能模型精度。其余 P3 项为防御性改进建议。

类型 数量
🔴 阻塞 0
🟡 建议 3

💬 仅评论

likedislike
不准确?
atomgit-bot
atomgit-bot6月23日进行代码检视1
autofuse/att/gen_model_info/expr_gen/arg_list_reorder.cpp
@@ -18,0 +98,4 @@
98+ return 0U;
99+ }
100+ return dividend / divisor + ((dividend % divisor) == 0U ? 0U : 1U);
101+}
atomgit-bot
atomgit-bot6月23日评论:

🔵 Low Priority

CeilDiv(arg_list_reorder.cpp 第 96–101 行)在 divisor == 0U 时直接返回 0U,不记录任何错误。这会在调用方意外传入 0 除数时静默产生 0 结果,可能掩盖数据问题。当前所有调用点在传入 divisor 前均已校验(例如 HasSmallTailLargeReduceTile 第 296 行检查 tensor->data_type_size == 0U),但该函数作为工具函数,静默吞错的设计不够安全,未来新增调用点可能遗漏校验。

建议:在 divisor==0 分支增加错误日志,并返回一个标示异常的上界值(例如 UINT32_MAX),或者使用 assert/GE_ASSERT_TRUE 确保调用方不会传入零除数。

likedislike
不准确?
atomgit-bot
atomgit-bot6月23日进行代码检视1
autofuse/att/generator/solver_pass_gen/axes_reorder_solver/axes_reorder_solver_gen.cpp
@@ -1056,0 +1096,4 @@
1096+ }
1097+ }
1098+ return runtime_expr_str;
1099+}
atomgit-bot
atomgit-bot6月23日评论:

🔵 Low Priority

GenRuntimeCompoundExprValue(axes_reorder_solver_gen.cpp 第 1079–1099 行)使用字符串 find + replace 的方式将临时占位名(__runtime_expr_N__)替换为运行时求值表达式。该方案依赖占位名不会出现在任何替换值的字符串表示中。如果某个 GenRuntimeExprValue(arg) 的返回值恰好包含形如 __runtime_expr_M__ 的子串,则可能发生错误替换,导致生成的 C++ 代码不正确。

虽然当前占位名格式(带序号 + 双下划线)与正常 Expr 名称冲突概率极低,但随着系统演进,存在潜在风险。建议使用更安全的替换方式,例如按 token 边界替换,或在替换前对占位名做更唯一的编码(如加入随机后缀或使用不合法 C++ 标识符字符)。

建议:考虑在占位名中加入不会出现在正常 Expr 名中的字符(如 $),或用一次 regex_replace 整体替换,避免先后替换时的子串匹配问题。

likedislike
不准确?
atomgit-bot
atomgit-bot6月23日进行代码检视1
autofuse/v35/att/api_perf_register/ascendc_api_perf/reduce_api_perf_v2.cpp
@@ -24,6 +24,7 @@ namespace {
2424constexpr int64_t kOneStride = 32768;
2525constexpr int64_t kVectorBlockBytes = 256;
2626constexpr int64_t kB64RegTraitNumTwoElements = 64;
27+constexpr int64_t kReduceVfHeadCost = 64;
atomgit-bot
atomgit-bot6月23日评论:

🟡 Medium Priority

BuildRepeatedVfGroupCost(第 420 行)使用硬编码常量 kReduceVfHeadCost = 64(第 27 行)作为 VF 头开销。但系统中已有的 VF 头开销获取途径 VfPerfUtils::GetVFHeadCost() 返回 PerfParamTableV2::GetVectorFunctionHeadCost() 的值 20(perf_param_v2.cpp 第 418–420 行)。旧代码中,tail 路径通过 VfOpCost / GetVfGroupCost 间接使用了 20 这个值(每次调用加 20),新代码的 body builder(BuildVfGroupBodyCost、BuildB64VfGroupBodyCost)完全去掉了 per-group 的 20,改为在 BuildRepeatedVfGroupCost 中一次性加 64。

具体影响:对于 tail repeats=N 的场景,旧模型为 N * (20 * group_count + body_instructions),新模型为 1 * 64 + N * body_instructions。这意味着 per-instruction-group 的 20 cycles 头开销完全消失,取而代之的是一次性的 64 cycles。如果 20 和 64 对应的是不同层级的开销(per-instruct-group vs per-VF-invocation),需要确认 body cost 是否应当保留 per-group 的 20。目前 main-fold 路径仍使用 GetVfGroupCost(含 20),而所有 tail 路径均使用新模型(含 64),两者不一致。

触发条件:任何 Reduce tail 场景(AR tail inplace、RA tree reduce tail、RA B64 const tree cost、RA normal aligned/unaligned tail)。

建议:确认 kReduceVfHeadCost = 64 的语义:如果 64 是 Reduce VF 调用级头开销,应将其统一定义到 arch_param 或通过 PerfParamTable 查询,避免硬编码;同时确认 body builder 中是否需要恢复 per-instruct-group 的 GetVectorFunctionHeadCost()(20)。如果 20 本身就不该在 body 中出现,那 main-fold 路径也需同步调整,消除两者不一致。

likedislike
不准确?
wangxiaotian995成员
6月23日 评论:

/approve

likedislike
CANN-robotCANN-robot成员
6月23日 添加了label:approved
CANN-robotCANN-robot成员
6月23日 解决了最后一个问题
CANN-robotCANN-robot成员
6月23日 合入了pull request