Pull Request已成功合入, 合并人@CANN-robot
(感谢 Tian_1122 的贡献)变更摘要
本 PR 为 foreach 系列算子补齐整型(int16/int8/uint8)UT 覆盖,围绕整型路径的溢出回绕(wrap)语义展开:为 ForeachSubList 新增整型用例(tiling key 5/7/8 走 wrap kernel),ForeachAddScalar 整型输入改为全量程随机以触发溢出路径,ForeachExp/ForeachExpm1 按"整进浮出"语义补充 DT_FLOAT 推导用例并适配 kernel UT;同时将 UT tiling 辅助函数 DivideUbMemory1 的 BINARY_LIST 分支对齐生产端 NeedCastCompute 的 UB 划分,并修复 sub_list 比对脚本的整型精确比较与退出码问题。
主要改动
- 新增
ForeachSubList整型用例:test_foreach_sub_list.cpp新增 int16/int8/uint8 三个用例(tiling key 5/7/8,走foreach_sub_listwrap kernel),对应gen_data.py增加整型 dtype 映射,采用全量程随机输入、int64精算后astype折叠生成 golden。 ForeachAddScalar数据生成改随机:add_scalar_data/gen_data.py将整型输入由固定值 10 改为全量程随机,触发溢出回绕路径,golden 同样按int64精算后折叠。ForeachExp/ForeachExpm1整进浮出适配: 新增infer_dtype_int_to_floatinfershape 用例验证整型输入推导DT_FLOAT输出;kernel UT 输出 tensor 改按float32分配/比对,gen_data.py为整型写入float32golden 与占位输入文件。- tiling 辅助函数对齐生产端 UB 划分:
foreach_abs_tiling_function.h的DivideUbMemory1BINARY_LIST 分支将 dataType 5/7/8(int16/int8/uint8)纳入与 bf16 相同的额外中间缓冲及 32 字节对齐处理,避免 wrap kernel 缓冲需求超出 UB 预算。 - 修复
sub_list比对脚本:compare_data.py对整型采用精确比较(precision=0),并让process()返回比对结果,修复退出码恒错的问题。


代码审查
审查结论
发现汇总
- P2 × 1
foreach/foreach_sub_list/tests/ut/op_kernel/sub_list_data/compare_data.py:59—process()已改为返回比对结果, 但__main__仍然丢弃返回值, C++ 侧system()也不检查返回码, PR 声称修复的"退出码恒错"问题实际未生效; 比对失败时脚本仍以退出码 0 结束, 新增的 int16/int8/uint8 精确比对在 CI 中会被静默放行。
各文件审查确认
| 文件 | 审查结果 |
|---|---|
foreach_abs_tiling_function.h |
无问题。BINARY_LIST 分支把 dataType 4/5/7/8 纳入 temp-casting 划分, 与生产端 foreach_tiling_class.h 的 NeedCastCompute(DT_BF16/DT_INT16/DT_INT8/DT_UINT8) 及 BYTE_BLOCK_FOR_BF16 对齐逻辑逐条一致; dataType 编号映射(4=bf16、5=int16、7=int8、8=uint8)与各测试 Init() 及 GetTilingN() 一致 |
add_scalar_data/gen_data.py |
无问题。整型全量程随机 + int64 精算后 astype 折叠, 与 kernel 二进制补码回绕语义一致; 区间 (-32768,32768)/(-128,128)/(0,256) 均为 [lo, hi) 正确取值 |
test_foreach_exp_infershape.cpp |
无问题。新用例结构与既有 infer_dtype_test_1 完全同构, 仅输入改为 int16/int8/uint8、断言 DT_FLOAT, 符合整进浮出语义 |
exp_data/gen_data.py |
无问题。整型输入提升 float32 计算, golden 写 float32_golden_*, 占位 float32_input_* 与输出缓冲分配配套; 输入值 -5..4 在 int 与 float32 间精确可表示, 与 kernel 输入一致 |
test_foreach_exp.cpp |
无问题。输出 tensor 改 float32、补上原先遗漏的 GmFree(x1), 比对改 compare_data.py 'float32', 与 golden 文件命名配套 |
test_foreach_expm1_infershape.cpp |
无问题。同 exp infershape, 结构正确 |
expm1_data/gen_data.py |
无问题。整型路径 golden 用 np.expm1 提升 float32, 语义正确 |
test_foreach_expm1.cpp |
无问题。输出改 float32、比对改 float32, 与 exp 侧改动一致, x1 释放原本已有 |
sub_list_data/compare_data.py |
发现问题(P2) — 退出码修复不完整, 见上方 finding |
sub_list_data/gen_data.py |
无问题。int16/int8/uint8 全量程随机, golden 用 int64 精算(含 input2*scale 先算)后 astype 折叠, 与 kernel 的分步回绕在模算术意义下逐位一致 |
test_foreach_sub_list.cpp |
无问题。新增三例与既有 float/int32 用例结构完全一致(标量 x3 为 float 3.0、tiling key 5/7/8、BINARY_LIST opCode 3), 无内存泄漏或越界 |
总体风险判断
本 PR 为纯测试补充, 变更整体质量较高: golden 回绕语义(int64 精算 + astype 折叠)、UT tiling 与生产端 NeedCastCompute 的对齐、exp/expm1 整进浮出的数据/缓冲适配均自洽且可复核, 未发现会误判或引入错误的行为。唯一实质问题在于 sub_list compare_data.py 声明的退出码修复未真正落地——比对失败仍无法通过退出码暴露, 影响该 PR 新增整型用例在 CI 上的失败可发现性, 属中等优先级, 修复成本低(一行 sys.exit + C++ 侧校验)。整体风险可控, 建议修复后合入。
| 类型 | 数量 |
|---|---|
| 🔴 阻塞 | 0 |
| 🟡 建议 | 1 |
💬 仅评论


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| foreach | ✅ 陈昊文, 任如海 (2/2) | ✅ 陈昊文 (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
Tian_1122, thanks for your pull request. All authors of the commits have signed the CLA. 👍


The MR can not be merged, because of CodeReview discussion not resolved
If you want to solve this problem, you can click here to do it in the FAQs.


Pull Request 已合并或已关闭。
If you want to solve this problem, you can click here to do it in the FAQs.


描述
新增整型路径缺少对应 UT 覆盖。本 PR 为 foreach 系列算子(foreach_exp、foreach_expm1、foreach_sub_list、foreach_add_scalar 等)补充整型(int16/int8/uint8)单元测试覆盖:扩展公共 tiling 的 UB 划分以支持整型输入所需的中间缓冲,测试数据生成改为全量程随机以触发溢出回绕路径,并针对 ForeachExp/ForeachExpm1 的"整进浮出"(输出 float32)语义调整 kernel 测试输出与 golden 生成,同时新增 host 侧 infer_datatype 整型转浮点用例及 foreach_sub_list 的整型 kernel 用例。
关联的Issue
#5618
测试
文档更新
类型标签
AI/Agent生成声明