已合并
fix(foreach): 四算子整型语义对齐 PyTorch — AddScalarV2/SubListV2 溢出回绕 + Exp/Expm1 整进浮出 #9890
Tian_1122创建于 29 天前
fix(foreach): 四算子整型语义对齐 PyTorch — AddScalarV2/SubListV2 溢出回绕 + Exp/Expm1 整进浮出 #9890
已合并
Tian_1122创建于 29 天前
Tian_1122
Tian_1122
29 天前

描述

int16/int8/uint8 新增类型下, AddScalarV2/SubListV2 整数溢出饱和钳位(竞品为二进制补码回绕);Exp/Expm1 整进整出+rint取整+饱和钳位(竞品整进浮出fp32+IEEE饱和), 常规取值即有差异(如 uint8 exp(3)=20 vs 20.085)。
修复:

  • AddScalarV2/SubListV2: 参照 !9151(AddListV2) 方案, 新增各自私有 wrap kernel —int16 域直接计算(硬件原生模2^16回绕), int8/uint8 经half提升int16精算后位掩码折叠低8位
  • Exp/Expm1: 新增公共模板 foreach_utils/op_kernel/foreach_int_to_float_unary.h(骨架+函数指针复用),整数输入提升fp32计算后直写float32输出

关联的Issue

关联Issue #5618

测试

库上相关ut、st、atk精度测试均验证通过

文档更新

foreach/foreach_exp/README.md
foreach/foreach_exp/docs/aclnnForeachExp.md
foreach/foreach_expm1/README.md
foreach/foreach_expm1/docs/aclnnForeachExpm1.md

类型标签

AI/Agent生成声明

likedislike
Pull Request已成功合入, 合并人@CANN-robot
(感谢 Tian_1122 的贡献)
Tian_1122Tian_1122
29 天前 创建了 pull request,commit 8d67fda7 1. Use git rebase command to rebase locally 2. Deselect the squash merge option of the current merge request 3. Use a new branch based on the target branch to squash the following is error message while CodeHub try to do squash for you: merge: could not auto-merge due to conflicts
CANN-robotCANN-robot成员
29 天前 添加了label:cann-cla/no
CANN-robotCANN-robot成员
29 天前 添加了label:stat/needs-squash
CANN-robot
CANN-robot成员
29 天前 评论:

Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.


PR Approval Progress

✅ Congratulations! All modules have met the lgtm and approve requirements.

Module Approval Details

module lgtm status approve status
*/*/README.md ✅ 任如海, 查建青, 杜慧萍 (3/2) ✅ 杜慧萍 (1/1)
*/*/docs/acl*.md ✅ 杜慧萍, 任如海, 查建青 (3/2) ✅ 杜慧萍 (1/1)
*/*/op_graph/*_proto.h ✅ 王永光, 任如海, 查建青 (3/2) ✅ 王永光 (1/1)
*/*/op_host/*_def.cpp ✅ 查建青, 王永光, 任如海 (3/2) ✅ 王永光 (1/1)
foreach ✅ 查建青, 任如海 (2/2) ✅ 查建青 (1/1)

💡 Tip:

  • Committer can comment /approve or /lgtm
  • Commenting /approve implies both code review (lgtm) and intent to merge (approve)

CLA Signature Pass

Tian_1122, thanks for your pull request. All authors of the commits have signed the CLA. 👍

likedislike
CANN-robotCANN-robot成员
29 天前 将zengjuan,chaotang233,kevin_huang1234,crystalhu,yangyang016,fanqirui,wangrui_,renruhai,su-yueming,zhajianqing123,tang-lei01,chenqi317,wang-xing001,liubo75,tangweiwei2,chenxingyu18,liujie12345678,pingchuantang,Chen_HaoWen,qianzehong,wangyongguang设为评审人
此处折叠了140条消息 查看更多
Chen_HaoWen成员25 天前进行代码检视2
foreach/foreach_sub_list/op_kernel/foreach_sub_list_wrap.h
@@ -0,0 +25,4 @@
25+constexpr int32_t SUB_LIST_WRAP_BUFFER_NUM = 2;
26+ 
27+template <typename T, int32_t bufferNum = SUB_LIST_WRAP_BUFFER_NUM>
28+class ForeachSubListWrap {
Chen_HaoWen25 天前评论:

新增的 ForeachSubListWrap(int16/int8/uint8 专属路径)目前没有任何 UT 覆盖:tests/ut/op_kernel/test_foreach_sub_list.cpp 只有 float32/float16/int32/bfloat16 用例,sub_list_data/gen_data.py 的 d_type_dict 也不含 int16/int8/uint8。建议参考 foreach_add_list 的用例补充三个 dtype 的测试,golden 可利用 numpy astype 的回绕特性构造(astype 溢出行为与本 kernel 的回绕语义一致)。

likedislike
Tian_1122
Tian_1122
23 天前 评论:

已于PR_10170和PR_10204修复合入

Chen_HaoWen成员25 天前进行代码检视2
foreach/foreach_add_scalar/op_kernel/foreach_add_scalar_wrap.h
@@ -0,0 +174,4 @@
174+ PipeBarrier<PIPE_V>();
175+ if constexpr (std::is_same_v<T, int16_t>) {
176+ // int16 域: x + scalar, 溢出按二进制补码回绕
177+ Adds(outLocal, inLocal, scalarVal, dataCount);
Chen_HaoWen25 天前评论:

当前 add_scalar 的 UT(add_scalar_data/gen_data.py)输入为 np.ones(shape) * 10、scalar=3,结果为 13,不会触发任何溢出——饱和钳位与回绕两种语义在该输入下输出完全相同,本 PR 的核心改动(溢出回绕)实际没有被验证到。建议将输入范围扩大到 dtype 边界(int16 取 ±32768 邻域、int8 取 127/-128 邻域、uint8 取 0/255 邻域),golden 按回绕语义生成。

likedislike
Tian_1122
Tian_1122
23 天前 评论:

已于PR_10170和PR_10204修复合入

Chen_HaoWen成员25 天前进行代码检视2
foreach/foreach_utils/op_host/foreach_infershape.cpp
@@ -74,2 +74,4 @@
7474}
7575 
76+// 整进浮出: 整数输入(int16/int8/uint8)推导输出 DT_FLOAT, 浮点输入输出同 dtype
77+static ge::graphStatus InferDataType4ForeachIntToFloat(gert::InferDataTypeContext* context)
Chen_HaoWen25 天前评论:

新增的 InferDataType4ForeachIntToFloat(int16/int8/uint8 → DT_FLOAT 推导)没有对应的 infershape UT:tests/ut/op_host/test_foreach_exp_infershape.cpp 目前仅覆盖 FP16→FP16。建议补充 int 输入推导 DT_FLOAT 的用例,并考虑覆盖同一 tensor 列表内 int/float 混合 dtype 的场景,foreach_expm1 的 infershape UT 同步补充。

likedislike
Tian_1122
Tian_1122
23 天前 评论:

已于PR_10170和PR_10204修复合入

Chen_HaoWen成员25 天前进行代码检视2
foreach/foreach_add_scalar/op_kernel/foreach_add_scalar_wrap.h
@@ -0,0 +77,4 @@
77+ ParseTilingData(tilingData);
78+ 
79+ inScalarGM.SetGlobalBuffer((__gm__ DTYPE_SCALAR*)scalar, 1);
80+ scalarVal = static_cast<int16_t>(inScalarGM.GetValue(0));
Chen_HaoWen25 天前评论:

生产路径上 int16/int8/uint8 变体的 scalar 为 DT_INT32(binary.json 变体定义 + aclnn 侧 ConvertToTensor(scalar, DataType::DT_INT32)),而 UT 中以 float 位模式写入 scalar(((float)x3) = 3)且 CMake 传入 -DDTYPE_SCALAR=float 编译,与真实部署路径不一致(当前 3.0f 与 int32 的 3 截断后碰巧相同)。建议 UT 的 int 用例按 int32 写入 scalar,对齐真实变体的编译宏与数据。

likedislike
Tian_1122
Tian_1122
23 天前 评论:

已于PR_10170和PR_10204修复合入

Chen_HaoWen成员25 天前进行代码检视2
foreach/foreach_sub_list/op_kernel/foreach_sub_list_wrap.h
@@ -0,0 +199,4 @@
199+ PipeBarrier<PIPE_V>();
200+ if constexpr (std::is_same_v<T, int16_t>) {
201+ // int16 域: x1 - alpha * x2, 下溢按二进制补码回绕
202+ Muls(inLocal2, inLocal2, alphaVal, dataCount);
Chen_HaoWen25 天前评论:

确认下:int16 路径的 Muls 依赖硬件 int16 乘法按模 2^16 回绕(而非饱和),PyTorch 对齐语义是否成立取决于这一点。该假设与已合入的 foreach_add_list_wrap 相同,但建议在 PR 描述或测试记录中补充 NPU 实测佐证(选取 alpha*x2 超出 int16 表示范围的用例,如 alpha=300、x2=127)。

likedislike
Tian_1122
Tian_1122
23 天前 评论:

已于PR_10170和PR_10204修复合入