Pull Request已成功合入, 合并人@CANN-robot
(感谢 Tian_1122 的贡献)Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
For more, you also can visit HICANN.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| */*/README.md | ✅ 任如海, 查建青, 杜慧萍 (3/2) | ✅ 杜慧萍 (1/1) |
| */*/docs/acl*.md | ✅ 杜慧萍, 任如海, 查建青 (3/2) | ✅ 杜慧萍 (1/1) |
| */*/op_graph/*_proto.h | ✅ 王永光, 任如海, 查建青 (3/2) | ✅ 王永光 (1/1) |
| */*/op_host/*_def.cpp | ✅ 查建青, 王永光, 任如海 (3/2) | ✅ 王永光 (1/1) |
| foreach | ✅ 查建青, 任如海 (2/2) | ✅ 查建青 (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
Tian_1122, thanks for your pull request. All authors of the commits have signed the CLA. 👍


新增的 ForeachSubListWrap(int16/int8/uint8 专属路径)目前没有任何 UT 覆盖:tests/ut/op_kernel/test_foreach_sub_list.cpp 只有 float32/float16/int32/bfloat16 用例,sub_list_data/gen_data.py 的 d_type_dict 也不含 int16/int8/uint8。建议参考 foreach_add_list 的用例补充三个 dtype 的测试,golden 可利用 numpy astype 的回绕特性构造(astype 溢出行为与本 kernel 的回绕语义一致)。


当前 add_scalar 的 UT(add_scalar_data/gen_data.py)输入为 np.ones(shape) * 10、scalar=3,结果为 13,不会触发任何溢出——饱和钳位与回绕两种语义在该输入下输出完全相同,本 PR 的核心改动(溢出回绕)实际没有被验证到。建议将输入范围扩大到 dtype 边界(int16 取 ±32768 邻域、int8 取 127/-128 邻域、uint8 取 0/255 邻域),golden 按回绕语义生成。


新增的 InferDataType4ForeachIntToFloat(int16/int8/uint8 → DT_FLOAT 推导)没有对应的 infershape UT:tests/ut/op_host/test_foreach_exp_infershape.cpp 目前仅覆盖 FP16→FP16。建议补充 int 输入推导 DT_FLOAT 的用例,并考虑覆盖同一 tensor 列表内 int/float 混合 dtype 的场景,foreach_expm1 的 infershape UT 同步补充。


生产路径上 int16/int8/uint8 变体的 scalar 为 DT_INT32(binary.json 变体定义 + aclnn 侧 ConvertToTensor(scalar, DataType::DT_INT32)),而 UT 中以 float 位模式写入 scalar(((float)x3) = 3)且 CMake 传入 -DDTYPE_SCALAR=float 编译,与真实部署路径不一致(当前 3.0f 与 int32 的 3 截断后碰巧相同)。建议 UT 的 int 用例按 int32 写入 scalar,对齐真实变体的编译宏与数据。


确认下:int16 路径的 Muls 依赖硬件 int16 乘法按模 2^16 回绕(而非饱和),PyTorch 对齐语义是否成立取决于这一点。该假设与已合入的 foreach_add_list_wrap 相同,但建议在 PR 描述或测试记录中补充 NPU 实测佐证(选取 alpha*x2 超出 int16 表示范围的用例,如 alpha=300、x2=127)。


描述
int16/int8/uint8 新增类型下, AddScalarV2/SubListV2 整数溢出饱和钳位(竞品为二进制补码回绕);Exp/Expm1 整进整出+rint取整+饱和钳位(竞品整进浮出fp32+IEEE饱和), 常规取值即有差异(如 uint8 exp(3)=20 vs 20.085)。
修复:
关联的Issue
关联Issue #5618
测试
库上相关ut、st、atk精度测试均验证通过
文档更新
foreach/foreach_exp/README.md
foreach/foreach_exp/docs/aclnnForeachExp.md
foreach/foreach_expm1/README.md
foreach/foreach_expm1/docs/aclnnForeachExpm1.md
类型标签
AI/Agent生成声明