已合并
fix: support quantized NPU flip dispatch #36070
hz893创建于 5月19日
fix: support quantized NPU flip dispatch #36070
已合并
Pull Request已成功合入, 合并人@ascend-robot
(感谢 hz893 的贡献)ascend-robot
5月19日 评论:
5月19日 评论:
Thanks for your pull-request.
The full list of commands accepted by me can be found at here。
You can get sig-info at here
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| repo-Ascend/pytorch | ✅ chujinjin, hbhu_bin (2/2) | ✅ chujinjin (1/1) |
| test | ✅ chujinjin, hbhu_bin (2/2) | ✅ chujinjin (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
hz893, thanks for your pull request. All authors of the commits have signed the CLA. 👍


5月19日 添加了label:ascend-cla/yes
ascend-robot
5月19日 评论:
5月19日 评论:
当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.11.0 | ||
| v2.10.0 | ||
| v2.9.0 | ||
| v2.7.1 | ||
| v2.12.0 |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


ascend-robot
5月19日 评论:
5月19日 评论:
Ascend docs pipeline is running...


此处折叠了44条消息 查看更多
6月5日 添加了label:approvedlgtm
6月5日 合入了pull request
ascend-robot
6月5日 评论:
6月5日 评论:
流水线 pytorch_gitcode_PR_multiVersion#10031 [ commitID:65d694e7 ] 已完成


【合入来源】
【修改方案】
QuantizedPrivateUse1的 codegen 注册中补充aten::flip,使量化 NPU tensor 能命中 torch_npu 的 quantized helper。quantized_fliphelper:per-tensor 量化场景对int_repr()调用普通 NPUaten::flip,复用现有op_plugin::flip -> aclnnFlip数据翻转路径,再用原 scale/zero_point 重建 affine quantized tensor。Setting strides is possible only on uniformly quantized tensor报错。【资料变更】
不涉及。
【接口变更】
不涉及。
【功能验证】
bash ci/build.sh --python=3.11,编译成功并生成 wheel。python -m pytest test_shape_ops.py -v -k test_flip_npu_float32,结果:1 passed。python -m pytest --import-mode=importlib test/test_shape_ops.py -v -k "test_flip_per_channel_quantized_error or test_flip_npu_float32",结果:2 passed。flipbackward 与 CPU 在dims=(0,)、(1,)、(0, 1)、()下 forward/grad 均一致;验证 quantized flip 的 autograd 状态和错误行为与 CPU 一致。git diff --check,无异常。【CheckList】