已关闭
[Performance]: background_matting性能优化 #4601
伦创建于 27 天前关闭于 17 天前
27 天前 添加了label:performance
27 天前 添加了label:triage-review
TorchNPU-Bot
27 天前 评论:
27 天前 评论:
issue待分派,添加triage-review标签


27 天前 关联了pull request:fix(npu): preserve native spatial ops for matting
27 天前 添加了label:bot-triaged;删除了label:triage-review
27 天前 删除了关联的pull request:fix(npu): preserve native spatial ops for matting
TorchNPU-Bot
27 天前 评论:
27 天前 评论:
检测到当前 issue 已关联 PR,自动添加标签:bot-triaged


27 天前 关联了pull request:fix(npu): preserve native spatial ops for matting
27 天前 修改了issue 的描述
27 天前 修改了issue 的描述
27 天前 修改了issue 的描述
17 天前 关闭了 issue
17 天前 issue状态由 TODO 改变为 DONE
17 天前 添加了label:resolved
提交提案之前,请先检索仓库内是否已有相同的提案,如已有请在同一提案中进行讨论。
性能优化具体描述
场景
Background_Mattingtorch.compile(..., backend="inductor", options={"npu_backend": "triton_experimental"})问题现象
模型中的反射填充和双线性插值在 NPU
triton_experimental后端会沿用通用 Inductor decomposition:reflection_pad2d/reflection_pad2d_backward被展开为通用索引计算;upsample_bilinear2d/upsample_bilinear2d_backward被展开为通用插值计算。上述算子已有 NPU 原生实现。通用分解路径会在大尺寸空间特征图上增加索引计算、访存和 kernel 调度开销,未能利用原生算子的实现与性能优势。
优化方案
triton_experimental后端移除上述四个算子的 decomposition,使其保留为 NPU 原生算子。upsample_bilinear2d.vec补充 dispatcher,兼容output_size和scale_factors,并转发到原生default重载。implicit_fallbacks=False时注册四个 native fallback,避免因算子没有上游 lowering 而编译失败。对应修复 PR:Ascend/pytorch#44957
验收标准
triton_experimental生成代码中走 NPU 原生调用,不再展开为通用 decomposition;F.pad(mode="reflect")和F.interpolate(mode="bilinear")的前反向结果与 eager 一致;F.interpolate的size、scale_factor两种入口均通过;implicit_fallbacks=False下可正常编译、执行;性能劣化说明
这是算子执行路径层面的性能劣化:单个原生空间算子被展开为多段通用索引/逐点计算,在
Background_Matting的大尺寸特征图上放大了额外计算、HBM 访存和 kernel 调度成本,前向与反向均受影响。当前没有同一软件栈、同一输入下可归因于这四个算子的稳定端到端 A/B 数据,因此不填写未经验证的加速数字。本次修复通过生成代码检查确认执行路径由通用 decomposition 切换为 NPU 原生算子,并以定向前反向测试保证功能正确性;端到端收益以后续统一性能流水线结果为准。
其他相关讨论
reflection_pad2d、upsample_bilinear2d及其 backward 的 NPUtriton_experimental执行路径。d4b2b35a0339a9ec35f49a0e9a4977c5b0d2be95。Ran 4 tests in 2.971s - OK。环境信息
欢迎加入社区,感谢您对社区的贡献 🎉!