Pull Request已成功合入, 合并人@ascend-robot
(感谢 Erwinnn 的贡献)变更摘要
本 PR 主要聚焦 NPU Fast Launch 计划化提交的最终定型优化:在 plan promotion 成功后,将生成 wrapper 的调用槽位从 BoundFastLaunch 替换为新的 FinalizedFastLaunch,使稳态调用不再重复经过 BoundFastLaunch.__call__、_canonical_args 等参数规范化路径;同时 C++ 侧在 MakeFastLaunchPlan 阶段将固定参数(fixed args)与静态 grid 预写入 packedArgsTemplate,并让 SubmitLaunch 复用 at_npu::native::OpCommand::RunOpApiV2 的 callable 入队接口,避免为每次 planned hit 构造通用 OpCommand、导出 ExecuteParas 和创建 COMPILE_AND_EXECUTE payload。PR 描述中的实测收益为 fast launch 相比 base 约 2.4 ms,优化后进一步增量约 0.76 ms。
主要改动
SubmitLaunch复用RunOpApiV2入队接口:csrc/bindings.cpp中将提交入口从构造通用零 IO 的OpCommand改为at_npu::native::OpCommand::RunOpApiV2(plan.kernelName, launchCall),kernel handle、packed args、Block Num、stream 等均已在进入SubmitLaunch()前准备完毕,避免重复构造 payload。- 新增
FinalizedFastLaunch稳态调用入口:bind.py中新增该类并安装到生成的 call slot,稳态调用直接执行self.direct(args, stream=stream);在benchmark_run、profiler 开启、best launcher 变更或存在 launch hooks 时回退到BoundFastLaunch,出错时按backend_submitted/stable走负向安装或_fallback。 - C++ plan 预打包固定参数与静态 grid:
FastLaunchPlan新增runtimeArgCount、packedArgsTemplate、staticBlockNum、hasStaticGrid字段;MakeFastLaunchPlan增加runtime_arg_count/fixed_args/static_grid入参,将固定参数与 FFTS 地址预写入模板(固定 tensor 参数被拒绝),并新增_npu_inductor_fast_launch_static_with_plan绑定与PackStaticLaunch,静态 grid 场景跳过 grid 解析直接使用预计算的staticBlockNum。 PlannedFastLaunch拆分运行期与固定参数:backend.py中PlannedFastLaunch新增runtime_arg_count、fixed_args、static_grid、untimed_static_launch字段,__call__按runtime_arg_count校验参数数量,动态 grid 时把fixed_args一并传给get_grid;_constant_grid通过ast.literal_eval解析 launcher 的_npu_fast_launch_grid_exprs识别静态 grid。- 导出
FinalizedFastLaunch:__init__.py在__getattr__与__all__中补充导出FinalizedFastLaunch,并同步更新bind.py的__all__。


Thanks for your pull-request.
The full list of commands accepted by me can be found at here.
You can get sig-info at here.
You can self-configure the PR merge rules for this repository. For more details, please refer to Here.
PR Approval Progress
✅ Congratulations! All modules have met the lgtm and approve requirements.
Module Approval Details
| module | lgtm status | approve status |
|---|---|---|
| torch_npu/_inductor | ✅ TonyYA, crazyDannyBoy (2/2) | ✅ crazyDannyBoy (1/1) |
💡 Tip:
- Committer can comment
/approveor/lgtm- Commenting
/approveimplies both code review (lgtm) and intent to merge (approve)
CLA Signature Pass
jiabiao_o, thanks for your pull request. All authors of the commits have signed the CLA. 👍


当前仓库存在以下 保护分支 :
| Protected Branch | Version | Release |
|---|---|---|
| master | ||
| v2.9.0 | ||
| v2.10.0 | ||
| v2.11.0 | ||
| v2.12.0 | ||
| v2.7.1 | ||
| v2.7.1-26.1.0 | ||
| v2.9.0-26.1.0 | ||
| v2.12.0-26.1.0 | ||
| v2.11.0-26.1.0 | ||
| v2.10.0-26.1.0 | ||
| ci-test |
评论 /sync <branch1> <branch2> ... 可将当前 PR 修改同步到其它分支(创建同步 PR):
a) 如果当前 PR 是 Open 状态,同步操作将延迟到 PR 被合并时执行
b) 如果当前 PR 已经 Merged,将立即执行同步操作
注意:
- /sync 命令可以指定同步到多个分支,仅最后一个 /sync 命令生效
- 如果创建的同步 PR 不正确,可通过向同步 PR 的源分支提交轻量级 PR 完善,或使用 /close 命令关闭


ascend docs pipeline is running...


✅ 跳过 docs ci 检查,没有需要检查的文档文件


| 阶段 | 任务名 | 状态 | 详情 |
|---|---|---|---|
| 编译构建 | Build_X86 | ✅ | >>> |
| Build_ARM | ✅ | >>> | |
| Build_X86_torchair | 🛑 | >>> | |
| Build_ARM_torchair | 🛑 | >>> | |
| patch_test | 🛑 | >>> | |
| Build_X86_213 | ✅ | >>> | |
| Build_ARM_213 | ✅ | >>> | |
| 恶意代码检查 | Antipoison | ✅ | >>> |
| 编码安全与规范检查 | codecheck_pre-commit | ✅ | >>> |
| check_error | ✅ | >>> | |
| lintrunner | ✅ | >>> | |
| 开源片段检查 | SCA | ✅ | >>> |
| 开发者测试 | UT_ARM_A3_Part_01 | 🛑 | >>> |
| UT_ARM_A3_Part_02 | 🛑 | >>> | |
| UT_ARM_A2_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_03 | ✅ | >>> | |
| UT_inductor_Part_01 | 🛑 | >>> | |
| UT_inductor_Part_02 | 🛑 | >>> | |
| UT_inductor_Part_03 | 🛑 | >>> | |
| UT_inductor_Part_04 | 🛑 | >>> | |
| UT_DIST_ARM_Part_01 | 🛑 | >>> | |
| UT_DIST_ARM_Part_02 | 🛑 | >>> | |
| UT_DIST_ARM_Part_03 | 🛑 | >>> | |
| UT_DIST_ARM_Part_04 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_01 | ✅ | >>> | |
| UT_ARM_A2_Select_Part_02 | ✅ | >>> | |
| UT_ARM_A2_Part_213 | ✅ | >>> | |
| UT_inductor_Part_213 | 🛑 | >>> | |
| UT_DIST_ARM_Part_213 | 🛑 | >>> | |
| UT_ARM_A2_Select_Part_213 | ✅ | >>> | |
| 流水线 | PR-pipeline_pytorch | ✅ | >>> |
- compile、compile_inductor、compile_torchair : 运行流水线
- retry : 重试流水线所有失败子任务
- retry <任务名> : 仅重试指定失败子任务
- stop : 停止流水线


/lgtm


/lgtm
/approve




【合入来源】
【修改方案】
SubmitLaunch()时已经准备好:因此将通用opcommand入口改为OpCommand::RunOpApiV2(plan.kernelName, launchCall);复用仓库已有的 callable 入队接口,避免为每次 planned hit 构造通用 OpCommand、导出 ExecuteParas 和创建 COMPILE_AND_EXECUTE payload。
BoundFastLaunch替换为:
FinalizedFastLaunch稳态调用不再重复经过:
【资料变更】
【接口变更】
【功能验证】
free time实测收益如下:
base:13.974 ms
fastlaunch:11.574 ms (收益:2.4 ms)
fastlaunch优化后:19.814 ms (收益:3.16ms 增量0.76 ms)
【CheckList】
【issues】
https://gitcode.com/Ascend/pytorch/issues/4345