已合并
【PR】: 新增gather+add融合用例 #1529
【PR】: 新增gather+add融合用例 #1529
已合并
zzq创建于 7月28日
共 5 个文件变更+166-28
@@ -4,12 +4,15 @@
4 4 
5使用 `torch.compile` 完成 PyTorch 网络下的算子融合。5使用 `torch.compile` 完成 PyTorch 网络下的算子融合。
6 6 
7-当前包含以下两个用例:7+当前包含以下三个用例:
8 8 
9- `add + ge`:将加法和比较算子融合为一个算子;9- `add + ge`:将加法和比较算子融合为一个算子;
10-- `mul + reducesum`:将乘法和求和归约算子融合为一个算子。10+- `mul + reducesum`:将乘法和求和归约算子融合为一个算子;
11+- `gather + add`:构造索引取数和逐元素加法图模式
11 12 
12-两个用例均开启 NPU Profiling,可通过生成的性能分析文件查看融合后的 Kernel。13+注:当前暂不支持gather融合能力,等待[ issue175 ](https://gitcode.com/cann/graph-autofusion/issues/175)这个issue完成后gather可以和add进行融合。
14+ 
15+三个用例均开启 NPU Profiling,可通过生成的性能分析文件查看融合后的 Kernel。
13 16 
14## 目录结构17## 目录结构
15 18 
@@ -21,10 +24,14 @@ pytorch
21│ ├── README.md24│ ├── README.md
22│ ├── README_en.md25│ ├── README_en.md
23│ └── af_add_ge.py # 融合 add + ge26│ └── af_add_ge.py # 融合 add + ge
24-└── af_reduce27+├── af_reduce
28+│ ├── README.md
29+│ ├── README_en.md
30+│ └── af_mul_reducesum.py # 融合 mul + reducesum
31+└── af_gather
25 ├── README.md32 ├── README.md
26 ├── README_en.md33 ├── README_en.md
27- └── af_mul_reducesum.py # 融合 mul + reducesum34+ └── af_gather_add.py # gather + add 图模式
28```35```
29 36 
30## 前置说明37## 前置说明
@@ -67,6 +74,13 @@ cd af_reduce
67python af_mul_reducesum.py74python af_mul_reducesum.py
68```75```
69 76 
77+### gather + add 图模式
78+ 
79+```bash
80+cd af_gather
81+python af_gather_add.py
82+```
83+ 
70## 预期执行结果84## 预期执行结果
71 85 
72程序执行完成后,当前目录下会生成 `profiling` 目录。86程序执行完成后,当前目录下会生成 `profiling` 目录。
@@ -89,3 +103,4 @@ op_summary_时间戳.csv
89 103 
90- [Autofuse 简介与快速上手](../../README.md)104- [Autofuse 简介与快速上手](../../README.md)
91- [Profiling 性能分析工具指南](https://hiascend.com/document/redirect/CannCommunityToolProfiling)105- [Profiling 性能分析工具指南](https://hiascend.com/document/redirect/CannCommunityToolProfiling)
106+- [精度调试工具指南](https://hiascend.com/document/redirect/CannCommunityToolAccucacy)
@@ -1,15 +1,18 @@
1-# PyTorch Inductor + AscendC Example Demonstration1+# PyTorch Inductor Examples
2 2 
3## Description3## Description
4 4 
5-Use the AscendC backend of `torch.compile` to perform operator fusion for PyTorch networks.5+These examples demonstrate how to use `torch.compile` to perform operator fusion in PyTorch models.
6 6 
7-The following two examples are currently included:7+The following three examples are currently provided:
8 8 
9-- `add + ge`: fuses the addition and comparison operators into a single operator;9+* `add + ge`: Fuses the addition and comparison operators into a single operator.
10-- `mul + reducesum`: fuses the multiplication and sum-reduction operators into a single operator.10+* `mul + reducesum`: Fuses the multiplication and sum-reduction operators into a single operator.
11+* `gather + add`: Constructs a graph pattern containing index gathering and element-wise addition.
11 12 
12-NPU Profiling is enabled for both examples. The generated profiling files can be used to view the fused Kernel.13+> **Note:** Gather fusion is not currently supported. After [issue175](https://gitcode.com/cann/graph-autofusion/issues/175) is resolved, `gather` will be able to fuse with `add`.
14+ 
15+NPU Profiling is enabled in all three examples. You can inspect the generated profiling data to view the fused kernels.
13 16 
14## Directory Structure17## Directory Structure
15 18 
@@ -20,35 +23,40 @@ pytorch
20├── af_pointwise23├── af_pointwise
21│ ├── README.md24│ ├── README.md
22│ ├── README_en.md25│ ├── README_en.md
23-│ └── af_add_ge.py # Fuse add + ge26+│ └── af_add_ge.py # Fuses add + ge
24-└── af_reduce27+├── af_reduce
28+│ ├── README.md
29+│ ├── README_en.md
30+│ └── af_mul_reducesum.py # Fuses mul + reducesum
31+└── af_gather
25 ├── README.md32 ├── README.md
26 ├── README_en.md33 ├── README_en.md
27- └── af_mul_reducesum.py # Fuse mul + reducesum34+ └── af_gather_add.py # Constructs the gather + add graph pattern
28```35```
29 36 
30## Prerequisites37## Prerequisites
31 38 
32-Before running the examples, carefully read the [PyTorch Environment Installation Guide](../../../docs/env_install/pytorch/env_pytorch.md) and complete the following steps:39+Before running these examples, carefully read the [PyTorch Environment Installation Guide](../../../docs/env_install/pytorch/env_pytorch.md) and complete the following steps:
33 40 
34-1. CANN version `9.0.0` or later is required. Install the Toolkit and OPS packages correctly through [CANN Quick Installation](https://www.hiascend.com/cann/download?versionId=745&ids=d802%2Ch0501%2Ch0602%2Ch0701). For details, see the [Installation Guide](../../../docs/zh/quick_install.md).41+1. Ensure that the CANN package version is `9.0.0` or later. Install the toolkit and ops packages correctly by using [CANN Quick Installation](https://www.hiascend.com/cann/download?versionId=745&ids=d802%2Ch0501%2Ch0602%2Ch0701). For more information, see the [Installation Guide](../../../docs/zh/quick_install.md).
35-2. `torch_npu` version `2.9.0` or later is required. You can use the [Quick Environment Installation Script](../../../scripts/env_install/pytorch/setup_torch_npu_daily.sh) to quickly install the Python environment and `torch_npu`.42+ 
43+2. Ensure that the `torch_npu` version is `2.9.0` or later. You can use the [environment quick installation script](../../../scripts/env_install/pytorch/setup_torch_npu_daily.sh) to quickly install the Python environment and `torch_npu`.
36 44 
37## Setting Environment Variables45## Setting Environment Variables
38 46 
39-Run the following commands each time you open a new terminal:47+Run the following commands whenever you open a new terminal:
40 48 
41```bash49```bash
42-# Activate the environment.50+# Activate the Python environment.
43source /mnt/workspace/env/venv/torch210_daily/bin/activate51source /mnt/workspace/env/venv/torch210_daily/bin/activate
44 52 
45-# Set the CANN installation path according to the actual installation location.53+# Set the CANN installation path based on the actual installation location.
46export CANN_INSTALL_PATH=/home/developer/Ascend54export CANN_INSTALL_PATH=/home/developer/Ascend
47 55 
48# Load CANN environment variables.56# Load CANN environment variables.
49source $CANN_INSTALL_PATH/cann/set_env.sh57source $CANN_INSTALL_PATH/cann/set_env.sh
50 58 
51-# Assume the example runs on device 0.59+# Assume that the examples run on device 0.
52export ASCEND_DEVICE_ID=060export ASCEND_DEVICE_ID=0
53```61```
54 62 
@@ -68,25 +76,33 @@ cd af_reduce
68python af_mul_reducesum.py76python af_mul_reducesum.py
69```77```
70 78 
79+### gather + add Graph Pattern
80+ 
81+```bash
82+cd af_gather
83+python af_gather_add.py
84+```
85+ 
71## Expected Results86## Expected Results
72 87 
73-After the program finishes running, a `profiling` directory is generated in the current directory.88+After an example finishes running, a `profiling` directory is generated in the current directory.
74 89 
75-Operator execution details can be viewed in the following directory:90+You can view operator execution details in the following directory:
76 91 
77```text92```text
78-profiling/PROF_timestamp/mindstudio_profiler_output93+profiling/PROF_<timestamp>/mindstudio_profiler_output
79```94```
80 95 
81Open the following file:96Open the following file:
82 97 
83```text98```text
84-op_summary_timestamp.csv99+op_summary_<timestamp>.csv
85```100```
86 101 
87-If the operator list contains a Kernel whose name starts with `autofused_`, the related operators have been successfully fused into a single fused operator.102+If the operator list contains a kernel whose name starts with `autofused_`, the related operators have been successfully fused into a single fused operator.
88 103 
89## References104## References
90 105 
91-- [Autofuse Introduction and Quick Start](../../README.md)106+* [Autofuse Overview and Quick Start](../../README.md)
92-- [Profiling Performance Analysis Tool Guide](https://hiascend.com/document/redirect/CannCommunityToolProfiling)107+* [Profiling Performance Analysis Tool Guide](https://hiascend.com/document/redirect/CannCommunityToolProfiling)
108+* [Accuracy Debugging Tool Guide](https://hiascend.com/document/redirect/CannCommunityToolAccucacy)
@@ -0,0 +1,15 @@
1+# autofuse 用例演示
2+ 
3+## 用例功能:
4+ 
5+本用例构造 gather + add 图模式:先通过 `torch.gather(x, 1, indices)` 沿第 1 轴取数,再将取数结果与同形状张量 `y` 逐元素相加。其中 add 属于 elewise 计算。
6+ 
7+## 执行命令
8+ 
9+```bash
10+python3 af_gather_add.py
11+```
12+ 
13+## 预期执行结果
14+ 
15+当前目录下的 profiling 目录下,有生成的 profiling 文件。其中 PROF_000001_时间戳xx/mindstudio_profiler_output 下面,可以打开 op_summary_时间戳xx.csv 文件,查看算子执行详情。若当前 PyTorch、torch_npu 和 AscendC 后端支持 Gather 前端转换,取数与加法应表现为一个名称以 autofused_ 开头的 kernel,且不再出现独立的 GatherElementsV2 和独立的 add;若仍出现独立 GatherElementsV2,则说明 Gather 已回退到 ACLNN,本用例在当前环境未发生融合。
@@ -0,0 +1,15 @@
1+# Autofuse Use Case Demonstration
2+ 
3+## Use Case Function:
4+ 
5+This example constructs a gather + add graph pattern. It first selects values along dimension 1 through `torch.gather(x, 1, indices)`, and then performs an element-wise addition with tensor `y` of the same shape. The add operation is an elewise computation.
6+ 
7+## Execution Command
8+ 
9+```bash
10+python3 af_gather_add.py
11+```
12+ 
13+## Expected Execution Result
14+ 
15+Profiling files are generated in the profiling directory under the current directory. Open op_summary_timestampxx.csv under PROF_000001_timestampxx/mindstudio_profiler_output to view operator execution details. If the current PyTorch, torch_npu, and AscendC backend versions support Gather frontend conversion, the indexed selection and addition should appear as one kernel whose name starts with autofused_, without a standalone GatherElementsV2 or add. If a standalone GatherElementsV2 is still present, Gather has fallen back to ACLNN and this pattern is not fused in the current environment.
@@ -0,0 +1,77 @@
1+#!/usr/bin/env python3
2+# -*- coding: utf-8 -*-
3+# ----------------------------------------------------------------------------------------------------------------------
4+# Copyright (c) 2026 Huawei Technologies Co., Ltd.
5+# This program is free software, you can redistribute it and/or modify it under the terms and conditions of
6+# CANN Open Software License Agreement Version 2.0 (the "License").
7+# Please refer to the License for details. You may not use this file except in compliance with the License.
8+# THIS SOFTWARE IS PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED,
9+# INCLUDING BUT NOT LIMITED TO NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE.
10+# See LICENSE in the root of the software repository for the full text of the License.
11+# ----------------------------------------------------------------------------------------------------------------------
12+# 导包
13+import torch
14+import torch_npu
15+import torch.nn as nn
16+ 
17+# ===== 1. 昇腾 NPU 配置 =====
18+DEVICE = "npu:0" # 假设使用0卡
19+torch.npu.set_device(DEVICE)
20+ 
21+ 
22+# ===== 2. 构造简单模型 =====
23+class MyModel(nn.Module):
24+ def __init__(self):
25+ super().__init__()
26+ 
27+ def forward(self, x, indices, y):
28+ result = torch.add(torch.gather(x, 1, indices), y)
29+ return result
30+ 
31+ 
32+# ===== 3. inductor + 昇腾NPU自动融合后端 =====
33+model = MyModel().to(DEVICE)
34+model = torch.compile(
35+ model,
36+ dynamic=False,
37+ fullgraph=True,
38+ options={"npu_backend": "ascendc"},
39+)
40+ 
41+# ===== 4. 创建输入 =====
42+x = torch.randn(128, 50, device=DEVICE)
43+indices = torch.randint(0, 50, (128, 25), dtype=torch.int64, device=DEVICE)
44+y = torch.randn(128, 25, device=DEVICE)
45+ 
46+# ===== 5. 执行 =====
47+model.eval()
48+ 
49+# 开启 profiling
50+experimental_config = torch_npu.profiler._ExperimentalConfig(
51+ export_type=[torch_npu.profiler.ExportType.Text],
52+ profiler_level=torch_npu.profiler.ProfilerLevel.Level2,
53+ msprof_tx=False,
54+ aic_metrics=torch_npu.profiler.AiCMetrics.PipeUtilization,
55+ l2_cache=False,
56+ op_attr=False,
57+ data_simplification=False,
58+ record_op_args=False,
59+ gc_detect_threshold=None,
60+)
61+ 
62+with torch_npu.profiler.profile(
63+ activities=[
64+ torch_npu.profiler.ProfilerActivity.CPU,
65+ torch_npu.profiler.ProfilerActivity.NPU,
66+ ],
67+ on_trace_ready=torch_npu.profiler.tensorboard_trace_handler("./profiling"),
68+ record_shapes=True,
69+ profile_memory=False,
70+ with_stack=False,
71+ with_modules=False,
72+ with_flops=False,
73+ experimental_config=experimental_config,
74+) as prof:
75+ # 跑 100 step
76+ for _ in range(100):
77+ result = model(x, indices, y)