已合并
【PR】: 新增gather+add融合用例 #1529
zzq创建于 7月28日
【PR】: 新增gather+add融合用例 #1529
已合并
共 5 个文件变更+166-28
| @@ -4,12 +4,15 @@ | |||
| 4 | 4 | ||
| 5 | 使用 `torch.compile` 完成 PyTorch 网络下的算子融合。 | 5 | 使用 `torch.compile` 完成 PyTorch 网络下的算子融合。 |
| 6 | 6 | ||
| 7 | -当前包含以下两个用例: | 7 | +当前包含以下三个用例: |
| 8 | 8 | ||
| 9 | - `add + ge`:将加法和比较算子融合为一个算子; | 9 | - `add + ge`:将加法和比较算子融合为一个算子; |
| 10 | -- `mul + reducesum`:将乘法和求和归约算子融合为一个算子。 | 10 | +- `mul + reducesum`:将乘法和求和归约算子融合为一个算子; |
| 11 | +- `gather + add`:构造索引取数和逐元素加法图模式 | ||
| 11 | 12 | ||
| 12 | -两个用例均开启 NPU Profiling,可通过生成的性能分析文件查看融合后的 Kernel。 | 13 | +注:当前暂不支持gather融合能力,等待[ issue175 ](https://gitcode.com/cann/graph-autofusion/issues/175)这个issue完成后gather可以和add进行融合。 |
| 14 | + | ||
| 15 | +三个用例均开启 NPU Profiling,可通过生成的性能分析文件查看融合后的 Kernel。 | ||
| 13 | 16 | ||
| 14 | ## 目录结构 | 17 | ## 目录结构 |
| 15 | 18 | ||
| @@ -21,10 +24,14 @@ pytorch | |||
| 21 | │ ├── README.md | 24 | │ ├── README.md |
| 22 | │ ├── README_en.md | 25 | │ ├── README_en.md |
| 23 | │ └── af_add_ge.py # 融合 add + ge | 26 | │ └── af_add_ge.py # 融合 add + ge |
| 24 | -└── af_reduce | 27 | +├── af_reduce |
| 28 | +│ ├── README.md | ||
| 29 | +│ ├── README_en.md | ||
| 30 | +│ └── af_mul_reducesum.py # 融合 mul + reducesum | ||
| 31 | +└── af_gather | ||
| 25 | ├── README.md | 32 | ├── README.md |
| 26 | ├── README_en.md | 33 | ├── README_en.md |
| 27 | - └── af_mul_reducesum.py # 融合 mul + reducesum | 34 | + └── af_gather_add.py # gather + add 图模式 |
| 28 | ``` | 35 | ``` |
| 29 | 36 | ||
| 30 | ## 前置说明 | 37 | ## 前置说明 |
| @@ -67,6 +74,13 @@ cd af_reduce | |||
| 67 | python af_mul_reducesum.py | 74 | python af_mul_reducesum.py |
| 68 | ``` | 75 | ``` |
| 69 | 76 | ||
| 77 | +### gather + add 图模式 | ||
| 78 | + | ||
| 79 | +```bash | ||
| 80 | +cd af_gather | ||
| 81 | +python af_gather_add.py | ||
| 82 | +``` | ||
| 83 | + | ||
| 70 | ## 预期执行结果 | 84 | ## 预期执行结果 |
| 71 | 85 | ||
| 72 | 程序执行完成后,当前目录下会生成 `profiling` 目录。 | 86 | 程序执行完成后,当前目录下会生成 `profiling` 目录。 |
| @@ -89,3 +103,4 @@ op_summary_时间戳.csv | |||
| 89 | 103 | ||
| 90 | - [Autofuse 简介与快速上手](../../README.md) | 104 | - [Autofuse 简介与快速上手](../../README.md) |
| 91 | - [Profiling 性能分析工具指南](https://hiascend.com/document/redirect/CannCommunityToolProfiling) | 105 | - [Profiling 性能分析工具指南](https://hiascend.com/document/redirect/CannCommunityToolProfiling) |
| 106 | +- [精度调试工具指南](https://hiascend.com/document/redirect/CannCommunityToolAccucacy) | ||
| @@ -1,15 +1,18 @@ | |||
| 1 | -# PyTorch Inductor + AscendC Example Demonstration | 1 | +# PyTorch Inductor Examples |
| 2 | 2 | ||
| 3 | ## Description | 3 | ## Description |
| 4 | 4 | ||
| 5 | -Use the AscendC backend of `torch.compile` to perform operator fusion for PyTorch networks. | 5 | +These examples demonstrate how to use `torch.compile` to perform operator fusion in PyTorch models. |
| 6 | 6 | ||
| 7 | -The following two examples are currently included: | 7 | +The following three examples are currently provided: |
| 8 | 8 | ||
| 9 | -- `add + ge`: fuses the addition and comparison operators into a single operator; | 9 | +* `add + ge`: Fuses the addition and comparison operators into a single operator. |
| 10 | -- `mul + reducesum`: fuses the multiplication and sum-reduction operators into a single operator. | 10 | +* `mul + reducesum`: Fuses the multiplication and sum-reduction operators into a single operator. |
| 11 | +* `gather + add`: Constructs a graph pattern containing index gathering and element-wise addition. | ||
| 11 | 12 | ||
| 12 | -NPU Profiling is enabled for both examples. The generated profiling files can be used to view the fused Kernel. | 13 | +> **Note:** Gather fusion is not currently supported. After [issue175](https://gitcode.com/cann/graph-autofusion/issues/175) is resolved, `gather` will be able to fuse with `add`. |
| 14 | + | ||
| 15 | +NPU Profiling is enabled in all three examples. You can inspect the generated profiling data to view the fused kernels. | ||
| 13 | 16 | ||
| 14 | ## Directory Structure | 17 | ## Directory Structure |
| 15 | 18 | ||
| @@ -20,35 +23,40 @@ pytorch | |||
| 20 | ├── af_pointwise | 23 | ├── af_pointwise |
| 21 | │ ├── README.md | 24 | │ ├── README.md |
| 22 | │ ├── README_en.md | 25 | │ ├── README_en.md |
| 23 | -│ └── af_add_ge.py # Fuse add + ge | 26 | +│ └── af_add_ge.py # Fuses add + ge |
| 24 | -└── af_reduce | 27 | +├── af_reduce |
| 28 | +│ ├── README.md | ||
| 29 | +│ ├── README_en.md | ||
| 30 | +│ └── af_mul_reducesum.py # Fuses mul + reducesum | ||
| 31 | +└── af_gather | ||
| 25 | ├── README.md | 32 | ├── README.md |
| 26 | ├── README_en.md | 33 | ├── README_en.md |
| 27 | - └── af_mul_reducesum.py # Fuse mul + reducesum | 34 | + └── af_gather_add.py # Constructs the gather + add graph pattern |
| 28 | ``` | 35 | ``` |
| 29 | 36 | ||
| 30 | ## Prerequisites | 37 | ## Prerequisites |
| 31 | 38 | ||
| 32 | -Before running the examples, carefully read the [PyTorch Environment Installation Guide](../../../docs/env_install/pytorch/env_pytorch.md) and complete the following steps: | 39 | +Before running these examples, carefully read the [PyTorch Environment Installation Guide](../../../docs/env_install/pytorch/env_pytorch.md) and complete the following steps: |
| 33 | 40 | ||
| 34 | -1. CANN version `9.0.0` or later is required. Install the Toolkit and OPS packages correctly through [CANN Quick Installation](https://www.hiascend.com/cann/download?versionId=745&ids=d802%2Ch0501%2Ch0602%2Ch0701). For details, see the [Installation Guide](../../../docs/zh/quick_install.md). | 41 | +1. Ensure that the CANN package version is `9.0.0` or later. Install the toolkit and ops packages correctly by using [CANN Quick Installation](https://www.hiascend.com/cann/download?versionId=745&ids=d802%2Ch0501%2Ch0602%2Ch0701). For more information, see the [Installation Guide](../../../docs/zh/quick_install.md). |
| 35 | -2. `torch_npu` version `2.9.0` or later is required. You can use the [Quick Environment Installation Script](../../../scripts/env_install/pytorch/setup_torch_npu_daily.sh) to quickly install the Python environment and `torch_npu`. | 42 | + |
| 43 | +2. Ensure that the `torch_npu` version is `2.9.0` or later. You can use the [environment quick installation script](../../../scripts/env_install/pytorch/setup_torch_npu_daily.sh) to quickly install the Python environment and `torch_npu`. | ||
| 36 | 44 | ||
| 37 | ## Setting Environment Variables | 45 | ## Setting Environment Variables |
| 38 | 46 | ||
| 39 | -Run the following commands each time you open a new terminal: | 47 | +Run the following commands whenever you open a new terminal: |
| 40 | 48 | ||
| 41 | ```bash | 49 | ```bash |
| 42 | -# Activate the environment. | 50 | +# Activate the Python environment. |
| 43 | source /mnt/workspace/env/venv/torch210_daily/bin/activate | 51 | source /mnt/workspace/env/venv/torch210_daily/bin/activate |
| 44 | 52 | ||
| 45 | -# Set the CANN installation path according to the actual installation location. | 53 | +# Set the CANN installation path based on the actual installation location. |
| 46 | export CANN_INSTALL_PATH=/home/developer/Ascend | 54 | export CANN_INSTALL_PATH=/home/developer/Ascend |
| 47 | 55 | ||
| 48 | # Load CANN environment variables. | 56 | # Load CANN environment variables. |
| 49 | source $CANN_INSTALL_PATH/cann/set_env.sh | 57 | source $CANN_INSTALL_PATH/cann/set_env.sh |
| 50 | 58 | ||
| 51 | -# Assume the example runs on device 0. | 59 | +# Assume that the examples run on device 0. |
| 52 | export ASCEND_DEVICE_ID=0 | 60 | export ASCEND_DEVICE_ID=0 |
| 53 | ``` | 61 | ``` |
| 54 | 62 | ||
| @@ -68,25 +76,33 @@ cd af_reduce | |||
| 68 | python af_mul_reducesum.py | 76 | python af_mul_reducesum.py |
| 69 | ``` | 77 | ``` |
| 70 | 78 | ||
| 79 | +### gather + add Graph Pattern | ||
| 80 | + | ||
| 81 | +```bash | ||
| 82 | +cd af_gather | ||
| 83 | +python af_gather_add.py | ||
| 84 | +``` | ||
| 85 | + | ||
| 71 | ## Expected Results | 86 | ## Expected Results |
| 72 | 87 | ||
| 73 | -After the program finishes running, a `profiling` directory is generated in the current directory. | 88 | +After an example finishes running, a `profiling` directory is generated in the current directory. |
| 74 | 89 | ||
| 75 | -Operator execution details can be viewed in the following directory: | 90 | +You can view operator execution details in the following directory: |
| 76 | 91 | ||
| 77 | ```text | 92 | ```text |
| 78 | -profiling/PROF_timestamp/mindstudio_profiler_output | 93 | +profiling/PROF_<timestamp>/mindstudio_profiler_output |
| 79 | ``` | 94 | ``` |
| 80 | 95 | ||
| 81 | Open the following file: | 96 | Open the following file: |
| 82 | 97 | ||
| 83 | ```text | 98 | ```text |
| 84 | -op_summary_timestamp.csv | 99 | +op_summary_<timestamp>.csv |
| 85 | ``` | 100 | ``` |
| 86 | 101 | ||
| 87 | -If the operator list contains a Kernel whose name starts with `autofused_`, the related operators have been successfully fused into a single fused operator. | 102 | +If the operator list contains a kernel whose name starts with `autofused_`, the related operators have been successfully fused into a single fused operator. |
| 88 | 103 | ||
| 89 | ## References | 104 | ## References |
| 90 | 105 | ||
| 91 | -- [Autofuse Introduction and Quick Start](../../README.md) | 106 | +* [Autofuse Overview and Quick Start](../../README.md) |
| 92 | -- [Profiling Performance Analysis Tool Guide](https://hiascend.com/document/redirect/CannCommunityToolProfiling) | 107 | +* [Profiling Performance Analysis Tool Guide](https://hiascend.com/document/redirect/CannCommunityToolProfiling) |
| 108 | +* [Accuracy Debugging Tool Guide](https://hiascend.com/document/redirect/CannCommunityToolAccucacy) | ||
| @@ -0,0 +1,15 @@ | |||
| 1 | +# autofuse 用例演示 | ||
| 2 | + | ||
| 3 | +## 用例功能: | ||
| 4 | + | ||
| 5 | +本用例构造 gather + add 图模式:先通过 `torch.gather(x, 1, indices)` 沿第 1 轴取数,再将取数结果与同形状张量 `y` 逐元素相加。其中 add 属于 elewise 计算。 | ||
| 6 | + | ||
| 7 | +## 执行命令 | ||
| 8 | + | ||
| 9 | +```bash | ||
| 10 | +python3 af_gather_add.py | ||
| 11 | +``` | ||
| 12 | + | ||
| 13 | +## 预期执行结果 | ||
| 14 | + | ||
| 15 | +当前目录下的 profiling 目录下,有生成的 profiling 文件。其中 PROF_000001_时间戳xx/mindstudio_profiler_output 下面,可以打开 op_summary_时间戳xx.csv 文件,查看算子执行详情。若当前 PyTorch、torch_npu 和 AscendC 后端支持 Gather 前端转换,取数与加法应表现为一个名称以 autofused_ 开头的 kernel,且不再出现独立的 GatherElementsV2 和独立的 add;若仍出现独立 GatherElementsV2,则说明 Gather 已回退到 ACLNN,本用例在当前环境未发生融合。 | ||
| @@ -0,0 +1,15 @@ | |||
| 1 | +# Autofuse Use Case Demonstration | ||
| 2 | + | ||
| 3 | +## Use Case Function: | ||
| 4 | + | ||
| 5 | +This example constructs a gather + add graph pattern. It first selects values along dimension 1 through `torch.gather(x, 1, indices)`, and then performs an element-wise addition with tensor `y` of the same shape. The add operation is an elewise computation. | ||
| 6 | + | ||
| 7 | +## Execution Command | ||
| 8 | + | ||
| 9 | +```bash | ||
| 10 | +python3 af_gather_add.py | ||
| 11 | +``` | ||
| 12 | + | ||
| 13 | +## Expected Execution Result | ||
| 14 | + | ||
| 15 | +Profiling files are generated in the profiling directory under the current directory. Open op_summary_timestampxx.csv under PROF_000001_timestampxx/mindstudio_profiler_output to view operator execution details. If the current PyTorch, torch_npu, and AscendC backend versions support Gather frontend conversion, the indexed selection and addition should appear as one kernel whose name starts with autofused_, without a standalone GatherElementsV2 or add. If a standalone GatherElementsV2 is still present, Gather has fallen back to ACLNN and this pattern is not fused in the current environment. | ||
| @@ -0,0 +1,77 @@ | |||
| 1 | +#!/usr/bin/env python3 | ||
| 2 | +# -*- coding: utf-8 -*- | ||
| 3 | +# ---------------------------------------------------------------------------------------------------------------------- | ||
| 4 | +# Copyright (c) 2026 Huawei Technologies Co., Ltd. | ||
| 5 | +# This program is free software, you can redistribute it and/or modify it under the terms and conditions of | ||
| 6 | +# CANN Open Software License Agreement Version 2.0 (the "License"). | ||
| 7 | +# Please refer to the License for details. You may not use this file except in compliance with the License. | ||
| 8 | +# THIS SOFTWARE IS PROVIDED ON AN "AS IS" BASIS, WITHOUT WARRANTIES OF ANY KIND, EITHER EXPRESS OR IMPLIED, | ||
| 9 | +# INCLUDING BUT NOT LIMITED TO NON-INFRINGEMENT, MERCHANTABILITY, OR FITNESS FOR A PARTICULAR PURPOSE. | ||
| 10 | +# See LICENSE in the root of the software repository for the full text of the License. | ||
| 11 | +# ---------------------------------------------------------------------------------------------------------------------- | ||
| 12 | +# 导包 | ||
| 13 | +import torch | ||
| 14 | +import torch_npu | ||
| 15 | +import torch.nn as nn | ||
| 16 | + | ||
| 17 | +# ===== 1. 昇腾 NPU 配置 ===== | ||
| 18 | +DEVICE = "npu:0" # 假设使用0卡 | ||
| 19 | +torch.npu.set_device(DEVICE) | ||
| 20 | + | ||
| 21 | + | ||
| 22 | +# ===== 2. 构造简单模型 ===== | ||
| 23 | +class MyModel(nn.Module): | ||
| 24 | + def __init__(self): | ||
| 25 | + super().__init__() | ||
| 26 | + | ||
| 27 | + def forward(self, x, indices, y): | ||
| 28 | + result = torch.add(torch.gather(x, 1, indices), y) | ||
| 29 | + return result | ||
| 30 | + | ||
| 31 | + | ||
| 32 | +# ===== 3. inductor + 昇腾NPU自动融合后端 ===== | ||
| 33 | +model = MyModel().to(DEVICE) | ||
| 34 | +model = torch.compile( | ||
| 35 | + model, | ||
| 36 | + dynamic=False, | ||
| 37 | + fullgraph=True, | ||
| 38 | + options={"npu_backend": "ascendc"}, | ||
| 39 | +) | ||
| 40 | + | ||
| 41 | +# ===== 4. 创建输入 ===== | ||
| 42 | +x = torch.randn(128, 50, device=DEVICE) | ||
| 43 | +indices = torch.randint(0, 50, (128, 25), dtype=torch.int64, device=DEVICE) | ||
| 44 | +y = torch.randn(128, 25, device=DEVICE) | ||
| 45 | + | ||
| 46 | +# ===== 5. 执行 ===== | ||
| 47 | +model.eval() | ||
| 48 | + | ||
| 49 | +# 开启 profiling | ||
| 50 | +experimental_config = torch_npu.profiler._ExperimentalConfig( | ||
| 51 | + export_type=[torch_npu.profiler.ExportType.Text], | ||
| 52 | + profiler_level=torch_npu.profiler.ProfilerLevel.Level2, | ||
| 53 | + msprof_tx=False, | ||
| 54 | + aic_metrics=torch_npu.profiler.AiCMetrics.PipeUtilization, | ||
| 55 | + l2_cache=False, | ||
| 56 | + op_attr=False, | ||
| 57 | + data_simplification=False, | ||
| 58 | + record_op_args=False, | ||
| 59 | + gc_detect_threshold=None, | ||
| 60 | +) | ||
| 61 | + | ||
| 62 | +with torch_npu.profiler.profile( | ||
| 63 | + activities=[ | ||
| 64 | + torch_npu.profiler.ProfilerActivity.CPU, | ||
| 65 | + torch_npu.profiler.ProfilerActivity.NPU, | ||
| 66 | + ], | ||
| 67 | + on_trace_ready=torch_npu.profiler.tensorboard_trace_handler("./profiling"), | ||
| 68 | + record_shapes=True, | ||
| 69 | + profile_memory=False, | ||
| 70 | + with_stack=False, | ||
| 71 | + with_modules=False, | ||
| 72 | + with_flops=False, | ||
| 73 | + experimental_config=experimental_config, | ||
| 74 | +) as prof: | ||
| 75 | + # 跑 100 step | ||
| 76 | + for _ in range(100): | ||
| 77 | + result = model(x, indices, y) | ||