已合并
docs: add comprehensive English documentation set #76
LiuZonggu创建于 8月1日
docs: add comprehensive English documentation set #76
已合并
LiuZonggu创建于 8月1日
从已删除 :docs/eng-ver合入到cann/ops-rasmaster
共 93 个文件变更+5091-0
@@ -0,0 +1 @@
1+# Changelog
@@ -0,0 +1,60 @@
1+# Contribution Guide
2+ 
3+This project welcomes developers to experience and participate in contributions. Before participating in community contributions, please see [cann-community](https://gitcode.com/cann/community) to understand the code of conduct, sign the CLA agreement, and understand the contribution process of the source code repository.
4+ 
5+Developers need to pay attention to the following points when preparing local code and submitting PRs:
6+ 
7+1. When submitting a PR, please carefully fill in the business background, purpose, solution, and other information of this PR according to the PR template.
8+2. If your modification is not a simple bug fix, but involves adding new features, new interfaces, new configuration parameters, or modifying code flow, please be sure to discuss the solution through an Issue first to avoid your code being rejected. If you are not sure whether this modification can be classified as a "simple bug fix", you can also discuss the solution by submitting an Issue.
9+ 
10+Developer contribution scenarios mainly include:
11+ 
12+- Operator Bug Fix
13+ 
14+ If you discover certain operator bugs in this project and want to fix them, we welcome you to create a new Issue for feedback and tracking.
15+ 
16+ You can create a new `Bug-Report|Bug Report` type Issue according to the [Submit Issue/Handle Issue Task](https://gitcode.com/cann/community#提交Issue处理Issue任务) guide to describe the bug, and then enter "/assign" or "/assign @yourself" in the comment box to assign this Issue to you for processing.
17+ 
18+- Operator Optimization
19+ 
20+ If you have generalization enhancement/performance optimization ideas for certain operator implementations in this project and want to implement these optimization points, we welcome you to contribute operator optimizations.
21+ 
22+ You can create a new `Requirement|Feature Request` type Issue according to the [Submit Issue/Handle Issue Task](https://gitcode.com/cann/community#提交Issue处理Issue任务) guide to explain the optimization points and provide your design solution, and then enter "/assign" or "/assign @yourself" in the comment box to assign this Issue to you for tracking optimization.
23+ 
24+- Contribute New Operators
25+ 
26+ If you have a brand new operator that you want to design and implement based on NPU, we welcome you to propose new ideas and designs in an Issue.
27+ 
28+ You can create a new `Requirement|Feature Request` type Issue according to the [Submit Issue/Handle Issue Task](https://gitcode.com/cann/community#提交Issue处理Issue任务) guide to provide the new operator description and design solution. Project members will communicate and confirm with you, and provide a suitable `contrib` directory classification for your operator under the `experimental` directory. You can contribute the new operator to the corresponding directory.
29+ 
30+ At the same time, you need to comment "/assign" or "/assign @yourself" in the submitted Issue to claim this Issue and subsequently complete the new operator submission.
31+ 
32+ The deliverables for new operators are usually quite numerous. You can refer to the following list to check the minimum deliverable set, where `${op_name}` indicates the new operator name.
33+ ```
34+ ${op_class} # operator classification
35+ ├── ${op_name} # operator name
36+ │ ├── op_host # operator definition, Tiling, InferShape related implementation
37+ │ │ ├── ${op_name}_def.cpp # operator definition file
38+ │ │ ├── ${op_name}_tiling.cpp # operator Tiling implementation file
39+ │ │ └── CMakeLists.txt
40+ │ ├── op_kernel # operator Kernel directory
41+ │ │ ├── ${op_name}.cpp
42+ │ │ ├── ${op_name}.h
43+ │ │ ├── ${op_name}_tiling_data.h
44+ │ │ ├── ${op_name}_tiling_key.h
45+ │ │ └── CMakeLists.txt
46+ │ ├── CMakeLists.txt # operator compilation configuration file, keep the original file
47+ │ └── README.md # operator description document
48+ ```
49+ 
50+- Document Correction
51+ 
52+ If you discover certain operator document description errors in this project, we welcome you to create a new Issue for feedback and correction.
53+ 
54+ You can create a new `Documentation|Documentation Feedback` type Issue according to the [Submit Issue/Handle Issue Task](https://gitcode.com/cann/community#提交Issue处理Issue任务) guide to point out the problems in the corresponding document, and then enter "/assign" or "/assign @yourself" in the comment box to assign this Issue to you to correct the corresponding document description.
55+ 
56+- Help Solve Others' Issues
57+ 
58+ If you have suitable solutions for problems encountered by others in the community, we welcome you to comment and communicate in the Issue to help others solve problems and pain points, and jointly optimize usability.
59+ 
60+ If the corresponding Issue requires code modification, you can enter "/assign" or "/assign @yourself" in the Issue comment box to assign this Issue to you for tracking and assisting in solving the problem.
@@ -0,0 +1,51 @@
1+# ops-ras
2+ 
3+## 🔥Latest News
4+ 
5+- [2026/07] The ops-ras project was first released.
6+ 
7+## 🚀Overview
8+ 
9+ops-ras is the security and RAS (Reliability, Availability and Serviceability) operator library in the [CANN](https://hiascend.com/software/cann) (Compute Architecture for Neural Networks) operator library, providing reliability, availability, and maintainability capabilities, including security, encryption, and RAS-related operators. "ras" is derived from the initials of these three core characteristics. The operator library architecture is shown below:
10+ 
11+<img src="docs/zh/figures/architecture.png" alt="Architecture Diagram" width="700px" height="320px">
12+ 
13+## 📌Version Compatibility
14+ 
15+The source code of this project will be released along with the CANN software version. For the correspondence between CANN software versions and project tags, refer to the relevant version descriptions in the [release repository](https://gitcode.com/cann/release-management).
16+Note that to ensure smooth custom development of your source code, select the matching CANN version and Gitcode tag source code. Using the master branch may pose version mismatch risks.
17+ 
18+## 🛠️Environment Setup
19+ 
20+[Environment Deployment](docs/zh/install/quick_install.md) is the prerequisite for experiencing the capabilities of this project. Please complete the NPU driver installation, CANN package installation, and so on to ensure the environment is normal.
21+ 
22+## ⬇️Source Code Download
23+ 
24+After the environment is ready, download the branch source code matching the CANN version. The general command is as follows. Replace `${tag_version}` with the branch tag name. Take the 9.0.0 branch source code download as an example:
25+ 
26+```bash
27+# General command: git clone -b ${tag_version} https://gitcode.com/cann/ops-ras.git
28+git clone -b 9.0.0 https://gitcode.com/cann/ops-ras.git
29+```
30+ 
31+> Note: If the matching branch source code already exists in the environment, **you can skip this step**. For example, CANNLab provides the source code matching the latest CANN version by default.
32+ 
33+## 📖Learning Tutorials
34+ 
35+- [Quick Start](docs/QUICKSTART.md): Quickly experience the core basic capabilities of the project from scratch, covering source code compilation, operator invocation, development, debugging, and other operations.
36+- [Advanced Tutorials](docs/README.md): If you need a deeper understanding of the project's compilation and deployment, operator invocation, development, debugging and tuning, and other capabilities, please refer to the documentation center for detailed guidance.
37+ 
38+## 💬Related Information
39+ 
40+- [Directory Structure](docs/zh/install/dir_structure.md)
41+- [Contribution Guide](CONTRIBUTING.md)
42+- [Security Statement](SECURITY.md)
43+- [License](LICENSE)
44+- [Affiliated SIG](https://gitcode.com/cann/community/tree/master/CANN/sigs/ops-basic)
45+ 
46+-----
47+PS: The functions and documentation of this project are being continuously updated and improved. We recommend that you follow the latest version.
48+ 
49+- **Issue Feedback**: Submit issues through GitCode [Issues](https://gitcode.com/cann/ops-ras/issues).
50+- **Community Interaction**: Participate in discussions through GitCode [Discussions](https://gitcode.com/cann/ops-ras/discussions).
51+- **Technical Column**: Access technical articles through GitCode [Wiki](https://gitcode.com/cann/ops-ras/wiki), such as serialized tutorials and best practices.
@@ -0,0 +1,65 @@
1+# Security Statement
2+ 
3+## Running User Recommendations
4+ 
5+Based on security considerations, we do not recommend using root or other administrator type accounts to execute any commands. Follow the principle of minimum permissions.
6+ 
7+## File Permission Control
8+ 
9+- We recommend that users set the running system umask value to 0027 or above on the host machine (including the host machine) and in the container to ensure that the default maximum permission for new folders is 750 and the default maximum permission for new files is 640.
10+- We recommend that users take security measures such as permission control for sensitive content such as personal privacy data, business assets, source files, and various files saved during operator development. For example, for project installation directory permission control and input public data file permission control, the set permissions should refer to [A-File (Folder) Permission Control Recommended Maximum Values in Various Scenarios](#a-file-folder-permission-control-recommended-maximum-values-in-various-scenarios).
11+- When the operator runs, it may cache operator compilation files, which are stored in the `kernel_meta_*` folder under the running directory to speed up subsequent operator invocation. Users can perform permission control on the generated related files as needed.
12+- Users need to perform permission control during installation and use. We recommend referring to [A-File (Folder) Permission Control Recommended Maximum Values in Various Scenarios](#a-file-folder-permission-control-recommended-maximum-values-in-various-scenarios) for file permission reference settings.
13+ 
14+## Build Security Statement
15+ 
16+When compiling and installing this project from source code, you need to compile it yourself. During the compilation process, some intermediate files will be generated. We recommend that you perform permission control on the intermediate files after compilation to ensure file security.
17+ 
18+## Running Security Statement
19+ 
20+- We recommend that users write corresponding operator invocation scripts based on the running environment resource status. If the operator invocation script does not match the resource status, such as the space used for generating input data or benchmark calculation results exceeding the memory capacity limit, or the script saving data locally exceeding the disk space size, it may cause errors and lead to unexpected process exit.
21+- When the operator runs abnormally, it will exit the process and print error information. We recommend locating the specific error cause based on the error prompt, including setting operator synchronous execution, viewing log files, and other methods.
22+- When the operator is invoked through [PyTorch](https://gitee.com/ascend/pytorch), running errors may occur due to version mismatch. For details, please refer to [PyTorch Security Statement](https://gitee.com/ascend/pytorch#%E5%AE%89%E5%85%A8%E5%A3%B0%E6%98%8E).
23+ 
24+## Public Network Address Statement
25+ 
26+The public network addresses contained in this project code are declared as follows:
27+ 
28+| Type | Open Source Code Address | File Name | Public Network IP Address/Public Network URL Address/Domain Name/Email Address/Compressed File Address | Usage Description |
29+| :------------: |:------------------------------------------------------------------------------------------:|:----------------------------------------------------------| :---------------------------------------------------------- |:-----------------------------------------|
30+| Dependency | Not involved | cmake/third_party/makeself-fetch.cmake | [https://gitcode.com/cann-src-third-party/makeself/releases/download/release-2.5.0-patch1.0/makeself-release-2.5.0-patch1.tar.gz](https://gitcode.com/cann-src-third-party/makeself/releases/download/release-2.5.0-patch1.0/makeself-release-2.5.0-patch1.tar.gz) | Download makeself source code from gitcode, used as compilation dependency |
31+| Dependency | Not involved | cmake/third_party/nlohmann_json.cmake | [https://gitcode.com/cann-src-third-party/json/releases/download/v3.11.3/include.zip](https://gitcode.com/cann-src-third-party/json/releases/download/v3.11.3/include.zip) | Download json source code from gitcode, used as compilation dependency |
32+| Dependency | Not involved | cmake/third_party/gtest.cmake | [https://gitcode.com/cann-src-third-party/googletest/releases/download/v1.14.0/googletest-1.14.0.tar.gz](https://gitcode.com/cann-src-third-party/googletest/releases/download/v1.14.0/googletest-1.14.0.tar.gz) | Download googletest source code from gitcode, used as compilation dependency |
33+| Dependency | Not involved | cmake/third_party/eigen.cmake | [https://gitcode.com/cann-src-third-party/eigen/releases/download/5.0.0-h0.trunk/eigen-5.0.0.tar.gz](https://gitcode.com/cann-src-third-party/eigen/releases/download/5.0.0-h0.trunk/eigen-5.0.0.tar.gz) | Download eigen source code from gitcode, used as compilation dependency |
34+| Dependency | Not involved | ops-ras/install_deps.sh | [https://apt.kitware.com/keys/kitware-archive-latest.asc](https://apt.kitware.com/keys/kitware-archive-latest.asc) | Download install_deps source code from gitcode, used as compilation dependency |
35+| Dependency | Not involved | ops-ras/install_deps.sh | [https://apt.kitware.com/ubuntu/](https://apt.kitware.com/ubuntu/) | Download install_deps source code from gitcode, used as compilation dependency |
36+| Dependency | Not involved | cmake | [https://apt.kitware.com/keys/kitware-archive-latest.asc](https://apt.kitware.com/keys/kitware-archive-latest.asc) | Download cmake software from kitware, used as compilation dependency |
37+| Dependency | Not involved | cmake | [https://apt.kitware.com/ubuntu/](https://apt.kitware.com/ubuntu/) | Download cmake software from kitware, used as compilation dependency |
38+ 
39+## Vulnerability Mechanism Description
40+ 
41+[Vulnerability Management](https://gitcode.com/cann/community/blob/master/security/security.md)
42+ 
43+## Appendix
44+ 
45+### A-File (Folder) Permission Control Recommended Maximum Values in Various Scenarios
46+ 
47+| Type | Linux Permission Reference Maximum Value |
48+| -------------- | --------------- |
49+| User Home Directory | 750 (rwxr-x---) |
50+| Program Files (including script files, library files, etc.) | 550 (r-xr-x---) |
51+| Program File Directory | 550 (r-xr-x---) |
52+| Configuration File | 640 (rw-r-----) |
53+| Configuration File Directory | 750 (rwxr-x---) |
54+| Log File (recording completed or archived) | 440 (r--r-----) |
55+| Log File (currently recording) | 640 (rw-r-----) |
56+| Log File Directory | 750 (rwxr-x---) |
57+| Debug File | 640 (rw-r-----) |
58+| Debug File Directory | 750 (rwxr-x---) |
59+| Temporary File Directory | 750 (rwxr-x---) |
60+| Maintenance Upgrade File Directory | 770 (rwxrwx---) |
61+| Business Data File | 640 (rw-r-----) |
62+| Business Data File Directory | 750 (rwxr-x---) |
63+| Key Component, Private Key, Certificate, Ciphertext File Directory | 700 (rwx-----) |
64+| Key Component, Private Key, Certificate, Encrypted Ciphertext | 600 (rw-------) |
65+| Encryption/Decryption Interface, Encryption/Decryption Script | 500 (r-x------) |
@@ -0,0 +1,102 @@
1+# Documentation Contribution Guide
2+ 
3+We welcome your contributions to the project documentation. High-quality documentation is crucial for project success. This guide will help you efficiently submit documentation that meets the standards.
4+ 
5+## Contribution Scope
6+ 
7+We welcome any contributions that can improve documentation quality, including but not limited to:
8+ 
9+- Correction and Improvement: Fix typos, grammar errors, incorrect code examples, outdated information, or broken links.
10+ 
11+- Clarification and Optimization: Make descriptions clearer and easier to understand, optimize sentence structure, and supplement background knowledge.
12+ 
13+- Content Supplement: Add usage examples, API documentation, frequently asked questions (FAQ), best practices, or warning descriptions for existing features.
14+ 
15+- New Content Creation: Write new chapters or tutorials for newly added features, such as operator README, API introduction documents, and so on. If you have questions, we recommend creating an Issue for discussion first.
16+ 
17+- Localization Translation: Help us translate or proofread documents in other languages.
18+ 
19+- Style and Navigation: Improve the layout, readability, and navigation structure of the documentation website.
20+ 
21+## Contribution Process
22+ 
23+1. **Preparation Work**
24+ 
25+ - Determine the Task: If there are documentation issues, you can create new Issues. We recommend using the label category `[Documentation|文档反馈]` and providing a detailed description. Based on the existing Issues list, determine the documentation issues to be resolved.
26+ - Claim the Task: Comment `/assign @yourself` under the corresponding Issue to indicate that you will handle it and avoid duplicate work.
27+ 
28+2. **Document Modification**
29+ 
30+ - Select Branch: Please download the source code from the master or other Tag branches to the local machine.
31+ - Follow Format:
32+ - This project recommends using **Markdown format**.
33+ - Follow the existing writing style of the project.
34+ - Put static resources such as images in the corresponding directory. For example, images are generally in the `figures` folder under the docs directory. You can adjust them yourself in special cases.
35+ - Careful Addition and Deletion: When modifying content, please try to maintain the original line width and line break conventions.
36+ 
37+3. **Submit Changes**
38+ 
39+ - Atomic Commit: Each commit should focus on an independent modification. For example, "Fix spelling errors in xx guide" and "Update example code in API reference" should be submitted separately.
40+ 
41+ - Write Clear Commit Messages:
42+ 
43+ ```text
44+ Brief description (no more than 50 characters)
45+ 
46+ If necessary, provide a more detailed description here. Explain the reason and content of the modification, rather than what specifically was changed (the code itself will show).
47+ Associated Issue: #123
48+ ```
49+ 
50+4. **Initiate Pull Request**
51+ 
52+ - Target Branch: Please merge the PR into the target branch of the project.
53+ - Title and Description:
54+ - PR Title: Should clearly summarize the modification, for example: `[Docs] Fix configuration example in quick start`.
55+ - PR Description: Detailed explanation of your changes, motivation, and associated Issues (use Closes #123 or Fixes #456).
56+ - Preview Check: Please check the document effect in local or online browsing in advance to ensure that the rendering meets expectations.
57+ - Wait for Review: Maintainers will review and may propose modification suggestions. Please follow up on the discussion in a timely manner.
58+ 
59+## Writing Standards
60+ 
61+Before developers write project documentation, please be sure to read the following standards first. If you have questions, you are welcome to make suggestions at any time!
62+ 
63+- Prerequisites: Please first learn the unified writing standards provided by the CANN organization. For details, see [CANN Document Writing Standards](https://gitcode.com/cann/community/blob/master/contributor/docs/document_writing_specs.md).
64+ 
65+ - Document Content Requirements: Introduce the required and optional document deliverables in the project.
66+ - Directory Structure Standards: Introduce the principles of directory division, such as Chinese and English management.
67+ - Content Element Standards: Introduce rules for different writing elements, such as file naming, titles, fonts, images, code blocks, links, and so on.
68+ 
69+- Precautions:
70+ 
71+ In addition to the above writing rules, you also need to pay attention to the following:
72+ 
73+ - Tone: Use a friendly, professional, and neutral tone. For beginners, avoid unnecessary jargon.
74+ - Terminology: Maintain terminology consistency (such as uniformly using "click" instead of "single click"). Please refer to the project terminology table (if available).
75+ - Code Examples:
76+ - Ensure that all code examples are runnable and tested.
77+ - Provide sufficient context and explanation.
78+ - Indicate the environment or prerequisites required for code running.
79+ - Punctuation and Format:
80+ - When mixing Chinese and English, use full-width punctuation. Punctuation marks must conform to the Chinese/English context.
81+ - Use appropriate hierarchy for titles (#, ##, ###).
82+ - Use lists and tables to organize complex information.
83+ - Links: Use descriptive link text, avoid "click here", and ensure that link resources are authentic and reliable.
84+ - Images:
85+ - Common Formats: We recommend the png format. Try to keep the style consistent with existing images.
86+ - Resolution and Clarity: Must be clear and of moderate size. Avoid blurring or excessive compression.
87+ - File Size: We do not recommend that a single image exceeds 10M.
88+ - Copyright: For all quoted images, literature, and other resources, please ensure compliance.
89+ 
90+## Get Help
91+ 
92+If you have any questions during the contribution process:
93+ 
94+1. Check Existing Documentation: If there are problems with templates or standards, please first check the existing guides, API documentation, or README of the project.
95+2. Initiate Discussion: You can create a new Issue or leave a message directly in the relevant Issue or PR.
96+ 
97+## Document Templates
98+ 
99+The key documents involved in operator deliverables mainly include the following. For specific writing formats and content requirements, please refer to the templates.
100+ 
101+- [Operator README Document Template](https://gitcode.com/cann/ops-ras/wiki/%E7%AE%97%E5%AD%90README%E6%96%87%E6%A1%A3%E6%A8%A1%E6%9D%BF)
102+- [aclnn API Document Template](https://gitcode.com/cann/ops-ras/wiki/aclnn%20API%E6%96%87%E6%A1%A3%E6%A8%A1%E6%9D%BF)
@@ -0,0 +1,273 @@
1+# Quick Start: Based on ops-ras Repository
2+ 
3+## Usage Notice
4+ 
5+This guide aims to help you quickly get started with CANN and the `ops-ras` operator repository. To help you quickly understand the entire process of operator development, we will use the **AddExample** operator as a practical object. Its source code is located in `ops-ras/examples/add_example`. The operation process is as follows:
6+ 
7+1. **[Prerequisites](../README.md)**: Complete the environment setup and source code download by referring to the project README. The process is not repeated here. For the quick start scenario, **CANNLab or Docker deployment is recommended** for simple operation.
8+ 
9+ > **Note**: The CANNLab or Docker environment provides the latest version of the CANN package by default. If you need to experience the latest capabilities of the master branch, you can manually set up the environment.
10+ 
11+2. **[Compilation and Running](#i-compilation-and-running)**: Compile the custom operator package and install it to achieve quick operator invocation.
12+ 
13+3. **[Operator Development](#ii-operator-development)**: Experience the complete loop of development, compilation, and verification by modifying the existing operator Kernel.
14+ 
15+4. **[Operator Debugging](#iii-operator-debugging)**: Master the methods of operator printing and performance collection.
16+ 
17+5. **[Operator Verification](#iv-operator-verification)**: Learn how to modify operator example samples to verify the functional correctness of operators under different inputs.
18+ 
19+## I. Compilation and Running
20+ 
21+The purpose of this stage is to **quickly experience the project standard process** and verify whether the environment can successfully perform operator source code compilation, packaging, installation, and running.
22+ 
23+### 1. Enter the Project Source Code
24+ 
25+- CANNLab Cloud Development Environment:
26+ 
27+ The latest CANN package matching project source code is provided by default. Enter the source code directory and replace `${gitCode_id}` with the developer's personal gitCode account.
28+ 
29+ ```bash
30+ cd /mnt/workspace/gitCode/${gitCode_id}/ops-ras
31+ ```
32+ 
33+- Non-CANNLab Cloud Development Environment:
34+ 
35+ According to the correspondence between source code and CANN versions in the [release repository](https://gitcode.com/cann/release-management), execute the following command to download the source code. Replace `${tag_version}` with the target branch tag, for example, 9.0.0.
36+ 
37+ ```bash
38+ git clone -b ${tag_version} https://gitcode.com/cann/ops-ras.git && cd ops-ras
39+ ```
40+ 
41+> Note: If you need to switch the source code branch version, refer to the following guidance.
42+>
43+> 1. Execute `git branch` in the source code directory to query the current source code version.
44+> 2. Execute `git checkout ${tag_version}` in the source code directory to switch to the target branch source code. Ensure that the source code matches the CANN version. If the source code already exists, execute `git pull` to pull the latest source code.
45+ 
46+### 2. Compile the AddExample Operator
47+ 
48+This guide uses **single operator compilation** by default: only the target operator is built, the compilation time is short, and it is suitable for quick start and daily development. The general command format: `bash build.sh --pkg --soc=<chip version> --ops=<operator name>`.
49+ 
50+> If you need to compile the entire operator library (omit `--ops`), see [build Parameter Description](zh/install/build.md).
51+ 
52+Taking the AddExample operator as an example, the compilation command is as follows:
53+ 
54+```bash
55+bash build.sh --pkg --soc=${soc_version} --ops=add_example -j16
56+```
57+ 
58+For the value of `${soc_version}`, visit the [CANN Download Center](https://www.hiascend.com/cann/download) and query the hardware product name according to the page prompts. The corresponding `${soc_version}` values for product names are as follows. Please pass the parameter according to the actual scenario.
59+ 
60+- Atlas A2 Training Series Products/Atlas A2 Inference Series Products: `ascend910b`
61+- Atlas A3 Training Series Products/Atlas A3 Inference Series Products: `ascend910_93`
62+- Ascend 950 Series Products: `ascend950`
63+ 
64+If the following information is prompted, the compilation is successful.
65+ 
66+```bash
67+Self-extractable archive "cann-ops-ras-custom_linux-${arch}.run" successfully created.
68+```
69+ 
70+After successful compilation, the run package is stored in the build_out directory under the project root directory.
71+ 
72+### 3. Install the AddExample Operator Package
73+ 
74+```bash
75+./build_out/cann-ops-ras-*linux*.run
76+```
77+ 
78+`AddExample` is installed in the ```${ASCEND_HOME_PATH}/opp/vendors``` path. ```${ASCEND_HOME_PATH}``` indicates the CANN software installation directory.
79+ 
80+### 4. Configure Environment Variables
81+ 
82+Add the path of the custom operator package to the environment variables to ensure that it can be found at runtime.
83+ 
84+```bash
85+export LD_LIBRARY_PATH=${ASCEND_HOME_PATH}/opp/vendors/custom_ras/op_api/lib:${LD_LIBRARY_PATH}
86+```
87+ 
88+### 5. Quick Verification: Run Operator Sample
89+ 
90+The general running command format: `bash build.sh --run_example <operator name> <running mode> <package mode>`.
91+ 
92+Taking AddExample as an example, it provides a simple operator sample `add_example/examples/test_aclnn_add_example.cpp`. Run this sample to verify whether the operator function is normal.
93+ 
94+```bash
95+bash build.sh --run_example add_example eager cust --vendor_name=custom
96+```
97+ 
98+Expected output: Print the addition calculation result of the operator `AddExample`, indicating that the operator has been successfully deployed and executed correctly.
99+ 
100+```bash
101+mean result[0] is: 2.000000
102+mean result[1] is: 2.000000
103+mean result[2] is: 2.000000
104+mean result[3] is: 2.000000
105+mean result[4] is: 2.000000
106+mean result[5] is: 2.000000
107+mean result[6] is: 2.000000
108+mean result[7] is: 2.000000
109+...
110+```
111+ 
112+## II. Operator Development
113+ 
114+The purpose of this stage is to try **modifying the kernel function code** for the successfully running AddExample operator.
115+ 
116+### 1. Modify Kernel Implementation
117+ 
118+Find the core kernel implementation file of the AddExample operator `ops-ras/examples/add_example/op_kernel/add_example.h`, and try to change the Add operation in the operator to a Mul operation:
119+ 
120+```cpp
121+__aicore__ inline void AddExample<T>::Compute(int32_t progress)
122+{
123+ AscendC::LocalTensor<T> xLocal = inputQueueX.DeQue<T>();
124+ AscendC::LocalTensor<T> yLocal = inputQueueY.DeQue<T>();
125+ AscendC::LocalTensor<T> zLocal = outputQueueZ.AllocTensor<T>();
126+ // === Replace Add with Mul here ===
127+ // AscendC::Add(zLocal, xLocal, yLocal, tileLength_);
128+ AscendC::Mul(zLocal, xLocal, yLocal, tileLength_);
129+ outputQueueZ.EnQue<T>(zLocal);
130+ inputQueueX.FreeTensor(xLocal);
131+ inputQueueY.FreeTensor(yLocal);
132+}
133+```
134+ 
135+### 2. Compile and Verify
136+ 
137+Repeat the steps in the [Compilation and Running](#i-compilation-and-running) section:
138+ 
139+1. **Recompile**:
140+ 
141+ First return to the project root directory. The compilation command is as follows:
142+ 
143+ ```bash
144+ bash build.sh --pkg --soc=${soc_version} --ops=add_example -j16
145+ ```
146+ 
147+ > **Note**: Please fill in `${soc_version}` according to the actual chip model. The value method is the same as described in [Compile the AddExample Operator](#2-compile-the-addexample-operator).
148+ 
149+2. **Reinstall**:
150+ 
151+ ```bash
152+ ./build_out/cann-ops-ras-*linux*.run
153+ ```
154+ 
155+3. **Re-verify**:
156+ 
157+ ```bash
158+ bash build.sh --run_example add_example eager cust --vendor_name=custom
159+ ```
160+ 
161+4. **Success Sign**: The output result becomes the multiplication result.
162+ 
163+ ```bash
164+ mean result[0] is: 1.000000
165+ mean result[1] is: 1.000000
166+ mean result[2] is: 1.000000
167+ mean result[3] is: 1.000000
168+ mean result[4] is: 1.000000
169+ mean result[5] is: 1.000000
170+ mean result[6] is: 1.000000
171+ mean result[7] is: 1.000000
172+ ...
173+ ```
174+ 
175+## III. Operator Debugging
176+ 
177+This stage takes AddExample as an example to add printing in the operator and collect operator performance data for subsequent problem analysis and positioning.
178+ 
179+### 1. Printing
180+ 
181+If the operator has execution failure, precision abnormality, or other problems, add printing for problem analysis and positioning.
182+ 
183+Please modify the code in `examples/add_example/op_kernel/add_example.h`.
184+ 
185+* **printf**
186+ 
187+ This interface supports printing Scalar type data, such as integers, character types, Boolean types, and so on. For detailed introduction, see "Operator Debugging API > printf" in "[Ascend C API](https://hiascend.com/document/redirect/CannCommunityAscendCApi)".
188+ 
189+ ```c++
190+ blockLength_ = (tilingData->totalLength + AscendC::GetBlockNum() - 1) / AscendC::GetBlockNum();
191+ tileNum_ = tilingData->tileNum;
192+ tileLength_ = ((blockLength_ + tileNum_ - 1) / tileNum_ / BUFFER_NUM) ?
193+ ((blockLength_ + tileNum_ - 1) / tileNum_ / BUFFER_NUM) : 1;
194+ // Print the current kernel calculation Block length
195+ AscendC::PRINTF("Tiling blockLength is %llu\n", blockLength_);
196+ ```
197+ 
198+* **DumpTensor**
199+ 
200+ This interface supports dumping the content of the specified Tensor, and also supports printing custom additional information, such as the current line number. For detailed introduction, see "Operator Debugging API > DumpTensor" in "[Ascend C API](https://hiascend.com/document/redirect/CannCommunityAscendCApi)".
201+ 
202+ ```c++
203+ AscendC::LocalTensor<T> zLocal = outputQueueZ.DeQue<T>();
204+ // Print zLocal Tensor information
205+ DumpTensor(zLocal, 0, 128);
206+ ```
207+ 
208+### 2. Performance Collection
209+ 
210+When the operator function verification is correct, you can collect operator performance data through the `msprof` tool.
211+ 
212+- **Generate Executable File**
213+ 
214+ Call the example sample of the AddExample operator to generate an executable file (test_aclnn_add_example), which is located in the project `ops-ras/build` directory.
215+ 
216+ ```bash
217+ bash build.sh --run_example add_example eager cust --vendor_name=custom
218+ ```
219+ 
220+- **Collect Performance Data**
221+ 
222+ Enter the AddExample operator executable file directory `ops-ras/build/` and execute the following command:
223+ 
224+ ```bash
225+ msprof --application="./test_aclnn_add_example"
226+ ```
227+ 
228+The collection result is in the project `ops-ras/build/` directory. After the msprof command is executed, it will automatically parse and export the performance data result file. For detailed content, see [msprof](https://www.hiascend.com/document/detail/zh/mindstudio/82RC1/T&ITools/Profiling/atlasprofiling_16_0110.html#ZH-CN_TOPIC_0000002504160251).
229+ 
230+## IV. Operator Verification
231+ 
232+This stage verifies the functional correctness of the operator in multiple scenarios by modifying the input data of the AddExample operator example sample.
233+ 
234+### 1. Modify Test Input
235+ 
236+Find and edit the `ops-ras/examples/add_example/examples/test_aclnn_add_example.cpp` of `AddExample`, and modify the shape and numerical values of the input tensor.
237+ 
238+**Modify Input/Output Data**: Modify the shape information of input and output, as well as the initialization data, and construct the corresponding input and output tensors.
239+ 
240+```c++
241+int main() {
242+ // ... initialization code ...
243+ 
244+ // === ① Modify selfX input ===
245+ // Before modification: shape = {32, 4, 4, 4}, all values are 1
246+ // After modification: change input shape to {8, 8, 8, 8}, and fill with different test data
247+ std::vector<int64_t> selfXShape = {8, 8, 8, 8};
248+ std::vector<float> selfXHostData(4096); // 4096 = 8 * 8 * 8 *8
249+ // You can use a loop to fill more distinguishable data, such as an increasing sequence
250+ for (int i = 0; i < 4096; ++i) {
251+ selfXHostData[i] = static_cast<float>(i % 10); // Fill with cyclic values of 0-9
252+ }
253+ // === ② Refer to selfX, similarly modify selfY and selfZ inputs ===
254+ 
255+ // ... subsequent execution code ...
256+}
257+```
258+ 
259+### 2. Recompile and Verify
260+ 
261+1. Since only the example test code is modified, there is no need to recompile the operator package.
262+ 
263+2. Re-execute the verification command:
264+ 
265+ ```bash
266+ bash build.sh --run_example add_example eager cust --vendor_name=custom
267+ ```
268+ 
269+3. Observe whether the operator output result meets expectations.
270+ 
271+## Conclusion
272+ 
273+After experiencing the above processes, you have basically completed the operator development process. If you want to further contribute new operators or learn more advanced development, debugging, and other skills, please visit the project README to learn about [Advanced Tutorials](../README.md#学习教程) and [Contribution Guide](../README.md#相关信息), and so on.
@@ -0,0 +1,72 @@
1+# Documentation Center
2+ 
3+## Directory Structure
4+ 
5+The Docs directory structure is described as follows:
6+ 
7+```text
8+├── zh
9+ ├── context # Public documents, such as terminology, basic concepts, and so on
10+ ├── debug # Operator debugging guidance documents
11+ │ ├── cann_sim.md
12+ │ ├── op_debug_prof.md
13+ │ └── ...
14+ ├── develop # Operator development guidance documents
15+ │ ├── aicore_develop_guide.md
16+ │ ├── aicpu_develop_guide.md
17+ │ ├── cross_platform_migration_guide.md
18+ │ ├── graph_develop_guide.md
19+ │ └── ...
20+ ├── figures # Image directory
21+ ├── install # Environment installation and compilation guidance documents
22+ │ ├── build.md
23+ │ ├── compile.md
24+ │ ├── dir_structure.md
25+ │ ├── quick_install.md
26+ │ └── ...
27+ ├── invocation # Operator invocation guidance documents (including aclnn invocation, graph mode invocation, and so on)
28+ │ ├── quick_op_invocation.md
29+ │ ├── op_invocation.md
30+ │ └── ...
31+ ├── op_api_list.md # Complete operator interface list (aclnn)
32+ ├── op_list.md # Complete operator list
33+├── CONTRIBUTING_DOCS.md # Documentation contribution guide
34+├── QUICKSTART.md # Quick start
35+└── README.md
36+```
37+ 
38+## Advanced Tutorials
39+ 
40+### Guide Documents
41+ 
42+| Document | Description |
43+| ------------------------------------------------------------ | ------------------------------------------------------------ |
44+| [Source Code Build Guide](zh/install/compile.md) | Introduces different source code build methods and verification methods in online and offline scenarios. |
45+| [Operator Invocation Guide](zh/invocation/quick_op_invocation.md) | Introduces the method of invoking operator samples and different operator invocation methods (such as aclnn/graph, and so on). |
46+| [Standard Operator Development Guide](zh/develop/aicore_develop_guide.md) | Introduces how to define operator prototypes and implement Tiling and Kernel based on standard engineering. Such operators are called "standard operators".<br>Standard operators support aclnn and graph mode invocation. |
47+| [Simple Operator Development Guide](../examples/fast_kernel_launch_example/README.md) | Introduces how to implement fast_kernel_launch based on simple engineering, that is, the `<<<>>>` method. Such operators are called "simple operators".<br>Simple operators only support PyTorch invocation. |
48+| [Operator Debugging and Tuning](zh/debug/op_debug_prof.md) | Introduces common operator function debugging and performance tuning methods (such as data collection and simulation pipeline, and so on). |
49+ 
50+### API Documents
51+ 
52+| Document | Description |
53+| ----------------------- | ---------------------- |
54+| [Operator List](zh/op_list.md) | Introduces the list of all operators included in the project. |
55+| [aclnn List](zh/op_api_list.md) | Introduces the list of all operator aclnn APIs included in the project. To facilitate users to invoke operators on the Host side, C language APIs are provided, that is, APIs with the aclnn prefix. |
56+ 
57+### Tool Documents
58+ 
59+| Document | Description |
60+| ----------------------- | ---------------------- |
61+| [Simulator Tool](zh/debug/cann_sim.md) | A SoC-level simulation tool for operator development scenarios, used to analyze the precision and performance data of AI tasks running on the AI simulator at each stage. |
62+ 
63+### More Documents
64+ 
65+- Sample Documents: Refer to the operator samples in the [cann-samples](https://gitcode.com/cann/cann-samples) repository.
66+ 
67+## Appendix
68+ 
69+| Document | Description |
70+| ----------------------------------- | ------------------------------------------------------------ |
71+| [Operator Basic Concepts](zh/context/basic_concept.md) | Introduces basic concepts and terminology in the operator domain, such as quantization/sparse, data types, data formats, and so on. |
72+| [build Parameter Description](zh/install/build.md) | Introduces the functions and parameter values of the build.sh script in this project, including source code compilation, operator invocation, debugging, and so on. |
The file is empty
@@ -0,0 +1,161 @@
1+# aclnn Return Codes
2+ 
3+When calling aclnn APIs, common interface return codes are shown in [Table 1](#table1).
4+For abnormal status code values, you can use the aclGetRecentErrMsg interface (refer to [ACL API (C)](https://www.hiascend.com/document/detail/en/canncommercial/latest/API/appdevgapi/aclcppdevg_03_0004.html)) to obtain exception information. You can troubleshoot the problem based on the error message or contact technical support.
5+ 
6+**Table 1** Return Status Codes
7+ 
8+<a name="table1"></a>
9+<table><thead align="left"><tr><th class="cellrowborder" valign="top" width="30.543054305430545%" id="mcps1.2.4.1.1"><p>Status Code Name</p>
10+</th>
11+<th class="cellrowborder" valign="top" width="15.971597159715973%" id="mcps1.2.4.1.2"><p>Status Code Value</p>
12+</th>
13+<th class="cellrowborder" valign="top" width="53.48534853485349%" id="mcps1.2.4.1.3"><p>Status Code Description</p>
14+</th>
15+</tr>
16+</thead>
17+<tbody><tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_SUCCESS</p>
18+</td>
19+<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>0</p>
20+</td>
21+<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>Success.</p>
22+</td>
23+</tr>
24+<tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_PARAM_NULLPTR</p>
25+</td>
26+<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>161001</p>
27+</td>
28+<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>Parameter validation error, illegal nullptr exists in parameters.</p>
29+</td>
30+</tr>
31+<tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_PARAM_INVALID</p>
32+</td>
33+<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>161002</p>
34+</td>
35+<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>Parameter validation error, such as two input data types not satisfying the input type promotion relationship.</p>
36+</td>
37+</tr>
38+<tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_RUNTIME_ERROR</p>
39+</td>
40+<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>361001</p>
41+</td>
42+<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>API internally calls npu runtime interface abnormally.</p>
43+</td>
44+</tr>
45+<tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_XXX</p>
46+</td>
47+<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>561xxx</p>
48+</td>
49+<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>API internal exception occurred.</p>
50+ 
51+</td>
52+</tr>
53+</tbody>
54+</table>
55+ 
56+For more information about ACLNN_ERR_INNER_XXX status codes, see [Table 2](#table2).
57+ 
58+**Table 2** Exception Status Codes
59+ 
60+<a name="table2"></a>
61+<table><thead align="left"><tr><th class="cellrowborder" valign="top" width="30.183018301830185%" id="mcps1.2.4.1.1"><p>Status Code Name</p>
62+</th>
63+<th class="cellrowborder" valign="top" width="16.521652165216523%" id="mcps1.2.4.1.2"><p>Status Code Value</p>
64+</th>
65+<th class="cellrowborder" valign="top" width="53.295329532953296%" id="mcps1.2.4.1.3"><p>Status Code Description</p>
66+</th>
67+</tr>
68+</thead>
69+<tbody><tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER</p>
70+</td>
71+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561000</p>
72+</td>
73+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal exception occurred.</p>
74+</td>
75+</tr>
76+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_INFERSHAPE_ERROR</p>
77+</td>
78+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561001</p>
79+</td>
80+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal output shape deduction error occurred.</p>
81+</td>
82+</tr>
83+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_TILING_ERROR</p>
84+</td>
85+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561002</p>
86+</td>
87+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal tiling for npu kernel exception occurred.</p>
88+</td>
89+</tr>
90+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_FIND_KERNEL_ERROR</p>
91+</td>
92+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561003</p>
93+</td>
94+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal npu kernel lookup exception (possibly because operator binary package is not installed).</p>
95+</td>
96+</tr>
97+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_CREATE_EXECUTOR</p>
98+</td>
99+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561101</p>
100+</td>
101+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal aclOpExecutor creation failed (possibly due to operating system exception).</p>
102+</td>
103+</tr>
104+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_NOT_TRANS_EXECUTOR</p>
105+</td>
106+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561102</p>
107+</td>
108+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal uniqueExecutor ReleaseTo not called.</p>
109+</td>
110+</tr>
111+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_NULLPTR</p>
112+</td>
113+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561103</p>
114+</td>
115+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, nullptr exception appeared.</p>
116+</td>
117+</tr>
118+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_WRONG_ATTR_INFO_SIZE</p>
119+</td>
120+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561104</p>
121+</td>
122+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, operator attribute count exception.</p>
123+</td>
124+</tr>
125+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_KEY_CONFILICT</p>
126+</td>
127+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561105</p>
128+</td>
129+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, operator kernel matching hash key conflict.</p>
130+</td>
131+</tr>
132+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_INVALID_IMPL_MODE</p>
133+</td>
134+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561106</p>
135+</td>
136+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, operator implementation mode parameter error.</p>
137+</td>
138+</tr>
139+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_OPP_PATH_NOT_FOUND</p>
140+</td>
141+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561107</p>
142+</td>
143+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, environment variable ASCEND_OPP_PATH to be configured not detected.</p>
144+</td>
145+</tr>
146+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_LOAD_JSON_FAILED</p>
147+</td>
148+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561108</p>
149+</td>
150+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, failed to load operator information json file in operator kernel library.</p>
151+</td>
152+</tr>
153+<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_JSON_VALUE_NOT_FOUND</p>
154+</td>
155+<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561109</p>
156+</td>
157+<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, failed to load a field in operator information json file in operator kernel library.</p>
158+</td>
159+</tr>
160+</tbody>
161+</table>
@@ -0,0 +1,12 @@
1+# Basic Concepts
2+ 
3+ - [Two Phase API](./two_phase_api.md)
4+ - [Data Structure](./data_structure.md)
5+ - [Data Type](./data_type.md)
6+ - [Data Format](./data_format.md)
7+ - [Non Contiguous Tensor](./non_contiguous_tensor.md)
8+ - [Broadcast Relationship](./broadcast_relationship.md)
9+ - [Decuction Relationship](./decuction_relationship.md)
10+ - [Conversion Relationship](./conversion_relationship.md)
11+ - [Quant more Introduction](./quant_more_introduction.md)
12+ - [Sparse Mode Introduction](./sparse_mode_introduction.md)
@@ -0,0 +1,55 @@
1+# Broadcast Relationships
2+ 
3+## Broadcast Concept
4+ 
5+Broadcast describes how operators handle tensors (or arrays) of different shapes during computation. In most cases, tensors (or arrays) of different shapes are allowed to automatically expand their shapes during element operations to make their dimensions compatible. Usually, smaller tensors (or arrays) are "broadcast" to larger tensors (or arrays).
6+ 
7+Currently, many CANN operator API parameter shapes support broadcasting, which can appropriately improve calculation efficiency and reduce memory usage (especially in large-scale data scenarios). For more detailed broadcast technology introduction, refer to the [NumPy](https://numpy.org/doc/stable/user/basics.broadcasting.html) official website.
8+ 
9+## Broadcast Rules
10+ 
11+When performing broadcast calculations, you generally need to understand the following rules:
12+ 
13+- Rule 1: If the number of dimensions between arrays is inconsistent, all arrays align to the array with the longest shape, and the insufficient part of the shape is padded with 1 on the **left** until the number of dimensions is the same.
14+ 
15+ > Note:
16+ > - Example 1: Number of Dimensions refers to the dimension count of the tensor (or array) corresponding to the shape. For example, x.shape=(1,1,2,4), the number of dimensions is 4.
17+ > - Example 2: For example, when calculating a+b, where a.shape=(2, 2, 3) and b.shape=(2, 3), array b will be broadcast to b.shape=(1, 2, 3).
18+ 
19+- Rule 2: If the number of dimensions between arrays is consistent, and a certain dimension of an array is 1, then the array with dimension 1 will be stretched to match the corresponding dimension shape of the other array.
20+ 
21+ > Note:
22+ > In this scenario, you only need to ensure broadcasting in a certain dimension. For example, when calculating a+b, where a.shape=(1, 3) and b.shape=(3, 1), both arrays will be broadcast to a.shape=(3, 3) and b.shape=(3, 3).
23+ 
24+- Rule 3: If the number of dimensions between arrays is inconsistent, and neither has a dimension equal to 1, an error will be reported.
25+ 
26+Based on the above rules, the broadcast process generally first expands dimensions according to **Rule 1**, and then stretches the shape according to **Rule 2**. Specific examples are as follows:
27+ 
28+```text
29+Assuming a.shape=(2,2,3), values look like:
30+[[[1 2 3],[4 5 6]],
31+ [[1 2 3],[4 5 6]]]
32+Assuming b.shape=(2,3), values look like:
33+[[1 2 3],
34+ [-1 -2 -3]]
35+According to Rule 1, expand dimensions, b.shape=(1,2,3), values are:
36+[[[1 2 3],
37+ [-1 -2 -3]]]
38+According to Rule 2, stretch shape, b.shape=(2,2,3), values are:
39+[[[1 2 3],[-1 -2 -3]],
40+ [[1 2 3],[-1 -2 -3]]]
41+Calculate a+b, actual result is:
42+ [[[2 4 6],[3 3 3]],
43+ [[2 4 6],[3 3 3]]]
44+```
45+ 
46+## Limitations
47+ 
48+When the data types of two inputs a and b that satisfy the broadcast relationship, or the deduced data types, are among COMPLEX64, COMPLEX128, DOUBLE, INT16, UINT16, or UINT64, in addition to satisfying the above broadcast rules, the following conditions must also be met, otherwise the broadcast will fail and cause the operator execution to report an error.
49+ 
50+Condition: The merged dimension of consecutive axes that need broadcasting and consecutive axes that do not need broadcasting must be less than 6.
51+ 
52+Examples:
53+ 
54+- When a.shape=(5, 1, 5, 1, 5, 1) and b.shape=(5, 5, 5, 5, 5, 5), there are no axes that need to be merged, the final dimension is 6, and the broadcast reports an error.
55+- When a.shape=(5, 1, 5, 5, 1, 1) and b.shape=(5, 5, 5, 5, 5, 5), broadcasting is not needed in dimensions 2 and 3, and broadcasting is needed in dimensions 4 and 5. They are merged separately and continuously, and the merged dimension is 4, so the broadcast succeeds.
@@ -0,0 +1,143 @@
1+# Compilation and Running Examples
2+ 
3+## Prerequisites
4+ 
5+- If you need to compile and execute operator APIs, ensure that the basic environment has been set up, including driver, firmware, CANN software package, ops package, etc.
6+- For the operator API calling process and compilation and running operations, refer to [Application Development (C&C++)](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/programug/acldevg/aclcppdevg_000006.html) under "Single Operator Invocation > Single Operator API Execution > Calling aclnn Interface Example Code".
7+ 
8+## Pre-compilation Preparation
9+ 
10+This chapter takes the development and runtime environment co-location scenario as an example, that is, the machine with AI processor serves as both the development environment and the runtime environment. In this scenario, code development and code running are on the same machine. Here we take the **AddMatMul operator** as an example. The calling logic, process, and compilation script of other operators are roughly the same as the AddMatMul operator. Please modify the API calling script (*.cpp) and compilation script (CMakeLists) according to the actual situation.
11+ 
12+- **Example Code**
13+ 
14+ The AddMatMul operator implements tensor addition operation, and the calculation formula is: out = β * self + α * (mat1 @ mat2). You can obtain the example code from the "Calling Example" section in [aclnnAddmm&aclnnInplaceAddmm.md](../../../matmul/mat_mul_v3/docs/aclnnAddmm&aclnnInplaceAddmm.md) and name the code file "**test\_addmm.cpp**".
15+ 
16+- **CMakeLists File**
17+ 
18+ The CMake file example is as follows. Please modify according to the actual situation:
19+ 
20+ ```bash
21+ # Copyright (c) Huawei Technologies Co., Ltd. 2025. All rights reserved.
22+ 
23+ # CMake lowest version requirement
24+ cmake_minimum_required(VERSION 3.14)
25+ 
26+ # Set project name
27+ project(ACLNN_EXAMPLE)
28+ 
29+ # Compile options
30+ add_compile_options(-std=c++11)
31+ 
32+ # Set compilation options
33+ set(CMAKE_RUNTIME_OUTPUT_DIRECTORY "./bin")
34+ set(CMAKE_CXX_FLAGS_DEBUG "-fPIC -O0 -g -Wall")
35+ set(CMAKE_CXX_FLAGS_RELEASE "-fPIC -O2 -Wall")
36+ 
37+ # Set executable file name (such as opapi_test) and specify the directory where the operator file *.cpp to be run is located
38+ add_executable(opapi_test
39+ test_addmm.cpp)
40+ 
41+ # Set ASCEND_PATH (CANN software package directory, please modify according to the actual path) and INCLUDE_BASE_DIR (header file directory)
42+ if(NOT "$ENV{ASCEND_CUSTOM_PATH}" STREQUAL "")
43+ set(ASCEND_PATH $ENV{ASCEND_CUSTOM_PATH})
44+ else()
45+ set(ASCEND_PATH "/usr/local/Ascend/cann")
46+ endif()
47+ set(INCLUDE_BASE_DIR "${ASCEND_PATH}/include")
48+ include_directories(
49+ ${INCLUDE_BASE_DIR}
50+ ${INCLUDE_BASE_DIR}/aclnn
51+ )
52+ 
53+ # Set linked library file path
54+ target_link_libraries(opapi_test PRIVATE
55+ ${ASCEND_PATH}/lib64/libascendcl.so
56+ ${ASCEND_PATH}/lib64/libnnopbase.so
57+ ${ASCEND_PATH}/lib64/libopapi_math.so
58+ ${ASCEND_PATH}/lib64/libopapi_nn.so)
59+ 
60+ # The executable file is in the bin directory under the CMakeLists file directory
61+ install(TARGETS opapi_test DESTINATION ${CMAKE_RUNTIME_OUTPUT_DIRECTORY})
62+ ```
63+ 
64+ For operators that combine collective communication and MatMul calculation, and run in parallel, they are collectively called MC2 operators (communication-computation fusion operators), including AllGatherMatmul, AlltoAllAllGatherBatchMatMul, BatchMatMulReduceScatterAlltoAll, MatmulAllReduce, MatmulAllReduceAddRmsNorm, MatmulReduceScatter, etc. When calling such operator APIs, multi-threading and HCCL (Huawei Collective Communication Library) are generally involved. Therefore, the CMake file needs to additionally import the following content, otherwise compilation will fail.
65+ 
66+ ```bash
67+ # Set linked library file path
68+ find_package(Threads REQUIRED)
69+ target_link_libraries(opapi_test PRIVATE
70+ ${ASCEND_PATH}/lib64/libascendcl.so
71+ ${ASCEND_PATH}/lib64/libnnopbase.so
72+ ${ASCEND_PATH}/lib64/libopapi_math.so
73+ ${ASCEND_PATH}/lib64/libopapi_nn.so
74+ ${ASCEND_PATH}/lib64/libhccl.so # Collective communication library file
75+ ${CMAKE_THREAD_LIBS_INIT}) # Library file that multi-threading depends on
76+ ```
77+ 
78+ Where "find_package(Threads REQUIRED)" is a CMake command used to find the thread library, which can automatically link the header files or indirectly dependent library files that the thread library depends on.
79+ 
80+## Compilation and Running
81+ 
82+ 1. Prepare the operator calling code (*.cpp) and compilation script (CMakeLists.txt) in advance.
83+ 2. Configure environment variables.
84+ 
85+ After installing the CANN software, log in to the environment as the CANN runtime user and execute the following command to make the environment variables effective.
86+ 
87+ ```bash
88+ source ${INSTALL_DIR}/set_env.sh
89+ ```
90+ 
91+ Where ${INSTALL_DIR} is the storage path after CANN software installation. Please replace according to the actual situation.
92+ 3. Compile and run.
93+ - Enter the directory where CMakeLists.txt is located and execute the following command to create a new build directory to store the generated compilation files.
94+ 
95+ ```bash
96+ mkdir -p build
97+ ```
98+ 
99+ - Enter the build directory, execute the cmake command to compile, and then execute the make command to generate the executable file.
100+ 
101+ ```bash
102+ cd build
103+ cmake ../ -DCMAKE_CXX_COMPILER=g++ -DCMAKE_SKIP_RPATH=TRUE
104+ make
105+ ```
106+ 
107+ After successful compilation, the opapi\_test executable file will be generated in the bin folder under the build directory.
108+ 
109+ - Enter the bin directory and run the executable file opapi_test.
110+ 
111+ ```bash
112+ cd bin
113+ ./opapi_test
114+ ```
115+ 
116+ Taking the running result of the AddMatMul operator as an example, the result after running is shown below:
117+ 
118+ ```bash
119+ result[0] is: 1.200000
120+ result[1] is: 2.200000
121+ result[2] is: 3.200000
122+ result[3] is: 5.400000
123+ result[4] is: 6.400000
124+ result[5] is: 7.400000
125+ result[6] is: 9.600000
126+ result[7] is: 10.600000
127+ ```
128+ 
129+ If the execution result reports an error and the expected result does not appear, you can use the aclGetRecentErrMsg interface to obtain the specific error information.
130+ Example of obtaining exception information when calling aclnnAddmmGetWorkspaceSize fails:
131+ 
132+ ```bash
133+ // self is nullptr
134+ ret = aclnnAddmmGetWorkspaceSize(self, mat1, mat2, beta, alpha, out, cubeMathType, &workspaceSize, &executor);
135+ CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnAddmmGetWorkspaceSize failed. ERROR: %d\n[ERROR msg]%s", ret, aclGetRecentErrMsg()); return ret);
136+ ```
137+ 
138+ The above null pointer construction problem obtains error information as shown below:
139+ 
140+ ```bash
141+ aclnnAddmmGetWorkspaceSize failed. ERROR: 161001
142+ [ERROR msg][PID:xxxx] xxx(timesamp) AclNN_Parameter_Error(EZ1001): Expected a proper Tensor but got null for argument addmmTennsor.self.
143+ ```
@@ -0,0 +1,15 @@
1+# Type Conversion Relationships
2+ 
3+When the **output aclTensor data type** of an API (such as aclnnAdd, aclnnMul, etc.) is inconsistent with the **calculation type after input data type promotion**, the API internally converts the calculation result to the data type corresponding to the output type.
4+ 
5+Data type conversion must satisfy the following rules. Conversions that do not satisfy the rules cannot be performed, and parameter validation will fail when calling the API.
6+ 
7+ - Floating-point types: ACL\_FLOAT16, ACL\_FLOAT, ACL\_DOUBLE, ACL\_BF16.
8+ - Integer types: ACL\_INT8, ACL\_UINT8, ACL\_INT16, ACL\_UINT16, ACL\_INT32, ACL\_UINT32, ACL\_INT64, ACL\_UINT64.
9+ - Complex types: ACL\_COMPLEX64, ACL\_COMPLEX128.
10+ - Conversions between integer types are supported, as well as conversions to floating-point and complex types.
11+ - Conversions between floating-point types are supported, as well as conversions to complex types.
12+ - Conversions between complex types are supported.
13+ - BOOL supports conversion to integer, floating-point, and complex types.
14+ 
15+Except for the above scenarios, other conversion scenarios are not supported.
@@ -0,0 +1,36 @@
1+# Data Formats
2+ 
3+Data format (format) is used to describe the business semantics of the axes of a multi-dimensional Tensor, representing the physical layout format of data, such as 1D, 2D, 3D, 4D, 5D, and so on. Generally, CNN (Convolutional Neural Networks) APIs require specific formats to be described.
4+ 
5+For the **full range of data formats** supported by aclTensor, refer to [ACL API (C)](https://www.hiascend.com/document/detail/en/canncommercial/latest/API/appdevgapi/aclcppdevg_03_0004.html) under "Data Types and Their Operation Interfaces > aclFormat".
6+ 
7+For an introduction to **data format layout principles**, refer to [Ascend C Operator Development Guide](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/programug/Ascendcopdevg/atlas_ascendc_map_10_0002.html) under "Concept Principles and Terminology > Neural Networks and Operators > Data Layout Formats".
8+ 
9+## Usage Instructions
10+ 
11+Currently, most operator APIs support the ND data format. For example, the aclnnAdd interface indicates that the supported data format is ND (that is, the rule of low-dimensional priority continuous layout for multi-dimensional Tensors). For aclnnConvolution, which is a CNN-type API, the input aclTensor is required to be set with a format that has business semantics, rather than the ND format. Such operators need to know the business semantics in the Tensor during the calculation process to perform the corresponding computation. For example, in 2D convolution, you need to know the correspondence between the Batch dimension, Channel dimension, Height dimension, Width dimension, and the Tensor dimensions.
12+ 
13+>**Note:**
14+>
15+>- For the parameter description of two-stage interfaces, to simplify the description, **the original data format "ACL\_FORMAT\_XXXX_" is abbreviated as "_XXXX_"**.
16+>- The meaning of each dimension in the data format: N (Batch) represents the batch size, H (Height) represents the feature map height, W (Width) represents the feature map width, C (Channels) represents the feature map channels, D (Depth) represents the feature map depth, L (Length) represents the feature map length.
17+ 
18+## Common Data Formats
19+ 
20+When creating an aclTensor through the **aclCreateTensor** interface, you need to set the data format according to the API business requirements. The **supported data formats** are:
21+ 
22+ACL\_FORMAT\_ND, ACL\_FORMAT\_NCHW, ACL\_FORMAT\_NHWC, ACL\_FORMAT\_HWCN, ACL\_FORMAT\_NDHWC, ACL\_FORMAT\_NCDHW, ACL\_FORMAT\_NC, ACL\_FORMAT\_NCL.
23+ 
24+For non-ND Tensors, the Tensor dimension requirements are consistent with the format description. For example:
25+ 
26+- 5D Tensor: Requires ACL\_FORMAT\_NCDHW, ACL\_FORMAT\_NDHWC, or ACL\_FORMAT\_ND (if the API parameter description does not indicate support for ND, setting the ND format will result in an API validation error).
27+- 4D Tensor: Requires ACL\_FORMAT\_NCHW, ACL\_FORMAT\_NHWC, ACL\_FORMAT\_HWCN, or ACL\_FORMAT\_ND.
28+- 3D Tensor: Requires ACL\_FORMAT\_NCL or ACL\_FORMAT\_ND.
29+- 2D Tensor: Requires ACL\_FORMAT\_NC or ACL\_FORMAT\_ND.
30+- Other dimension Tensors: Require ACL\_FORMAT\_ND.
31+ 
32+## Private Data Formats
33+ 
34+In addition to the common data formats mentioned above, there are other data formats, such as ACL\_FORMAT\_NC1HWC0, ACL\_FORMAT\_FRACTAL\_Z, ACL\_FORMAT\_NC1HWC0\_C04, ACL\_FORMAT\_FRACTAL\_NZ, ACL\_FORMAT\_NDC1HWC0, ACL\_FORMAT\_FRACTAL\_Z\_3D, and so on.
35+ 
36+These formats are private formats of the NPU. Currently, most aclnn APIs do not support these formats. If an individual API declares supported data formats, refer to the actual description of that API.
@@ -0,0 +1,79 @@
1+# Data Structures
2+ 
3+This chapter provides the basic data structures required for calling CANN operator APIs. **Developers do not need to focus on their internal implementation and can use them directly.**
4+ 
5+Note that these basic data structures can be created through the "Public Interfaces" section in [Operator Library Interface](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/API/aolapi/operatorlist_00001.html), such as aclCreateTensor.
6+ 
7+- **aclTensor**
8+ 
9+ A structure defined by the framework to manage and store tensor data (such as multi-dimensional data like vectors and matrices). You can create this object through the **aclCreateTensor** interface.
10+ 
11+ ```bash
12+ typedef struct aclTensor aclTensor
13+ ```
14+ 
15+- **aclScalar**
16+ 
17+ A structure defined by the framework to manage and store scalar data (that is, a single value). You can create this object through the **aclCreateScalar** interface.
18+ 
19+ ```bash
20+ typedef struct aclScalar aclScalar
21+ ```
22+ 
23+- **aclIntArray**
24+ 
25+ An array structure defined by the framework to manage and store integer data. You can create this object through the **aclCreateIntArray** interface.
26+ 
27+ ```bash
28+ typedef struct aclIntArray aclIntArray
29+ ```
30+ 
31+- **aclFloatArray**
32+ 
33+ An array structure defined by the framework to manage and store float32 data. You can create this object through the **aclCreateFloatArray** interface.
34+ 
35+ ```bash
36+ typedef struct aclFloatArray aclFloatArray
37+ ```
38+ 
39+- **aclBoolArray**
40+ 
41+ An array structure defined by the framework to manage and store boolean data. You can create this object through the **aclCreateBoolArray** interface.
42+ 
43+ ```bash
44+ typedef struct aclBoolArray aclBoolArray
45+ ```
46+ 
47+- **aclTensorList**
48+ 
49+ An array structure defined by the framework to manage and store multiple tensor data. You can create this object through the **aclCreateTensorList** interface.
50+ 
51+ ```bash
52+ typedef struct aclTensorList aclTensorList
53+ ```
54+ 
55+- **aclScalarList**
56+ 
57+ An array structure defined by the framework to manage and store scalar data. You can create this object through the **aclCreateScalarList** interface.
58+ 
59+ ```bash
60+ typedef struct aclScalarList aclScalarList
61+ ```
62+ 
63+- **aclOpExecutor**
64+ 
65+ An executor data structure defined by the framework, which is a container used to execute operator calculations.
66+ 
67+ Typically, when calling the first-stage interface aclxxXxxGetWorkspaceSize, the framework automatically creates an aclOpExecutor; after calling the second-stage interface aclxxXxx, the object is automatically released.
68+ 
69+ ```bash
70+ typedef struct aclOpExecutor aclOpExecutor
71+ ```
72+ 
73+- **aclrtStream**
74+ 
75+ A stream processing data structure defined by the framework, used to manage and maintain the execution order of some asynchronous operations.
76+ 
77+ ```bash
78+ typedef void *aclrtStream
79+ ```
@@ -0,0 +1,37 @@
1+# Data Types
2+ 
3+When creating an aclTensor through the **aclCreateTensor** interface, refer to [ACL API (C)](https://www.hiascend.com/document/detail/en/canncommercial/latest/API/appdevgapi/aclcppdevg_03_0004.html) for the full list of supported data types under "Data Types and Their Operation Interfaces > aclDataType".
4+ 
5+For the parameter description of two-stage interfaces, the supported data types will use the following abbreviated forms for convenience.
6+ 
7+**Table 1** Data Type Abbreviations
8+ 
9+| Original Data Type | Abbreviation (case-insensitive) |
10+| :---------------: | :----------------------: |
11+| ACL_FLOAT | FLOAT or FLOAT32 |
12+| ACL_FLOAT16 | FLOAT16 |
13+| ACL_INT8 | INT8 |
14+| ACL_INT32 | INT32 |
15+| ACL_UINT8 | UINT8 |
16+| ACL_INT16 | INT16 |
17+| ACL_UINT16 | UINT16 |
18+| ACL_UINT32 | UINT32 |
19+| ACL_INT64 | INT64 |
20+| ACL_UINT64 | UINT64 |
21+| ACL_DOUBLE | DOUBLE or FLOAT64 |
22+| ACL_BOOL | BOOL |
23+| ACL_STRING | STRING |
24+| ACL_COMPLEX64 | COMPLEX64 |
25+| ACL_COMPLEX128 | COMPLEX128 |
26+| ACL_BF16 | BF16 or BFLOAT16 |
27+| ACL_INT4 | INT4 |
28+| ACL_UINT1 | UINT1 |
29+| ACL_COMPLEX32 | COMPLEX32 |
30+| ACL_HIFLOAT8 | HIFLOAT8 |
31+| ACL_FLOAT8_E5M2 | FLOAT8_E5M2 |
32+| ACL_FLOAT8_E4M3FN | FLOAT8_E4M3FN |
33+| ACL_FLOAT8_E8M0 | FLOAT8_E8M0 |
34+| ACL_FLOAT6_E3M2 | FLOAT6_E3M2 |
35+| ACL_FLOAT6_E2M3 | FLOAT6_E2M3 |
36+| ACL_FLOAT4_E2M1 | FLOAT4_E2M1 |
37+| ACL_FLOAT4_E1M2 | FLOAT4_E1M2 |
@@ -0,0 +1,81 @@
1+# Type Promotion Relationships
2+ 
3+## Promotion Rules
4+ 
5+When the **input aclTensor data types** of an API (such as aclnnAdd, aclnnMul, etc.) are inconsistent, the API internally deduces a data type and converts the input data to that data type for calculation.
6+ 
7+For the data types supported by aclTensor, refer to [Data Types](./data_type.md). Some of these types satisfy the following promotion rules, and the promotion principle is similar to PyTorch's [Type Promotion](https://pytorch.org/docs/stable/tensor_attributes.html#type-promotion-doc).
8+ 
9+> Note:
10+>
11+> - For convenience of description, the data types used in the table are **abbreviated forms**, representing: ACL\_FLOAT(f32), ACL\_FLOAT16(f16), ACL\_DOUBLE(f64), ACL\_BF16(bf16), ACL\_INT8(s8), ACL\_UINT8(u8), ACL\_INT16(s16), ACL\_UINT16(u16), ACL\_INT32(s32), ACL\_UINT32(u32), ACL\_INT64(s64), ACL\_UINT64(u64), ACL\_BOOL(bool), ACL\_COMPLEX32(c32), ACL\_COMPLEX64(c64), ACL\_COMPLEX128(c128).
12+> - The table header and the leftmost column represent the two input data types to be deduced, and the corresponding position in the table represents the deduced data type.
13+> - The cross mark (×) in the table indicates that these two types cannot perform promotion calculation.
14+ 
15+**Table 1** Data Type Promotion Relationships
16+ 
17+| Data Type | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | bool | c32 | c64 | c128 |
18+| :------: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: |
19+| **f32** | f32 | f32 | f64 | f32 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c64 | c64 | c128 |
20+| **f16** | f32 | f16 | f64 | f32 | f16 | f16 | f16 | × | f16 | × | f16 | × | f16 | c32 | c64 | c128 |
21+| **f64** | f64 | f64 | f64 | f64 | f64 | f64 | f64 | × | f64 | × | f64 | × | f64 | c128 | c128 | c128 |
22+| **bf16** | f32 | f32 | f64 | bf16 | bf16 | bf16 | bf16 | × | bf16 | × | bf16 | × | bf16 | c32 | c64 | c128 |
23+| **s8** | f32 | f16 | f64 | bf16 | s8 | s16 | s16 | × | s32 | × | s64 | × | s8 | c32 | c64 | c128 |
24+| **u8** | f32 | f16 | f64 | bf16 | s16 | u8 | s16 | × | s32 | × | s64 | × | u8 | c32 | c64 | c128 |
25+| **s16** | f32 | f16 | f64 | bf16 | s16 | s16 | s16 | × | s32 | × | s64 | × | s16 | c32 | c64 | c128 |
26+| **u16** | × | × | × | × | × | × | × | u16 | × | × | × | × | × | × | × | × |
27+| **s32** | f32 | f16 | f64 | bf16 | s32 | s32 | s32 | × | s32 | × | s64 | × | s32 | c32 | c64 | c128 |
28+| **u32** | × | × | × | × | × | × | × | × | × | u32 | × | × | × | × | × | × |
29+| **s64** | f32 | f16 | f64 | bf16 | s64 | s64 | s64 | × | s64 | × | s64 | × | s64 | c32 | c64 | c128 |
30+| **u64** | × | × | × | × | × | × | × | × | × | × | × | u64 | × | × | × | × |
31+| **bool** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | × | s32 | × | s64 | × | bool | c32 | c64 | c128 |
32+| **c32** | c64 | c32 | c128 | c32 | c32 | c32 | c32 | × | c32 | × | c32 | × | c32 | c32 | c64 | c128 |
33+| **c64** | c64 | c64 | c128 | c64 | c64 | c64 | c64 | × | c64 | × | c64 | × | c64 | c64 | c64 | c128 |
34+| **c128** | c128 | c128 | c128 | c128 | c128 | c128 | c128 | × | c128 | × | c128 | × | c128 | c128 | c128 | c128 |
35+ 
36+## Promotion Examples
37+ 
38+- When calling the aclnnAdd interface, if the data types of the input parameters are inconsistent, one is float16 and one is float32, the API internally converts the float16 data type to float32 data type and then performs the calculation.
39+- When calling the aclnnAdd interface, if the data types of the input parameters are inconsistent, one is float32 and one is bool, the API internally converts the bool data type to float32 data type and then performs the calculation.
40+ 
41+## TensorScalar Promotion Relationships
42+ 
43+### Promotion Rules
44+ 
45+When the **input Tensor data type** and **input Scalar data type** of an API (such as aclnnAdds, aclnnMuls, and so on) are inconsistent, the API internally deduces a data type and converts the input data to that data type for calculation.
46+ 
47+For the data types supported by aclTensor, refer to [Data Types](./data_type.md). Some of these types satisfy the following promotion rules, and the promotion principle is similar to PyTorch's [Type Promotion](https://pytorch.org/docs/stable/tensor_attributes.html#type-promotion-doc).
48+ 
49+The type promotion rules are as follows:
50+ 
51+> Note:
52+>
53+> - For convenience of description, the data types used in the table are abbreviated forms, representing: ACL\_FLOAT(f32), ACL\_FLOAT16(f16), ACL\_DOUBLE(f64), ACL\_BF16(bf16), ACL\_INT8(s8), ACL\_UINT8(u8), ACL\_INT16(s16), ACL\_UINT16(u16), ACL\_INT32(s32), ACL\_UINT32(u32), ACL\_INT64(s64), ACL\_UINT64(u64), ACL\_BOOL(bool), ACL\_COMPLEX32(c32), ACL\_COMPLEX64(c64), ACL\_COMPLEX128(c128).
54+> - The table header represents the input Tensor data type to be deduced, and the leftmost column represents the input Scalar data type to be deduced. The corresponding position in the table represents the deduced data type.
55+> - The cross mark (×) in the table indicates that these two types cannot perform promotion calculation.
56+ 
57+**Table 2** Data Type Promotion Relationships
58+ 
59+| Data Type | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | bool | c32 | c64 | c128 |
60+| :------: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: |
61+| **f32** | f32 | f16 | f64 | bf16 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c32 | c64 | c128 |
62+| **f16** | f32 | f16 | f64 | bf16 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c32 | c64 | c128 |
63+| **f64** | f32 | f16 | f64 | bf16 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c128 | c128 | c128 |
64+| **bf16** | f32 | f16 | f64 | bf16 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c32 | c64 | c128 |
65+| **s8** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | s8 | c32 | c64 | c128 |
66+| **u8** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | u8 | c32 | c64 | c128 |
67+| **s16** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | s16 | c32 | c64 | c128 |
68+| **u16** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | × | c32 | c64 | c128 |
69+| **s32** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | s32 | c32 | c64 | c128 |
70+| **u32** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | × | c32 | c64 | c128 |
71+| **s64** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | s64 | c32 | c64 | c128 |
72+| **u64** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | × | c32 | c64 | c128 |
73+| **bool** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | bool | c32 | c64 | c128 |
74+| **c32** | c64 | c32 | c128 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c32 | c64 | c128 |
75+| **c64** | c64 | c32 | c128 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c32 | c64 | c128 |
76+| **c128** | c64 | c32 | c128 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c32 | c64 | c128 |
77+ 
78+### Promotion Examples
79+ 
80+ - If the input Tensor data type is float16 and the input Scalar data type is float32, the API internally converts the input Scalar float32 data type to float16 data type and then performs the calculation.
81+ - If the input Tensor data type is bool and the input Scalar data type is float32, the API internally converts the input Tensor bool data type to float32 data type and then performs the calculation.
@@ -0,0 +1,38 @@
1+# Non-contiguous Tensor
2+ 
3+Currently, most operator APIs support "**non-contiguous Tensor**" as input aclTensor, that is, a Tensor can be represented by (shape, strides, offset).
4+ 
5+Note: You can create an aclTensor through the "Public Interfaces > aclCreateTensor" section in [Operator Library Interface](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/API/aolapi/operatorlist_00001.html).
6+ 
7+## Example 1
8+ 
9+For example, consider a Tensor with shape=(6, 5), strides=(10, 1), and offset=22. Its memory layout is as follows:
10+> a<sub>0,0</sub> , a<sub>0,1</sub> , a<sub>0,2</sub> , a<sub>0,3</sub> , a<sub>0,4</sub> , a<sub>0,5</sub> , a<sub>0,6</sub> , a<sub>0,7</sub> , a<sub>0,8</sub> , a<sub>0,9</sub>
11+> a<sub>1,0</sub> , a<sub>1,1</sub> , a<sub>1,2</sub> , a<sub>1,3</sub> , a<sub>1,4</sub> , a<sub>1,5</sub> , a<sub>1,6</sub> , a<sub>1,7</sub> , a<sub>1,8</sub> , a<sub>1,9</sub>
12+> a<sub>2,0</sub> , a<sub>2,1</sub> , **a<sub>2,2</sub> , a<sub>2,3</sub> , a<sub>2,4</sub> , a<sub>2,5</sub> , a<sub>2,6</sub>** , a<sub>2,7</sub> , a<sub>2,8</sub> , a<sub>2,9</sub>
13+> a<sub>3,0</sub> , a<sub>3,1</sub> , **a<sub>3,2</sub> , a<sub>3,3</sub> , a<sub>3,4</sub> , a<sub>3,5</sub> , a<sub>3,6</sub>** , a<sub>3,7</sub> , a<sub>3,8</sub> , a<sub>3,9</sub>
14+> a<sub>4,0</sub> , a<sub>4,1</sub> , **a<sub>4,2</sub> , a<sub>4,3</sub> , a<sub>4,4</sub> , a<sub>4,5</sub> , a<sub>4,6</sub>** , a<sub>4,7</sub> , a<sub>4,8</sub> , a<sub>4,9</sub>
15+> a<sub>5,0</sub> , a<sub>5,1</sub> , **a<sub>5,2</sub> , a<sub>5,3</sub> , a<sub>5,4</sub> , a<sub>5,5</sub> , a<sub>5,6</sub>** , a<sub>5,7</sub> , a<sub>5,8</sub> , a<sub>5,9</sub>
16+> a<sub>6,0</sub> , a<sub>6,1</sub> , **a<sub>6,2</sub> , a<sub>6,3</sub> , a<sub>6,4</sub> , a<sub>6,5</sub> , a<sub>6,6</sub>** , a<sub>6,7</sub> , a<sub>6,8</sub> , a<sub>6,9</sub>
17+> a<sub>7,0</sub> , a<sub>7,1</sub> , **a<sub>7,2</sub> , a<sub>7,3</sub> , a<sub>7,4</sub> , a<sub>7,5</sub> , a<sub>7,6</sub>** , a<sub>7,7</sub> , a<sub>7,8</sub> , a<sub>7,9</sub>
18+> a<sub>8,0</sub> , a<sub>8,1</sub> , a<sub>8,2</sub> , a<sub>8,3</sub> , a<sub>8,4</sub> , a<sub>8,5</sub> , a<sub>8,6</sub> , a<sub>8,7</sub> , a<sub>8,8</sub> , a<sub>8,9</sub>
19+> a<sub>9,0</sub> , a<sub>9,1</sub> , a<sub>9,2</sub> , a<sub>9,3</sub> , a<sub>9,4</sub> , a<sub>9,5</sub> , a<sub>9,6</sub> , a<sub>9,7</sub> , a<sub>9,8</sub> , a<sub>9,9</sub>
20+ 
21+That is, the Tensor is laid out in the dark positions shown above. This complete Tensor is non-contiguous in memory layout. Strides describe the interval between two adjacent elements in the Tensor dimension. If the stride in dimension 1 is 1, that dimension is contiguous; if the stride in dimension 0 is 10, then adjacent elements are separated by 10 elements, which is non-contiguous. Offset represents the offset of the first element of this Tensor relative to addr.
22+ 
23+## Example 2
24+ 
25+For example, consider a Tensor with shape=(4, 3), strides=(20, 2), and offset=22. Its memory layout is as follows:
26+ 
27+> a<sub>0,0</sub> , a<sub>0,1</sub> , a<sub>0,2</sub> , a<sub>0,3</sub> , a<sub>0,4</sub> , a<sub>0,5</sub> , a<sub>0,6</sub> , a<sub>0,7</sub> , a<sub>0,8</sub> , a<sub>0,9</sub>
28+> a<sub>1,0</sub> , a<sub>1,1</sub> , a<sub>1,2</sub> , a<sub>1,3</sub> , a<sub>1,4</sub> , a<sub>1,5</sub> , a<sub>1,6</sub> , a<sub>1,7</sub> , a<sub>1,8</sub> , a<sub>1,9</sub>
29+> a<sub>2,0</sub> , a<sub>2,1</sub> , **a<sub>2,2</sub>** , a<sub>2,3</sub> , **a<sub>2,4</sub>** , a<sub>2,5</sub> , **a<sub>2,6</sub>** , a<sub>2,7</sub> , a<sub>2,8</sub> , a<sub>2,9</sub>
30+> a<sub>3,0</sub> , a<sub>3,1</sub> , a<sub>3,2</sub> , a<sub>3,3</sub> , a<sub>3,4</sub> , a<sub>3,5</sub> , a<sub>3,6</sub> , a<sub>3,7</sub> , a<sub>3,8</sub> , a<sub>3,9</sub>
31+> a<sub>4,0</sub> , a<sub>4,1</sub> , **a<sub>4,2</sub>** , a<sub>4,3</sub> , **a<sub>4,4</sub>** , a<sub>4,5</sub> , **a<sub>4,6</sub>** , a<sub>4,7</sub> , a<sub>4,8</sub> , a<sub>4,9</sub>
32+> a<sub>5,0</sub> , a<sub>5,1</sub> , a<sub>5,2</sub> , a<sub>5,3</sub> , a<sub>5,4</sub> , a<sub>5,5</sub> , a<sub>5,6</sub> , a<sub>5,7</sub> , a<sub>5,8</sub> , a<sub>5,9</sub>
33+> a<sub>6,0</sub> , a<sub>6,1</sub> , **a<sub>6,2</sub>** , a<sub>6,3</sub> , **a<sub>6,4</sub>** , a<sub>6,5</sub> , **a<sub>6,6</sub>** , a<sub>6,7</sub> , a<sub>6,8</sub> , a<sub>6,9</sub>
34+> a<sub>7,0</sub> , a<sub>7,1</sub> , a<sub>7,2</sub> , a<sub>7,3</sub> , a<sub>7,4</sub> , a<sub>7,5</sub> , a<sub>7,6</sub> , a<sub>7,7</sub> , a<sub>7,8</sub> , a<sub>7,9</sub>
35+> a<sub>8,0</sub> , a<sub>8,1</sub> , **a<sub>8,2</sub>** , a<sub>8,3</sub> , **a<sub>8,4</sub>** , a<sub>8,5</sub> , **a<sub>8,6</sub>** , a<sub>8,7</sub> , a<sub>8,8</sub> , a<sub>8,9</sub>
36+> a<sub>9,0</sub> , a<sub>9,1</sub> , a<sub>9,2</sub> , a<sub>9,3</sub> , a<sub>9,4</sub> , a<sub>9,5</sub> , a<sub>9,6</sub> , a<sub>9,7</sub> , a<sub>9,8</sub> , a<sub>9,9</sub>
37+ 
38+That is, the Tensor is laid out in the dark positions shown above. This complete Tensor is non-contiguous in memory layout. Strides describe the interval between two adjacent elements in the Tensor dimension. If the stride in dimension 1 is 2, that dimension has an interval of 1 element; if the stride in dimension 0 is 20, then adjacent elements are separated by 20 elements, which is non-contiguous. Offset represents the offset of the first element of this Tensor relative to addr.
@@ -0,0 +1,59 @@
1+# Quantization Introduction
2+ 
3+Quantization is widely used in deep learning models, especially during inference. Through quantization, models can run more efficiently on hardware, reducing computational resource consumption and accelerating the inference process, while also reducing the model's storage requirements.
4+ 
5+CANN operator quantization refers to the computational process of converting the input tensors of matrix (cube) operators such as Matmul in neural networks from high-bit to low-bit representation, while generating the corresponding quantization parameter scale. After the low-bit cube computation is complete, the quantization parameter scale can convert the low-bit values back to high-bit values, thereby ensuring the correctness of the overall computation results (the effect is approximately equivalent to directly using high-bit computation) and effectively improving computational efficiency.
6+ 
7+- Static quantization: Uses predetermined quantization parameters for quantization. In inference scenarios, the quantization of weights generally uses static quantization, which provides better operator performance.
8+- Dynamic quantization: Uses input data to compute quantization parameters online for quantization. In inference scenarios, the quantization of activations generally uses dynamic quantization, which better adapts to data variations and provides higher precision. In training scenarios, dynamic quantization is also generally used to improve quantization precision. Note that dynamic quantization results in slightly worse operator performance because quantization parameters are generated online.
9+ 
10+## Quantization Modes
11+ 
12+Quantization modes (also known as quantization granularity) refer to the use of different quantization computation levels for different input tensors of an operator. Common quantization computation modes include:
13+ 
14+> Description:
15+>
16+> - The m, n, and k variables represent the sizes of different axes for tensor computation.
17+> - The left matrix and right matrix refer to the two input tensors used for matrix multiplication computation in cube operators. Generally, the left matrix represents the activation and the right matrix represents the weight. Interpret and use them according to the actual situation.
18+ 
19+- Pertensor quantization (abbreviated as T quantization): The quantization target can be either the left matrix or the right matrix. Each tensor shares the same quantization parameter.
20+ 
21+ Assuming the left matrix shape is (m, k) and the right matrix shape is (k, n), where k is the reduce axis, the shape of the generated quantization parameter is (1, ).
22+ 
23+ ![Schematic diagram](../figures/pertensor_quantization.png)
24+ 
25+- Perchannel quantization (abbreviated as C quantization): The quantization target is the right matrix. Each channel uses an independent quantization parameter.
26+ 
27+ Assuming the right matrix shape is (k, n), where k is the reduce axis, the shape of the generated quantization parameter is (n, ).
28+ 
29+ ![Schematic diagram](../figures/perchannel_quantization.png)
30+ 
31+- Pertoken quantization (abbreviated as K quantization): The quantization target is the left matrix. Each token uses an independent quantization parameter.
32+ 
33+ Assuming the left matrix shape is (m, k), where k is the reduce axis, the shape of the generated quantization parameter is (m, ).
34+ 
35+ ![Schematic diagram](../figures/pertoken_quantization.png)
36+ 
37+- Pergroup quantization (abbreviated as G quantization): The quantization target can be either the left matrix or the right matrix. Data is grouped on the reduce axis, and each group uses an independent quantization parameter.
38+ - Assuming the left matrix shape is (m, k), where k is the reduce axis, data is grouped on the k axis with a group size of gs, and the shape of the generated quantization parameter is (m, k/gs).
39+ - Assuming the right matrix shape is (k, n), where k is the reduce axis, data is grouped on the k axis with a group size of gs, and the shape of the generated quantization parameter is (k/gs, n).
40+ 
41+ ![Schematic diagram](../figures/pergroup_quantization.png)
42+ 
43+- Perblock quantization (abbreviated as B quantization): The quantization target can be either the left matrix or the right matrix. Data is divided into blocks on all axes, and each block uses an independent quantization parameter.
44+ 
45+ - Assuming the left matrix shape is (m, k), where k is the reduce axis, data is grouped on the m and k axes by (bs, bs) blocks, where bs is the block size. The shape of the generated quantization parameter is (m/bs, k/bs).
46+ - Assuming the right matrix shape is (k, n), where k is the reduce axis, data is grouped on the k and n axes by (bs, bs) blocks, where bs is the block size. The shape of the generated quantization parameter is (k/bs, n/bs).
47+ 
48+ ![Schematic diagram](../figures/perblock_quantization.png)
49+ 
50+## Common Combined Quantization
51+ 
52+- Full quantization: Generally refers to the mode in which both the left and right matrices are quantized, including:
53+ - Pertensor-perchannel quantization mode (abbreviated as T-C quantization mode)
54+ - Pertoken-perchannel quantization mode (abbreviated as K-C quantization mode)
55+ - Pergroup-perblock quantization mode (abbreviated as G-B quantization mode)
56+ - Pertensor-perchannel-pergroup quantization mode (abbreviated as T-CG quantization mode)
57+ - Perblock-perblock quantization mode (abbreviated as B-B quantization mode)
58+- Pseudo-quantization: Generally refers to the mode in which only the weight matrix is quantized, including the perchannel quantization mode (abbreviated as C quantization mode).
59+- MX quantization: Essentially Microscaling quantization, which maintains model precision at extremely low bits (such as 1 bit) by dynamically adjusting scaling factors. Here it refers to the pergroup-pergroup quantization mode (abbreviated as G-G quantization mode), which is a special case where the quantization parameter type is FLOAT8_E8M0 and the group size is 32.
@@ -0,0 +1,152 @@
1+# sparseMode Introduction
2+ 
3+In the large model field, sparseMode (sparse mode) usually refers to the sparsity design of parameters or activations in the model architecture or calculation formula, as opposed to the dense mode (DenseMode).
4+ 
5+This section introduces common sparseModes and their corresponding scenario descriptions.
6+ 
7+| sparseMode | Meaning | Note |
8+| ---------- | --------------------- | ------------------ |
9+| 0 | defaultMask mode. | - |
10+| 1 | allMask mode. | - |
11+| 2 | leftUpCausal mode. | - |
12+| 3 | rightDownCausal mode. | - |
13+| 4 | band mode. | - |
14+| 5 | prefix non-compressed mode. | Not supported in varlen scenarios. |
15+| 6 | prefix compressed mode. | - |
16+| 7 | varlen outer slice scenario, rightDownCausal mode. | Only supported in varlen scenarios. |
17+| 8 | varlen outer slice scenario, leftUpCausal mode. | Only supported in varlen scenarios. |
18+ 
19+The working principle of attenMask is to mask the value of the query (Q) and key (K) transpose matrix product at the position where Mask is True, as shown below:
20+ 
21+![Schematic](../figures/QK_transpose_diagram.png)
22+ 
23+The $QK^T$ matrix will be masked at the position where attenMask is True, with the following effect:
24+ 
25+![Schematic](../figures/masked_QK_diagram.png)
26+ 
27+## sparseMode=0
28+ 
29+When sparseMode is 0, it represents the defaultMask mode.
30+ 
31+- No mask passed: If attenMask is not passed, no mask operation is performed. attenMask takes the value None, and preTokens and nextTokens values are ignored. The Masked $QK^T$ matrix is shown below:
32+ 
33+ ![Schematic](../figures/sparsemode_0_masked_matrix.png)
34+ 
35+- nextTokens is 0, preTokens is greater than or equal to Sq, indicating a causal scenario sparse. attenMask should pass a lower triangular matrix. At this time, the part between preTokens and nextTokens needs to be calculated. The Masked $QK^T$ matrix is shown below:
36+ 
37+ ![Schematic](../figures/sparsemode_0_masked_matrix_1.png)
38+ 
39+ attenMask should pass a lower triangular matrix, as shown below:
40+ 
41+ ![Schematic](../figures/attenmask_lower_triangle.png)
42+ 
43+- preTokens is less than Sq, nextTokens is less than Skv, and both are greater than or equal to 0, indicating a band scenario. At this time, the part between preTokens and nextTokens needs to be calculated. The Masked $QK^T$ matrix is shown below:
44+ 
45+ ![Schematic](../figures/sparsemode_0_masked_matrix_2.png)
46+ 
47+ attenMask should pass a band-shaped matrix, as shown below:
48+ 
49+ ![Schematic](../figures/attenmask_band_matrix.png)
50+ 
51+- nextTokens is negative. Taking preTokens=9, nextTokens=-3 as an example, the part between preTokens and nextTokens needs to be calculated. The Masked $QK^T$ is shown below:
52+ 
53+ **Note: When nextTokens is negative, preTokens must be greater than or equal to the absolute value of nextTokens, and the absolute value of nextTokens must be less than Skv.**
54+ 
55+ ![Schematic](../figures/sparsemode_0_masked_matrix_3.png)
56+ 
57+- preTokens is negative. Taking nextTokens=7, preTokens=-3 as an example, the part between preTokens and nextTokens needs to be calculated. The Masked $QK^T$ is shown below:
58+ 
59+ **Note: When preTokens is negative, nextTokens must be greater than or equal to the absolute value of preTokens, and the absolute value of preTokens must be less than Sq.**
60+ 
61+ ![Schematic](../figures/sparsemode_0_masked_matrix_4.png)
62+ 
63+## sparseMode=1
64+ 
65+When sparseMode is 1, it represents allMask, that is, passing the complete attenMask matrix.
66+ 
67+In this scenario, nextTokens and preTokens values are ignored. The Masked $QK^T$ matrix is shown below:
68+ 
69+![Schematic](../figures/sparsemode_1_masked_matrix.png)
70+ 
71+## sparseMode=2
72+ 
73+When sparseMode is 2, it represents the leftUpCausal mode mask, corresponding to the lower triangular scenario divided by the upper-left vertex (parameter starting point is the upper-left corner).
74+ 
75+In this scenario, preTokens and nextTokens values are ignored. The Masked $QK^T$ matrix is shown below:
76+ 
77+![Schematic](../figures/sparsemode_2_masked_matrix.png)
78+ 
79+The passed attenMask is an optimized compressed lower triangular matrix (2048\*2048). The compressed lower triangular matrix is shown below (same below):
80+ 
81+![Schematic](../figures/attenmask_compressed_lower_triangle.png)
82+ 
83+## sparseMode=3
84+ 
85+When sparseMode is 3, it represents the rightDownCausal mode mask, corresponding to the lower triangular scenario divided by the lower-right vertex (parameter starting point is the lower-right corner).
86+ 
87+In this scenario, preTokens and nextTokens values are ignored. attenMask is an optimized compressed lower triangular matrix (2048\*2048). The Masked $QK^T$ matrix is shown below:
88+ 
89+![Schematic](../figures/sparsemode_3_masked_matrix.png)
90+ 
91+## sparseMode=4
92+ 
93+When sparseMode is 4, it represents the band scenario, that is, calculating the part between preTokens and nextTokens. The parameter starting point is the lower-right corner, and there must be an intersection between preTokens and nextTokens. attenMask is an optimized compressed lower triangular matrix (2048\*2048). The Masked $QK^T$ matrix is shown below:
94+ 
95+![Schematic](../figures/sparsemode_4_masked_matrix.png)
96+ 
97+## sparseMode=5
98+ 
99+When sparseMode is 5, it represents the prefix non-compressed scenario, that is, adding a matrix with length Sq and width N to the left on the basis of rightDownCausal. The value of N is obtained from the optional input prefix. For example, the figure below shows prefix passing array [4,5] in batch=2 scenario. The N value of each batch axis can be different. The parameter starting point is the upper-left corner.
100+ 
101+In this scenario, preTokens and nextTokens values are ignored. The attenMask matrix data format must be BNSS or B1SS. The Masked $QK^T$ matrix is shown below:
102+ 
103+![Schematic](../figures/sparsemode_5_masked_matrix.png)
104+ 
105+attenMask should pass a matrix as shown below:
106+ 
107+![Schematic](../figures/attenmask_matrix.png)
108+ 
109+## sparseMode=6
110+ 
111+When sparseMode is 6, it represents the prefix compressed scenario, that is, in the prefix scenario, attenMask is an optimized compressed lower triangular + rectangular matrix (3072\*2048): the upper part is a [2048, 2048] lower triangular matrix, and the lower part is a [1024, 2048] rectangular matrix. The left half of the rectangular matrix is all 0, and the right half is all 1. attenMask should pass a matrix as shown below. In this scenario, preTokens and nextTokens values are ignored.
112+ 
113+![Schematic](../figures/sparsemode_6_masked_matrix.png)
114+ 
115+## sparseMode=7
116+ 
117+When sparseMode is 7, it indicates a varlen and long sequence outer slice scenario (that is, long sequences are multi-card sliced by query sequence length in the model script). You need to ensure that the scenario using sparseMode 3 was used before outer slicing. In the current mode, you need to set preTokens and nextTokens (starting point is the lower-right vertex), and you need to ensure that the parameters are correct, otherwise there will be precision issues.
118+ 
119+The Masked $QK^T$ matrix is shown below. In the second batch, the query is sliced, and the key and value are not sliced. The 4x6 mask matrix is sliced into 2x6 and 2x6 masks, which are calculated on card 1 and card 2 respectively:
120+ 
121+- The last mask block of card 1 is a band-type mask. Configure preTokens=6 (ensure it is greater than or equal to the last Skv), nextTokens=-2. actual_seq_qlen should pass {3,5}, and actual_seq_kvlen should pass {3,9}.
122+- The mask type of card 2 remains unchanged after slicing. sparseMode is 3. actual_seq_qlen should pass {2,7,11}, and actual_seq_kvlen should pass {6,11,15}.
123+ 
124+![Schematic](../figures/sparsemode_7_masked_matrix.png)
125+ 
126+**Note**:
127+ 
128+- sparseMode=7, band represents the sparse type of the last non-empty tensor Batch. If there is only one batch, you need to configure parameters according to the band mode requirements. For sparseMode=7, you need to input a 2048x2048 lower triangular mask as the input of this fusion operator.
129+- The sparse parameters of the band mode generated based on sparseMode=3 outer slicing should meet the following conditions:
130+ - preTokens >= last_Skv.
131+ - last_Sq-last_Skv <= nextTokens <= 0.
132+ - The current mode does not support the optional input pse.
133+- The non-band mode batch should satisfy: Sq <= Skv.
134+ 
135+## sparseMode=8
136+ 
137+When sparseMode is 8, it indicates a varlen and long sequence outer slice scenario. You need to ensure that the scenario using sparseMode 2 was used before outer slicing. In the current mode, you need to set preTokens and nextTokens (starting point is the lower-right vertex), and you need to ensure that the parameters are correct, otherwise there will be precision issues.
138+ 
139+The Masked $QK^T$ matrix is shown below. In the second batch, the query is sliced, and the key and value are not sliced. The 5x4 mask matrix is sliced into 2x4 and 3x4 masks, which are calculated on card 1 and card 2 respectively:
140+ 
141+- The mask type of card 1 remains unchanged after slicing. sparseMode is 2. actual_seq_qlen should pass {3,5}, and actual_seq_kvlen should pass {3,7}.
142+- The first mask block of card 2 is a band-type mask. Configure preTokens=4 (ensure it is greater than or equal to the first Skv), nextTokens=1. actual_seq_qlen should pass {3,8,12}, and actual_seq_kvlen should pass {4,9,13}.
143+ 
144+![Schematic](../figures/sparsemode_8_masked_matrix.png)
145+ 
146+**Note**:
147+ 
148+- sparseMode=8, band represents the sparse type of the first non-empty tensor Batch. If there is only one batch, you need to configure parameters according to the band mode requirements. For sparseMode=8, you need to input a 2048x2048 lower triangular mask as the input of this fusion operator.
149+- The sparse parameters of the band mode generated based on sparseMode=2 outer slicing should meet the following conditions:
150+ - preTokens >= first_Skv.
151+ - nextTokens >= first_Sq - first_Skv, configure according to the actual situation.
152+ - The current mode does not support the optional input pse.
@@ -0,0 +1,23 @@
1+# Two-stage Interface
2+ 
3+When calling an operator API based on the single-operator API execution method, it is usually divided into "two stages", with the following pattern:
4+ 
5+```Cpp
6+aclnnStatus aclxxXxxGetWorkspaceSize(const aclTensor *src, ..., aclTensor *out, ..., uint64_t *workspaceSize, aclOpExecutor **executor);
7+aclnnStatus aclxxXxx(void *workspace, uint64_t workspaceSize, aclOpExecutor *executor, aclrtStream stream);
8+```
9+ 
10+You must first call the first-stage interface aclxxXxxGetWorkspaceSize to calculate how much workspace memory is required during this API call. After obtaining the calculated workspaceSize, apply for NPU memory according to the workspaceSize, and then call the second-stage interface aclxxXxx to execute the calculation.
11+ 
12+Here, "aclxx" represents the operator interface prefix, such as aclnn; and "Xxx" represents the corresponding operator type, such as the Add operator.
13+ 
14+> Note:
15+>
16+> - workspace refers to the temporary memory required by the API to complete the calculation on the AI processor, in addition to input/output.
17+> - The second-stage interface aclxxXxx(...) cannot be called repeatedly. The following calling method will cause an exception:
18+>
19+> ```Cpp
20+> aclxxXxxGetWorkspaceSize(...)
21+> aclxxXxx(...)
22+> aclxxXxx(...)
23+> ```
@@ -0,0 +1,249 @@
1+# Introduction
2+ 
3+CANN Simulator is a SoC-level chip simulation tool designed for operator development scenarios. It analyzes the accuracy and performance data (such as instruction execution status) of AI tasks running on the AI simulator at each stage. This tool helps users perform deep performance tuning, enabling developers to obtain verification results and performance feedback nearly consistent with real chips even when real chips are unavailable or chip resources are scarce.
4+ 
5+# Main Functions
6+ 
7+This tool maintains binary compatibility with on-board execution (the same kernel can be executed on both the simulator and the AI processor). The main uses are as follows:
8+ 
9+* Accuracy simulation: Outputs bit-level accuracy results, helping users complete operator accuracy verification.
10+* Performance simulation: Outputs instruction pipeline diagrams, helping users identify operator performance bottlenecks.
11+ 
12+# Preparation Before Use
13+ 
14+## Usage Constraints
15+ 
16+* Recommended tool environment configuration: CPU with 16 cores or more, memory of 32 GB or more.
17+* All paths mentioned in this document must ensure that the running user has read or read-write permissions.
18+* For security and minimal permissions, it is recommended to use regular user permissions to execute this tool. Avoid using root or other high-privilege accounts.
19+* This tool depends on the CANN software package. Before using it, install the CANN software package. Driver and firmware installation is not required. Execute the CANN set_env.sh environment variable file through the source command. For security, do not modify the environment variables involved in set_env.sh after executing the source command.
20+* Users should follow the principle of least privilege. For example, files input to the tool must not be writable by other users. In some more stringent security scenarios, ensure that input files are not writable by group users.
21+* This tool is a development tool and is not recommended for use in production environments.
22+* The simulation function of the tool only supports single-card scenarios and cannot simulate multi-card environments. Only card 0 can be set in the code. Modifying the visible card number will cause simulation failure.
23+* The simulation environment only supports AI Core computation-type operators (MC2 and HCCL type operators are not supported).
24+* The CANN Simulator tool is currently in the early-access version stage and only supports the Ascend950PR chip. It is recommended that the simulator running environment be configured with a 16-core CPU and 32 GB or more memory.
25+* ARM environment simulation is not supported at this time.
26+ 
27+## Environment Preparation
28+ 
29+CANN Simulator is integrated in the CANN toolkit package. Complete the software package installation by following [Environment Deployment](../context/quick_install.md).
30+ 
31+# Quick Start
32+ 
33+The following uses [add_examples](../../../examples/add_example/) as an example to describe operator simulation in detail.
34+ 
35+## Operator Compilation
36+ 
37+* Complete the add_example operator compilation and installation by following [Operator Invocation](../invocation/quick_op_invocation.md).
38+ 
39+```bash
40+# Note: Enter the project root directory and execute the following compilation command. The command is for reference only. For details, refer to the operator invocation instructions.
41+bash build.sh --pkg --soc=Ascend950 --vendor_name=custom --ops=add_example
42+# Install the custom operator package
43+./build_out/cann-ops-nn-${vendor_name}_linux-${arch}.run
44+```
45+ 
46+* Complete the compilation of test_aclnn_add_example.cpp by following [aclnn Invocation](../invocation/op_invocation.md#aclnn-invocation), and generate the executable file test_aclnn_add_example.
47+ 
48+## Execute Simulation Command
49+ 
50+```bash
51+cannsim record ./test_aclnn_add_example -s Ascend950 --gen-report
52+```
53+ 
54+The simulation tool execution log files are in the examples/add_example/examples/build/bin/cannsim_* directory. The execution log file is:
55+ 
56+```bash
57+cannsim.log
58+```
59+ 
60+From the simulation tool log file, you can see the print information in the sample:
61+ 
62+```bash
63+add_example first input[0] is: 1.000000, second input[0] is: 1.000000, result[0] is: 2.000000
64+add_example first input[1] is: 1.000000, second input[1] is: 1.000000, result[1] is: 2.000000
65+add_example first input[2] is: 1.000000, second input[2] is: 1.000000, result[2] is: 2.000000
66+add_example first input[3] is: 1.000000, second input[3] is: 1.000000, result[3] is: 2.000000
67+add_example first input[4] is: 1.000000, second input[4] is: 1.000000, result[4] is: 2.000000
68+add_example first input[5] is: 1.000000, second input[5] is: 1.000000, result[5] is: 2.000000
69+add_example first input[6] is: 1.000000, second input[6] is: 1.000000, result[6] is: 2.000000
70+```
71+ 
72+## View Performance Pipeline
73+ 
74+The simulation performance pipeline files are in the `examples/add_example/examples/build/bin/cannsim_*/report` directory of this project. The pipeline-related file is:
75+ 
76+```bash
77+trace_core0.json
78+```
79+ 
80+Enter "chrome://tracing" in the Chrome browser and drag the generated instruction pipeline diagram file (trace_core0.json) to the blank area to open it. For specific parameter descriptions, refer to the "Simulation Result Analysis" section.
81+ 
82+# Simulation Execution Instructions
83+ 
84+## Command Function
85+ 
86+Execute the application in the simulation environment.
87+ 
88+## Command Format
89+ 
90+cannsim record [options] user_app --user-options
91+ 
92+## Parameter Description
93+ 
94+Table 1 Simulation Execution Parameter Description
95+ 
96+|Parameter|Required/Optional|Description|
97+| --- | --- | --- |
98+|-s or --soc-version [options] parameter | Required | Specify the target chip version for simulation (for example: Ascend950).|
99+|-o or --output [options] parameter | Optional| The path where the generated files are stored. It can be configured as an absolute path or a relative path, and the user executing the tool must have read-write permissions. If the path is not specified, data is saved in the current directory by default.|
100+|-g or --gen-report [options] parameter | Optional | Enable automatic analysis after simulation completion and generate an analysis report. By default, automatic analysis is not enabled.|
101+|user_app|Required|Operator executable file.|
102+|--user-options|Optional|Running parameters of the operator executable file.|
103+ 
104+## Usage Example
105+ 
106+1. Complete operator development and compilation.
107+2. Execute the simulation command. Refer to the following usage examples:
108+ 
109+ ```text
110+ Method 1: Enable simulation and save the output to the ./output directory. /path/to/app is the operator program.
111+ $ cannsim record /path/to/app -o ./output -s Ascend950
112+ 
113+ Method 2: Enable simulation and generate a report for subsequent performance analysis.
114+ $ cannsim record /path/to/app -o ./output -s Ascend950 --gen-report
115+ ```
116+ 
117+3. After the command completes, a folder named "cannsim_{timestamp}_${user_app}" is generated in the default path or the specified "output" directory. The structure example is as follows:
118+ 
119+ ```text
120+ ├─cannsim_{timestamp}_${user_app}
121+ ├── cannsim.log
122+ ```
123+ 
124+4. You can obtain the operator execution results and compare the accuracy. The results are displayed in cannsim.log. An example is as follows:
125+ 
126+ The following output is only an example of the AscendC single-operator direct invocation accuracy comparison result. It may vary slightly depending on the version. Please refer to the actual output.
127+ 
128+ ```bash
129+ INFO:root:[INFO] compare data case[ case001]
130+ INFO:root:---------------RESULT---------------
131+ INFO:root:['case_name', 'wrong_num', 'total_num', 'result', 'task_duration']
132+ INFO:root:[' case001', 0, 65536, 'Success']
133+ ```
134+ 
135+5. View the operator instruction pipeline diagram. Refer to the simulation result analysis section.
136+ 
137+# Simulation Result Analysis Instructions
138+ 
139+## Command Function
140+ 
141+Generate a visualized instruction pipeline diagram.
142+ 
143+## Command Format
144+ 
145+cannsim report [options]
146+ 
147+## Parameter Description
148+ 
149+Table 1 Simulation Result Analysis Parameter Description
150+ 
151+|Parameter | Required/Optional | Description|
152+| --- | --- | --- |
153+|-e or --export [options] parameter | Required | The original result file directory. It must be specified as the result directory generated after simulation execution, pointing to the cannsim_{timestamp}_${user_app} level. It can be configured as an absolute path or a relative path, and the tool execution user must have read-write permissions.|
154+|-o or --output [options] parameter | Optional | The analysis result output directory. It can be configured as an absolute path or a relative path, and the execution user must have read-write permissions. If the path is not specified, data is saved in the current directory by default. If the generated result file has the same name as an existing file, the existing file is overwritten.|
155+|-n or --core-id [options] parameter | Optional | Specify the core ID for generating the instruction pipeline. If not specified, the pipeline for core 0 is generated by default. The configuration format is as follows: To generate pipelines for all cores, configure 'all'. To specify a core ID range, for example: '0-1'. To specify a single core ID, for example: '5'.|
156+ 
157+## Usage Example
158+ 
159+1. Execute operator simulation by following the simulation execution instructions, and compare the output example to ensure the corresponding results are correct.
160+2. Execute the simulation result analysis command. Refer to the following execution example.
161+ 
162+ ```bash
163+ Generate a performance analysis report in the current directory (default: analyze only core 0)
164+ cannsim report -e /path/to/cannsim_{timestamp}_${user_app}
165+ 
166+ Generate performance analysis reports for core 0, core 1, core 11, and core 12 in the specified directory
167+ cannsim report -e /path/to/cannsim_{timestamp}_${user_app} -o /path/to/report -n '0-1, 11-12'
168+ ```
169+ 
170+3. After the command execution completes, the corresponding pipeline files are generated in the output configured directory. The file format is JSON. The output result example is as follows:
171+ 
172+ ```bash
173+ trace_core0.json
174+ trace_core1.json
175+ ...
176+ ```
177+ 
178+4. View simulation results
179+ Enter "chrome://tracing" in the Chrome browser and drag the generated instruction pipeline diagram file (trace.json) to the blank area to open it. Use keyboard shortcuts (W: zoom in, S: zoom out, A: move left, D: move right) to view the results.
180+ ![Instruction Pipeline Diagram](../figures/Instruction_pipeline.png)
181+ 
182+ Table 2 Key Field Description
183+ 
184+ |Field Name|Field Meaning|
185+ | --- | --- |
186+ |VECTOR|Vector computation unit.|
187+ |SCALAR|Scalar computation unit.|
188+ |Cube|Matrix multiplication computation unit.|
189+ |MTE1|Data transfer pipeline; data transfer direction: L1 ->{L0A/L0B, UBUF}.|
190+ |MTE2|Data transfer pipeline; data transfer direction: {DDR/GM, L2} ->{L1, L0A/B, UBUF}.|
191+ |MTE3|Data transfer pipeline; data transfer direction: UBUF -> {DDR/GM, L2, L1}, L1->{DDR/L2}.|
192+ |FIXP|Data transfer pipeline; data transfer direction: FIXPIPE L0C -> OUT/L1.|
193+ |FLOWCTRL|Control flow instruction.|
194+ |ICACHELOAD|View ICache misses.|
195+ 
196+# Query Help Information
197+ 
198+## Command Function
199+ 
200+Query tool help information.
201+ 
202+## Command Format
203+ 
204+Query tool help information:
205+ 
206+```bash
207+cannsim --help
208+```
209+ 
210+Query tool record subcommand help information:
211+ 
212+```bash
213+cannsim record --help
214+```
215+ 
216+Query tool report subcommand help information:
217+ 
218+```bash
219+cannsim report --help
220+```
221+ 
222+## Parameter Description
223+ 
224+None
225+ 
226+## Usage Example
227+ 
228+1. Log in to the Host-side server.
229+2. Execute the following command.
230+ 
231+ ```bash
232+ cannsim --help
233+ ```
234+ 
235+## Output Description
236+ 
237+```bash
238+usage: cannsim [-h] {record,report} ...
239+ 
240+Command-line tool for performance simulation analysis on Ascend hardware.
241+ 
242+positional arguments:
243+ {record,report} Available commands
244+ record Run user application in AscendOps simulation environment
245+ report Generate performance analysis reports
246+ 
247+options:
248+ -h, --help show this help message and exit
249+```
@@ -0,0 +1,190 @@
1+# Operator Debugging and Tuning
2+ 
3+## Debugging and Troubleshooting (AI Core Operators)
4+ 
5+If an operator execution failure or accuracy anomaly occurs during operator execution, you can print information at each stage, such as Kernel intermediate results, for problem analysis and troubleshooting.
6+ 
7+### 1. Host-Side Log Acquisition Method
8+ 
9+* **plog acquisition**
10+ 
11+ After program execution completes, you can view the logs by default in "$HOME/ascendc/log". The host log file storage path is as follows:
12+ 
13+ ```bash
14+ $HOME/ascend/log/debug/plog/plog-pid_*.log
15+ ```
16+ 
17+ Enable the environment variable ASCEND_SLOG_PRINT_TO_STDOUT to display log output directly on the screen (1: enable screen display, 0: disable screen display). The configuration example is as follows:
18+ 
19+ ```bash
20+ export ASCEND_SLOG_PRINT_TO_STDOUT=1
21+ ```
22+ 
23+ For log-related information, refer to [Log Reference](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/maintenref/logreference/logreference_0001.html). For environment variable information, refer to [Environment Variable Reference](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/maintenref/envvar/envref_07_0001.html).
24+ 
25+* **aclnn exception error message acquisition**
26+ 
27+ Obtain exception information during aclnn interface invocation through the aclGetRecentErrMsg interface (refer to [acl API (C)](https://www.hiascend.com/document/detail/en/canncommercial/latest/API/appdevgapi/aclcppdevg_03_0004.html)). The usage method is as follows:
28+ 
29+ ```bash
30+ printf(aclGetRecentErrMsg());
31+ ```
32+ 
33+ The printed error message example is as follows:
34+ 
35+ ```bash
36+ [PID:646612] 2026-01-24-11:53:44.671.727 AclNN_Parameter_Error(EZ1001): Expected a proper Tensor but got null for argument addmmTensor.self.
37+ ```
38+ 
39+### 2. Kernel Debugging
40+ 
41+Common debugging methods are as follows:
42+ 
43+* **printf**
44+ 
45+ This interface supports printing Scalar-type data, such as integers, characters, and Boolean values. For detailed information, refer to [Ascend C API](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/API/ascendcopapi/atlasascendc_api_07_0003.html) in "Operator Debugging API > printf".
46+ 
47+ ```c++
48+ blockLength_ = tilingData->totalLength / AscendC::GetBlockNum();
49+ tileNum_ = tilingData->tileNum;
50+ tileLength_ = blockLength_ / tileNum_ / BUFFER_NUM;
51+ // Print the current core computation Block length
52+ AscendC::PRINTF("Tiling blockLength is %llu\n", blockLength_);
53+ ```
54+ 
55+* **DumpTensor**
56+ 
57+ This interface supports dumping the content of a specified Tensor and also supports printing custom additional information, such as the current line number. For detailed information, refer to [Ascend C API](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/API/ascendcopapi/atlasascendc_api_07_0003.html) in "Operator Debugging API > DumpTensor".
58+ 
59+ ```c++
60+ AscendC::LocalTensor<T> zLocal = outputQueueZ.DeQue<T>();
61+ // Print zLocal Tensor information
62+ DumpTensor(zLocal, 0, 128);
63+ AscendC::DataCopy(outputGMZ[progress * tileLength_], zLocal, tileLength_);
64+ ```
65+ 
66+For troubleshooting in complex scenarios, such as operator hangs or GM/UB access out-of-bounds, you can use **step-by-step debugging**. For specific operations, refer to the [msDebug](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/devaids/optool/docs/en/quick_start/msdebug_quick_start.md) operator debugging tool.
67+ 
68+## Debugging and Troubleshooting (AI CPU Operators)
69+ 
70+If an operator execution failure or accuracy anomaly occurs during operator execution, you can print information at each stage, such as Kernel intermediate results, for problem analysis and troubleshooting.
71+ 
72+### 1. Host-Side Log Acquisition Method
73+ 
74+ Refer to the AI Core operator [Host-Side Log Acquisition Method](#1-host-side-log-acquisition-method)
75+ 
76+### 2. Kernel Debugging
77+ 
78+Common debugging methods are as follows:
79+ 
80+* **KERNEL_LOG macro**
81+ 
82+ You can print log information during operator execution through the following macros, including DEBUG, INFO, WARN, and ERROR level logs.
83+ 
84+ ```Cpp
85+ KERNEL_LOG_DEBUG(fmt, ...) // The fmt parameter represents the format control string
86+ KERNEL_LOG_INFO(fmt, ...)
87+ KERNEL_LOG_WARN(fmt, ...)
88+ KERNEL_LOG_ERROR(fmt, ...) // ERROR level logs are printed by default
89+ ```
90+ 
91+ To print logs at non-ERROR levels, you need to configure the environment variable `ASCEND_GLOBAL_LOG_LEVEL` in advance. For specific usage, refer to [Environment Variable Reference](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/maintenref/envvar/envref_07_0001.html).
92+ 
93+ The printing example is as follows:
94+ 
95+ ```c++
96+ Tensor* input0 = ctx.Input(kFirstInputIndex);
97+ Tensor* input1 = ctx.Input(kSecondInputIndex);
98+ Tensor* output = ctx.Output(0);
99+ 
100+ if (input0 == nullptr || input1 == nullptr || output == nullptr) {
101+ // Print error information
102+ KERNEL_LOG_ERROR("Invalid argument");
103+ return kParamInvalid;
104+ }
105+ 
106+ int64_t num_elements = input0->NumElements();
107+ // Print the number of input elements
108+ KERNEL_LOG_INFO("Num of elements is %ld", data_size);
109+ ```
110+ 
111+## Performance Tuning
112+ 
113+### Method 1 (For Atlas A2/A3 Series Products)
114+ 
115+If execution accuracy degradation or abnormal memory usage occurs during operator execution, you can use the [msProf](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/devaids/optool/docs/en/quick_start/msopprof_quick_start.md) performance analysis tool to analyze the operator's performance metrics at each execution stage (such as throughput, memory usage, and latency), thereby identifying the root cause and performing targeted optimization.
116+ 
117+This chapter uses the [AddExample custom operator](../../../examples/add_example/) as an example to introduce the two commonly used methods in operator tuning: on-board performance collection and pipeline simulation. By collecting the on-board running pipeline metrics of the operator, you can analyze the operator's Bound scenario. Understanding the simulation pipeline diagram helps optimize the operator's internal pipeline.
118+ 
119+1. Prerequisites.
120+ 
121+ After completing operator development and compilation, assuming the aclnn interface invocation method is used, the generated operator executable file (test_aclnn_add_example) is located in the `examples/add_example/examples/build/bin/` directory of this project.
122+ 
123+2. Collect performance data.
124+ 
125+ When you need to collect the on-board running pipeline metrics of the operator, navigate to the directory where the operator executable file is located and execute the following command:
126+ 
127+ ```bash
128+ msprof op ./test_aclnn_add_example
129+ ```
130+ 
131+ The collection results are in the `examples/add_example/examples/build/bin/OPPROF_*` directory of this project. After collection completes, the following information is printed:
132+ 
133+ ``` text
134+ Op Name: AddExample_a1532827238e1555db7b997c7bce2928_high_performance_1
135+ Op Type: vector
136+ Task Duration(us): 97.861954
137+ Block Dim: 8
138+ Mix Block Dim:
139+ Device Id: 0
140+ Pid: 2776181
141+ Current Freq: 1800
142+ Rated Freq: 1800
143+ ```
144+ 
145+ Task Duration is the current operator Kernel execution time, and Block Dim is the current operator execution core count.
146+ 
147+ For detailed pipeline metrics of the operator, refer to the `ArithmeticUtilization` file under `OPPROF_*`, which contains the proportion of each pipeline. For specific descriptions, refer to the [msProf](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/devaids/optool/docs/en/quick_start/msopprof_quick_start.md) section "Performance Data Files > msprof op > ArithmeticUtilization (cube and vector type instruction latency and proportion)".
148+ 
149+3. Collect simulation pipeline diagrams.
150+ 
151+ Before using the msProf tool for operator simulation tuning, execute the following command to configure the environment variable.
152+ 
153+ ```bash
154+ export LD_LIBRARY_PATH=${INSTALL_DIR}/tools/simulator/Ascendxxxyy/lib:$LD_LIBRARY_PATH
155+ ```
156+ 
157+ Modify the above environment variable according to the actual CANN software package installation path and AI processor model.
158+ 
159+ Then navigate to the directory where the operator executable file is located and execute the following command:
160+ 
161+ ```bash
162+ msprof op simulator --output=$PWD/pipeline_auto --kernel-name"AddExample" ./test_aclnn_add_example
163+ ```
164+ 
165+ The collection results are in the `$PWD/pipeline_auto/OPPROF_**` directory of this project.
166+ The pipeline-related file path is `OPPROF**/simulator/visualize_data.bin`, which can be viewed using the [mindStudio Insight](https://www.hiascend.com/document/detail/en/mindstudio/latest/visualization_tool/MindStudioInsight/docs/en/user_guide/overview.md) tool.
167+ 
168+### Method 2 (For Ascend 950PR)
169+ 
170+If execution accuracy degradation or abnormal memory usage occurs during operator development, you can use the [CANN Simulator](./cann_simulator.md) simulation tool to analyze the operator's instruction pipeline situation, thereby identifying the root cause and performing targeted optimization.
171+ 
172+This chapter uses the [AddExample custom operator](../../../examples/add_example/) as an example to introduce the use of the simulation tool. It describes how to perform accuracy and performance tuning through the simulation tool.
173+ 
174+1. Prerequisites.
175+ 
176+ After completing operator development and compilation, assuming the aclnn interface invocation method is used, the generated operator executable file (test_aclnn_add_example) is located in the `examples/add_example/examples/build/bin/` directory of this project.
177+ 
178+2. Execute the simulation command to generate simulation data.
179+ 
180+ ```text
181+ cannsim record ./test_aclnn_add_example -s Ascend950 --gen-report
182+ ```
183+ 
184+ The simulation results are in the `examples/add_example/examples/build/bin/cannsim_*` directory of this project. The pipeline-related file is:
185+ 
186+ ```text
187+ trace_core0.json
188+ ```
189+ 
190+3. Enter "chrome://tracing" in the Chrome browser and drag the generated instruction pipeline diagram file (trace_core0.json) to the blank area to open it. For specific parameter descriptions, refer to the [Simulation Result Analysis](./cann_simulator.md#simulation-result-analysis-instructions) section in CANN Simulator.
@@ -0,0 +1,247 @@
1+# AI CPU Operator Development Guide
2+ 
3+## Overview
4+ 
5+> Note:
6+>
7+> 1. For basic concepts and AI CPU interfaces involved in operator development, refer to [TBE & AI CPU Operator Development](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/others/tbeaicpudevg/atlasopdev_10_0001.html) for detailed information.
8+> 2. AI CPU operators are developed using the C++ language and run on the AI CPU hardware unit.
9+> 3. build.sh: The commands involved in operator development can be viewed through `bash build.sh --help`. For function parameter descriptions, refer to [build Parameter Description](../context/build.md).
10+ 
11+This development guide uses the `AddExample` operator as an example to introduce the new operator development process and the deliverables involved. For complete sample code, visit the project `examples` directory.
12+ 
13+1. [Project Creation](#project-creation): Before developing an operator, complete the environment deployment and create the operator directory for subsequent operator compilation and deployment.
14+ 
15+2. [Operator Definition](#operator-definition): Determine the operator functionality and prototype definition.
16+ 
17+3. [Kernel Implementation](#kernel-implementation): Implement the Device-side operator kernel function.
18+ 
19+4. [aclnn Adaptation](#aclnn-adaptation): Custom operators are recommended to use the aclnn interface for invocation, which requires completing binary publishing in advance. **If you use graph mode to invoke the operator**, refer to the [Graph Mode Adaptation Guide](./graph_develop_guide.md).
20+ 
21+5. [Compilation and Deployment](#compilation-and-deployment): Complete the compilation and installation of the custom operator through the project compilation script.
22+ 
23+6. [Operator Verification](#operator-verification): Verify the custom operator functionality through common operator invocation methods.
24+ 
25+## Project Creation
26+ 
27+**1. Environment Deployment**
28+ 
29+Before developing an operator, complete the basic environment setup by following [Environment Deployment](../context/quick_install.md).
30+ 
31+**2. Directory Creation**
32+ 
33+Directory creation is an important step in operator development, providing a unified directory structure and file organization for subsequent code writing, compilation, and debugging.
34+ 
35+You can quickly create the operator directory through `build.sh`. Enter the project root directory and execute the following command:
36+ 
37+```bash
38+# Create the specified operator directory, for example: bash build.sh --genop=activation/op_example
39+# ${op_class} represents the operator type, such as the activation class.
40+# ${op_name} represents the lowercase underscore form of the operator name. For example, the `AddExample` operator corresponds to add_example. New operators must not have the same name as existing operators.
41+bash build.sh --genop_aicpu=${op_class}/${op_name}
42+```
43+ 
44+After the command executes successfully, the following message appears:
45+ 
46+```bash
47+Create the AI CPU initial directory for ${op_name} under ${op_class} success
48+```
49+ 
50+After creation, the directory structure is as follows:
51+ 
52+```text
53+${op_name} # Replace with the lowercase underscore form of the actual operator name
54+├── examples # Operator invocation samples
55+│ └── test_aclnn_${op_name}.cpp # Operator aclnn invocation sample
56+├── op_host # Host-side implementation
57+│ └── ${op_name}_infershape.cpp # InferShape implementation, implementing operator shape inference, inferring the output shape at runtime
58+├── op_kernel_aicpu # Device-side Kernel implementation
59+│ ├── ${op_name}_aicpu.cpp # Kernel entry file, containing the main function and scheduling logic
60+│ ├── ${op_name}_aicpu.h # Kernel header file, including function declarations, structure definitions, and logic implementation
61+│ └── ${op_name}.json # Operator information library, defining basic operator information such as name, input/output, and data types
62+├── tests # UT implementation
63+│ └── ut # Kernel/aclnn UT implementation
64+└── CMakeLists.txt # Operator Cmakelist entry
65+```
66+ 
67+If `${op_class}` is a new operator category, you need to additionally add `${op_class}` to the `OP_CATEGORY_LIST` in `cmake/variables.cmake`; otherwise, normal compilation is not possible.
68+ 
69+## Operator Definition
70+ 
71+Operator definition requires two deliverables: `README.md` and `${op_name}.json`
72+ 
73+**Deliverable 1: README.md**
74+ 
75+Before developing an operator, determine the functionality and computation logic of the target operator.
76+ 
77+For an example of the custom `AddExample` operator, refer to [AddExample Operator Description](../../../examples/add_example_aicpu/README.md).
78+ 
79+**Deliverable 2: ${op_name}.json**
80+ 
81+Operator information library.
82+ 
83+For an example of the custom `AddExample` operator, refer to [AddExample Operator Information Library](../../../examples/add_example_aicpu/op_kernel_aicpu/add_example.json).
84+ 
85+## Kernel Implementation
86+ 
87+### Kernel Introduction
88+ 
89+Kernel is the core part of an operator executed on the NPU. The Kernel implementation includes the following steps:
90+ 
91+```mermaid
92+graph LR
93+H([Operator Class Declaration]) -->A([Compute Function Implementation])
94+A -->B([Register Operator])
95+```
96+ 
97+### Code Implementation
98+ 
99+Kernel requires two deliverables: `${op_name}_aicpu.cpp` and `${op_name}_aicpu.h`
100+ 
101+**Deliverable 1: ${op_name}_aicpu.h**
102+ 
103+Operator class declaration
104+ 
105+The first step of Kernel implementation is to declare the operator class in the header file `op_kernel_aicpu/${op_name}_aicpu.h`. The operator class must inherit from the CpuKernel base class.
106+For detailed implementation, refer to [add_example_aicpu.h](../../../examples/add_example_aicpu/op_kernel_aicpu/add_example_aicpu.h).
107+ 
108+```CPP
109+// 1. Operator class declaration
110+// Include the AI CPU base library header file
111+#include "cpu_kernel.h"
112+// Define the namespace aicpu (fixed, do not modify), and define the operator Compute implementation function
113+namespace aicpu {
114+// The operator class inherits from the CpuKernel base class
115+class AddExampleCpuKernel : public CpuKernel {
116+ public:
117+ ~AddExampleCpuKernel() = default;
118+ // Declare the Compute function (needs to be overridden); the parameter CpuKernelContext is the context of CPUKernel, including operator input, output, and attribute information
119+ uint32_t Compute(CpuKernelContext &ctx) override;
120+};
121+} // namespace aicpu
122+```
123+ 
124+**Deliverable 2: ${op_name}_aicpu.cpp**
125+ 
126+Compute function implementation and AI CPU operator registration
127+ 
128+Obtain the input/output Tensor information and perform validity checks, then implement the core computation logic (such as the addition operation), and set the computation result to the output Tensor.
129+ 
130+For detailed implementation, refer to [add_example_aicpu.cpp](../../../examples/add_example_aicpu/op_kernel_aicpu/add_example_aicpu.cpp).
131+ 
132+```C++
133+// 2. Compute function implementation
134+#include "add_example_aicpu.h"
135+ 
136+namespace {
137+// Operator name
138+const char* const kAddExample = "AddExample";
139+const uint32_t kParamInvalid = 1;
140+} // namespace
141+ 
142+// Define the namespace aicpu
143+namespace aicpu {
144+// Implement the Compute function of the custom operator class
145+uint32_t AddExampleCpuKernel::Compute(CpuKernelContext& ctx) {
146+ // Obtain the input tensor from CpuKernelContext
147+ Tensor* input0 = ctx.Input(0);
148+ Tensor* input1 = ctx.Input(1);
149+ // Obtain the output tensor from CpuKernelContext
150+ Tensor* output = ctx.Output(0);
151+ 
152+ // Perform basic validation on the tensor; check for null pointers
153+ if (input0 == nullptr || input1 == nullptr || output == nullptr) {
154+ return kParamInvalid;
155+ }
156+ 
157+ // Obtain the data type of the input tensor
158+ auto data_type = static_cast<DataType>(input0->GetDataType());
159+ // Obtain the data address of the input tensor, for example, the input data type is int32
160+ auto input0_data = reinterpret_cast<int32_t*>(input0->GetData());
161+ // Obtain the tensor shape
162+ auto input0_shape = input->GetTensorShape();
163+ 
164+ // Obtain the data address of the output tensor, for example, the output data type is int32
165+ auto y = reinterpret_cast<int32_t*>(output->GetData());
166+ 
167+ // The AddCompute function executes the corresponding computation based on the input type.
168+ // Since C++ does not natively support half-precision floating-point types, you can use the third-party library Eigen (version 3.3.9 recommended) for representation.
169+ switch (data_type) {
170+ case DT_FLOAT:
171+ return AddCompute<float>(...);
172+ case DT_INT32:
173+ return AddCompute<int32>(...);
174+ ....
175+ default : return PARAM_INVALID;
176+ }
177+}
178+ 
179+// 3. Register the operator Kernel implementation for the framework to obtain the Compute function of the operator Kernel.
180+REGISTER_CPU_KERNEL(kAddExample, AddExampleCpuKernel);
181+} // namespace aicpu
182+```
183+ 
184+## aclnn Adaptation
185+ 
186+After operator development and compilation are completed, the aclnn interface (a set of C-based APIs) is automatically generated. No additional configuration is required. You can directly invoke the aclnn interface in your application to call the operator.
187+ 
188+## Compilation and Deployment
189+ 
190+After operator development is completed, compile the operator project to generate a custom operator installation package *.run. The specific operations are as follows:
191+ 
192+1. **Preparation.**
193+ 
194+ Complete the basic environment setup by following [Project Creation](#project-creation), and check whether the operator development deliverables are complete and in the corresponding operator category directory.
195+ 
196+2. **Compile the custom operator package.**
197+ 
198+ Using the `AddExample` operator as an example, assuming the development deliverables are in the `examples` directory, the complete code is in the [add_example](../../../examples/add_example_aicpu) directory.
199+ 
200+ ```bash
201+ # Compile the specified operator, for example: bash build.sh --pkg --ops=add_example
202+ bash build.sh --pkg --soc=${soc_version} --vendor_name=${vendor_name} --ops=${op_list} [--experimental]
203+ ```
204+ 
205+ - --soc: ${soc_version} represents the NPU model. For Atlas A2 series products, use "ascend910b" (default). For Atlas A3 series products, use "ascend910_93". For Ascend 950PR/Ascend 950DT products, use "ascend950".
206+ - --vendor_name (optional): ${vendor_name} represents the name of the custom operator package to build. The default name is custom.
207+ - --ops (optional): ${op_list} represents the operators to compile. If not specified, all operators are compiled by default. The format is "--ops=add_example".
208+ - --experimental (optional): If the operator being compiled is a contributed operator, configure --experimental.
209+ 
210+ If the following message appears, the compilation is successful:
211+ 
212+ ```bash
213+ Self-extractable archive "cann-ops-nn-${vendor_name}-linux.${arch}.run" successfully created.
214+ ```
215+ 
216+3. **Install the custom operator package.**
217+ 
218+ ```bash
219+ # Install the run package
220+ ./build_out/cann-ops-nn-${vendor_name}-linux.${arch}.run
221+ ```
222+ 
223+ The custom operator package is installed in the `${ASCEND_HOME_PATH}/opp/vendors` path. `${ASCEND_HOME_PATH}` represents the CANN software installation directory, which can be configured in the environment variable in advance.
224+ 
225+4. **(Optional) Uninstall the custom operator package.**
226+ 
227+ After the custom operator package is installed, an `uninstall.sh` script is generated in the `${ASCEND_HOME_PATH}/opp/vendors/${vendor_name}_nn/scripts` directory. You can uninstall the custom operator package through this script. The command is as follows:
228+ 
229+ ```bash
230+ bash ${ASCEND_HOME_PATH}/opp/vendors/${vendor_name}_nn/scripts/uninstall.sh
231+ ```
232+ 
233+## Operator Verification
234+ 
235+Before verifying the operator, ensure that the environment variables are configured. The command is as follows:
236+ 
237+```bash
238+export LD_LIBRARY_PATH=${ASCEND_HOME_PATH}/opp/vendors/${vendor_name}_nn/op_api/lib:${LD_LIBRARY_PATH}
239+```
240+ 
241+- **UT Verification**
242+ 
243+ During operator development, you can quickly verify through UT verification (such as Kernel).
244+ 
245+- **aclnn Invocation Verification**
246+ 
247+ After the developed operator is compiled and deployed, you can verify the functionality through the aclnn method. For the method, refer to [Operator Invocation Methods](../invocation/op_invocation.md).
@@ -0,0 +1,466 @@
1+# Cross-Platform Migration Guide for Operators
2+ 
3+This guide describes the key adaptation points and solutions for migrating operators across multiple platforms. Taking the migration of operators from the Atlas A2 series to the Ascend 950 series as an example, it compares hardware architecture differences and related adaptation points, and provides relevant operator adaptation samples.
4+ 
5+## I. Hardware Architecture and Specification Parameter Comparison
6+ 
7+### Atlas A2 Series Hardware Architecture
8+ 
9+<div align="center">
10+ <img src="../figures/AtlasA2_hardware_architecture.png" width="900" alt="Atlas A2 Hardware Architecture" />
11+</div>
12+ 
13+### Ascend 950 Series Hardware Architecture
14+ 
15+<div align="center">
16+ <img src="../figures/Ascend950_hardware_architecture.png" width="900" alt="Ascend 950 Hardware Architecture" />
17+</div>
18+ 
19+### Intergenerational Specification Parameter Comparison
20+ 
21+Multiple product models are typically divided based on different application scenarios, processes, or hardware configurations. Each model may have certain differences in performance, resource configuration, and other aspects. For ease of explanation and direct comparison, this section selects representative configurations as parameter display and difference analysis objects. For other related adjustments, refer to the actual manual or official release.
22+ 
23+<table>
24+ <tr>
25+ <th colspan="2" style="width: 25%;">Specification Item</th>
26+ <th style="width:37.5%;">Atlas A2</th>
27+ <th style="width:37.5%;">Ascend 950</th>
28+ </tr>
29+ <tr>
30+ <td rowspan="4">AICore</td>
31+ <td>Core Count</td>
32+ <td>24</td>
33+ <td>32</td>
34+ </tr>
35+ <tr>
36+ <td>Frequency</td>
37+ <td>1.8</td>
38+ <td>1.65</td>
39+ </tr>
40+ <tr>
41+ <td>Cube Computing Power</td>
42+ <td>353T/376T @BF16,FP16</td>
43+ <td>426T@BF16,FP16 757T@FP8,HIFP8,MXFP8,INT8 1514T@MXFP4</td>
44+ </tr>
45+ <tr>
46+ <td>Vector Computing Power (FP16)</td>
47+ <td>23.5T</td>
48+ <td>54T</td>
49+ </tr>
50+ <tr>
51+ <td rowspan="2">Memory</td>
52+ <td>Memory Capacity (GB)</td>
53+ <td>64</td>
54+ <td>128</td>
55+ </tr>
56+ <tr>
57+ <td>Memory Bandwidth</td>
58+ <td>1.6TB/s</td>
59+ <td>1.6TB/s</td>
60+ </tr>
61+</table>
62+ 
63+## II. Adaptation Points Introduced by Hardware Capability Changes
64+ 
65+<table>
66+ <tr>
67+ <th style="width: 25%;">Hardware Unit</th>
68+ <th style="width:35%;">Hardware Capability Change</th>
69+ <th style="width:40%;">Typical Impact Scope</th>
70+ </tr>
71+ <tr>
72+ <td rowspan="5">Data Transfer Unit</td>
73+ <td>Removed the data path from L1 to GM</td>
74+ <td>Kernels that rely on L1 directly writing back to GM must be changed to use the L1 to UB to GM or L0C/FIXPIPE to GM path. Related DataCopy links, event synchronization, and buffer planning need to be adjusted.</td>
75+ </tr>
76+ <tr>
77+ <td>Removed the data paths from GM to L0A and L0B</td>
78+ <td>The direct GM to L0A/L0B connection is no longer available. Use GM to L1 to L0A/L0B instead. The L1 tiling strategy and MTE1/MTE2 pipelines need to be restructured.</td>
79+ </tr>
80+ <tr>
81+ <td>ND DMA flexible data transfer, supporting in-line ND to NZ conversion</td>
82+ <td>ND2NZ/DN2NZ can be used to complete format conversion during the MTE2 stage, reducing intermediate buffers and format conversion overhead. Pay attention to stride, alignment, and NZ shape mapping.</td>
83+ </tr>
84+ <tr>
85+ <td>Supports efficient Cube-to-Vector internal data paths: L1 to UB, L0C to UB, FIXP to UB</td>
86+ <td>Intermediate accumulation/activation/fusion (such as K-split accumulation and post-processing) can be performed on the UB side, reducing GM round trips. The corresponding synchronization and pipeline partitioning need to be adjusted.</td>
87+ </tr>
88+ <tr>
89+ <td>Introduced the collective communication accelerator CCU1.0</td>
90+ <td>For communication-computation fusion operators, adjust HcclServerType in Eager mode, and switch to the CCU series GE interfaces in Graph mode.</td>
91+ </tr>
92+ <tr>
93+ <td rowspan="3">Compute Unit</td>
94+ <td>Vector now supports the Regbase paradigm</td>
95+ <td>The memory access patterns, alignment methods, and register count assumptions that originally relied on Membase need to be re-examined. Templates and tiling may need to be updated to the Regbase version.</td>
96+ </tr>
97+ <tr>
98+ <td>Cube no longer supports int4_t</td>
99+ <td>All operators using int4_t need to switch to supported data types (such as int8) and update the quantization calculation logic.</td>
100+ </tr>
101+ <tr>
102+ <td>4:2 sparse matrix computation is not supported</td>
103+ <td>Kernels that originally relied on the 4:2 sparse feature for acceleration need to be changed to dense or other supported sparse strategies, and the performance expectation description needs to be updated.</td>
104+ </tr>
105+ <tr>
106+ <td rowspan="1">Storage Unit</td>
107+ <td>Local Buffer memory improvements: Cube L0C 256 KB, Vector UB 256 KB</td>
108+ <td>Larger L0C/UB allows for increasing the basic block size and double buffering capacity, reducing the number of K-split and block-split rounds. The L1/L0/UB ratio and tile size need to be re-evaluated.</td>
109+ </tr>
110+ <tr>
111+ <td rowspan="2">Other</td>
112+ <td>Performance optimization for multiple cores simultaneously accessing the same Global Memory address</td>
113+ <td>Templates related to matrix multiplication operators can be optimized.</td>
114+ </tr>
115+ <tr>
116+ <td>SIMT</td>
117+ <td>With the introduction of SIMT, thread-level parallelism can be used to handle branching and irregular computation, but it requires adaptation of thread partitioning, shared memory, and synchronization semantics. Some Vector implementations can be migrated to SIMT versions.</td>
118+ </tr>
119+</table>
120+ 
121+## III. Recommended Migration Steps
122+ 
123+1. Confirm whether the compute units (Cube/Vector) involved in the operator and the supported data types of the corresponding units differ between platforms.
124+2. Confirm whether the data transfer units involved (ND-to-NZ, GM-to-Lx, collective communication, and so on) differ between platforms.
125+3. Modify item by item according to the hardware capability change points (Vector architecture, Cube supported data types, L1/L0/UB size, CCU communication, and so on).
126+4. Refer to the operator migration samples to adjust or supplement the Atlas A2/Ascend 950 branching logic.
127+ 
128+## IV. Operator Migration Samples
129+ 
130+### Cube Matrix Computation Operators
131+ 
132+#### Global Memory Same-Address Access Conflict Optimization
133+ 
134+The Ascend 950 hardware introduces a new feature for parallel processing of same-address requests, eliminating the need to specifically avoid same-address access conflicts in various multi-core scenarios. During migration, the multi-core strategy designed for "offset-based conflict avoidance" on Atlas A2 can be simplified to a more regular sliding window template (such as row-group windowing with column-wise back-and-forth scanning), reducing invalid offsets and redundant address transformations. In practice, it is recommended to first retain the original tile size with the goal of functional equivalence, and then gradually relax the multi-core constraints. Observe key metrics such as MAC utilization, MTE2 utilization, and L2 hit rate based on profiling data to confirm whether the template adjustment brings stable benefits.
135+ 
136+<div align="center">
137+ <img src="../figures/SWAT_sliding_window_template.png" width="900" alt="SWAT Sliding Window Template" />
138+</div>
139+ 
140+#### Tile Size Adjustment
141+ 
142+On Atlas A2, the L0C size is 128 KB, while on Ascend 950 it is increased to 256 KB. This means that a single instance can carry a larger accumulation result block. During migration, prioritize increasing the tile block partitioning granularity or the single-round processing depth in the K direction to reduce the number of block and K-split rounds, thereby reducing loop control and data transfer overhead. At the same time, rebalance the L1/L0/UB capacity budget to avoid pipeline breakpoints caused by L0C expansion squeezing A/B/scale buffering.
143+ 
144+### Vector Computation Operators
145+ 
146+#### SIMT
147+ 
148+The Ascend 950 series introduces a new SIMT unit. SIMT has significant advantages over SIMD in handling irregular discrete access, and is suitable for scenarios with discontinuous addresses, large variations in memory access span, and inconsistent branch paths (such as scatter/gather, index reordering, and sparse updates).
149+ 
150+During migration, it is recommended to prioritize identifying operator sub-processes that are "memory-access-dominated" and have "low vectorization efficiency." If the original SIMD implementation has a large number of mask branches, a high proportion of invalid lanes, or requires complex address assembly, that part can be rewritten to the SIMT path, which typically reduces control overhead and improves effective memory access throughput.
151+ 
152+In practice, focus on the following points: first, the thread task partitioning must match the data sparsity to avoid extremely unbalanced thread loads; second, reduce pipeline idle time caused by high-frequency random memory access, and try to complete index regularization and bucketing upstream; third, decouple boundary processing from the main path to avoid introducing too many branches in hot loops. After migration, it is recommended to compare the "pure SIMD implementation" and the "SIMD + SIMT hybrid implementation" and select the optimal strategy based on data distribution, rather than fixing a single path.
153+ 
154+**Taking the gather_v2 operator as an example: SIMD vs. SIMT implementation comparison**
155+ 
156+The gather_v2 operator performs gather based on the last axis after axis merging. Therefore, the template selection basis is: use the SIMT template when the last axis is less than or equal to 2048, and use the SIMD template when the last axis is greater than 2048. This is because when the last axis is small, multiple discontinuous small block addresses need to be accessed discretely, and SIMT is more efficient. The following compares the core differences between the two implementations:
157+ 
158+**1. Programming Model Differences**
159+ 
160+The SIMD implementation uses the traditional vectorized programming model, requiring explicit management of UB buffers and pipeline queues:
161+ 
162+```cpp
163+// SIMD: Uses queue mechanism to manage data buffering
164+TQueBind<QuePosition::VECIN, QuePosition::VECOUT, BUFFER_NUM> inQueue_;
165+TBuf<QuePosition::VECCALC> indexBuf_;
166+ 
167+// SIMD: Row-by-row processing, explicit data transfer and synchronization
168+for (int64_t j = 0; j < rows; j++) {
169+ INDICES_T index = GetIndex(yIdx, indiceEndIdx); // Scalar index read
170+ int64_t xIndex = index * tilingData_->innerSize;
171+ DataCopyPad(xLocal[j * colsAlign], xGm[offset], dataCoptExtParams, dataCopyPadExtParams); // Batch continuous data transfer in
172+}
173+inQueue_.EnQue<int8_t>(xLocal); // Enqueue for output
174+```
175+ 
176+The SIMT implementation uses a thread-level parallelism model, where each thread independently processes elements:
177+ 
178+```cpp
179+// SIMT: Uses thread-level parallelism, no explicit buffer management required
180+__simt_vf__ LAUNCH_BOUND(2048) void GatherSimt(...) {
181+ for (INDEX_SIZE_T index = Simt::GetThreadIdx();
182+ index < currentCoreElements;
183+ index += Simt::GetThreadNum()) { // Thread jump-style parallelism
184+ // Each thread independently computes a single-point index and accesses memory
185+ INDEX_SIZE_T gatherI = Simt::UintDiv(yIndex, m0, shift0);
186+ INDICES_T indicesValue = indices[gatherI]; // Directly access GM based on the single-point index gatherI
187+ y[yIndex] = idxOutOfBound ? 0 : x[xIndex]; // Directly write back to GM
188+ }
189+}
190+```
191+ 
192+**2. Memory Access Pattern Differences**
193+ 
194+| Feature | SIMD Implementation | SIMT Implementation |
195+|------|----------|----------|
196+| Data Access | Explicit data transfer to UB through DataCopyPad | Threads directly access GM through `__gm__` pointers |
197+| Buffer Management | AllocTensor/EnQue/DeQue/FreeTensor required | No explicit buffer required; hardware manages automatically |
198+| Synchronization Mechanism | Explicit event synchronization (HardEvent::MTE2_V, etc.) | Implicit synchronization between threads |
199+ 
200+**3. Applicable Scenario Differences**
201+ 
202+SIMD is suitable for scenarios with continuous access to large blocks of addresses, efficiently processing continuous data through vectorized instructions.
203+ 
204+SIMT is suitable for discrete memory access, with threads processing in parallel.
205+ 
206+#### Regbase
207+ 
208+The Ascend 950 series introduces the Regbase programming paradigm. Compared with the traditional Membase (Vector API) programming, Regbase is closer to the underlying hardware register operations and provides finer-grained vectorization control capabilities.
209+ 
210+**Features**
211+ 
212+- Uses the underlying APIs under the `AscendC::MicroAPI` namespace
213+- Directly operates on registers `RegTensor<T>` instead of explicitly managing UB buffer queues
214+- Implements flexible element-level mask control through `MaskReg`
215+ 
216+**Comparison with the Membase Programming Model**
217+ 
218+| Feature | Membase (Traditional Vector API) | Regbase (MicroAPI) |
219+|------|---------------------------|---------------------|
220+| Data Carrier | `LocalTensor<T>` + Queue mechanism | `RegTensor<T>` register |
221+| Memory Management | Explicit Alloc/EnQue/DeQue/Free | Register auto-allocation |
222+| Mask Control | Function parameter control | `MaskReg` register control |
223+| Data Transfer | `DataCopy`/`DataCopyPad` | `MicroAPI::DataCopy` + distribution mode |
224+ 
225+**Code Examples**
226+ 
227+```cpp
228+__simd_vf__ __aicore__ void GenIndexBuf(ubuf int32_t* helpAddr, int32_t colFactor)
229+{
230+ // Declare register tensors
231+ AscendC::MicroAPI::RegTensor<int32_t> v0;
232+ AscendC::MicroAPI::RegTensor<int32_t> v1;
233+ AscendC::MicroAPI::RegTensor<int32_t> vd1;
234+ 
235+ // Create a full mask
236+ AscendC::MicroAPI::MaskReg preg =
237+ AscendC::MicroAPI::CreateMask<int32_t, AscendC::MicroAPI::MaskPattern::ALL>();
238+ 
239+ // Duplicate scalar to register
240+ AscendC::MicroAPI::Duplicate(v1, colFactor, preg);
241+ // Generate sequence [0, 1, 2, ...]
242+ AscendC::MicroAPI::Arange(v0, 0);
243+ // Vector operations
244+ AscendC::MicroAPI::Div(vd1, v0, v1, preg);
245+ AscendC::MicroAPI::Mul(vd2, vd1, v1, preg);
246+ AscendC::MicroAPI::Sub(vd3, v0, vd2, preg);
247+ // Write register data back to UB
248+ AscendC::MicroAPI::DataCopy(helpAddr, vd3, preg);
249+}
250+```
251+ 
252+```cpp
253+// Dynamic mask: Handle tail incomplete data
254+__simd_vf__ __aicore__ void GatherProcess(ubuf int8_t* curYAddr, uint16_t repeatimes, uint16_t computeSize)
255+{
256+ MicroAPI::RegTensor<int8_t> vregTemp;
257+ MicroAPI::MaskReg preg;
258+ 
259+ for (uint16_t r = 0; r < repeatTimes; r++) {
260+ // Update mask based on the number of remaining elements
261+ preg = MicroAPI::UpdateMask<int8_t>(sreg);
262+ // Create an address offset register
263+ MicroAPI::AddrReg offset = MicroAPI::CreateAddrReg<int8_t>(r, computeSize);
264+ MicroAPI::DataCopy(vregTemp, curXAddr, offset);
265+ // Store data with mask
266+ MicroAPI::DataCopy(curYAddr, vregTemp, offset, preg);
267+ }
268+}
269+```
270+ 
271+```cpp
272+// Data aggregation
273+__VEC_SCOPE__
274+{
275+ MicroAPI::RegTensor<uint32_t> indicesReg;
276+ MicroAPI::RegTensor<int32_t> vd0;
277+ 
278+ for (uint16_t indices = 0; indices < indicesLoopNum; indices++) {
279+ // Load indices (E2B distribution mode: broadcast scalar to vector)
280+ MicroAPI::DataCopy<uint32_t, MicroAPI::LoadDist::DIST_E2B_B32>(indicesReg, indicesAddr);
281+ // Gather data aggregation based on indices
282+ MicroAPI::DataCopyGather(vd0, curXAddr, indicesReg, preg);
283+ // Data block copy output
284+ MicroAPI::DataCopy<int32_t, MicroAPI::DataCopyMode::DATA_BLOCK_COPY>(
285+ curYAddr, vd0, blockStride, preg);
286+ }
287+}
288+```
289+ 
290+**Key Regbase API Descriptions**
291+ 
292+| API Category | API Name | Function Description |
293+|---------|---------|----------|
294+| Register Type | `RegTensor<T>` | Vector register tensor type |
295+| Mask Type | `MaskReg` | Mask register type |
296+| Mask Creation | `CreateMask<T, Pattern>()` | Create a mask (ALL/HALF and other patterns) |
297+| Mask Update | `UpdateMask<T>(count)` | Dynamically update the mask based on the number of remaining elements |
298+| Scalar Operation | `Duplicate(reg, val, mask)` | Duplicate a scalar value to all elements of a register |
299+| Sequence Generation | `Arange(reg, start)` | Generate a continuous sequence |
300+| Arithmetic Operations | `Add/Sub/Mul/Div(dst, src1, src2, mask)` | Vector arithmetic operations |
301+| Scalar Operations | `Adds/Muls(dst, src, scalar, mask)` | Vector and scalar operations |
302+| Type Conversion | `Cast<DT, ST>(dst, src, mask)` | Data type conversion |
303+| Comparison Operations | `Compare<T, CMPMODE>(mask, src1, src2, pred)` | Vector comparison generates a mask |
304+| Data Load | `DataCopy<T, LoadDist>(reg, addr)` | Load from UB to register |
305+| Data Store | `DataCopy<T>(addr, reg, mask)` | Store from register to UB |
306+| Gather | `DataCopyGather(dst, base, indices, mask)` | Collect data based on indices |
307+| Address Offset | `CreateAddrReg<T>(loop, stride)` | Create a loop address offset register |
308+ 
309+**LoadDist (Distribution Mode) Descriptions**
310+ 
311+| Mode | Description | Typical Use |
312+|------|------|----------|
313+| `DIST_NORM` | Normal continuous load | Continuous data processing |
314+| `DIST_UNPACK_B16` | 16-bit unpack load | FP16/BF16 to FP32 conversion |
315+| `DIST_BRC_B32/B16` | Broadcast load | Scalar scale broadcast |
316+| `DIST_E2B_B32` | Scalar to vector broadcast | Index value broadcast |
317+ 
318+**Migration Suggestions**
319+ 
320+1. Scenarios suitable for Regbase: Cases requiring fine-grained control of register allocation, complex mask logic, and Gather/Scatter memory access patterns.
321+2. Scenarios to retain Membase: Simple continuous data transfer and computation, double-buffered pipelines.
322+3. Hybrid use: Combine both paradigms within the same operator, using Regbase to handle core computation logic and Membase to manage data transfer.
323+ 
324+### Cube-Vector Fusion Operators
325+ 
326+#### MTE Data Transfer Path Changes
327+ 
328+The Ascend 950 new architecture introduces direct connection paths between UB-to-L1 and L0C-to-UB, enabling fast transfer of matrix computation data. This aims to simplify CV fusion operator development and improve performance.
329+<div align="center">
330+ <img src="../figures/Ascend950_CV_passthrough_link.png" width="700" alt="Ascend 950 New CV Passthrough Link" />
331+</div>
332+ 
333+**Matrix Move-In**
334+ 
335+Enable the UB-to-L1 (UB2L1) direct connection path. Through the DataCopy interface, the vector computation results of fusion operators can be directly moved into L1.
336+ 
337+**Matrix Move-Out**
338+ 
339+Enable the L0C-to-UB (L0C2UB) direct connection path. Through the DataCopy interface, the matrix computation results of fusion operators can be directly moved into UB for subsequent vector computation.
340+ 
341+For K-split or multi-stage fusion scenarios, the "L0C moved back to GM and then read back to UB" approach can be changed to "L0C directly to UB for accumulation/post-processing," reducing GM round-trip bandwidth pressure and latency. During migration, it is recommended to place intermediate result merging and activation/quantization pre-processing on the UB side, and explicitly sort out the event synchronization order of MTE1/MTE2/MTE3 and compute units to ensure continuous cross-unit pipelines, avoiding data visibility or synchronization timing issues introduced by the new paths. For the key enabling interface definitions, refer to:
342+ 
343+```cpp
344+// 1. New: The move-in interface adds UB2L1 Nd2Nz move-in, supporting the form where both Src and Dst are LocalTensor
345+template <typename T>
346+__aicore__ inline void DataCopy(const LocalTensor<T>& dst, const LocalTensor<T>& src, const Nd2NzParams& intriParams);
347+ 
348+// 2. New: The move-out interface adds L0C2UB move-out, supporting direct move-out from L0C to UB, supporting the form where both Src and Dst are LocalTensor
349+template <typename T, typename U, const FixpipeConfig& config = CFG_ROW_MAJOR>
350+__aicore__ inline void Fixpipe(const LocalTensor<T>& dst, const LocalTensor<U>& src, const FixpipeParamsC310<config.format>& intriParams);
351+template <CO2Layout format = CO2Layout::ROW_MAJOR>
352+struct FixpipeParamsC310 {
353+ // ...
354+ uint8_t dualDstCtl = 0;
355+};
356+ 
357+// 3. Capability enhancement: The cross-core synchronization interface adds mode 3
358+template <uint8_t modeId, pipe_t pipe>
359+__aicore__ inline void CrossCoreSetFlag(uint16_t flagId)
360+template <uint8_t modeId = 0, pipe_t pipe = PIPE_S>
361+__aicore__ inline void CrossCoreWaitFlag(uint16_t flagId)
362+ 
363+```
364+ 
365+#### Cross-Core Synchronization Semaphore Matching
366+ 
367+`CrossCoreSetFlag` and `CrossCoreWaitFlag` are cross-core synchronization semaphore interfaces, widely used for data dependency and collaborative control between multiple cores. They essentially decouple and orderly advance data processing stages between different AICores in the form of "semaphores," and are commonly used in scenarios such as pipeline control, double-buffer switching, and cross-core collaboration.
368+ 
369+- `CrossCoreSetFlag`: After the current core (or thread) completes data processing for a certain stage, it actively sets the specified flag signal to inform the dependent party (generally another core or downstream pipeline stage) that this stage is complete and the subsequent process can continue.
370+- `CrossCoreWaitFlag`: The current core (or thread) needs to wait for a certain flag signal to be set (that is, the dependent data or event is complete). After detecting the flag, it continues to execute downward.
371+ 
372+The essence of this semaphore mechanism is to ensure a consistent synchronization sequence between multiple threads or pipeline stages, preventing hardware exceptions such as data races or deadlocks caused by resources not being ready or dependencies not being completed. For detailed interface descriptions, refer to the official documentation: [CrossCoreSetFlag and CrossCoreWaitFlag Cross-Core Synchronization Interface Details](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/900beta1/API/ascendcopapi/atlasascendc_api_07_0273.html).
373+ 
374+On Ascend 950, the numbers of `CrossCoreWaitFlag` and `CrossCoreSetFlag` must strictly match, and it is recommended to design them in pairs within the same synchronization semantic domain, following the "produce first, then consume" order. On Atlas A2, if there are redundant `CrossCoreSetFlag` semaphores between operators, HWTS performs special handling to clear the counter. The Ascend 950 series, to reduce hardware overhead, no longer relies on this type of fallback mechanism, requiring that the synchronization semaphores within a single operator kernel match one-to-one. Otherwise, a deterministic hang will occur.
375+ 
376+During migration, focus on troubleshooting the following issues: first, abnormal branch early returns that cause only `Set` to be executed without the corresponding `Wait` (or vice versa); second, multi-stage pipelines reusing the same `flagId` but with overlapping lifecycles, causing "cross-stage crosstalk"; third, conditionally triggered synchronization within loops but with unaligned loop boundaries, resulting in inconsistent iteration counts. The above issues may be masked on Atlas A2 but directly exposed as blocking timeouts or deadlocks on Ascend 950. For operators with complex cross-core pipelines, first build a minimal dataset for single-stage verification, and then gradually add double buffering and multiple stages to reduce the complexity of locating synchronization issues.
377+ 
378+### Collective Communication Operators
379+ 
380+Ascend 950 introduces the collective communication accelerator CCU1.0, which reduces memory access requirements and scheduling latency. To effectively utilize this feature, the cross-chip communication method of operators is changed from AICPU on A2 to CCU communication.
381+ 
382+**Eager Mode**
383+ 
384+In the second-stage interface of the aclnn two-phase interface, specify the collective communication type for the operator executor aclOpExecutor.
385+ 
386+Taking the [MatmulAllReduce](https://gitcode.com/cann/ops-transformer/tree/master/mc2/matmul_all_reduce) operator migration as an example:
387+Set the NnopbaseSetHcclServerType enum value. For A2, it is NNOPBASE_HCCL_SERVER_AICPU, and for 950, it is NNOPBASE_HCCL_SERVER_TYPE_CCU.
388+ 
389+```CPP
390+// ...
391+aclnnStatus aclnnMatmulAllReduce(
392+ void* workspace, uint64_t workspaceSize, aclOpExecutor* executor, const aclrtStream stream)
393+{
394+ // ...
395+ if (NnopbaseSetHcclServerType) {
396+ if (op::GetCurrentPlatformInfo().GetCurNpuArch() == NpuArch::DAV_3510) {
397+ NnopbaseSetHcclServerType(executor, NnopbaseHcclServerType::NNOPBASE_HCCL_SERVER_TYPE_CCU);
398+ }
399+ }
400+ // ...
401+ return ACLNN_SUCCESS;
402+}
403+```
404+ 
405+**Graph Mode**
406+ 
407+1. In the CalcParamFunc callback interface used for resource computation and application, which involves auxiliary stream-related information, differentiate the collective communication type of the auxiliary stream for the GE context context.
408+2. In the GenerateTask callback interface used for setting custom tasks and parameter customization on the main stream and auxiliary streams, differentiate between the two sets of GE KernelLaunch interfaces, and call the AICPU communication or CCU communication creation and customization processes respectively.
409+ 
410+For the static graph GE side, the task type for creating communication tasks is aicpu kfc server + kfc_stream for A2, and ccu server + ccu_stream for 950. The related code file is: [matmul_all_reduce_gen_task.cpp](https://gitcode.com/cann/ops-transformer/blob/master/mc2/matmul_all_reduce/op_graph/matmul_all_reduce_gen_task.cpp)
411+ 
412+```CPP
413+// ...
414+ge::Status MatmulAllReduceCalcParamFunc(gert::ExeResGenerationContext *context)
415+{
416+ if (Mc2GenTaskOpsUtils::IsTargetPlatformNpuArch(context->GetNodeName(), NPUARCH_A5)) {
417+ // 950
418+ return Mc2GenTaskOpsUtils::CommonKFCMc2CalcParamFunc(context, "ccu server", "ccu_stream");
419+ }
420+ // A2
421+ return Mc2GenTaskOpsUtils::CommonKFCMc2CalcParamFunc(context, "aicpu kfc server", "kfc_stream");
422+}
423+// ...
424+```
425+ 
426+The static graph GenTask invocation interfaces differ, and the processes are different. The related code file is: [matmul_all_reduce_gen_task.cpp](https://gitcode.com/cann/ops-transformer/blob/master/mc2/matmul_all_reduce/op_graph/matmul_all_reduce_gen_task.cpp)
427+ 
428+```CPP
429+// ...
430+// A2
431+ge::Status MatmulAllReduceGenTaskOpsUtils::MatmulAllReduceGenTaskCallback(
432+ const gert::ExeResGenerationContext *context, std::vector<std::vector<uint8_t>>& tasks) {
433+ // ...
434+ // aicpu task
435+ ge::KernelLaunchInfo aicpu_task =
436+ ge::KernelLaunchInfo::CreateAicpuKfcTask(context, SO_NAME.c_str(), KERNEL_NAME_V1.c_str());
437+ // ...
438+}
439+ 
440+// 950
441+ge::Status Mc2Arch35GenTaskOpsUtils::Mc2Arch35GenTaskCallBack(const gert::ExeResGenerationContext *context, std::vector<std::vector<uint8_t>> &tasks) {
442+ // ...
443+ // ccu task
444+ ge::KernelLaunchInfo ccuTask = ge::KernelLaunchInfo::CreateCcuTask(context, ccuGroups);
445+ // ...
446+}
447+ 
448+ge::Status MatmulAllReduceGenTaskFunc(const gert::ExeResGenerationContext *context, std::vector<std::vector<uint8_t>> &tasks)
449+{
450+ if (Mc2GenTaskOpsUtils::IsTargetPlatformNpuArch(context->GetNodeName(), NPUARCH_A5)) {
451+ // 950
452+ return Mc2Arch35GenTaskOpsUtils::Mc2Arch35GenTaskCallBack(context, tasks);
453+ }
454+ // A2
455+ return MatmulAllReduceGenTaskOpsUtils::MatmulAllReduceGenTaskCallback(context, tasks);
456+}
457+// ...
458+```
459+ 
460+## V. Common Issues and Performance Tuning Suggestions (FAQ/Performance Tips)
461+ 
462+If the performance of an operator on Ascend 950 does not improve but declines, prioritize the following checks:
463+ 
464+1. Whether the Atlas A2 offset-based multi-core template is still being used.
465+2. Whether CCU communication is not enabled and AICPU is still being used.
466+3. Whether the tiling still uses the Atlas A2 L1/L0/UB partitioning strategy, resulting in the larger on-chip cache of Ascend 950 not being fully utilized.
@@ -0,0 +1,118 @@
1+# Graph Mode Adaptation Guide
2+ 
3+## Overview
4+ 
5+If a custom operator needs to run in graph mode, the overall process is consistent with the operator development guide ([AI Core Operator Development Guide](aicore_develop_guide.md)/[AI CPU Operator Development Guide](aicpu_develop_guide.md)). Note that **aclnn adaptation is not required**; only the following deliverable adaptations are needed.
6+ 
7+```text
8+${op_name} # Replace with the lowercase underscore form of the actual operator name
9+├── op_host # Host-side implementation
10+│ └── ${op_name}_infershape.cpp # InferShape implementation, implementing operator shape inference, inferring the output shape at runtime
11+├── op_graph # Graph fusion-related implementation
12+│ ├── CMakeLists.txt # op_graph side cmakelist file
13+│ ├── ${op_name}_graph_infer.cpp # InferDataType file, implementing operator type inference, inferring the output dataType at runtime
14+└── └── ${op_name}_proto.h # Operator prototype definition, used for graph optimization and fusion phase operator identification
15+```
16+ 
17+This document uses the `AddExample` operator (assuming it is an AI Core operator) as an example to introduce the implementation of graph mode deliverables. The implementation for AI CPU operators entering the graph is basically similar. For complete code, refer to the `add_example` and `add_example_aicpu` directories under `examples`.
18+ 
19+## Shape and DataType Inference
20+ 
21+Graph mode requires two deliverables: `${op_name}_graph_infer.cpp` and `${op_name}_infershape.cpp`
22+ 
23+**Deliverable 1: ${op_name}_infershape.cpp**
24+ 
25+The InferShape function infers the output shape based on the input shape.
26+ 
27+The example is as follows. For the complete code of the `AddExample` operator, refer to [add_example_infershape.cpp](../../../examples/add_example/op_host/add_example_infershape.cpp) under `examples/add_example/op_host`.
28+ 
29+```C++
30+// The AddExample operator logic is adding two numbers, so the output shape is consistent with the input shape
31+static ge::graphStatus InferShapeAddExample(gert::InferShapeContext* context)
32+{
33+ ....
34+ // Obtain the input shape
35+ const gert::Shape* xShape = context->GetInputShape(IDX_0);
36+ // Obtain the output shape
37+ gert::Shape* yShape = context->GetOutputShape(IDX_0);
38+ // Obtain the input DimNum
39+ auto xShapeSize = xShape->GetDimNum();
40+ // Set the output DimNum
41+ yShape->SetDimNum(xShapeSize);
42+ // Set the input Dim values to the output one by one
43+ for (size_t i = 0; i < xShapeSize; i++) {
44+ int64_t dim = xShape->GetDim(i);
45+ yShape->SetDim(i, dim);
46+ }
47+ ....
48+}
49+// InferShape registration
50+IMPL_OP_INFERSHAPE(AddExample).InferShape(InferShapeAddExample);
51+```
52+ 
53+**Deliverable 2: ${op_name}_graph_infer.cpp**
54+ 
55+The InferDataType function infers the output DataType based on the input DataType. The example is as follows.
56+ 
57+```C++
58+// The AddExample operator logic is adding two numbers, so the output dataType is consistent with the input dataType
59+static ge::graphStatus InferDataTypeAddExample(gert::InferDataTypeContext* context)
60+{
61+ ....
62+ // Obtain the input dataType
63+ ge::DataType sizeDtype = context->GetInputDataType(IDX_0);
64+ // Set the input dataType to the output
65+ context->SetOutputDataType(IDX_0, sizeDtype);
66+ ....
67+}
68+ 
69+// Register InferDataType
70+IMPL_OP(AddExample).InferDataType(InferDataTypeAddExample);
71+```
72+ 
73+## Operator Prototype Configuration
74+ 
75+Graph mode invocation requires registering the operator prototype into [Graph Engine](https://www.hiascend.com/eng/cann/graph-engine) (abbreviated as GE) so that GE can identify the input, output, and attribute information of this type of operator. Registration is completed through the `REG_OP` interface. Developers need to define basic information such as the operator input, output tensor types, and quantities.
76+ 
77+Common tensor/attribute data type examples are as follows:
78+ 
79+|Tensor Type|Attribute Type|Example|
80+|-----|------|-----|
81+|int64|/|DT_INT64|
82+|int32|/|DT_INT32|
83+|int16|/|DT_INT16|
84+|int8|/|DT_INT8|
85+|double|/|DT_DOUBLE|
86+|float32|/|DT_FLOAT|
87+|float16|/|DT_FLOAT16|
88+|bfloat16|/|DT_BF16|
89+|complex128|/|DT_COMPLEX128|
90+|complex64|/|DT_COMPLEX64|
91+|complex32|/|DT_COMPLEX32|
92+|/|int|Int|
93+|/|bool|Bool|
94+|/|string|String|
95+|/|float|Float|
96+|/|list|ListInt|
97+ 
98+Basic information is as follows:
99+ 
100+|Input/Output|Keyword|Example|
101+|-----|------|-----|
102+|Required input|INPUT|.INPUT(${name}, TensorType({input_dtype}))|
103+|Optional input|OPTIONAL_INPUT|.OPTIONAL_INPUT(${name}, TensorType({optional_input_dtype}))|
104+|Required attribute|REQUIRED_ATTR|.REQUIRED_ATTR(${name}, ${dtype})|
105+|Optional attribute|ATTR|.ATTR(${name}, ${dtype}, ${default_value})|
106+|Output|OUTPUT|.OUTPUT(${name}, TensorType({output_dtype}))|
107+ 
108+The sample code below shows how to register the `AddExample` operator:
109+ 
110+```CPP
111+REG_OP(AddExample)
112+ .INPUT(x1, TensorType({DT_FLOAT}))
113+ .INPUT(x2, TensorType({DT_FLOAT}))
114+ .OUTPUT(y, TensorType({DT_FLOAT}))
115+ .OP_END_FACTORY_REG(AddExample)
116+```
117+ 
118+For complete code, refer to [add_example_proto.h](../../../examples/add_example/op_graph/add_example_proto.h) under the `examples/add_example/op_graph` directory.
@@ -0,0 +1,70 @@
1+# build Parameter Description
2+ 
3+## Introduction
4+ 
5+build.sh is the build script of this project, located in the project root directory by default. Its function is to automatically compile, link, and configure the source code, and finally generate executable files, library files, or other target files that can be installed or run directly. Specifically, the script configures different parameters to achieve multiple functions, including building multiple target libraries (such as libophost_nn.so), compiling operator packages, executing unit tests, etc.
6+ 
7+## Usage
8+ 
9+1. **Configure Environment Variables**
10+ 
11+ Complete the basic environment setup by referring to [Environment Deployment](../context/quick_install.md).
12+ 
13+ ```bash
14+ # Default path installation, taking root user as an example
15+ source /usr/local/Ascend/cann/set_env.sh
16+ ```
17+ 
18+2. **Build Command Format**
19+ 
20+ Taking the compile operator package command as an example, the format is as follows, where `--vendor_name` and `--ops` are optional in this scenario.
21+ 
22+ ```bash
23+ bash build.sh --pkg --soc=${soc_version} [--vendor_name=${vendor_name}] [--ops=${op_list}]
24+ ```
25+ 
26+ For the meaning of all parameters, refer to the parameter description section below. Choose the appropriate parameters according to the actual situation.
27+ 
28+## Parameter Description
29+ 
30+build.sh supports multiple functions. You can view all function parameters through the following command.
31+ 
32+```bash
33+bash build.sh --help
34+```
35+ 
36+| Parameter Name | Optional/Required | Parameter Description |
37+|------------------|--------|-----------------------------------------------------------------------------|
38+| -j${n} | Optional | Specifies the number of compilation threads. ${n} is the specific number of threads. The default value is 8 (such as -j8). If the number of threads exceeds the number of CPU cores, it will be automatically adjusted to the number of CPU cores. |
39+| -v | Optional | View CMake compilation configuration information. |
40+| -O${n} | Optional | Specifies the compilation optimization level. Supports O0/O1/O2/O3 (such as -O3). ${n} is the optimization level identifier. |
41+| -u | Optional | Enables unit test (UT) compilation mode and compiles all UT targets. |
42+| --help, -h | Optional | Prints script usage help information. |
43+| --ops | Optional | Specifies the operators to be compiled, such as mat_mul_v3, mse_loss. Multiple operators are separated by English commas ",". Cannot be used with --ophost and --opapi at the same time. |
44+| --soc | Optional | Specifies the NPU model. Only 1 NPU model is supported per compilation. |
45+| --jit | Optional | In the static graph scenario, when compiling the `cann-${soc_name}-ops-nn_${cann_version}_linux-${arch}.run` package, you do not need to compile the operator binary files (the graph runtime will compile online). You can configure this option to improve compilation speed. |
46+| --static | Optional | When configured, it means generating a static library file, including libcann_nn_static.a and aclnn interface header files. Combined with the --pkg parameter, it generates a static library compressed package.|
47+| --vendor_name | Optional | Specifies the name of the custom operator package. The default value is custom. |
48+| --build-type | Optional | Enables debug mode. Optional types: Release/Debug. The default is Release. When the value is Debug, it cannot be used with --mssanitizer, --oom, --dump_cce at the same time |
49+| --debug | Optional | Enables debug mode. |
50+| --cov | Optional | Reserved parameter, developers do not need to pay attention for now. |
51+| --noexec | Optional | Only compiles the unit test binary file without automatically executing the compiled UT executable file. |
52+| --opkernel | Optional | Compiles the binary kernel. |
53+| --pkg | Optional | Generates the installation package. Cannot be used with -u (UT mode) or --ophost, --opapi at the same time. |
54+| --asan | Optional | Enables host-side ASAN (AddressSanitizer) memory detection function. |
55+| --valgrind | Optional | Reserved parameter, developers do not need to pay attention for now. |
56+| --make_clean | Optional | Executes basic cleanup operations (cleans compilation products). The script exits after execution. |
57+| --make_clean_all | Optional | Executes complete cleanup operations (deletes all compilation-related files). The script exits after execution. |
58+| --ophost | Optional | Compiles the libophost_nn.so library. Cannot be used with --pkg, --ops at the same time. |
59+| --opapi | Optional | Compiles the libopapi_nn.so library. Cannot be used with --pkg, --ops at the same time. |
60+| --run_example | Optional | Compiles the sample of the specified operator and mode and executes the compiled executable file. Use --run_example --help to view the usage. |
61+| --genop | Optional | Creates the AI Core custom operator initial directory. |
62+| --genop_aicpu | Optional | Creates the AI CPU custom operator initial directory. |
63+| --experimental | Optional | Compiles user operators in the experimental directory. |
64+| --mssanitizer | Optional | Enables kernel-side mssanitizer memory detection function. |
65+| --oom | Optional | Enables kernel-side oom memory detection function. |
66+| --dump_cce | Optional | Enables kernel-side dump precompiled file function. |
67+| --cann_3rd_lib_path| Optional | The directory where third-party libraries are stored in the offline compilation scenario. |
68+| --simulator | Optional | Used in combination with --run_example to enable simulator mode to execute --run_example tasks. In simulator mode, the corresponding simulator library will be linked according to soc_version. |
69+| --bisheng_flags | Optional | Specifies the BiSheng compiler compilation parameters. Multiple compilation parameters are separated by English commas ",". Cannot be used with --mssanitizer, --oom, --dump_cce at the same time. |
70+| --kernel_template_input | Optional | Specifies the tilingKey template when compiling the kernel. Only one template can be specified. Used with --ops and only one operator can be specified. It will not compile the binary files of other operators that this operator depends on. |