已合并
docs: add comprehensive English documentation set #76
LiuZonggu创建于 8月1日
docs: add comprehensive English documentation set #76
已合并
从已删除 :docs/eng-ver合入到cann/ops-rasmaster
共 93 个文件变更+5091-0
| @@ -0,0 +1,60 @@ | |||
| 1 | +# Contribution Guide | ||
| 2 | + | ||
| 3 | +This project welcomes developers to experience and participate in contributions. Before participating in community contributions, please see [cann-community](https://gitcode.com/cann/community) to understand the code of conduct, sign the CLA agreement, and understand the contribution process of the source code repository. | ||
| 4 | + | ||
| 5 | +Developers need to pay attention to the following points when preparing local code and submitting PRs: | ||
| 6 | + | ||
| 7 | +1. When submitting a PR, please carefully fill in the business background, purpose, solution, and other information of this PR according to the PR template. | ||
| 8 | +2. If your modification is not a simple bug fix, but involves adding new features, new interfaces, new configuration parameters, or modifying code flow, please be sure to discuss the solution through an Issue first to avoid your code being rejected. If you are not sure whether this modification can be classified as a "simple bug fix", you can also discuss the solution by submitting an Issue. | ||
| 9 | + | ||
| 10 | +Developer contribution scenarios mainly include: | ||
| 11 | + | ||
| 12 | +- Operator Bug Fix | ||
| 13 | + | ||
| 14 | + If you discover certain operator bugs in this project and want to fix them, we welcome you to create a new Issue for feedback and tracking. | ||
| 15 | + | ||
| 16 | + You can create a new `Bug-Report|Bug Report` type Issue according to the [Submit Issue/Handle Issue Task](https://gitcode.com/cann/community#提交Issue处理Issue任务) guide to describe the bug, and then enter "/assign" or "/assign @yourself" in the comment box to assign this Issue to you for processing. | ||
| 17 | + | ||
| 18 | +- Operator Optimization | ||
| 19 | + | ||
| 20 | + If you have generalization enhancement/performance optimization ideas for certain operator implementations in this project and want to implement these optimization points, we welcome you to contribute operator optimizations. | ||
| 21 | + | ||
| 22 | + You can create a new `Requirement|Feature Request` type Issue according to the [Submit Issue/Handle Issue Task](https://gitcode.com/cann/community#提交Issue处理Issue任务) guide to explain the optimization points and provide your design solution, and then enter "/assign" or "/assign @yourself" in the comment box to assign this Issue to you for tracking optimization. | ||
| 23 | + | ||
| 24 | +- Contribute New Operators | ||
| 25 | + | ||
| 26 | + If you have a brand new operator that you want to design and implement based on NPU, we welcome you to propose new ideas and designs in an Issue. | ||
| 27 | + | ||
| 28 | + You can create a new `Requirement|Feature Request` type Issue according to the [Submit Issue/Handle Issue Task](https://gitcode.com/cann/community#提交Issue处理Issue任务) guide to provide the new operator description and design solution. Project members will communicate and confirm with you, and provide a suitable `contrib` directory classification for your operator under the `experimental` directory. You can contribute the new operator to the corresponding directory. | ||
| 29 | + | ||
| 30 | + At the same time, you need to comment "/assign" or "/assign @yourself" in the submitted Issue to claim this Issue and subsequently complete the new operator submission. | ||
| 31 | + | ||
| 32 | + The deliverables for new operators are usually quite numerous. You can refer to the following list to check the minimum deliverable set, where `${op_name}` indicates the new operator name. | ||
| 33 | + ``` | ||
| 34 | + ${op_class} # operator classification | ||
| 35 | + ├── ${op_name} # operator name | ||
| 36 | + │ ├── op_host # operator definition, Tiling, InferShape related implementation | ||
| 37 | + │ │ ├── ${op_name}_def.cpp # operator definition file | ||
| 38 | + │ │ ├── ${op_name}_tiling.cpp # operator Tiling implementation file | ||
| 39 | + │ │ └── CMakeLists.txt | ||
| 40 | + │ ├── op_kernel # operator Kernel directory | ||
| 41 | + │ │ ├── ${op_name}.cpp | ||
| 42 | + │ │ ├── ${op_name}.h | ||
| 43 | + │ │ ├── ${op_name}_tiling_data.h | ||
| 44 | + │ │ ├── ${op_name}_tiling_key.h | ||
| 45 | + │ │ └── CMakeLists.txt | ||
| 46 | + │ ├── CMakeLists.txt # operator compilation configuration file, keep the original file | ||
| 47 | + │ └── README.md # operator description document | ||
| 48 | + ``` | ||
| 49 | + | ||
| 50 | +- Document Correction | ||
| 51 | + | ||
| 52 | + If you discover certain operator document description errors in this project, we welcome you to create a new Issue for feedback and correction. | ||
| 53 | + | ||
| 54 | + You can create a new `Documentation|Documentation Feedback` type Issue according to the [Submit Issue/Handle Issue Task](https://gitcode.com/cann/community#提交Issue处理Issue任务) guide to point out the problems in the corresponding document, and then enter "/assign" or "/assign @yourself" in the comment box to assign this Issue to you to correct the corresponding document description. | ||
| 55 | + | ||
| 56 | +- Help Solve Others' Issues | ||
| 57 | + | ||
| 58 | + If you have suitable solutions for problems encountered by others in the community, we welcome you to comment and communicate in the Issue to help others solve problems and pain points, and jointly optimize usability. | ||
| 59 | + | ||
| 60 | + If the corresponding Issue requires code modification, you can enter "/assign" or "/assign @yourself" in the Issue comment box to assign this Issue to you for tracking and assisting in solving the problem. | ||
| @@ -0,0 +1,51 @@ | |||
| 1 | +# ops-ras | ||
| 2 | + | ||
| 3 | +## 🔥Latest News | ||
| 4 | + | ||
| 5 | +- [2026/07] The ops-ras project was first released. | ||
| 6 | + | ||
| 7 | +## 🚀Overview | ||
| 8 | + | ||
| 9 | +ops-ras is the security and RAS (Reliability, Availability and Serviceability) operator library in the [CANN](https://hiascend.com/software/cann) (Compute Architecture for Neural Networks) operator library, providing reliability, availability, and maintainability capabilities, including security, encryption, and RAS-related operators. "ras" is derived from the initials of these three core characteristics. The operator library architecture is shown below: | ||
| 10 | + | ||
| 11 | +<img src="docs/zh/figures/architecture.png" alt="Architecture Diagram" width="700px" height="320px"> | ||
| 12 | + | ||
| 13 | +## 📌Version Compatibility | ||
| 14 | + | ||
| 15 | +The source code of this project will be released along with the CANN software version. For the correspondence between CANN software versions and project tags, refer to the relevant version descriptions in the [release repository](https://gitcode.com/cann/release-management). | ||
| 16 | +Note that to ensure smooth custom development of your source code, select the matching CANN version and Gitcode tag source code. Using the master branch may pose version mismatch risks. | ||
| 17 | + | ||
| 18 | +## 🛠️Environment Setup | ||
| 19 | + | ||
| 20 | +[Environment Deployment](docs/zh/install/quick_install.md) is the prerequisite for experiencing the capabilities of this project. Please complete the NPU driver installation, CANN package installation, and so on to ensure the environment is normal. | ||
| 21 | + | ||
| 22 | +## ⬇️Source Code Download | ||
| 23 | + | ||
| 24 | +After the environment is ready, download the branch source code matching the CANN version. The general command is as follows. Replace `${tag_version}` with the branch tag name. Take the 9.0.0 branch source code download as an example: | ||
| 25 | + | ||
| 26 | +```bash | ||
| 27 | +# General command: git clone -b ${tag_version} https://gitcode.com/cann/ops-ras.git | ||
| 28 | +git clone -b 9.0.0 https://gitcode.com/cann/ops-ras.git | ||
| 29 | +``` | ||
| 30 | + | ||
| 31 | +> Note: If the matching branch source code already exists in the environment, **you can skip this step**. For example, CANNLab provides the source code matching the latest CANN version by default. | ||
| 32 | + | ||
| 33 | +## 📖Learning Tutorials | ||
| 34 | + | ||
| 35 | +- [Quick Start](docs/QUICKSTART.md): Quickly experience the core basic capabilities of the project from scratch, covering source code compilation, operator invocation, development, debugging, and other operations. | ||
| 36 | +- [Advanced Tutorials](docs/README.md): If you need a deeper understanding of the project's compilation and deployment, operator invocation, development, debugging and tuning, and other capabilities, please refer to the documentation center for detailed guidance. | ||
| 37 | + | ||
| 38 | +## 💬Related Information | ||
| 39 | + | ||
| 40 | +- [Directory Structure](docs/zh/install/dir_structure.md) | ||
| 41 | +- [Contribution Guide](CONTRIBUTING.md) | ||
| 42 | +- [Security Statement](SECURITY.md) | ||
| 43 | +- [License](LICENSE) | ||
| 44 | +- [Affiliated SIG](https://gitcode.com/cann/community/tree/master/CANN/sigs/ops-basic) | ||
| 45 | + | ||
| 46 | +----- | ||
| 47 | +PS: The functions and documentation of this project are being continuously updated and improved. We recommend that you follow the latest version. | ||
| 48 | + | ||
| 49 | +- **Issue Feedback**: Submit issues through GitCode [Issues](https://gitcode.com/cann/ops-ras/issues). | ||
| 50 | +- **Community Interaction**: Participate in discussions through GitCode [Discussions](https://gitcode.com/cann/ops-ras/discussions). | ||
| 51 | +- **Technical Column**: Access technical articles through GitCode [Wiki](https://gitcode.com/cann/ops-ras/wiki), such as serialized tutorials and best practices. | ||
| @@ -0,0 +1,65 @@ | |||
| 1 | +# Security Statement | ||
| 2 | + | ||
| 3 | +## Running User Recommendations | ||
| 4 | + | ||
| 5 | +Based on security considerations, we do not recommend using root or other administrator type accounts to execute any commands. Follow the principle of minimum permissions. | ||
| 6 | + | ||
| 7 | +## File Permission Control | ||
| 8 | + | ||
| 9 | +- We recommend that users set the running system umask value to 0027 or above on the host machine (including the host machine) and in the container to ensure that the default maximum permission for new folders is 750 and the default maximum permission for new files is 640. | ||
| 10 | +- We recommend that users take security measures such as permission control for sensitive content such as personal privacy data, business assets, source files, and various files saved during operator development. For example, for project installation directory permission control and input public data file permission control, the set permissions should refer to [A-File (Folder) Permission Control Recommended Maximum Values in Various Scenarios](#a-file-folder-permission-control-recommended-maximum-values-in-various-scenarios). | ||
| 11 | +- When the operator runs, it may cache operator compilation files, which are stored in the `kernel_meta_*` folder under the running directory to speed up subsequent operator invocation. Users can perform permission control on the generated related files as needed. | ||
| 12 | +- Users need to perform permission control during installation and use. We recommend referring to [A-File (Folder) Permission Control Recommended Maximum Values in Various Scenarios](#a-file-folder-permission-control-recommended-maximum-values-in-various-scenarios) for file permission reference settings. | ||
| 13 | + | ||
| 14 | +## Build Security Statement | ||
| 15 | + | ||
| 16 | +When compiling and installing this project from source code, you need to compile it yourself. During the compilation process, some intermediate files will be generated. We recommend that you perform permission control on the intermediate files after compilation to ensure file security. | ||
| 17 | + | ||
| 18 | +## Running Security Statement | ||
| 19 | + | ||
| 20 | +- We recommend that users write corresponding operator invocation scripts based on the running environment resource status. If the operator invocation script does not match the resource status, such as the space used for generating input data or benchmark calculation results exceeding the memory capacity limit, or the script saving data locally exceeding the disk space size, it may cause errors and lead to unexpected process exit. | ||
| 21 | +- When the operator runs abnormally, it will exit the process and print error information. We recommend locating the specific error cause based on the error prompt, including setting operator synchronous execution, viewing log files, and other methods. | ||
| 22 | +- When the operator is invoked through [PyTorch](https://gitee.com/ascend/pytorch), running errors may occur due to version mismatch. For details, please refer to [PyTorch Security Statement](https://gitee.com/ascend/pytorch#%E5%AE%89%E5%85%A8%E5%A3%B0%E6%98%8E). | ||
| 23 | + | ||
| 24 | +## Public Network Address Statement | ||
| 25 | + | ||
| 26 | +The public network addresses contained in this project code are declared as follows: | ||
| 27 | + | ||
| 28 | +| Type | Open Source Code Address | File Name | Public Network IP Address/Public Network URL Address/Domain Name/Email Address/Compressed File Address | Usage Description | | ||
| 29 | +| :------------: |:------------------------------------------------------------------------------------------:|:----------------------------------------------------------| :---------------------------------------------------------- |:-----------------------------------------| | ||
| 30 | +| Dependency | Not involved | cmake/third_party/makeself-fetch.cmake | [https://gitcode.com/cann-src-third-party/makeself/releases/download/release-2.5.0-patch1.0/makeself-release-2.5.0-patch1.tar.gz](https://gitcode.com/cann-src-third-party/makeself/releases/download/release-2.5.0-patch1.0/makeself-release-2.5.0-patch1.tar.gz) | Download makeself source code from gitcode, used as compilation dependency | | ||
| 31 | +| Dependency | Not involved | cmake/third_party/nlohmann_json.cmake | [https://gitcode.com/cann-src-third-party/json/releases/download/v3.11.3/include.zip](https://gitcode.com/cann-src-third-party/json/releases/download/v3.11.3/include.zip) | Download json source code from gitcode, used as compilation dependency | | ||
| 32 | +| Dependency | Not involved | cmake/third_party/gtest.cmake | [https://gitcode.com/cann-src-third-party/googletest/releases/download/v1.14.0/googletest-1.14.0.tar.gz](https://gitcode.com/cann-src-third-party/googletest/releases/download/v1.14.0/googletest-1.14.0.tar.gz) | Download googletest source code from gitcode, used as compilation dependency | | ||
| 33 | +| Dependency | Not involved | cmake/third_party/eigen.cmake | [https://gitcode.com/cann-src-third-party/eigen/releases/download/5.0.0-h0.trunk/eigen-5.0.0.tar.gz](https://gitcode.com/cann-src-third-party/eigen/releases/download/5.0.0-h0.trunk/eigen-5.0.0.tar.gz) | Download eigen source code from gitcode, used as compilation dependency | | ||
| 34 | +| Dependency | Not involved | ops-ras/install_deps.sh | [https://apt.kitware.com/keys/kitware-archive-latest.asc](https://apt.kitware.com/keys/kitware-archive-latest.asc) | Download install_deps source code from gitcode, used as compilation dependency | | ||
| 35 | +| Dependency | Not involved | ops-ras/install_deps.sh | [https://apt.kitware.com/ubuntu/](https://apt.kitware.com/ubuntu/) | Download install_deps source code from gitcode, used as compilation dependency | | ||
| 36 | +| Dependency | Not involved | cmake | [https://apt.kitware.com/keys/kitware-archive-latest.asc](https://apt.kitware.com/keys/kitware-archive-latest.asc) | Download cmake software from kitware, used as compilation dependency | | ||
| 37 | +| Dependency | Not involved | cmake | [https://apt.kitware.com/ubuntu/](https://apt.kitware.com/ubuntu/) | Download cmake software from kitware, used as compilation dependency | | ||
| 38 | + | ||
| 39 | +## Vulnerability Mechanism Description | ||
| 40 | + | ||
| 41 | +[Vulnerability Management](https://gitcode.com/cann/community/blob/master/security/security.md) | ||
| 42 | + | ||
| 43 | +## Appendix | ||
| 44 | + | ||
| 45 | +### A-File (Folder) Permission Control Recommended Maximum Values in Various Scenarios | ||
| 46 | + | ||
| 47 | +| Type | Linux Permission Reference Maximum Value | | ||
| 48 | +| -------------- | --------------- | | ||
| 49 | +| User Home Directory | 750 (rwxr-x---) | | ||
| 50 | +| Program Files (including script files, library files, etc.) | 550 (r-xr-x---) | | ||
| 51 | +| Program File Directory | 550 (r-xr-x---) | | ||
| 52 | +| Configuration File | 640 (rw-r-----) | | ||
| 53 | +| Configuration File Directory | 750 (rwxr-x---) | | ||
| 54 | +| Log File (recording completed or archived) | 440 (r--r-----) | | ||
| 55 | +| Log File (currently recording) | 640 (rw-r-----) | | ||
| 56 | +| Log File Directory | 750 (rwxr-x---) | | ||
| 57 | +| Debug File | 640 (rw-r-----) | | ||
| 58 | +| Debug File Directory | 750 (rwxr-x---) | | ||
| 59 | +| Temporary File Directory | 750 (rwxr-x---) | | ||
| 60 | +| Maintenance Upgrade File Directory | 770 (rwxrwx---) | | ||
| 61 | +| Business Data File | 640 (rw-r-----) | | ||
| 62 | +| Business Data File Directory | 750 (rwxr-x---) | | ||
| 63 | +| Key Component, Private Key, Certificate, Ciphertext File Directory | 700 (rwx-----) | | ||
| 64 | +| Key Component, Private Key, Certificate, Encrypted Ciphertext | 600 (rw-------) | | ||
| 65 | +| Encryption/Decryption Interface, Encryption/Decryption Script | 500 (r-x------) | | ||
| @@ -0,0 +1,102 @@ | |||
| 1 | +# Documentation Contribution Guide | ||
| 2 | + | ||
| 3 | +We welcome your contributions to the project documentation. High-quality documentation is crucial for project success. This guide will help you efficiently submit documentation that meets the standards. | ||
| 4 | + | ||
| 5 | +## Contribution Scope | ||
| 6 | + | ||
| 7 | +We welcome any contributions that can improve documentation quality, including but not limited to: | ||
| 8 | + | ||
| 9 | +- Correction and Improvement: Fix typos, grammar errors, incorrect code examples, outdated information, or broken links. | ||
| 10 | + | ||
| 11 | +- Clarification and Optimization: Make descriptions clearer and easier to understand, optimize sentence structure, and supplement background knowledge. | ||
| 12 | + | ||
| 13 | +- Content Supplement: Add usage examples, API documentation, frequently asked questions (FAQ), best practices, or warning descriptions for existing features. | ||
| 14 | + | ||
| 15 | +- New Content Creation: Write new chapters or tutorials for newly added features, such as operator README, API introduction documents, and so on. If you have questions, we recommend creating an Issue for discussion first. | ||
| 16 | + | ||
| 17 | +- Localization Translation: Help us translate or proofread documents in other languages. | ||
| 18 | + | ||
| 19 | +- Style and Navigation: Improve the layout, readability, and navigation structure of the documentation website. | ||
| 20 | + | ||
| 21 | +## Contribution Process | ||
| 22 | + | ||
| 23 | +1. **Preparation Work** | ||
| 24 | + | ||
| 25 | + - Determine the Task: If there are documentation issues, you can create new Issues. We recommend using the label category `[Documentation|文档反馈]` and providing a detailed description. Based on the existing Issues list, determine the documentation issues to be resolved. | ||
| 26 | + - Claim the Task: Comment `/assign @yourself` under the corresponding Issue to indicate that you will handle it and avoid duplicate work. | ||
| 27 | + | ||
| 28 | +2. **Document Modification** | ||
| 29 | + | ||
| 30 | + - Select Branch: Please download the source code from the master or other Tag branches to the local machine. | ||
| 31 | + - Follow Format: | ||
| 32 | + - This project recommends using **Markdown format**. | ||
| 33 | + - Follow the existing writing style of the project. | ||
| 34 | + - Put static resources such as images in the corresponding directory. For example, images are generally in the `figures` folder under the docs directory. You can adjust them yourself in special cases. | ||
| 35 | + - Careful Addition and Deletion: When modifying content, please try to maintain the original line width and line break conventions. | ||
| 36 | + | ||
| 37 | +3. **Submit Changes** | ||
| 38 | + | ||
| 39 | + - Atomic Commit: Each commit should focus on an independent modification. For example, "Fix spelling errors in xx guide" and "Update example code in API reference" should be submitted separately. | ||
| 40 | + | ||
| 41 | + - Write Clear Commit Messages: | ||
| 42 | + | ||
| 43 | + ```text | ||
| 44 | + Brief description (no more than 50 characters) | ||
| 45 | + | ||
| 46 | + If necessary, provide a more detailed description here. Explain the reason and content of the modification, rather than what specifically was changed (the code itself will show). | ||
| 47 | + Associated Issue: #123 | ||
| 48 | + ``` | ||
| 49 | + | ||
| 50 | +4. **Initiate Pull Request** | ||
| 51 | + | ||
| 52 | + - Target Branch: Please merge the PR into the target branch of the project. | ||
| 53 | + - Title and Description: | ||
| 54 | + - PR Title: Should clearly summarize the modification, for example: `[Docs] Fix configuration example in quick start`. | ||
| 55 | + - PR Description: Detailed explanation of your changes, motivation, and associated Issues (use Closes #123 or Fixes #456). | ||
| 56 | + - Preview Check: Please check the document effect in local or online browsing in advance to ensure that the rendering meets expectations. | ||
| 57 | + - Wait for Review: Maintainers will review and may propose modification suggestions. Please follow up on the discussion in a timely manner. | ||
| 58 | + | ||
| 59 | +## Writing Standards | ||
| 60 | + | ||
| 61 | +Before developers write project documentation, please be sure to read the following standards first. If you have questions, you are welcome to make suggestions at any time! | ||
| 62 | + | ||
| 63 | +- Prerequisites: Please first learn the unified writing standards provided by the CANN organization. For details, see [CANN Document Writing Standards](https://gitcode.com/cann/community/blob/master/contributor/docs/document_writing_specs.md). | ||
| 64 | + | ||
| 65 | + - Document Content Requirements: Introduce the required and optional document deliverables in the project. | ||
| 66 | + - Directory Structure Standards: Introduce the principles of directory division, such as Chinese and English management. | ||
| 67 | + - Content Element Standards: Introduce rules for different writing elements, such as file naming, titles, fonts, images, code blocks, links, and so on. | ||
| 68 | + | ||
| 69 | +- Precautions: | ||
| 70 | + | ||
| 71 | + In addition to the above writing rules, you also need to pay attention to the following: | ||
| 72 | + | ||
| 73 | + - Tone: Use a friendly, professional, and neutral tone. For beginners, avoid unnecessary jargon. | ||
| 74 | + - Terminology: Maintain terminology consistency (such as uniformly using "click" instead of "single click"). Please refer to the project terminology table (if available). | ||
| 75 | + - Code Examples: | ||
| 76 | + - Ensure that all code examples are runnable and tested. | ||
| 77 | + - Provide sufficient context and explanation. | ||
| 78 | + - Indicate the environment or prerequisites required for code running. | ||
| 79 | + - Punctuation and Format: | ||
| 80 | + - When mixing Chinese and English, use full-width punctuation. Punctuation marks must conform to the Chinese/English context. | ||
| 81 | + - Use appropriate hierarchy for titles (#, ##, ###). | ||
| 82 | + - Use lists and tables to organize complex information. | ||
| 83 | + - Links: Use descriptive link text, avoid "click here", and ensure that link resources are authentic and reliable. | ||
| 84 | + - Images: | ||
| 85 | + - Common Formats: We recommend the png format. Try to keep the style consistent with existing images. | ||
| 86 | + - Resolution and Clarity: Must be clear and of moderate size. Avoid blurring or excessive compression. | ||
| 87 | + - File Size: We do not recommend that a single image exceeds 10M. | ||
| 88 | + - Copyright: For all quoted images, literature, and other resources, please ensure compliance. | ||
| 89 | + | ||
| 90 | +## Get Help | ||
| 91 | + | ||
| 92 | +If you have any questions during the contribution process: | ||
| 93 | + | ||
| 94 | +1. Check Existing Documentation: If there are problems with templates or standards, please first check the existing guides, API documentation, or README of the project. | ||
| 95 | +2. Initiate Discussion: You can create a new Issue or leave a message directly in the relevant Issue or PR. | ||
| 96 | + | ||
| 97 | +## Document Templates | ||
| 98 | + | ||
| 99 | +The key documents involved in operator deliverables mainly include the following. For specific writing formats and content requirements, please refer to the templates. | ||
| 100 | + | ||
| 101 | +- [Operator README Document Template](https://gitcode.com/cann/ops-ras/wiki/%E7%AE%97%E5%AD%90README%E6%96%87%E6%A1%A3%E6%A8%A1%E6%9D%BF) | ||
| 102 | +- [aclnn API Document Template](https://gitcode.com/cann/ops-ras/wiki/aclnn%20API%E6%96%87%E6%A1%A3%E6%A8%A1%E6%9D%BF) | ||
| @@ -0,0 +1,273 @@ | |||
| 1 | +# Quick Start: Based on ops-ras Repository | ||
| 2 | + | ||
| 3 | +## Usage Notice | ||
| 4 | + | ||
| 5 | +This guide aims to help you quickly get started with CANN and the `ops-ras` operator repository. To help you quickly understand the entire process of operator development, we will use the **AddExample** operator as a practical object. Its source code is located in `ops-ras/examples/add_example`. The operation process is as follows: | ||
| 6 | + | ||
| 7 | +1. **[Prerequisites](../README.md)**: Complete the environment setup and source code download by referring to the project README. The process is not repeated here. For the quick start scenario, **CANNLab or Docker deployment is recommended** for simple operation. | ||
| 8 | + | ||
| 9 | + > **Note**: The CANNLab or Docker environment provides the latest version of the CANN package by default. If you need to experience the latest capabilities of the master branch, you can manually set up the environment. | ||
| 10 | + | ||
| 11 | +2. **[Compilation and Running](#i-compilation-and-running)**: Compile the custom operator package and install it to achieve quick operator invocation. | ||
| 12 | + | ||
| 13 | +3. **[Operator Development](#ii-operator-development)**: Experience the complete loop of development, compilation, and verification by modifying the existing operator Kernel. | ||
| 14 | + | ||
| 15 | +4. **[Operator Debugging](#iii-operator-debugging)**: Master the methods of operator printing and performance collection. | ||
| 16 | + | ||
| 17 | +5. **[Operator Verification](#iv-operator-verification)**: Learn how to modify operator example samples to verify the functional correctness of operators under different inputs. | ||
| 18 | + | ||
| 19 | +## I. Compilation and Running | ||
| 20 | + | ||
| 21 | +The purpose of this stage is to **quickly experience the project standard process** and verify whether the environment can successfully perform operator source code compilation, packaging, installation, and running. | ||
| 22 | + | ||
| 23 | +### 1. Enter the Project Source Code | ||
| 24 | + | ||
| 25 | +- CANNLab Cloud Development Environment: | ||
| 26 | + | ||
| 27 | + The latest CANN package matching project source code is provided by default. Enter the source code directory and replace `${gitCode_id}` with the developer's personal gitCode account. | ||
| 28 | + | ||
| 29 | + ```bash | ||
| 30 | + cd /mnt/workspace/gitCode/${gitCode_id}/ops-ras | ||
| 31 | + ``` | ||
| 32 | + | ||
| 33 | +- Non-CANNLab Cloud Development Environment: | ||
| 34 | + | ||
| 35 | + According to the correspondence between source code and CANN versions in the [release repository](https://gitcode.com/cann/release-management), execute the following command to download the source code. Replace `${tag_version}` with the target branch tag, for example, 9.0.0. | ||
| 36 | + | ||
| 37 | + ```bash | ||
| 38 | + git clone -b ${tag_version} https://gitcode.com/cann/ops-ras.git && cd ops-ras | ||
| 39 | + ``` | ||
| 40 | + | ||
| 41 | +> Note: If you need to switch the source code branch version, refer to the following guidance. | ||
| 42 | +> | ||
| 43 | +> 1. Execute `git branch` in the source code directory to query the current source code version. | ||
| 44 | +> 2. Execute `git checkout ${tag_version}` in the source code directory to switch to the target branch source code. Ensure that the source code matches the CANN version. If the source code already exists, execute `git pull` to pull the latest source code. | ||
| 45 | + | ||
| 46 | +### 2. Compile the AddExample Operator | ||
| 47 | + | ||
| 48 | +This guide uses **single operator compilation** by default: only the target operator is built, the compilation time is short, and it is suitable for quick start and daily development. The general command format: `bash build.sh --pkg --soc=<chip version> --ops=<operator name>`. | ||
| 49 | + | ||
| 50 | +> If you need to compile the entire operator library (omit `--ops`), see [build Parameter Description](zh/install/build.md). | ||
| 51 | + | ||
| 52 | +Taking the AddExample operator as an example, the compilation command is as follows: | ||
| 53 | + | ||
| 54 | +```bash | ||
| 55 | +bash build.sh --pkg --soc=${soc_version} --ops=add_example -j16 | ||
| 56 | +``` | ||
| 57 | + | ||
| 58 | +For the value of `${soc_version}`, visit the [CANN Download Center](https://www.hiascend.com/cann/download) and query the hardware product name according to the page prompts. The corresponding `${soc_version}` values for product names are as follows. Please pass the parameter according to the actual scenario. | ||
| 59 | + | ||
| 60 | +- Atlas A2 Training Series Products/Atlas A2 Inference Series Products: `ascend910b` | ||
| 61 | +- Atlas A3 Training Series Products/Atlas A3 Inference Series Products: `ascend910_93` | ||
| 62 | +- Ascend 950 Series Products: `ascend950` | ||
| 63 | + | ||
| 64 | +If the following information is prompted, the compilation is successful. | ||
| 65 | + | ||
| 66 | +```bash | ||
| 67 | +Self-extractable archive "cann-ops-ras-custom_linux-${arch}.run" successfully created. | ||
| 68 | +``` | ||
| 69 | + | ||
| 70 | +After successful compilation, the run package is stored in the build_out directory under the project root directory. | ||
| 71 | + | ||
| 72 | +### 3. Install the AddExample Operator Package | ||
| 73 | + | ||
| 74 | +```bash | ||
| 75 | +./build_out/cann-ops-ras-*linux*.run | ||
| 76 | +``` | ||
| 77 | + | ||
| 78 | +`AddExample` is installed in the ```${ASCEND_HOME_PATH}/opp/vendors``` path. ```${ASCEND_HOME_PATH}``` indicates the CANN software installation directory. | ||
| 79 | + | ||
| 80 | +### 4. Configure Environment Variables | ||
| 81 | + | ||
| 82 | +Add the path of the custom operator package to the environment variables to ensure that it can be found at runtime. | ||
| 83 | + | ||
| 84 | +```bash | ||
| 85 | +export LD_LIBRARY_PATH=${ASCEND_HOME_PATH}/opp/vendors/custom_ras/op_api/lib:${LD_LIBRARY_PATH} | ||
| 86 | +``` | ||
| 87 | + | ||
| 88 | +### 5. Quick Verification: Run Operator Sample | ||
| 89 | + | ||
| 90 | +The general running command format: `bash build.sh --run_example <operator name> <running mode> <package mode>`. | ||
| 91 | + | ||
| 92 | +Taking AddExample as an example, it provides a simple operator sample `add_example/examples/test_aclnn_add_example.cpp`. Run this sample to verify whether the operator function is normal. | ||
| 93 | + | ||
| 94 | +```bash | ||
| 95 | +bash build.sh --run_example add_example eager cust --vendor_name=custom | ||
| 96 | +``` | ||
| 97 | + | ||
| 98 | +Expected output: Print the addition calculation result of the operator `AddExample`, indicating that the operator has been successfully deployed and executed correctly. | ||
| 99 | + | ||
| 100 | +```bash | ||
| 101 | +mean result[0] is: 2.000000 | ||
| 102 | +mean result[1] is: 2.000000 | ||
| 103 | +mean result[2] is: 2.000000 | ||
| 104 | +mean result[3] is: 2.000000 | ||
| 105 | +mean result[4] is: 2.000000 | ||
| 106 | +mean result[5] is: 2.000000 | ||
| 107 | +mean result[6] is: 2.000000 | ||
| 108 | +mean result[7] is: 2.000000 | ||
| 109 | +... | ||
| 110 | +``` | ||
| 111 | + | ||
| 112 | +## II. Operator Development | ||
| 113 | + | ||
| 114 | +The purpose of this stage is to try **modifying the kernel function code** for the successfully running AddExample operator. | ||
| 115 | + | ||
| 116 | +### 1. Modify Kernel Implementation | ||
| 117 | + | ||
| 118 | +Find the core kernel implementation file of the AddExample operator `ops-ras/examples/add_example/op_kernel/add_example.h`, and try to change the Add operation in the operator to a Mul operation: | ||
| 119 | + | ||
| 120 | +```cpp | ||
| 121 | +__aicore__ inline void AddExample<T>::Compute(int32_t progress) | ||
| 122 | +{ | ||
| 123 | + AscendC::LocalTensor<T> xLocal = inputQueueX.DeQue<T>(); | ||
| 124 | + AscendC::LocalTensor<T> yLocal = inputQueueY.DeQue<T>(); | ||
| 125 | + AscendC::LocalTensor<T> zLocal = outputQueueZ.AllocTensor<T>(); | ||
| 126 | + // === Replace Add with Mul here === | ||
| 127 | + // AscendC::Add(zLocal, xLocal, yLocal, tileLength_); | ||
| 128 | + AscendC::Mul(zLocal, xLocal, yLocal, tileLength_); | ||
| 129 | + outputQueueZ.EnQue<T>(zLocal); | ||
| 130 | + inputQueueX.FreeTensor(xLocal); | ||
| 131 | + inputQueueY.FreeTensor(yLocal); | ||
| 132 | +} | ||
| 133 | +``` | ||
| 134 | + | ||
| 135 | +### 2. Compile and Verify | ||
| 136 | + | ||
| 137 | +Repeat the steps in the [Compilation and Running](#i-compilation-and-running) section: | ||
| 138 | + | ||
| 139 | +1. **Recompile**: | ||
| 140 | + | ||
| 141 | + First return to the project root directory. The compilation command is as follows: | ||
| 142 | + | ||
| 143 | + ```bash | ||
| 144 | + bash build.sh --pkg --soc=${soc_version} --ops=add_example -j16 | ||
| 145 | + ``` | ||
| 146 | + | ||
| 147 | + > **Note**: Please fill in `${soc_version}` according to the actual chip model. The value method is the same as described in [Compile the AddExample Operator](#2-compile-the-addexample-operator). | ||
| 148 | + | ||
| 149 | +2. **Reinstall**: | ||
| 150 | + | ||
| 151 | + ```bash | ||
| 152 | + ./build_out/cann-ops-ras-*linux*.run | ||
| 153 | + ``` | ||
| 154 | + | ||
| 155 | +3. **Re-verify**: | ||
| 156 | + | ||
| 157 | + ```bash | ||
| 158 | + bash build.sh --run_example add_example eager cust --vendor_name=custom | ||
| 159 | + ``` | ||
| 160 | + | ||
| 161 | +4. **Success Sign**: The output result becomes the multiplication result. | ||
| 162 | + | ||
| 163 | + ```bash | ||
| 164 | + mean result[0] is: 1.000000 | ||
| 165 | + mean result[1] is: 1.000000 | ||
| 166 | + mean result[2] is: 1.000000 | ||
| 167 | + mean result[3] is: 1.000000 | ||
| 168 | + mean result[4] is: 1.000000 | ||
| 169 | + mean result[5] is: 1.000000 | ||
| 170 | + mean result[6] is: 1.000000 | ||
| 171 | + mean result[7] is: 1.000000 | ||
| 172 | + ... | ||
| 173 | + ``` | ||
| 174 | + | ||
| 175 | +## III. Operator Debugging | ||
| 176 | + | ||
| 177 | +This stage takes AddExample as an example to add printing in the operator and collect operator performance data for subsequent problem analysis and positioning. | ||
| 178 | + | ||
| 179 | +### 1. Printing | ||
| 180 | + | ||
| 181 | +If the operator has execution failure, precision abnormality, or other problems, add printing for problem analysis and positioning. | ||
| 182 | + | ||
| 183 | +Please modify the code in `examples/add_example/op_kernel/add_example.h`. | ||
| 184 | + | ||
| 185 | +* **printf** | ||
| 186 | + | ||
| 187 | + This interface supports printing Scalar type data, such as integers, character types, Boolean types, and so on. For detailed introduction, see "Operator Debugging API > printf" in "[Ascend C API](https://hiascend.com/document/redirect/CannCommunityAscendCApi)". | ||
| 188 | + | ||
| 189 | + ```c++ | ||
| 190 | + blockLength_ = (tilingData->totalLength + AscendC::GetBlockNum() - 1) / AscendC::GetBlockNum(); | ||
| 191 | + tileNum_ = tilingData->tileNum; | ||
| 192 | + tileLength_ = ((blockLength_ + tileNum_ - 1) / tileNum_ / BUFFER_NUM) ? | ||
| 193 | + ((blockLength_ + tileNum_ - 1) / tileNum_ / BUFFER_NUM) : 1; | ||
| 194 | + // Print the current kernel calculation Block length | ||
| 195 | + AscendC::PRINTF("Tiling blockLength is %llu\n", blockLength_); | ||
| 196 | + ``` | ||
| 197 | + | ||
| 198 | +* **DumpTensor** | ||
| 199 | + | ||
| 200 | + This interface supports dumping the content of the specified Tensor, and also supports printing custom additional information, such as the current line number. For detailed introduction, see "Operator Debugging API > DumpTensor" in "[Ascend C API](https://hiascend.com/document/redirect/CannCommunityAscendCApi)". | ||
| 201 | + | ||
| 202 | + ```c++ | ||
| 203 | + AscendC::LocalTensor<T> zLocal = outputQueueZ.DeQue<T>(); | ||
| 204 | + // Print zLocal Tensor information | ||
| 205 | + DumpTensor(zLocal, 0, 128); | ||
| 206 | + ``` | ||
| 207 | + | ||
| 208 | +### 2. Performance Collection | ||
| 209 | + | ||
| 210 | +When the operator function verification is correct, you can collect operator performance data through the `msprof` tool. | ||
| 211 | + | ||
| 212 | +- **Generate Executable File** | ||
| 213 | + | ||
| 214 | + Call the example sample of the AddExample operator to generate an executable file (test_aclnn_add_example), which is located in the project `ops-ras/build` directory. | ||
| 215 | + | ||
| 216 | + ```bash | ||
| 217 | + bash build.sh --run_example add_example eager cust --vendor_name=custom | ||
| 218 | + ``` | ||
| 219 | + | ||
| 220 | +- **Collect Performance Data** | ||
| 221 | + | ||
| 222 | + Enter the AddExample operator executable file directory `ops-ras/build/` and execute the following command: | ||
| 223 | + | ||
| 224 | + ```bash | ||
| 225 | + msprof --application="./test_aclnn_add_example" | ||
| 226 | + ``` | ||
| 227 | + | ||
| 228 | +The collection result is in the project `ops-ras/build/` directory. After the msprof command is executed, it will automatically parse and export the performance data result file. For detailed content, see [msprof](https://www.hiascend.com/document/detail/zh/mindstudio/82RC1/T&ITools/Profiling/atlasprofiling_16_0110.html#ZH-CN_TOPIC_0000002504160251). | ||
| 229 | + | ||
| 230 | +## IV. Operator Verification | ||
| 231 | + | ||
| 232 | +This stage verifies the functional correctness of the operator in multiple scenarios by modifying the input data of the AddExample operator example sample. | ||
| 233 | + | ||
| 234 | +### 1. Modify Test Input | ||
| 235 | + | ||
| 236 | +Find and edit the `ops-ras/examples/add_example/examples/test_aclnn_add_example.cpp` of `AddExample`, and modify the shape and numerical values of the input tensor. | ||
| 237 | + | ||
| 238 | +**Modify Input/Output Data**: Modify the shape information of input and output, as well as the initialization data, and construct the corresponding input and output tensors. | ||
| 239 | + | ||
| 240 | +```c++ | ||
| 241 | +int main() { | ||
| 242 | + // ... initialization code ... | ||
| 243 | + | ||
| 244 | + // === ① Modify selfX input === | ||
| 245 | + // Before modification: shape = {32, 4, 4, 4}, all values are 1 | ||
| 246 | + // After modification: change input shape to {8, 8, 8, 8}, and fill with different test data | ||
| 247 | + std::vector<int64_t> selfXShape = {8, 8, 8, 8}; | ||
| 248 | + std::vector<float> selfXHostData(4096); // 4096 = 8 * 8 * 8 *8 | ||
| 249 | + // You can use a loop to fill more distinguishable data, such as an increasing sequence | ||
| 250 | + for (int i = 0; i < 4096; ++i) { | ||
| 251 | + selfXHostData[i] = static_cast<float>(i % 10); // Fill with cyclic values of 0-9 | ||
| 252 | + } | ||
| 253 | + // === ② Refer to selfX, similarly modify selfY and selfZ inputs === | ||
| 254 | + | ||
| 255 | + // ... subsequent execution code ... | ||
| 256 | +} | ||
| 257 | +``` | ||
| 258 | + | ||
| 259 | +### 2. Recompile and Verify | ||
| 260 | + | ||
| 261 | +1. Since only the example test code is modified, there is no need to recompile the operator package. | ||
| 262 | + | ||
| 263 | +2. Re-execute the verification command: | ||
| 264 | + | ||
| 265 | + ```bash | ||
| 266 | + bash build.sh --run_example add_example eager cust --vendor_name=custom | ||
| 267 | + ``` | ||
| 268 | + | ||
| 269 | +3. Observe whether the operator output result meets expectations. | ||
| 270 | + | ||
| 271 | +## Conclusion | ||
| 272 | + | ||
| 273 | +After experiencing the above processes, you have basically completed the operator development process. If you want to further contribute new operators or learn more advanced development, debugging, and other skills, please visit the project README to learn about [Advanced Tutorials](../README.md#学习教程) and [Contribution Guide](../README.md#相关信息), and so on. | ||
| @@ -0,0 +1,72 @@ | |||
| 1 | +# Documentation Center | ||
| 2 | + | ||
| 3 | +## Directory Structure | ||
| 4 | + | ||
| 5 | +The Docs directory structure is described as follows: | ||
| 6 | + | ||
| 7 | +```text | ||
| 8 | +├── zh | ||
| 9 | + ├── context # Public documents, such as terminology, basic concepts, and so on | ||
| 10 | + ├── debug # Operator debugging guidance documents | ||
| 11 | + │ ├── cann_sim.md | ||
| 12 | + │ ├── op_debug_prof.md | ||
| 13 | + │ └── ... | ||
| 14 | + ├── develop # Operator development guidance documents | ||
| 15 | + │ ├── aicore_develop_guide.md | ||
| 16 | + │ ├── aicpu_develop_guide.md | ||
| 17 | + │ ├── cross_platform_migration_guide.md | ||
| 18 | + │ ├── graph_develop_guide.md | ||
| 19 | + │ └── ... | ||
| 20 | + ├── figures # Image directory | ||
| 21 | + ├── install # Environment installation and compilation guidance documents | ||
| 22 | + │ ├── build.md | ||
| 23 | + │ ├── compile.md | ||
| 24 | + │ ├── dir_structure.md | ||
| 25 | + │ ├── quick_install.md | ||
| 26 | + │ └── ... | ||
| 27 | + ├── invocation # Operator invocation guidance documents (including aclnn invocation, graph mode invocation, and so on) | ||
| 28 | + │ ├── quick_op_invocation.md | ||
| 29 | + │ ├── op_invocation.md | ||
| 30 | + │ └── ... | ||
| 31 | + ├── op_api_list.md # Complete operator interface list (aclnn) | ||
| 32 | + ├── op_list.md # Complete operator list | ||
| 33 | +├── CONTRIBUTING_DOCS.md # Documentation contribution guide | ||
| 34 | +├── QUICKSTART.md # Quick start | ||
| 35 | +└── README.md | ||
| 36 | +``` | ||
| 37 | + | ||
| 38 | +## Advanced Tutorials | ||
| 39 | + | ||
| 40 | +### Guide Documents | ||
| 41 | + | ||
| 42 | +| Document | Description | | ||
| 43 | +| ------------------------------------------------------------ | ------------------------------------------------------------ | | ||
| 44 | +| [Source Code Build Guide](zh/install/compile.md) | Introduces different source code build methods and verification methods in online and offline scenarios. | | ||
| 45 | +| [Operator Invocation Guide](zh/invocation/quick_op_invocation.md) | Introduces the method of invoking operator samples and different operator invocation methods (such as aclnn/graph, and so on). | | ||
| 46 | +| [Standard Operator Development Guide](zh/develop/aicore_develop_guide.md) | Introduces how to define operator prototypes and implement Tiling and Kernel based on standard engineering. Such operators are called "standard operators".<br>Standard operators support aclnn and graph mode invocation. | | ||
| 47 | +| [Simple Operator Development Guide](../examples/fast_kernel_launch_example/README.md) | Introduces how to implement fast_kernel_launch based on simple engineering, that is, the `<<<>>>` method. Such operators are called "simple operators".<br>Simple operators only support PyTorch invocation. | | ||
| 48 | +| [Operator Debugging and Tuning](zh/debug/op_debug_prof.md) | Introduces common operator function debugging and performance tuning methods (such as data collection and simulation pipeline, and so on). | | ||
| 49 | + | ||
| 50 | +### API Documents | ||
| 51 | + | ||
| 52 | +| Document | Description | | ||
| 53 | +| ----------------------- | ---------------------- | | ||
| 54 | +| [Operator List](zh/op_list.md) | Introduces the list of all operators included in the project. | | ||
| 55 | +| [aclnn List](zh/op_api_list.md) | Introduces the list of all operator aclnn APIs included in the project. To facilitate users to invoke operators on the Host side, C language APIs are provided, that is, APIs with the aclnn prefix. | | ||
| 56 | + | ||
| 57 | +### Tool Documents | ||
| 58 | + | ||
| 59 | +| Document | Description | | ||
| 60 | +| ----------------------- | ---------------------- | | ||
| 61 | +| [Simulator Tool](zh/debug/cann_sim.md) | A SoC-level simulation tool for operator development scenarios, used to analyze the precision and performance data of AI tasks running on the AI simulator at each stage. | | ||
| 62 | + | ||
| 63 | +### More Documents | ||
| 64 | + | ||
| 65 | +- Sample Documents: Refer to the operator samples in the [cann-samples](https://gitcode.com/cann/cann-samples) repository. | ||
| 66 | + | ||
| 67 | +## Appendix | ||
| 68 | + | ||
| 69 | +| Document | Description | | ||
| 70 | +| ----------------------------------- | ------------------------------------------------------------ | | ||
| 71 | +| [Operator Basic Concepts](zh/context/basic_concept.md) | Introduces basic concepts and terminology in the operator domain, such as quantization/sparse, data types, data formats, and so on. | | ||
| 72 | +| [build Parameter Description](zh/install/build.md) | Introduces the functions and parameter values of the build.sh script in this project, including source code compilation, operator invocation, debugging, and so on. | | ||
The file is empty
| @@ -0,0 +1,161 @@ | |||
| 1 | +# aclnn Return Codes | ||
| 2 | + | ||
| 3 | +When calling aclnn APIs, common interface return codes are shown in [Table 1](#table1). | ||
| 4 | +For abnormal status code values, you can use the aclGetRecentErrMsg interface (refer to [ACL API (C)](https://www.hiascend.com/document/detail/en/canncommercial/latest/API/appdevgapi/aclcppdevg_03_0004.html)) to obtain exception information. You can troubleshoot the problem based on the error message or contact technical support. | ||
| 5 | + | ||
| 6 | +**Table 1** Return Status Codes | ||
| 7 | + | ||
| 8 | +<a name="table1"></a> | ||
| 9 | +<table><thead align="left"><tr><th class="cellrowborder" valign="top" width="30.543054305430545%" id="mcps1.2.4.1.1"><p>Status Code Name</p> | ||
| 10 | +</th> | ||
| 11 | +<th class="cellrowborder" valign="top" width="15.971597159715973%" id="mcps1.2.4.1.2"><p>Status Code Value</p> | ||
| 12 | +</th> | ||
| 13 | +<th class="cellrowborder" valign="top" width="53.48534853485349%" id="mcps1.2.4.1.3"><p>Status Code Description</p> | ||
| 14 | +</th> | ||
| 15 | +</tr> | ||
| 16 | +</thead> | ||
| 17 | +<tbody><tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_SUCCESS</p> | ||
| 18 | +</td> | ||
| 19 | +<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>0</p> | ||
| 20 | +</td> | ||
| 21 | +<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>Success.</p> | ||
| 22 | +</td> | ||
| 23 | +</tr> | ||
| 24 | +<tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_PARAM_NULLPTR</p> | ||
| 25 | +</td> | ||
| 26 | +<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>161001</p> | ||
| 27 | +</td> | ||
| 28 | +<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>Parameter validation error, illegal nullptr exists in parameters.</p> | ||
| 29 | +</td> | ||
| 30 | +</tr> | ||
| 31 | +<tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_PARAM_INVALID</p> | ||
| 32 | +</td> | ||
| 33 | +<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>161002</p> | ||
| 34 | +</td> | ||
| 35 | +<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>Parameter validation error, such as two input data types not satisfying the input type promotion relationship.</p> | ||
| 36 | +</td> | ||
| 37 | +</tr> | ||
| 38 | +<tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_RUNTIME_ERROR</p> | ||
| 39 | +</td> | ||
| 40 | +<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>361001</p> | ||
| 41 | +</td> | ||
| 42 | +<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>API internally calls npu runtime interface abnormally.</p> | ||
| 43 | +</td> | ||
| 44 | +</tr> | ||
| 45 | +<tr><td class="cellrowborder" valign="top" width="30.543054305430545%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_XXX</p> | ||
| 46 | +</td> | ||
| 47 | +<td class="cellrowborder" valign="top" width="15.971597159715973%" headers="mcps1.2.4.1.2 "><p>561xxx</p> | ||
| 48 | +</td> | ||
| 49 | +<td class="cellrowborder" valign="top" width="53.48534853485349%" headers="mcps1.2.4.1.3 "><p>API internal exception occurred.</p> | ||
| 50 | + | ||
| 51 | +</td> | ||
| 52 | +</tr> | ||
| 53 | +</tbody> | ||
| 54 | +</table> | ||
| 55 | + | ||
| 56 | +For more information about ACLNN_ERR_INNER_XXX status codes, see [Table 2](#table2). | ||
| 57 | + | ||
| 58 | +**Table 2** Exception Status Codes | ||
| 59 | + | ||
| 60 | +<a name="table2"></a> | ||
| 61 | +<table><thead align="left"><tr><th class="cellrowborder" valign="top" width="30.183018301830185%" id="mcps1.2.4.1.1"><p>Status Code Name</p> | ||
| 62 | +</th> | ||
| 63 | +<th class="cellrowborder" valign="top" width="16.521652165216523%" id="mcps1.2.4.1.2"><p>Status Code Value</p> | ||
| 64 | +</th> | ||
| 65 | +<th class="cellrowborder" valign="top" width="53.295329532953296%" id="mcps1.2.4.1.3"><p>Status Code Description</p> | ||
| 66 | +</th> | ||
| 67 | +</tr> | ||
| 68 | +</thead> | ||
| 69 | +<tbody><tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER</p> | ||
| 70 | +</td> | ||
| 71 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561000</p> | ||
| 72 | +</td> | ||
| 73 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal exception occurred.</p> | ||
| 74 | +</td> | ||
| 75 | +</tr> | ||
| 76 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_INFERSHAPE_ERROR</p> | ||
| 77 | +</td> | ||
| 78 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561001</p> | ||
| 79 | +</td> | ||
| 80 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal output shape deduction error occurred.</p> | ||
| 81 | +</td> | ||
| 82 | +</tr> | ||
| 83 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_TILING_ERROR</p> | ||
| 84 | +</td> | ||
| 85 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561002</p> | ||
| 86 | +</td> | ||
| 87 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal tiling for npu kernel exception occurred.</p> | ||
| 88 | +</td> | ||
| 89 | +</tr> | ||
| 90 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_FIND_KERNEL_ERROR</p> | ||
| 91 | +</td> | ||
| 92 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561003</p> | ||
| 93 | +</td> | ||
| 94 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal npu kernel lookup exception (possibly because operator binary package is not installed).</p> | ||
| 95 | +</td> | ||
| 96 | +</tr> | ||
| 97 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_CREATE_EXECUTOR</p> | ||
| 98 | +</td> | ||
| 99 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561101</p> | ||
| 100 | +</td> | ||
| 101 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal aclOpExecutor creation failed (possibly due to operating system exception).</p> | ||
| 102 | +</td> | ||
| 103 | +</tr> | ||
| 104 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_NOT_TRANS_EXECUTOR</p> | ||
| 105 | +</td> | ||
| 106 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561102</p> | ||
| 107 | +</td> | ||
| 108 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: API internal uniqueExecutor ReleaseTo not called.</p> | ||
| 109 | +</td> | ||
| 110 | +</tr> | ||
| 111 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_NULLPTR</p> | ||
| 112 | +</td> | ||
| 113 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561103</p> | ||
| 114 | +</td> | ||
| 115 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, nullptr exception appeared.</p> | ||
| 116 | +</td> | ||
| 117 | +</tr> | ||
| 118 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_WRONG_ATTR_INFO_SIZE</p> | ||
| 119 | +</td> | ||
| 120 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561104</p> | ||
| 121 | +</td> | ||
| 122 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, operator attribute count exception.</p> | ||
| 123 | +</td> | ||
| 124 | +</tr> | ||
| 125 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_KEY_CONFILICT</p> | ||
| 126 | +</td> | ||
| 127 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561105</p> | ||
| 128 | +</td> | ||
| 129 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, operator kernel matching hash key conflict.</p> | ||
| 130 | +</td> | ||
| 131 | +</tr> | ||
| 132 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_INVALID_IMPL_MODE</p> | ||
| 133 | +</td> | ||
| 134 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561106</p> | ||
| 135 | +</td> | ||
| 136 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, operator implementation mode parameter error.</p> | ||
| 137 | +</td> | ||
| 138 | +</tr> | ||
| 139 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_OPP_PATH_NOT_FOUND</p> | ||
| 140 | +</td> | ||
| 141 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561107</p> | ||
| 142 | +</td> | ||
| 143 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, environment variable ASCEND_OPP_PATH to be configured not detected.</p> | ||
| 144 | +</td> | ||
| 145 | +</tr> | ||
| 146 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_LOAD_JSON_FAILED</p> | ||
| 147 | +</td> | ||
| 148 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561108</p> | ||
| 149 | +</td> | ||
| 150 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, failed to load operator information json file in operator kernel library.</p> | ||
| 151 | +</td> | ||
| 152 | +</tr> | ||
| 153 | +<tr><td class="cellrowborder" valign="top" width="30.183018301830185%" headers="mcps1.2.4.1.1 "><p>ACLNN_ERR_INNER_JSON_VALUE_NOT_FOUND</p> | ||
| 154 | +</td> | ||
| 155 | +<td class="cellrowborder" valign="top" width="16.521652165216523%" headers="mcps1.2.4.1.2 "><p>561109</p> | ||
| 156 | +</td> | ||
| 157 | +<td class="cellrowborder" valign="top" width="53.295329532953296%" headers="mcps1.2.4.1.3 "><p>Internal exception: aclnn API internal exception occurred, failed to load a field in operator information json file in operator kernel library.</p> | ||
| 158 | +</td> | ||
| 159 | +</tr> | ||
| 160 | +</tbody> | ||
| 161 | +</table> | ||
| @@ -0,0 +1,12 @@ | |||
| 1 | +# Basic Concepts | ||
| 2 | + | ||
| 3 | + - [Two Phase API](./two_phase_api.md) | ||
| 4 | + - [Data Structure](./data_structure.md) | ||
| 5 | + - [Data Type](./data_type.md) | ||
| 6 | + - [Data Format](./data_format.md) | ||
| 7 | + - [Non Contiguous Tensor](./non_contiguous_tensor.md) | ||
| 8 | + - [Broadcast Relationship](./broadcast_relationship.md) | ||
| 9 | + - [Decuction Relationship](./decuction_relationship.md) | ||
| 10 | + - [Conversion Relationship](./conversion_relationship.md) | ||
| 11 | + - [Quant more Introduction](./quant_more_introduction.md) | ||
| 12 | + - [Sparse Mode Introduction](./sparse_mode_introduction.md) | ||
| @@ -0,0 +1,55 @@ | |||
| 1 | +# Broadcast Relationships | ||
| 2 | + | ||
| 3 | +## Broadcast Concept | ||
| 4 | + | ||
| 5 | +Broadcast describes how operators handle tensors (or arrays) of different shapes during computation. In most cases, tensors (or arrays) of different shapes are allowed to automatically expand their shapes during element operations to make their dimensions compatible. Usually, smaller tensors (or arrays) are "broadcast" to larger tensors (or arrays). | ||
| 6 | + | ||
| 7 | +Currently, many CANN operator API parameter shapes support broadcasting, which can appropriately improve calculation efficiency and reduce memory usage (especially in large-scale data scenarios). For more detailed broadcast technology introduction, refer to the [NumPy](https://numpy.org/doc/stable/user/basics.broadcasting.html) official website. | ||
| 8 | + | ||
| 9 | +## Broadcast Rules | ||
| 10 | + | ||
| 11 | +When performing broadcast calculations, you generally need to understand the following rules: | ||
| 12 | + | ||
| 13 | +- Rule 1: If the number of dimensions between arrays is inconsistent, all arrays align to the array with the longest shape, and the insufficient part of the shape is padded with 1 on the **left** until the number of dimensions is the same. | ||
| 14 | + | ||
| 15 | + > Note: | ||
| 16 | + > - Example 1: Number of Dimensions refers to the dimension count of the tensor (or array) corresponding to the shape. For example, x.shape=(1,1,2,4), the number of dimensions is 4. | ||
| 17 | + > - Example 2: For example, when calculating a+b, where a.shape=(2, 2, 3) and b.shape=(2, 3), array b will be broadcast to b.shape=(1, 2, 3). | ||
| 18 | + | ||
| 19 | +- Rule 2: If the number of dimensions between arrays is consistent, and a certain dimension of an array is 1, then the array with dimension 1 will be stretched to match the corresponding dimension shape of the other array. | ||
| 20 | + | ||
| 21 | + > Note: | ||
| 22 | + > In this scenario, you only need to ensure broadcasting in a certain dimension. For example, when calculating a+b, where a.shape=(1, 3) and b.shape=(3, 1), both arrays will be broadcast to a.shape=(3, 3) and b.shape=(3, 3). | ||
| 23 | + | ||
| 24 | +- Rule 3: If the number of dimensions between arrays is inconsistent, and neither has a dimension equal to 1, an error will be reported. | ||
| 25 | + | ||
| 26 | +Based on the above rules, the broadcast process generally first expands dimensions according to **Rule 1**, and then stretches the shape according to **Rule 2**. Specific examples are as follows: | ||
| 27 | + | ||
| 28 | +```text | ||
| 29 | +Assuming a.shape=(2,2,3), values look like: | ||
| 30 | +[[[1 2 3],[4 5 6]], | ||
| 31 | + [[1 2 3],[4 5 6]]] | ||
| 32 | +Assuming b.shape=(2,3), values look like: | ||
| 33 | +[[1 2 3], | ||
| 34 | + [-1 -2 -3]] | ||
| 35 | +According to Rule 1, expand dimensions, b.shape=(1,2,3), values are: | ||
| 36 | +[[[1 2 3], | ||
| 37 | + [-1 -2 -3]]] | ||
| 38 | +According to Rule 2, stretch shape, b.shape=(2,2,3), values are: | ||
| 39 | +[[[1 2 3],[-1 -2 -3]], | ||
| 40 | + [[1 2 3],[-1 -2 -3]]] | ||
| 41 | +Calculate a+b, actual result is: | ||
| 42 | + [[[2 4 6],[3 3 3]], | ||
| 43 | + [[2 4 6],[3 3 3]]] | ||
| 44 | +``` | ||
| 45 | + | ||
| 46 | +## Limitations | ||
| 47 | + | ||
| 48 | +When the data types of two inputs a and b that satisfy the broadcast relationship, or the deduced data types, are among COMPLEX64, COMPLEX128, DOUBLE, INT16, UINT16, or UINT64, in addition to satisfying the above broadcast rules, the following conditions must also be met, otherwise the broadcast will fail and cause the operator execution to report an error. | ||
| 49 | + | ||
| 50 | +Condition: The merged dimension of consecutive axes that need broadcasting and consecutive axes that do not need broadcasting must be less than 6. | ||
| 51 | + | ||
| 52 | +Examples: | ||
| 53 | + | ||
| 54 | +- When a.shape=(5, 1, 5, 1, 5, 1) and b.shape=(5, 5, 5, 5, 5, 5), there are no axes that need to be merged, the final dimension is 6, and the broadcast reports an error. | ||
| 55 | +- When a.shape=(5, 1, 5, 5, 1, 1) and b.shape=(5, 5, 5, 5, 5, 5), broadcasting is not needed in dimensions 2 and 3, and broadcasting is needed in dimensions 4 and 5. They are merged separately and continuously, and the merged dimension is 4, so the broadcast succeeds. | ||
| @@ -0,0 +1,143 @@ | |||
| 1 | +# Compilation and Running Examples | ||
| 2 | + | ||
| 3 | +## Prerequisites | ||
| 4 | + | ||
| 5 | +- If you need to compile and execute operator APIs, ensure that the basic environment has been set up, including driver, firmware, CANN software package, ops package, etc. | ||
| 6 | +- For the operator API calling process and compilation and running operations, refer to [Application Development (C&C++)](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/programug/acldevg/aclcppdevg_000006.html) under "Single Operator Invocation > Single Operator API Execution > Calling aclnn Interface Example Code". | ||
| 7 | + | ||
| 8 | +## Pre-compilation Preparation | ||
| 9 | + | ||
| 10 | +This chapter takes the development and runtime environment co-location scenario as an example, that is, the machine with AI processor serves as both the development environment and the runtime environment. In this scenario, code development and code running are on the same machine. Here we take the **AddMatMul operator** as an example. The calling logic, process, and compilation script of other operators are roughly the same as the AddMatMul operator. Please modify the API calling script (*.cpp) and compilation script (CMakeLists) according to the actual situation. | ||
| 11 | + | ||
| 12 | +- **Example Code** | ||
| 13 | + | ||
| 14 | + The AddMatMul operator implements tensor addition operation, and the calculation formula is: out = β * self + α * (mat1 @ mat2). You can obtain the example code from the "Calling Example" section in [aclnnAddmm&aclnnInplaceAddmm.md](../../../matmul/mat_mul_v3/docs/aclnnAddmm&aclnnInplaceAddmm.md) and name the code file "**test\_addmm.cpp**". | ||
| 15 | + | ||
| 16 | +- **CMakeLists File** | ||
| 17 | + | ||
| 18 | + The CMake file example is as follows. Please modify according to the actual situation: | ||
| 19 | + | ||
| 20 | + ```bash | ||
| 21 | + # Copyright (c) Huawei Technologies Co., Ltd. 2025. All rights reserved. | ||
| 22 | + | ||
| 23 | + # CMake lowest version requirement | ||
| 24 | + cmake_minimum_required(VERSION 3.14) | ||
| 25 | + | ||
| 26 | + # Set project name | ||
| 27 | + project(ACLNN_EXAMPLE) | ||
| 28 | + | ||
| 29 | + # Compile options | ||
| 30 | + add_compile_options(-std=c++11) | ||
| 31 | + | ||
| 32 | + # Set compilation options | ||
| 33 | + set(CMAKE_RUNTIME_OUTPUT_DIRECTORY "./bin") | ||
| 34 | + set(CMAKE_CXX_FLAGS_DEBUG "-fPIC -O0 -g -Wall") | ||
| 35 | + set(CMAKE_CXX_FLAGS_RELEASE "-fPIC -O2 -Wall") | ||
| 36 | + | ||
| 37 | + # Set executable file name (such as opapi_test) and specify the directory where the operator file *.cpp to be run is located | ||
| 38 | + add_executable(opapi_test | ||
| 39 | + test_addmm.cpp) | ||
| 40 | + | ||
| 41 | + # Set ASCEND_PATH (CANN software package directory, please modify according to the actual path) and INCLUDE_BASE_DIR (header file directory) | ||
| 42 | + if(NOT "$ENV{ASCEND_CUSTOM_PATH}" STREQUAL "") | ||
| 43 | + set(ASCEND_PATH $ENV{ASCEND_CUSTOM_PATH}) | ||
| 44 | + else() | ||
| 45 | + set(ASCEND_PATH "/usr/local/Ascend/cann") | ||
| 46 | + endif() | ||
| 47 | + set(INCLUDE_BASE_DIR "${ASCEND_PATH}/include") | ||
| 48 | + include_directories( | ||
| 49 | + ${INCLUDE_BASE_DIR} | ||
| 50 | + ${INCLUDE_BASE_DIR}/aclnn | ||
| 51 | + ) | ||
| 52 | + | ||
| 53 | + # Set linked library file path | ||
| 54 | + target_link_libraries(opapi_test PRIVATE | ||
| 55 | + ${ASCEND_PATH}/lib64/libascendcl.so | ||
| 56 | + ${ASCEND_PATH}/lib64/libnnopbase.so | ||
| 57 | + ${ASCEND_PATH}/lib64/libopapi_math.so | ||
| 58 | + ${ASCEND_PATH}/lib64/libopapi_nn.so) | ||
| 59 | + | ||
| 60 | + # The executable file is in the bin directory under the CMakeLists file directory | ||
| 61 | + install(TARGETS opapi_test DESTINATION ${CMAKE_RUNTIME_OUTPUT_DIRECTORY}) | ||
| 62 | + ``` | ||
| 63 | + | ||
| 64 | + For operators that combine collective communication and MatMul calculation, and run in parallel, they are collectively called MC2 operators (communication-computation fusion operators), including AllGatherMatmul, AlltoAllAllGatherBatchMatMul, BatchMatMulReduceScatterAlltoAll, MatmulAllReduce, MatmulAllReduceAddRmsNorm, MatmulReduceScatter, etc. When calling such operator APIs, multi-threading and HCCL (Huawei Collective Communication Library) are generally involved. Therefore, the CMake file needs to additionally import the following content, otherwise compilation will fail. | ||
| 65 | + | ||
| 66 | + ```bash | ||
| 67 | + # Set linked library file path | ||
| 68 | + find_package(Threads REQUIRED) | ||
| 69 | + target_link_libraries(opapi_test PRIVATE | ||
| 70 | + ${ASCEND_PATH}/lib64/libascendcl.so | ||
| 71 | + ${ASCEND_PATH}/lib64/libnnopbase.so | ||
| 72 | + ${ASCEND_PATH}/lib64/libopapi_math.so | ||
| 73 | + ${ASCEND_PATH}/lib64/libopapi_nn.so | ||
| 74 | + ${ASCEND_PATH}/lib64/libhccl.so # Collective communication library file | ||
| 75 | + ${CMAKE_THREAD_LIBS_INIT}) # Library file that multi-threading depends on | ||
| 76 | + ``` | ||
| 77 | + | ||
| 78 | + Where "find_package(Threads REQUIRED)" is a CMake command used to find the thread library, which can automatically link the header files or indirectly dependent library files that the thread library depends on. | ||
| 79 | + | ||
| 80 | +## Compilation and Running | ||
| 81 | + | ||
| 82 | + 1. Prepare the operator calling code (*.cpp) and compilation script (CMakeLists.txt) in advance. | ||
| 83 | + 2. Configure environment variables. | ||
| 84 | + | ||
| 85 | + After installing the CANN software, log in to the environment as the CANN runtime user and execute the following command to make the environment variables effective. | ||
| 86 | + | ||
| 87 | + ```bash | ||
| 88 | + source ${INSTALL_DIR}/set_env.sh | ||
| 89 | + ``` | ||
| 90 | + | ||
| 91 | + Where ${INSTALL_DIR} is the storage path after CANN software installation. Please replace according to the actual situation. | ||
| 92 | + 3. Compile and run. | ||
| 93 | + - Enter the directory where CMakeLists.txt is located and execute the following command to create a new build directory to store the generated compilation files. | ||
| 94 | + | ||
| 95 | + ```bash | ||
| 96 | + mkdir -p build | ||
| 97 | + ``` | ||
| 98 | + | ||
| 99 | + - Enter the build directory, execute the cmake command to compile, and then execute the make command to generate the executable file. | ||
| 100 | + | ||
| 101 | + ```bash | ||
| 102 | + cd build | ||
| 103 | + cmake ../ -DCMAKE_CXX_COMPILER=g++ -DCMAKE_SKIP_RPATH=TRUE | ||
| 104 | + make | ||
| 105 | + ``` | ||
| 106 | + | ||
| 107 | + After successful compilation, the opapi\_test executable file will be generated in the bin folder under the build directory. | ||
| 108 | + | ||
| 109 | + - Enter the bin directory and run the executable file opapi_test. | ||
| 110 | + | ||
| 111 | + ```bash | ||
| 112 | + cd bin | ||
| 113 | + ./opapi_test | ||
| 114 | + ``` | ||
| 115 | + | ||
| 116 | + Taking the running result of the AddMatMul operator as an example, the result after running is shown below: | ||
| 117 | + | ||
| 118 | + ```bash | ||
| 119 | + result[0] is: 1.200000 | ||
| 120 | + result[1] is: 2.200000 | ||
| 121 | + result[2] is: 3.200000 | ||
| 122 | + result[3] is: 5.400000 | ||
| 123 | + result[4] is: 6.400000 | ||
| 124 | + result[5] is: 7.400000 | ||
| 125 | + result[6] is: 9.600000 | ||
| 126 | + result[7] is: 10.600000 | ||
| 127 | + ``` | ||
| 128 | + | ||
| 129 | + If the execution result reports an error and the expected result does not appear, you can use the aclGetRecentErrMsg interface to obtain the specific error information. | ||
| 130 | + Example of obtaining exception information when calling aclnnAddmmGetWorkspaceSize fails: | ||
| 131 | + | ||
| 132 | + ```bash | ||
| 133 | + // self is nullptr | ||
| 134 | + ret = aclnnAddmmGetWorkspaceSize(self, mat1, mat2, beta, alpha, out, cubeMathType, &workspaceSize, &executor); | ||
| 135 | + CHECK_RET(ret == ACL_SUCCESS, LOG_PRINT("aclnnAddmmGetWorkspaceSize failed. ERROR: %d\n[ERROR msg]%s", ret, aclGetRecentErrMsg()); return ret); | ||
| 136 | + ``` | ||
| 137 | + | ||
| 138 | + The above null pointer construction problem obtains error information as shown below: | ||
| 139 | + | ||
| 140 | + ```bash | ||
| 141 | + aclnnAddmmGetWorkspaceSize failed. ERROR: 161001 | ||
| 142 | + [ERROR msg][PID:xxxx] xxx(timesamp) AclNN_Parameter_Error(EZ1001): Expected a proper Tensor but got null for argument addmmTennsor.self. | ||
| 143 | + ``` | ||
| @@ -0,0 +1,15 @@ | |||
| 1 | +# Type Conversion Relationships | ||
| 2 | + | ||
| 3 | +When the **output aclTensor data type** of an API (such as aclnnAdd, aclnnMul, etc.) is inconsistent with the **calculation type after input data type promotion**, the API internally converts the calculation result to the data type corresponding to the output type. | ||
| 4 | + | ||
| 5 | +Data type conversion must satisfy the following rules. Conversions that do not satisfy the rules cannot be performed, and parameter validation will fail when calling the API. | ||
| 6 | + | ||
| 7 | + - Floating-point types: ACL\_FLOAT16, ACL\_FLOAT, ACL\_DOUBLE, ACL\_BF16. | ||
| 8 | + - Integer types: ACL\_INT8, ACL\_UINT8, ACL\_INT16, ACL\_UINT16, ACL\_INT32, ACL\_UINT32, ACL\_INT64, ACL\_UINT64. | ||
| 9 | + - Complex types: ACL\_COMPLEX64, ACL\_COMPLEX128. | ||
| 10 | + - Conversions between integer types are supported, as well as conversions to floating-point and complex types. | ||
| 11 | + - Conversions between floating-point types are supported, as well as conversions to complex types. | ||
| 12 | + - Conversions between complex types are supported. | ||
| 13 | + - BOOL supports conversion to integer, floating-point, and complex types. | ||
| 14 | + | ||
| 15 | +Except for the above scenarios, other conversion scenarios are not supported. | ||
| @@ -0,0 +1,36 @@ | |||
| 1 | +# Data Formats | ||
| 2 | + | ||
| 3 | +Data format (format) is used to describe the business semantics of the axes of a multi-dimensional Tensor, representing the physical layout format of data, such as 1D, 2D, 3D, 4D, 5D, and so on. Generally, CNN (Convolutional Neural Networks) APIs require specific formats to be described. | ||
| 4 | + | ||
| 5 | +For the **full range of data formats** supported by aclTensor, refer to [ACL API (C)](https://www.hiascend.com/document/detail/en/canncommercial/latest/API/appdevgapi/aclcppdevg_03_0004.html) under "Data Types and Their Operation Interfaces > aclFormat". | ||
| 6 | + | ||
| 7 | +For an introduction to **data format layout principles**, refer to [Ascend C Operator Development Guide](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/programug/Ascendcopdevg/atlas_ascendc_map_10_0002.html) under "Concept Principles and Terminology > Neural Networks and Operators > Data Layout Formats". | ||
| 8 | + | ||
| 9 | +## Usage Instructions | ||
| 10 | + | ||
| 11 | +Currently, most operator APIs support the ND data format. For example, the aclnnAdd interface indicates that the supported data format is ND (that is, the rule of low-dimensional priority continuous layout for multi-dimensional Tensors). For aclnnConvolution, which is a CNN-type API, the input aclTensor is required to be set with a format that has business semantics, rather than the ND format. Such operators need to know the business semantics in the Tensor during the calculation process to perform the corresponding computation. For example, in 2D convolution, you need to know the correspondence between the Batch dimension, Channel dimension, Height dimension, Width dimension, and the Tensor dimensions. | ||
| 12 | + | ||
| 13 | +>**Note:** | ||
| 14 | +> | ||
| 15 | +>- For the parameter description of two-stage interfaces, to simplify the description, **the original data format "ACL\_FORMAT\_XXXX_" is abbreviated as "_XXXX_"**. | ||
| 16 | +>- The meaning of each dimension in the data format: N (Batch) represents the batch size, H (Height) represents the feature map height, W (Width) represents the feature map width, C (Channels) represents the feature map channels, D (Depth) represents the feature map depth, L (Length) represents the feature map length. | ||
| 17 | + | ||
| 18 | +## Common Data Formats | ||
| 19 | + | ||
| 20 | +When creating an aclTensor through the **aclCreateTensor** interface, you need to set the data format according to the API business requirements. The **supported data formats** are: | ||
| 21 | + | ||
| 22 | +ACL\_FORMAT\_ND, ACL\_FORMAT\_NCHW, ACL\_FORMAT\_NHWC, ACL\_FORMAT\_HWCN, ACL\_FORMAT\_NDHWC, ACL\_FORMAT\_NCDHW, ACL\_FORMAT\_NC, ACL\_FORMAT\_NCL. | ||
| 23 | + | ||
| 24 | +For non-ND Tensors, the Tensor dimension requirements are consistent with the format description. For example: | ||
| 25 | + | ||
| 26 | +- 5D Tensor: Requires ACL\_FORMAT\_NCDHW, ACL\_FORMAT\_NDHWC, or ACL\_FORMAT\_ND (if the API parameter description does not indicate support for ND, setting the ND format will result in an API validation error). | ||
| 27 | +- 4D Tensor: Requires ACL\_FORMAT\_NCHW, ACL\_FORMAT\_NHWC, ACL\_FORMAT\_HWCN, or ACL\_FORMAT\_ND. | ||
| 28 | +- 3D Tensor: Requires ACL\_FORMAT\_NCL or ACL\_FORMAT\_ND. | ||
| 29 | +- 2D Tensor: Requires ACL\_FORMAT\_NC or ACL\_FORMAT\_ND. | ||
| 30 | +- Other dimension Tensors: Require ACL\_FORMAT\_ND. | ||
| 31 | + | ||
| 32 | +## Private Data Formats | ||
| 33 | + | ||
| 34 | +In addition to the common data formats mentioned above, there are other data formats, such as ACL\_FORMAT\_NC1HWC0, ACL\_FORMAT\_FRACTAL\_Z, ACL\_FORMAT\_NC1HWC0\_C04, ACL\_FORMAT\_FRACTAL\_NZ, ACL\_FORMAT\_NDC1HWC0, ACL\_FORMAT\_FRACTAL\_Z\_3D, and so on. | ||
| 35 | + | ||
| 36 | +These formats are private formats of the NPU. Currently, most aclnn APIs do not support these formats. If an individual API declares supported data formats, refer to the actual description of that API. | ||
| @@ -0,0 +1,79 @@ | |||
| 1 | +# Data Structures | ||
| 2 | + | ||
| 3 | +This chapter provides the basic data structures required for calling CANN operator APIs. **Developers do not need to focus on their internal implementation and can use them directly.** | ||
| 4 | + | ||
| 5 | +Note that these basic data structures can be created through the "Public Interfaces" section in [Operator Library Interface](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/API/aolapi/operatorlist_00001.html), such as aclCreateTensor. | ||
| 6 | + | ||
| 7 | +- **aclTensor** | ||
| 8 | + | ||
| 9 | + A structure defined by the framework to manage and store tensor data (such as multi-dimensional data like vectors and matrices). You can create this object through the **aclCreateTensor** interface. | ||
| 10 | + | ||
| 11 | + ```bash | ||
| 12 | + typedef struct aclTensor aclTensor | ||
| 13 | + ``` | ||
| 14 | + | ||
| 15 | +- **aclScalar** | ||
| 16 | + | ||
| 17 | + A structure defined by the framework to manage and store scalar data (that is, a single value). You can create this object through the **aclCreateScalar** interface. | ||
| 18 | + | ||
| 19 | + ```bash | ||
| 20 | + typedef struct aclScalar aclScalar | ||
| 21 | + ``` | ||
| 22 | + | ||
| 23 | +- **aclIntArray** | ||
| 24 | + | ||
| 25 | + An array structure defined by the framework to manage and store integer data. You can create this object through the **aclCreateIntArray** interface. | ||
| 26 | + | ||
| 27 | + ```bash | ||
| 28 | + typedef struct aclIntArray aclIntArray | ||
| 29 | + ``` | ||
| 30 | + | ||
| 31 | +- **aclFloatArray** | ||
| 32 | + | ||
| 33 | + An array structure defined by the framework to manage and store float32 data. You can create this object through the **aclCreateFloatArray** interface. | ||
| 34 | + | ||
| 35 | + ```bash | ||
| 36 | + typedef struct aclFloatArray aclFloatArray | ||
| 37 | + ``` | ||
| 38 | + | ||
| 39 | +- **aclBoolArray** | ||
| 40 | + | ||
| 41 | + An array structure defined by the framework to manage and store boolean data. You can create this object through the **aclCreateBoolArray** interface. | ||
| 42 | + | ||
| 43 | + ```bash | ||
| 44 | + typedef struct aclBoolArray aclBoolArray | ||
| 45 | + ``` | ||
| 46 | + | ||
| 47 | +- **aclTensorList** | ||
| 48 | + | ||
| 49 | + An array structure defined by the framework to manage and store multiple tensor data. You can create this object through the **aclCreateTensorList** interface. | ||
| 50 | + | ||
| 51 | + ```bash | ||
| 52 | + typedef struct aclTensorList aclTensorList | ||
| 53 | + ``` | ||
| 54 | + | ||
| 55 | +- **aclScalarList** | ||
| 56 | + | ||
| 57 | + An array structure defined by the framework to manage and store scalar data. You can create this object through the **aclCreateScalarList** interface. | ||
| 58 | + | ||
| 59 | + ```bash | ||
| 60 | + typedef struct aclScalarList aclScalarList | ||
| 61 | + ``` | ||
| 62 | + | ||
| 63 | +- **aclOpExecutor** | ||
| 64 | + | ||
| 65 | + An executor data structure defined by the framework, which is a container used to execute operator calculations. | ||
| 66 | + | ||
| 67 | + Typically, when calling the first-stage interface aclxxXxxGetWorkspaceSize, the framework automatically creates an aclOpExecutor; after calling the second-stage interface aclxxXxx, the object is automatically released. | ||
| 68 | + | ||
| 69 | + ```bash | ||
| 70 | + typedef struct aclOpExecutor aclOpExecutor | ||
| 71 | + ``` | ||
| 72 | + | ||
| 73 | +- **aclrtStream** | ||
| 74 | + | ||
| 75 | + A stream processing data structure defined by the framework, used to manage and maintain the execution order of some asynchronous operations. | ||
| 76 | + | ||
| 77 | + ```bash | ||
| 78 | + typedef void *aclrtStream | ||
| 79 | + ``` | ||
| @@ -0,0 +1,37 @@ | |||
| 1 | +# Data Types | ||
| 2 | + | ||
| 3 | +When creating an aclTensor through the **aclCreateTensor** interface, refer to [ACL API (C)](https://www.hiascend.com/document/detail/en/canncommercial/latest/API/appdevgapi/aclcppdevg_03_0004.html) for the full list of supported data types under "Data Types and Their Operation Interfaces > aclDataType". | ||
| 4 | + | ||
| 5 | +For the parameter description of two-stage interfaces, the supported data types will use the following abbreviated forms for convenience. | ||
| 6 | + | ||
| 7 | +**Table 1** Data Type Abbreviations | ||
| 8 | + | ||
| 9 | +| Original Data Type | Abbreviation (case-insensitive) | | ||
| 10 | +| :---------------: | :----------------------: | | ||
| 11 | +| ACL_FLOAT | FLOAT or FLOAT32 | | ||
| 12 | +| ACL_FLOAT16 | FLOAT16 | | ||
| 13 | +| ACL_INT8 | INT8 | | ||
| 14 | +| ACL_INT32 | INT32 | | ||
| 15 | +| ACL_UINT8 | UINT8 | | ||
| 16 | +| ACL_INT16 | INT16 | | ||
| 17 | +| ACL_UINT16 | UINT16 | | ||
| 18 | +| ACL_UINT32 | UINT32 | | ||
| 19 | +| ACL_INT64 | INT64 | | ||
| 20 | +| ACL_UINT64 | UINT64 | | ||
| 21 | +| ACL_DOUBLE | DOUBLE or FLOAT64 | | ||
| 22 | +| ACL_BOOL | BOOL | | ||
| 23 | +| ACL_STRING | STRING | | ||
| 24 | +| ACL_COMPLEX64 | COMPLEX64 | | ||
| 25 | +| ACL_COMPLEX128 | COMPLEX128 | | ||
| 26 | +| ACL_BF16 | BF16 or BFLOAT16 | | ||
| 27 | +| ACL_INT4 | INT4 | | ||
| 28 | +| ACL_UINT1 | UINT1 | | ||
| 29 | +| ACL_COMPLEX32 | COMPLEX32 | | ||
| 30 | +| ACL_HIFLOAT8 | HIFLOAT8 | | ||
| 31 | +| ACL_FLOAT8_E5M2 | FLOAT8_E5M2 | | ||
| 32 | +| ACL_FLOAT8_E4M3FN | FLOAT8_E4M3FN | | ||
| 33 | +| ACL_FLOAT8_E8M0 | FLOAT8_E8M0 | | ||
| 34 | +| ACL_FLOAT6_E3M2 | FLOAT6_E3M2 | | ||
| 35 | +| ACL_FLOAT6_E2M3 | FLOAT6_E2M3 | | ||
| 36 | +| ACL_FLOAT4_E2M1 | FLOAT4_E2M1 | | ||
| 37 | +| ACL_FLOAT4_E1M2 | FLOAT4_E1M2 | | ||
| @@ -0,0 +1,81 @@ | |||
| 1 | +# Type Promotion Relationships | ||
| 2 | + | ||
| 3 | +## Promotion Rules | ||
| 4 | + | ||
| 5 | +When the **input aclTensor data types** of an API (such as aclnnAdd, aclnnMul, etc.) are inconsistent, the API internally deduces a data type and converts the input data to that data type for calculation. | ||
| 6 | + | ||
| 7 | +For the data types supported by aclTensor, refer to [Data Types](./data_type.md). Some of these types satisfy the following promotion rules, and the promotion principle is similar to PyTorch's [Type Promotion](https://pytorch.org/docs/stable/tensor_attributes.html#type-promotion-doc). | ||
| 8 | + | ||
| 9 | +> Note: | ||
| 10 | +> | ||
| 11 | +> - For convenience of description, the data types used in the table are **abbreviated forms**, representing: ACL\_FLOAT(f32), ACL\_FLOAT16(f16), ACL\_DOUBLE(f64), ACL\_BF16(bf16), ACL\_INT8(s8), ACL\_UINT8(u8), ACL\_INT16(s16), ACL\_UINT16(u16), ACL\_INT32(s32), ACL\_UINT32(u32), ACL\_INT64(s64), ACL\_UINT64(u64), ACL\_BOOL(bool), ACL\_COMPLEX32(c32), ACL\_COMPLEX64(c64), ACL\_COMPLEX128(c128). | ||
| 12 | +> - The table header and the leftmost column represent the two input data types to be deduced, and the corresponding position in the table represents the deduced data type. | ||
| 13 | +> - The cross mark (×) in the table indicates that these two types cannot perform promotion calculation. | ||
| 14 | + | ||
| 15 | +**Table 1** Data Type Promotion Relationships | ||
| 16 | + | ||
| 17 | +| Data Type | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | bool | c32 | c64 | c128 | | ||
| 18 | +| :------: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | | ||
| 19 | +| **f32** | f32 | f32 | f64 | f32 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c64 | c64 | c128 | | ||
| 20 | +| **f16** | f32 | f16 | f64 | f32 | f16 | f16 | f16 | × | f16 | × | f16 | × | f16 | c32 | c64 | c128 | | ||
| 21 | +| **f64** | f64 | f64 | f64 | f64 | f64 | f64 | f64 | × | f64 | × | f64 | × | f64 | c128 | c128 | c128 | | ||
| 22 | +| **bf16** | f32 | f32 | f64 | bf16 | bf16 | bf16 | bf16 | × | bf16 | × | bf16 | × | bf16 | c32 | c64 | c128 | | ||
| 23 | +| **s8** | f32 | f16 | f64 | bf16 | s8 | s16 | s16 | × | s32 | × | s64 | × | s8 | c32 | c64 | c128 | | ||
| 24 | +| **u8** | f32 | f16 | f64 | bf16 | s16 | u8 | s16 | × | s32 | × | s64 | × | u8 | c32 | c64 | c128 | | ||
| 25 | +| **s16** | f32 | f16 | f64 | bf16 | s16 | s16 | s16 | × | s32 | × | s64 | × | s16 | c32 | c64 | c128 | | ||
| 26 | +| **u16** | × | × | × | × | × | × | × | u16 | × | × | × | × | × | × | × | × | | ||
| 27 | +| **s32** | f32 | f16 | f64 | bf16 | s32 | s32 | s32 | × | s32 | × | s64 | × | s32 | c32 | c64 | c128 | | ||
| 28 | +| **u32** | × | × | × | × | × | × | × | × | × | u32 | × | × | × | × | × | × | | ||
| 29 | +| **s64** | f32 | f16 | f64 | bf16 | s64 | s64 | s64 | × | s64 | × | s64 | × | s64 | c32 | c64 | c128 | | ||
| 30 | +| **u64** | × | × | × | × | × | × | × | × | × | × | × | u64 | × | × | × | × | | ||
| 31 | +| **bool** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | × | s32 | × | s64 | × | bool | c32 | c64 | c128 | | ||
| 32 | +| **c32** | c64 | c32 | c128 | c32 | c32 | c32 | c32 | × | c32 | × | c32 | × | c32 | c32 | c64 | c128 | | ||
| 33 | +| **c64** | c64 | c64 | c128 | c64 | c64 | c64 | c64 | × | c64 | × | c64 | × | c64 | c64 | c64 | c128 | | ||
| 34 | +| **c128** | c128 | c128 | c128 | c128 | c128 | c128 | c128 | × | c128 | × | c128 | × | c128 | c128 | c128 | c128 | | ||
| 35 | + | ||
| 36 | +## Promotion Examples | ||
| 37 | + | ||
| 38 | +- When calling the aclnnAdd interface, if the data types of the input parameters are inconsistent, one is float16 and one is float32, the API internally converts the float16 data type to float32 data type and then performs the calculation. | ||
| 39 | +- When calling the aclnnAdd interface, if the data types of the input parameters are inconsistent, one is float32 and one is bool, the API internally converts the bool data type to float32 data type and then performs the calculation. | ||
| 40 | + | ||
| 41 | +## TensorScalar Promotion Relationships | ||
| 42 | + | ||
| 43 | +### Promotion Rules | ||
| 44 | + | ||
| 45 | +When the **input Tensor data type** and **input Scalar data type** of an API (such as aclnnAdds, aclnnMuls, and so on) are inconsistent, the API internally deduces a data type and converts the input data to that data type for calculation. | ||
| 46 | + | ||
| 47 | +For the data types supported by aclTensor, refer to [Data Types](./data_type.md). Some of these types satisfy the following promotion rules, and the promotion principle is similar to PyTorch's [Type Promotion](https://pytorch.org/docs/stable/tensor_attributes.html#type-promotion-doc). | ||
| 48 | + | ||
| 49 | +The type promotion rules are as follows: | ||
| 50 | + | ||
| 51 | +> Note: | ||
| 52 | +> | ||
| 53 | +> - For convenience of description, the data types used in the table are abbreviated forms, representing: ACL\_FLOAT(f32), ACL\_FLOAT16(f16), ACL\_DOUBLE(f64), ACL\_BF16(bf16), ACL\_INT8(s8), ACL\_UINT8(u8), ACL\_INT16(s16), ACL\_UINT16(u16), ACL\_INT32(s32), ACL\_UINT32(u32), ACL\_INT64(s64), ACL\_UINT64(u64), ACL\_BOOL(bool), ACL\_COMPLEX32(c32), ACL\_COMPLEX64(c64), ACL\_COMPLEX128(c128). | ||
| 54 | +> - The table header represents the input Tensor data type to be deduced, and the leftmost column represents the input Scalar data type to be deduced. The corresponding position in the table represents the deduced data type. | ||
| 55 | +> - The cross mark (×) in the table indicates that these two types cannot perform promotion calculation. | ||
| 56 | + | ||
| 57 | +**Table 2** Data Type Promotion Relationships | ||
| 58 | + | ||
| 59 | +| Data Type | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | bool | c32 | c64 | c128 | | ||
| 60 | +| :------: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | :--: | | ||
| 61 | +| **f32** | f32 | f16 | f64 | bf16 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c32 | c64 | c128 | | ||
| 62 | +| **f16** | f32 | f16 | f64 | bf16 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c32 | c64 | c128 | | ||
| 63 | +| **f64** | f32 | f16 | f64 | bf16 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c128 | c128 | c128 | | ||
| 64 | +| **bf16** | f32 | f16 | f64 | bf16 | f32 | f32 | f32 | × | f32 | × | f32 | × | f32 | c32 | c64 | c128 | | ||
| 65 | +| **s8** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | s8 | c32 | c64 | c128 | | ||
| 66 | +| **u8** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | u8 | c32 | c64 | c128 | | ||
| 67 | +| **s16** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | s16 | c32 | c64 | c128 | | ||
| 68 | +| **u16** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | × | c32 | c64 | c128 | | ||
| 69 | +| **s32** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | s32 | c32 | c64 | c128 | | ||
| 70 | +| **u32** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | × | c32 | c64 | c128 | | ||
| 71 | +| **s64** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | s64 | c32 | c64 | c128 | | ||
| 72 | +| **u64** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | × | c32 | c64 | c128 | | ||
| 73 | +| **bool** | f32 | f16 | f64 | bf16 | s8 | u8 | s16 | u16 | s32 | u32 | s64 | u64 | bool | c32 | c64 | c128 | | ||
| 74 | +| **c32** | c64 | c32 | c128 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c32 | c64 | c128 | | ||
| 75 | +| **c64** | c64 | c32 | c128 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c32 | c64 | c128 | | ||
| 76 | +| **c128** | c64 | c32 | c128 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c64 | c32 | c64 | c128 | | ||
| 77 | + | ||
| 78 | +### Promotion Examples | ||
| 79 | + | ||
| 80 | + - If the input Tensor data type is float16 and the input Scalar data type is float32, the API internally converts the input Scalar float32 data type to float16 data type and then performs the calculation. | ||
| 81 | + - If the input Tensor data type is bool and the input Scalar data type is float32, the API internally converts the input Tensor bool data type to float32 data type and then performs the calculation. | ||
| @@ -0,0 +1,38 @@ | |||
| 1 | +# Non-contiguous Tensor | ||
| 2 | + | ||
| 3 | +Currently, most operator APIs support "**non-contiguous Tensor**" as input aclTensor, that is, a Tensor can be represented by (shape, strides, offset). | ||
| 4 | + | ||
| 5 | +Note: You can create an aclTensor through the "Public Interfaces > aclCreateTensor" section in [Operator Library Interface](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/API/aolapi/operatorlist_00001.html). | ||
| 6 | + | ||
| 7 | +## Example 1 | ||
| 8 | + | ||
| 9 | +For example, consider a Tensor with shape=(6, 5), strides=(10, 1), and offset=22. Its memory layout is as follows: | ||
| 10 | +> a<sub>0,0</sub> , a<sub>0,1</sub> , a<sub>0,2</sub> , a<sub>0,3</sub> , a<sub>0,4</sub> , a<sub>0,5</sub> , a<sub>0,6</sub> , a<sub>0,7</sub> , a<sub>0,8</sub> , a<sub>0,9</sub> | ||
| 11 | +> a<sub>1,0</sub> , a<sub>1,1</sub> , a<sub>1,2</sub> , a<sub>1,3</sub> , a<sub>1,4</sub> , a<sub>1,5</sub> , a<sub>1,6</sub> , a<sub>1,7</sub> , a<sub>1,8</sub> , a<sub>1,9</sub> | ||
| 12 | +> a<sub>2,0</sub> , a<sub>2,1</sub> , **a<sub>2,2</sub> , a<sub>2,3</sub> , a<sub>2,4</sub> , a<sub>2,5</sub> , a<sub>2,6</sub>** , a<sub>2,7</sub> , a<sub>2,8</sub> , a<sub>2,9</sub> | ||
| 13 | +> a<sub>3,0</sub> , a<sub>3,1</sub> , **a<sub>3,2</sub> , a<sub>3,3</sub> , a<sub>3,4</sub> , a<sub>3,5</sub> , a<sub>3,6</sub>** , a<sub>3,7</sub> , a<sub>3,8</sub> , a<sub>3,9</sub> | ||
| 14 | +> a<sub>4,0</sub> , a<sub>4,1</sub> , **a<sub>4,2</sub> , a<sub>4,3</sub> , a<sub>4,4</sub> , a<sub>4,5</sub> , a<sub>4,6</sub>** , a<sub>4,7</sub> , a<sub>4,8</sub> , a<sub>4,9</sub> | ||
| 15 | +> a<sub>5,0</sub> , a<sub>5,1</sub> , **a<sub>5,2</sub> , a<sub>5,3</sub> , a<sub>5,4</sub> , a<sub>5,5</sub> , a<sub>5,6</sub>** , a<sub>5,7</sub> , a<sub>5,8</sub> , a<sub>5,9</sub> | ||
| 16 | +> a<sub>6,0</sub> , a<sub>6,1</sub> , **a<sub>6,2</sub> , a<sub>6,3</sub> , a<sub>6,4</sub> , a<sub>6,5</sub> , a<sub>6,6</sub>** , a<sub>6,7</sub> , a<sub>6,8</sub> , a<sub>6,9</sub> | ||
| 17 | +> a<sub>7,0</sub> , a<sub>7,1</sub> , **a<sub>7,2</sub> , a<sub>7,3</sub> , a<sub>7,4</sub> , a<sub>7,5</sub> , a<sub>7,6</sub>** , a<sub>7,7</sub> , a<sub>7,8</sub> , a<sub>7,9</sub> | ||
| 18 | +> a<sub>8,0</sub> , a<sub>8,1</sub> , a<sub>8,2</sub> , a<sub>8,3</sub> , a<sub>8,4</sub> , a<sub>8,5</sub> , a<sub>8,6</sub> , a<sub>8,7</sub> , a<sub>8,8</sub> , a<sub>8,9</sub> | ||
| 19 | +> a<sub>9,0</sub> , a<sub>9,1</sub> , a<sub>9,2</sub> , a<sub>9,3</sub> , a<sub>9,4</sub> , a<sub>9,5</sub> , a<sub>9,6</sub> , a<sub>9,7</sub> , a<sub>9,8</sub> , a<sub>9,9</sub> | ||
| 20 | + | ||
| 21 | +That is, the Tensor is laid out in the dark positions shown above. This complete Tensor is non-contiguous in memory layout. Strides describe the interval between two adjacent elements in the Tensor dimension. If the stride in dimension 1 is 1, that dimension is contiguous; if the stride in dimension 0 is 10, then adjacent elements are separated by 10 elements, which is non-contiguous. Offset represents the offset of the first element of this Tensor relative to addr. | ||
| 22 | + | ||
| 23 | +## Example 2 | ||
| 24 | + | ||
| 25 | +For example, consider a Tensor with shape=(4, 3), strides=(20, 2), and offset=22. Its memory layout is as follows: | ||
| 26 | + | ||
| 27 | +> a<sub>0,0</sub> , a<sub>0,1</sub> , a<sub>0,2</sub> , a<sub>0,3</sub> , a<sub>0,4</sub> , a<sub>0,5</sub> , a<sub>0,6</sub> , a<sub>0,7</sub> , a<sub>0,8</sub> , a<sub>0,9</sub> | ||
| 28 | +> a<sub>1,0</sub> , a<sub>1,1</sub> , a<sub>1,2</sub> , a<sub>1,3</sub> , a<sub>1,4</sub> , a<sub>1,5</sub> , a<sub>1,6</sub> , a<sub>1,7</sub> , a<sub>1,8</sub> , a<sub>1,9</sub> | ||
| 29 | +> a<sub>2,0</sub> , a<sub>2,1</sub> , **a<sub>2,2</sub>** , a<sub>2,3</sub> , **a<sub>2,4</sub>** , a<sub>2,5</sub> , **a<sub>2,6</sub>** , a<sub>2,7</sub> , a<sub>2,8</sub> , a<sub>2,9</sub> | ||
| 30 | +> a<sub>3,0</sub> , a<sub>3,1</sub> , a<sub>3,2</sub> , a<sub>3,3</sub> , a<sub>3,4</sub> , a<sub>3,5</sub> , a<sub>3,6</sub> , a<sub>3,7</sub> , a<sub>3,8</sub> , a<sub>3,9</sub> | ||
| 31 | +> a<sub>4,0</sub> , a<sub>4,1</sub> , **a<sub>4,2</sub>** , a<sub>4,3</sub> , **a<sub>4,4</sub>** , a<sub>4,5</sub> , **a<sub>4,6</sub>** , a<sub>4,7</sub> , a<sub>4,8</sub> , a<sub>4,9</sub> | ||
| 32 | +> a<sub>5,0</sub> , a<sub>5,1</sub> , a<sub>5,2</sub> , a<sub>5,3</sub> , a<sub>5,4</sub> , a<sub>5,5</sub> , a<sub>5,6</sub> , a<sub>5,7</sub> , a<sub>5,8</sub> , a<sub>5,9</sub> | ||
| 33 | +> a<sub>6,0</sub> , a<sub>6,1</sub> , **a<sub>6,2</sub>** , a<sub>6,3</sub> , **a<sub>6,4</sub>** , a<sub>6,5</sub> , **a<sub>6,6</sub>** , a<sub>6,7</sub> , a<sub>6,8</sub> , a<sub>6,9</sub> | ||
| 34 | +> a<sub>7,0</sub> , a<sub>7,1</sub> , a<sub>7,2</sub> , a<sub>7,3</sub> , a<sub>7,4</sub> , a<sub>7,5</sub> , a<sub>7,6</sub> , a<sub>7,7</sub> , a<sub>7,8</sub> , a<sub>7,9</sub> | ||
| 35 | +> a<sub>8,0</sub> , a<sub>8,1</sub> , **a<sub>8,2</sub>** , a<sub>8,3</sub> , **a<sub>8,4</sub>** , a<sub>8,5</sub> , **a<sub>8,6</sub>** , a<sub>8,7</sub> , a<sub>8,8</sub> , a<sub>8,9</sub> | ||
| 36 | +> a<sub>9,0</sub> , a<sub>9,1</sub> , a<sub>9,2</sub> , a<sub>9,3</sub> , a<sub>9,4</sub> , a<sub>9,5</sub> , a<sub>9,6</sub> , a<sub>9,7</sub> , a<sub>9,8</sub> , a<sub>9,9</sub> | ||
| 37 | + | ||
| 38 | +That is, the Tensor is laid out in the dark positions shown above. This complete Tensor is non-contiguous in memory layout. Strides describe the interval between two adjacent elements in the Tensor dimension. If the stride in dimension 1 is 2, that dimension has an interval of 1 element; if the stride in dimension 0 is 20, then adjacent elements are separated by 20 elements, which is non-contiguous. Offset represents the offset of the first element of this Tensor relative to addr. | ||
| @@ -0,0 +1,59 @@ | |||
| 1 | +# Quantization Introduction | ||
| 2 | + | ||
| 3 | +Quantization is widely used in deep learning models, especially during inference. Through quantization, models can run more efficiently on hardware, reducing computational resource consumption and accelerating the inference process, while also reducing the model's storage requirements. | ||
| 4 | + | ||
| 5 | +CANN operator quantization refers to the computational process of converting the input tensors of matrix (cube) operators such as Matmul in neural networks from high-bit to low-bit representation, while generating the corresponding quantization parameter scale. After the low-bit cube computation is complete, the quantization parameter scale can convert the low-bit values back to high-bit values, thereby ensuring the correctness of the overall computation results (the effect is approximately equivalent to directly using high-bit computation) and effectively improving computational efficiency. | ||
| 6 | + | ||
| 7 | +- Static quantization: Uses predetermined quantization parameters for quantization. In inference scenarios, the quantization of weights generally uses static quantization, which provides better operator performance. | ||
| 8 | +- Dynamic quantization: Uses input data to compute quantization parameters online for quantization. In inference scenarios, the quantization of activations generally uses dynamic quantization, which better adapts to data variations and provides higher precision. In training scenarios, dynamic quantization is also generally used to improve quantization precision. Note that dynamic quantization results in slightly worse operator performance because quantization parameters are generated online. | ||
| 9 | + | ||
| 10 | +## Quantization Modes | ||
| 11 | + | ||
| 12 | +Quantization modes (also known as quantization granularity) refer to the use of different quantization computation levels for different input tensors of an operator. Common quantization computation modes include: | ||
| 13 | + | ||
| 14 | +> Description: | ||
| 15 | +> | ||
| 16 | +> - The m, n, and k variables represent the sizes of different axes for tensor computation. | ||
| 17 | +> - The left matrix and right matrix refer to the two input tensors used for matrix multiplication computation in cube operators. Generally, the left matrix represents the activation and the right matrix represents the weight. Interpret and use them according to the actual situation. | ||
| 18 | + | ||
| 19 | +- Pertensor quantization (abbreviated as T quantization): The quantization target can be either the left matrix or the right matrix. Each tensor shares the same quantization parameter. | ||
| 20 | + | ||
| 21 | + Assuming the left matrix shape is (m, k) and the right matrix shape is (k, n), where k is the reduce axis, the shape of the generated quantization parameter is (1, ). | ||
| 22 | + | ||
| 23 | +  | ||
| 24 | + | ||
| 25 | +- Perchannel quantization (abbreviated as C quantization): The quantization target is the right matrix. Each channel uses an independent quantization parameter. | ||
| 26 | + | ||
| 27 | + Assuming the right matrix shape is (k, n), where k is the reduce axis, the shape of the generated quantization parameter is (n, ). | ||
| 28 | + | ||
| 29 | +  | ||
| 30 | + | ||
| 31 | +- Pertoken quantization (abbreviated as K quantization): The quantization target is the left matrix. Each token uses an independent quantization parameter. | ||
| 32 | + | ||
| 33 | + Assuming the left matrix shape is (m, k), where k is the reduce axis, the shape of the generated quantization parameter is (m, ). | ||
| 34 | + | ||
| 35 | +  | ||
| 36 | + | ||
| 37 | +- Pergroup quantization (abbreviated as G quantization): The quantization target can be either the left matrix or the right matrix. Data is grouped on the reduce axis, and each group uses an independent quantization parameter. | ||
| 38 | + - Assuming the left matrix shape is (m, k), where k is the reduce axis, data is grouped on the k axis with a group size of gs, and the shape of the generated quantization parameter is (m, k/gs). | ||
| 39 | + - Assuming the right matrix shape is (k, n), where k is the reduce axis, data is grouped on the k axis with a group size of gs, and the shape of the generated quantization parameter is (k/gs, n). | ||
| 40 | + | ||
| 41 | +  | ||
| 42 | + | ||
| 43 | +- Perblock quantization (abbreviated as B quantization): The quantization target can be either the left matrix or the right matrix. Data is divided into blocks on all axes, and each block uses an independent quantization parameter. | ||
| 44 | + | ||
| 45 | + - Assuming the left matrix shape is (m, k), where k is the reduce axis, data is grouped on the m and k axes by (bs, bs) blocks, where bs is the block size. The shape of the generated quantization parameter is (m/bs, k/bs). | ||
| 46 | + - Assuming the right matrix shape is (k, n), where k is the reduce axis, data is grouped on the k and n axes by (bs, bs) blocks, where bs is the block size. The shape of the generated quantization parameter is (k/bs, n/bs). | ||
| 47 | + | ||
| 48 | +  | ||
| 49 | + | ||
| 50 | +## Common Combined Quantization | ||
| 51 | + | ||
| 52 | +- Full quantization: Generally refers to the mode in which both the left and right matrices are quantized, including: | ||
| 53 | + - Pertensor-perchannel quantization mode (abbreviated as T-C quantization mode) | ||
| 54 | + - Pertoken-perchannel quantization mode (abbreviated as K-C quantization mode) | ||
| 55 | + - Pergroup-perblock quantization mode (abbreviated as G-B quantization mode) | ||
| 56 | + - Pertensor-perchannel-pergroup quantization mode (abbreviated as T-CG quantization mode) | ||
| 57 | + - Perblock-perblock quantization mode (abbreviated as B-B quantization mode) | ||
| 58 | +- Pseudo-quantization: Generally refers to the mode in which only the weight matrix is quantized, including the perchannel quantization mode (abbreviated as C quantization mode). | ||
| 59 | +- MX quantization: Essentially Microscaling quantization, which maintains model precision at extremely low bits (such as 1 bit) by dynamically adjusting scaling factors. Here it refers to the pergroup-pergroup quantization mode (abbreviated as G-G quantization mode), which is a special case where the quantization parameter type is FLOAT8_E8M0 and the group size is 32. | ||
| @@ -0,0 +1,152 @@ | |||
| 1 | +# sparseMode Introduction | ||
| 2 | + | ||
| 3 | +In the large model field, sparseMode (sparse mode) usually refers to the sparsity design of parameters or activations in the model architecture or calculation formula, as opposed to the dense mode (DenseMode). | ||
| 4 | + | ||
| 5 | +This section introduces common sparseModes and their corresponding scenario descriptions. | ||
| 6 | + | ||
| 7 | +| sparseMode | Meaning | Note | | ||
| 8 | +| ---------- | --------------------- | ------------------ | | ||
| 9 | +| 0 | defaultMask mode. | - | | ||
| 10 | +| 1 | allMask mode. | - | | ||
| 11 | +| 2 | leftUpCausal mode. | - | | ||
| 12 | +| 3 | rightDownCausal mode. | - | | ||
| 13 | +| 4 | band mode. | - | | ||
| 14 | +| 5 | prefix non-compressed mode. | Not supported in varlen scenarios. | | ||
| 15 | +| 6 | prefix compressed mode. | - | | ||
| 16 | +| 7 | varlen outer slice scenario, rightDownCausal mode. | Only supported in varlen scenarios. | | ||
| 17 | +| 8 | varlen outer slice scenario, leftUpCausal mode. | Only supported in varlen scenarios. | | ||
| 18 | + | ||
| 19 | +The working principle of attenMask is to mask the value of the query (Q) and key (K) transpose matrix product at the position where Mask is True, as shown below: | ||
| 20 | + | ||
| 21 | + | ||
| 22 | + | ||
| 23 | +The $QK^T$ matrix will be masked at the position where attenMask is True, with the following effect: | ||
| 24 | + | ||
| 25 | + | ||
| 26 | + | ||
| 27 | +## sparseMode=0 | ||
| 28 | + | ||
| 29 | +When sparseMode is 0, it represents the defaultMask mode. | ||
| 30 | + | ||
| 31 | +- No mask passed: If attenMask is not passed, no mask operation is performed. attenMask takes the value None, and preTokens and nextTokens values are ignored. The Masked $QK^T$ matrix is shown below: | ||
| 32 | + | ||
| 33 | +  | ||
| 34 | + | ||
| 35 | +- nextTokens is 0, preTokens is greater than or equal to Sq, indicating a causal scenario sparse. attenMask should pass a lower triangular matrix. At this time, the part between preTokens and nextTokens needs to be calculated. The Masked $QK^T$ matrix is shown below: | ||
| 36 | + | ||
| 37 | +  | ||
| 38 | + | ||
| 39 | + attenMask should pass a lower triangular matrix, as shown below: | ||
| 40 | + | ||
| 41 | +  | ||
| 42 | + | ||
| 43 | +- preTokens is less than Sq, nextTokens is less than Skv, and both are greater than or equal to 0, indicating a band scenario. At this time, the part between preTokens and nextTokens needs to be calculated. The Masked $QK^T$ matrix is shown below: | ||
| 44 | + | ||
| 45 | +  | ||
| 46 | + | ||
| 47 | + attenMask should pass a band-shaped matrix, as shown below: | ||
| 48 | + | ||
| 49 | +  | ||
| 50 | + | ||
| 51 | +- nextTokens is negative. Taking preTokens=9, nextTokens=-3 as an example, the part between preTokens and nextTokens needs to be calculated. The Masked $QK^T$ is shown below: | ||
| 52 | + | ||
| 53 | + **Note: When nextTokens is negative, preTokens must be greater than or equal to the absolute value of nextTokens, and the absolute value of nextTokens must be less than Skv.** | ||
| 54 | + | ||
| 55 | +  | ||
| 56 | + | ||
| 57 | +- preTokens is negative. Taking nextTokens=7, preTokens=-3 as an example, the part between preTokens and nextTokens needs to be calculated. The Masked $QK^T$ is shown below: | ||
| 58 | + | ||
| 59 | + **Note: When preTokens is negative, nextTokens must be greater than or equal to the absolute value of preTokens, and the absolute value of preTokens must be less than Sq.** | ||
| 60 | + | ||
| 61 | +  | ||
| 62 | + | ||
| 63 | +## sparseMode=1 | ||
| 64 | + | ||
| 65 | +When sparseMode is 1, it represents allMask, that is, passing the complete attenMask matrix. | ||
| 66 | + | ||
| 67 | +In this scenario, nextTokens and preTokens values are ignored. The Masked $QK^T$ matrix is shown below: | ||
| 68 | + | ||
| 69 | + | ||
| 70 | + | ||
| 71 | +## sparseMode=2 | ||
| 72 | + | ||
| 73 | +When sparseMode is 2, it represents the leftUpCausal mode mask, corresponding to the lower triangular scenario divided by the upper-left vertex (parameter starting point is the upper-left corner). | ||
| 74 | + | ||
| 75 | +In this scenario, preTokens and nextTokens values are ignored. The Masked $QK^T$ matrix is shown below: | ||
| 76 | + | ||
| 77 | + | ||
| 78 | + | ||
| 79 | +The passed attenMask is an optimized compressed lower triangular matrix (2048\*2048). The compressed lower triangular matrix is shown below (same below): | ||
| 80 | + | ||
| 81 | + | ||
| 82 | + | ||
| 83 | +## sparseMode=3 | ||
| 84 | + | ||
| 85 | +When sparseMode is 3, it represents the rightDownCausal mode mask, corresponding to the lower triangular scenario divided by the lower-right vertex (parameter starting point is the lower-right corner). | ||
| 86 | + | ||
| 87 | +In this scenario, preTokens and nextTokens values are ignored. attenMask is an optimized compressed lower triangular matrix (2048\*2048). The Masked $QK^T$ matrix is shown below: | ||
| 88 | + | ||
| 89 | + | ||
| 90 | + | ||
| 91 | +## sparseMode=4 | ||
| 92 | + | ||
| 93 | +When sparseMode is 4, it represents the band scenario, that is, calculating the part between preTokens and nextTokens. The parameter starting point is the lower-right corner, and there must be an intersection between preTokens and nextTokens. attenMask is an optimized compressed lower triangular matrix (2048\*2048). The Masked $QK^T$ matrix is shown below: | ||
| 94 | + | ||
| 95 | + | ||
| 96 | + | ||
| 97 | +## sparseMode=5 | ||
| 98 | + | ||
| 99 | +When sparseMode is 5, it represents the prefix non-compressed scenario, that is, adding a matrix with length Sq and width N to the left on the basis of rightDownCausal. The value of N is obtained from the optional input prefix. For example, the figure below shows prefix passing array [4,5] in batch=2 scenario. The N value of each batch axis can be different. The parameter starting point is the upper-left corner. | ||
| 100 | + | ||
| 101 | +In this scenario, preTokens and nextTokens values are ignored. The attenMask matrix data format must be BNSS or B1SS. The Masked $QK^T$ matrix is shown below: | ||
| 102 | + | ||
| 103 | + | ||
| 104 | + | ||
| 105 | +attenMask should pass a matrix as shown below: | ||
| 106 | + | ||
| 107 | + | ||
| 108 | + | ||
| 109 | +## sparseMode=6 | ||
| 110 | + | ||
| 111 | +When sparseMode is 6, it represents the prefix compressed scenario, that is, in the prefix scenario, attenMask is an optimized compressed lower triangular + rectangular matrix (3072\*2048): the upper part is a [2048, 2048] lower triangular matrix, and the lower part is a [1024, 2048] rectangular matrix. The left half of the rectangular matrix is all 0, and the right half is all 1. attenMask should pass a matrix as shown below. In this scenario, preTokens and nextTokens values are ignored. | ||
| 112 | + | ||
| 113 | + | ||
| 114 | + | ||
| 115 | +## sparseMode=7 | ||
| 116 | + | ||
| 117 | +When sparseMode is 7, it indicates a varlen and long sequence outer slice scenario (that is, long sequences are multi-card sliced by query sequence length in the model script). You need to ensure that the scenario using sparseMode 3 was used before outer slicing. In the current mode, you need to set preTokens and nextTokens (starting point is the lower-right vertex), and you need to ensure that the parameters are correct, otherwise there will be precision issues. | ||
| 118 | + | ||
| 119 | +The Masked $QK^T$ matrix is shown below. In the second batch, the query is sliced, and the key and value are not sliced. The 4x6 mask matrix is sliced into 2x6 and 2x6 masks, which are calculated on card 1 and card 2 respectively: | ||
| 120 | + | ||
| 121 | +- The last mask block of card 1 is a band-type mask. Configure preTokens=6 (ensure it is greater than or equal to the last Skv), nextTokens=-2. actual_seq_qlen should pass {3,5}, and actual_seq_kvlen should pass {3,9}. | ||
| 122 | +- The mask type of card 2 remains unchanged after slicing. sparseMode is 3. actual_seq_qlen should pass {2,7,11}, and actual_seq_kvlen should pass {6,11,15}. | ||
| 123 | + | ||
| 124 | + | ||
| 125 | + | ||
| 126 | +**Note**: | ||
| 127 | + | ||
| 128 | +- sparseMode=7, band represents the sparse type of the last non-empty tensor Batch. If there is only one batch, you need to configure parameters according to the band mode requirements. For sparseMode=7, you need to input a 2048x2048 lower triangular mask as the input of this fusion operator. | ||
| 129 | +- The sparse parameters of the band mode generated based on sparseMode=3 outer slicing should meet the following conditions: | ||
| 130 | + - preTokens >= last_Skv. | ||
| 131 | + - last_Sq-last_Skv <= nextTokens <= 0. | ||
| 132 | + - The current mode does not support the optional input pse. | ||
| 133 | +- The non-band mode batch should satisfy: Sq <= Skv. | ||
| 134 | + | ||
| 135 | +## sparseMode=8 | ||
| 136 | + | ||
| 137 | +When sparseMode is 8, it indicates a varlen and long sequence outer slice scenario. You need to ensure that the scenario using sparseMode 2 was used before outer slicing. In the current mode, you need to set preTokens and nextTokens (starting point is the lower-right vertex), and you need to ensure that the parameters are correct, otherwise there will be precision issues. | ||
| 138 | + | ||
| 139 | +The Masked $QK^T$ matrix is shown below. In the second batch, the query is sliced, and the key and value are not sliced. The 5x4 mask matrix is sliced into 2x4 and 3x4 masks, which are calculated on card 1 and card 2 respectively: | ||
| 140 | + | ||
| 141 | +- The mask type of card 1 remains unchanged after slicing. sparseMode is 2. actual_seq_qlen should pass {3,5}, and actual_seq_kvlen should pass {3,7}. | ||
| 142 | +- The first mask block of card 2 is a band-type mask. Configure preTokens=4 (ensure it is greater than or equal to the first Skv), nextTokens=1. actual_seq_qlen should pass {3,8,12}, and actual_seq_kvlen should pass {4,9,13}. | ||
| 143 | + | ||
| 144 | + | ||
| 145 | + | ||
| 146 | +**Note**: | ||
| 147 | + | ||
| 148 | +- sparseMode=8, band represents the sparse type of the first non-empty tensor Batch. If there is only one batch, you need to configure parameters according to the band mode requirements. For sparseMode=8, you need to input a 2048x2048 lower triangular mask as the input of this fusion operator. | ||
| 149 | +- The sparse parameters of the band mode generated based on sparseMode=2 outer slicing should meet the following conditions: | ||
| 150 | + - preTokens >= first_Skv. | ||
| 151 | + - nextTokens >= first_Sq - first_Skv, configure according to the actual situation. | ||
| 152 | + - The current mode does not support the optional input pse. | ||
| @@ -0,0 +1,23 @@ | |||
| 1 | +# Two-stage Interface | ||
| 2 | + | ||
| 3 | +When calling an operator API based on the single-operator API execution method, it is usually divided into "two stages", with the following pattern: | ||
| 4 | + | ||
| 5 | +```Cpp | ||
| 6 | +aclnnStatus aclxxXxxGetWorkspaceSize(const aclTensor *src, ..., aclTensor *out, ..., uint64_t *workspaceSize, aclOpExecutor **executor); | ||
| 7 | +aclnnStatus aclxxXxx(void *workspace, uint64_t workspaceSize, aclOpExecutor *executor, aclrtStream stream); | ||
| 8 | +``` | ||
| 9 | + | ||
| 10 | +You must first call the first-stage interface aclxxXxxGetWorkspaceSize to calculate how much workspace memory is required during this API call. After obtaining the calculated workspaceSize, apply for NPU memory according to the workspaceSize, and then call the second-stage interface aclxxXxx to execute the calculation. | ||
| 11 | + | ||
| 12 | +Here, "aclxx" represents the operator interface prefix, such as aclnn; and "Xxx" represents the corresponding operator type, such as the Add operator. | ||
| 13 | + | ||
| 14 | +> Note: | ||
| 15 | +> | ||
| 16 | +> - workspace refers to the temporary memory required by the API to complete the calculation on the AI processor, in addition to input/output. | ||
| 17 | +> - The second-stage interface aclxxXxx(...) cannot be called repeatedly. The following calling method will cause an exception: | ||
| 18 | +> | ||
| 19 | +> ```Cpp | ||
| 20 | +> aclxxXxxGetWorkspaceSize(...) | ||
| 21 | +> aclxxXxx(...) | ||
| 22 | +> aclxxXxx(...) | ||
| 23 | +> ``` | ||
| @@ -0,0 +1,249 @@ | |||
| 1 | +# Introduction | ||
| 2 | + | ||
| 3 | +CANN Simulator is a SoC-level chip simulation tool designed for operator development scenarios. It analyzes the accuracy and performance data (such as instruction execution status) of AI tasks running on the AI simulator at each stage. This tool helps users perform deep performance tuning, enabling developers to obtain verification results and performance feedback nearly consistent with real chips even when real chips are unavailable or chip resources are scarce. | ||
| 4 | + | ||
| 5 | +# Main Functions | ||
| 6 | + | ||
| 7 | +This tool maintains binary compatibility with on-board execution (the same kernel can be executed on both the simulator and the AI processor). The main uses are as follows: | ||
| 8 | + | ||
| 9 | +* Accuracy simulation: Outputs bit-level accuracy results, helping users complete operator accuracy verification. | ||
| 10 | +* Performance simulation: Outputs instruction pipeline diagrams, helping users identify operator performance bottlenecks. | ||
| 11 | + | ||
| 12 | +# Preparation Before Use | ||
| 13 | + | ||
| 14 | +## Usage Constraints | ||
| 15 | + | ||
| 16 | +* Recommended tool environment configuration: CPU with 16 cores or more, memory of 32 GB or more. | ||
| 17 | +* All paths mentioned in this document must ensure that the running user has read or read-write permissions. | ||
| 18 | +* For security and minimal permissions, it is recommended to use regular user permissions to execute this tool. Avoid using root or other high-privilege accounts. | ||
| 19 | +* This tool depends on the CANN software package. Before using it, install the CANN software package. Driver and firmware installation is not required. Execute the CANN set_env.sh environment variable file through the source command. For security, do not modify the environment variables involved in set_env.sh after executing the source command. | ||
| 20 | +* Users should follow the principle of least privilege. For example, files input to the tool must not be writable by other users. In some more stringent security scenarios, ensure that input files are not writable by group users. | ||
| 21 | +* This tool is a development tool and is not recommended for use in production environments. | ||
| 22 | +* The simulation function of the tool only supports single-card scenarios and cannot simulate multi-card environments. Only card 0 can be set in the code. Modifying the visible card number will cause simulation failure. | ||
| 23 | +* The simulation environment only supports AI Core computation-type operators (MC2 and HCCL type operators are not supported). | ||
| 24 | +* The CANN Simulator tool is currently in the early-access version stage and only supports the Ascend950PR chip. It is recommended that the simulator running environment be configured with a 16-core CPU and 32 GB or more memory. | ||
| 25 | +* ARM environment simulation is not supported at this time. | ||
| 26 | + | ||
| 27 | +## Environment Preparation | ||
| 28 | + | ||
| 29 | +CANN Simulator is integrated in the CANN toolkit package. Complete the software package installation by following [Environment Deployment](../context/quick_install.md). | ||
| 30 | + | ||
| 31 | +# Quick Start | ||
| 32 | + | ||
| 33 | +The following uses [add_examples](../../../examples/add_example/) as an example to describe operator simulation in detail. | ||
| 34 | + | ||
| 35 | +## Operator Compilation | ||
| 36 | + | ||
| 37 | +* Complete the add_example operator compilation and installation by following [Operator Invocation](../invocation/quick_op_invocation.md). | ||
| 38 | + | ||
| 39 | +```bash | ||
| 40 | +# Note: Enter the project root directory and execute the following compilation command. The command is for reference only. For details, refer to the operator invocation instructions. | ||
| 41 | +bash build.sh --pkg --soc=Ascend950 --vendor_name=custom --ops=add_example | ||
| 42 | +# Install the custom operator package | ||
| 43 | +./build_out/cann-ops-nn-${vendor_name}_linux-${arch}.run | ||
| 44 | +``` | ||
| 45 | + | ||
| 46 | +* Complete the compilation of test_aclnn_add_example.cpp by following [aclnn Invocation](../invocation/op_invocation.md#aclnn-invocation), and generate the executable file test_aclnn_add_example. | ||
| 47 | + | ||
| 48 | +## Execute Simulation Command | ||
| 49 | + | ||
| 50 | +```bash | ||
| 51 | +cannsim record ./test_aclnn_add_example -s Ascend950 --gen-report | ||
| 52 | +``` | ||
| 53 | + | ||
| 54 | +The simulation tool execution log files are in the examples/add_example/examples/build/bin/cannsim_* directory. The execution log file is: | ||
| 55 | + | ||
| 56 | +```bash | ||
| 57 | +cannsim.log | ||
| 58 | +``` | ||
| 59 | + | ||
| 60 | +From the simulation tool log file, you can see the print information in the sample: | ||
| 61 | + | ||
| 62 | +```bash | ||
| 63 | +add_example first input[0] is: 1.000000, second input[0] is: 1.000000, result[0] is: 2.000000 | ||
| 64 | +add_example first input[1] is: 1.000000, second input[1] is: 1.000000, result[1] is: 2.000000 | ||
| 65 | +add_example first input[2] is: 1.000000, second input[2] is: 1.000000, result[2] is: 2.000000 | ||
| 66 | +add_example first input[3] is: 1.000000, second input[3] is: 1.000000, result[3] is: 2.000000 | ||
| 67 | +add_example first input[4] is: 1.000000, second input[4] is: 1.000000, result[4] is: 2.000000 | ||
| 68 | +add_example first input[5] is: 1.000000, second input[5] is: 1.000000, result[5] is: 2.000000 | ||
| 69 | +add_example first input[6] is: 1.000000, second input[6] is: 1.000000, result[6] is: 2.000000 | ||
| 70 | +``` | ||
| 71 | + | ||
| 72 | +## View Performance Pipeline | ||
| 73 | + | ||
| 74 | +The simulation performance pipeline files are in the `examples/add_example/examples/build/bin/cannsim_*/report` directory of this project. The pipeline-related file is: | ||
| 75 | + | ||
| 76 | +```bash | ||
| 77 | +trace_core0.json | ||
| 78 | +``` | ||
| 79 | + | ||
| 80 | +Enter "chrome://tracing" in the Chrome browser and drag the generated instruction pipeline diagram file (trace_core0.json) to the blank area to open it. For specific parameter descriptions, refer to the "Simulation Result Analysis" section. | ||
| 81 | + | ||
| 82 | +# Simulation Execution Instructions | ||
| 83 | + | ||
| 84 | +## Command Function | ||
| 85 | + | ||
| 86 | +Execute the application in the simulation environment. | ||
| 87 | + | ||
| 88 | +## Command Format | ||
| 89 | + | ||
| 90 | +cannsim record [options] user_app --user-options | ||
| 91 | + | ||
| 92 | +## Parameter Description | ||
| 93 | + | ||
| 94 | +Table 1 Simulation Execution Parameter Description | ||
| 95 | + | ||
| 96 | +|Parameter|Required/Optional|Description| | ||
| 97 | +| --- | --- | --- | | ||
| 98 | +|-s or --soc-version [options] parameter | Required | Specify the target chip version for simulation (for example: Ascend950).| | ||
| 99 | +|-o or --output [options] parameter | Optional| The path where the generated files are stored. It can be configured as an absolute path or a relative path, and the user executing the tool must have read-write permissions. If the path is not specified, data is saved in the current directory by default.| | ||
| 100 | +|-g or --gen-report [options] parameter | Optional | Enable automatic analysis after simulation completion and generate an analysis report. By default, automatic analysis is not enabled.| | ||
| 101 | +|user_app|Required|Operator executable file.| | ||
| 102 | +|--user-options|Optional|Running parameters of the operator executable file.| | ||
| 103 | + | ||
| 104 | +## Usage Example | ||
| 105 | + | ||
| 106 | +1. Complete operator development and compilation. | ||
| 107 | +2. Execute the simulation command. Refer to the following usage examples: | ||
| 108 | + | ||
| 109 | + ```text | ||
| 110 | + Method 1: Enable simulation and save the output to the ./output directory. /path/to/app is the operator program. | ||
| 111 | + $ cannsim record /path/to/app -o ./output -s Ascend950 | ||
| 112 | + | ||
| 113 | + Method 2: Enable simulation and generate a report for subsequent performance analysis. | ||
| 114 | + $ cannsim record /path/to/app -o ./output -s Ascend950 --gen-report | ||
| 115 | + ``` | ||
| 116 | + | ||
| 117 | +3. After the command completes, a folder named "cannsim_{timestamp}_${user_app}" is generated in the default path or the specified "output" directory. The structure example is as follows: | ||
| 118 | + | ||
| 119 | + ```text | ||
| 120 | + ├─cannsim_{timestamp}_${user_app} | ||
| 121 | + ├── cannsim.log | ||
| 122 | + ``` | ||
| 123 | + | ||
| 124 | +4. You can obtain the operator execution results and compare the accuracy. The results are displayed in cannsim.log. An example is as follows: | ||
| 125 | + | ||
| 126 | + The following output is only an example of the AscendC single-operator direct invocation accuracy comparison result. It may vary slightly depending on the version. Please refer to the actual output. | ||
| 127 | + | ||
| 128 | + ```bash | ||
| 129 | + INFO:root:[INFO] compare data case[ case001] | ||
| 130 | + INFO:root:---------------RESULT--------------- | ||
| 131 | + INFO:root:['case_name', 'wrong_num', 'total_num', 'result', 'task_duration'] | ||
| 132 | + INFO:root:[' case001', 0, 65536, 'Success'] | ||
| 133 | + ``` | ||
| 134 | + | ||
| 135 | +5. View the operator instruction pipeline diagram. Refer to the simulation result analysis section. | ||
| 136 | + | ||
| 137 | +# Simulation Result Analysis Instructions | ||
| 138 | + | ||
| 139 | +## Command Function | ||
| 140 | + | ||
| 141 | +Generate a visualized instruction pipeline diagram. | ||
| 142 | + | ||
| 143 | +## Command Format | ||
| 144 | + | ||
| 145 | +cannsim report [options] | ||
| 146 | + | ||
| 147 | +## Parameter Description | ||
| 148 | + | ||
| 149 | +Table 1 Simulation Result Analysis Parameter Description | ||
| 150 | + | ||
| 151 | +|Parameter | Required/Optional | Description| | ||
| 152 | +| --- | --- | --- | | ||
| 153 | +|-e or --export [options] parameter | Required | The original result file directory. It must be specified as the result directory generated after simulation execution, pointing to the cannsim_{timestamp}_${user_app} level. It can be configured as an absolute path or a relative path, and the tool execution user must have read-write permissions.| | ||
| 154 | +|-o or --output [options] parameter | Optional | The analysis result output directory. It can be configured as an absolute path or a relative path, and the execution user must have read-write permissions. If the path is not specified, data is saved in the current directory by default. If the generated result file has the same name as an existing file, the existing file is overwritten.| | ||
| 155 | +|-n or --core-id [options] parameter | Optional | Specify the core ID for generating the instruction pipeline. If not specified, the pipeline for core 0 is generated by default. The configuration format is as follows: To generate pipelines for all cores, configure 'all'. To specify a core ID range, for example: '0-1'. To specify a single core ID, for example: '5'.| | ||
| 156 | + | ||
| 157 | +## Usage Example | ||
| 158 | + | ||
| 159 | +1. Execute operator simulation by following the simulation execution instructions, and compare the output example to ensure the corresponding results are correct. | ||
| 160 | +2. Execute the simulation result analysis command. Refer to the following execution example. | ||
| 161 | + | ||
| 162 | + ```bash | ||
| 163 | + Generate a performance analysis report in the current directory (default: analyze only core 0) | ||
| 164 | + cannsim report -e /path/to/cannsim_{timestamp}_${user_app} | ||
| 165 | + | ||
| 166 | + Generate performance analysis reports for core 0, core 1, core 11, and core 12 in the specified directory | ||
| 167 | + cannsim report -e /path/to/cannsim_{timestamp}_${user_app} -o /path/to/report -n '0-1, 11-12' | ||
| 168 | + ``` | ||
| 169 | + | ||
| 170 | +3. After the command execution completes, the corresponding pipeline files are generated in the output configured directory. The file format is JSON. The output result example is as follows: | ||
| 171 | + | ||
| 172 | + ```bash | ||
| 173 | + trace_core0.json | ||
| 174 | + trace_core1.json | ||
| 175 | + ... | ||
| 176 | + ``` | ||
| 177 | + | ||
| 178 | +4. View simulation results | ||
| 179 | + Enter "chrome://tracing" in the Chrome browser and drag the generated instruction pipeline diagram file (trace.json) to the blank area to open it. Use keyboard shortcuts (W: zoom in, S: zoom out, A: move left, D: move right) to view the results. | ||
| 180 | +  | ||
| 181 | + | ||
| 182 | + Table 2 Key Field Description | ||
| 183 | + | ||
| 184 | + |Field Name|Field Meaning| | ||
| 185 | + | --- | --- | | ||
| 186 | + |VECTOR|Vector computation unit.| | ||
| 187 | + |SCALAR|Scalar computation unit.| | ||
| 188 | + |Cube|Matrix multiplication computation unit.| | ||
| 189 | + |MTE1|Data transfer pipeline; data transfer direction: L1 ->{L0A/L0B, UBUF}.| | ||
| 190 | + |MTE2|Data transfer pipeline; data transfer direction: {DDR/GM, L2} ->{L1, L0A/B, UBUF}.| | ||
| 191 | + |MTE3|Data transfer pipeline; data transfer direction: UBUF -> {DDR/GM, L2, L1}, L1->{DDR/L2}.| | ||
| 192 | + |FIXP|Data transfer pipeline; data transfer direction: FIXPIPE L0C -> OUT/L1.| | ||
| 193 | + |FLOWCTRL|Control flow instruction.| | ||
| 194 | + |ICACHELOAD|View ICache misses.| | ||
| 195 | + | ||
| 196 | +# Query Help Information | ||
| 197 | + | ||
| 198 | +## Command Function | ||
| 199 | + | ||
| 200 | +Query tool help information. | ||
| 201 | + | ||
| 202 | +## Command Format | ||
| 203 | + | ||
| 204 | +Query tool help information: | ||
| 205 | + | ||
| 206 | +```bash | ||
| 207 | +cannsim --help | ||
| 208 | +``` | ||
| 209 | + | ||
| 210 | +Query tool record subcommand help information: | ||
| 211 | + | ||
| 212 | +```bash | ||
| 213 | +cannsim record --help | ||
| 214 | +``` | ||
| 215 | + | ||
| 216 | +Query tool report subcommand help information: | ||
| 217 | + | ||
| 218 | +```bash | ||
| 219 | +cannsim report --help | ||
| 220 | +``` | ||
| 221 | + | ||
| 222 | +## Parameter Description | ||
| 223 | + | ||
| 224 | +None | ||
| 225 | + | ||
| 226 | +## Usage Example | ||
| 227 | + | ||
| 228 | +1. Log in to the Host-side server. | ||
| 229 | +2. Execute the following command. | ||
| 230 | + | ||
| 231 | + ```bash | ||
| 232 | + cannsim --help | ||
| 233 | + ``` | ||
| 234 | + | ||
| 235 | +## Output Description | ||
| 236 | + | ||
| 237 | +```bash | ||
| 238 | +usage: cannsim [-h] {record,report} ... | ||
| 239 | + | ||
| 240 | +Command-line tool for performance simulation analysis on Ascend hardware. | ||
| 241 | + | ||
| 242 | +positional arguments: | ||
| 243 | + {record,report} Available commands | ||
| 244 | + record Run user application in AscendOps simulation environment | ||
| 245 | + report Generate performance analysis reports | ||
| 246 | + | ||
| 247 | +options: | ||
| 248 | + -h, --help show this help message and exit | ||
| 249 | +``` | ||
| @@ -0,0 +1,190 @@ | |||
| 1 | +# Operator Debugging and Tuning | ||
| 2 | + | ||
| 3 | +## Debugging and Troubleshooting (AI Core Operators) | ||
| 4 | + | ||
| 5 | +If an operator execution failure or accuracy anomaly occurs during operator execution, you can print information at each stage, such as Kernel intermediate results, for problem analysis and troubleshooting. | ||
| 6 | + | ||
| 7 | +### 1. Host-Side Log Acquisition Method | ||
| 8 | + | ||
| 9 | +* **plog acquisition** | ||
| 10 | + | ||
| 11 | + After program execution completes, you can view the logs by default in "$HOME/ascendc/log". The host log file storage path is as follows: | ||
| 12 | + | ||
| 13 | + ```bash | ||
| 14 | + $HOME/ascend/log/debug/plog/plog-pid_*.log | ||
| 15 | + ``` | ||
| 16 | + | ||
| 17 | + Enable the environment variable ASCEND_SLOG_PRINT_TO_STDOUT to display log output directly on the screen (1: enable screen display, 0: disable screen display). The configuration example is as follows: | ||
| 18 | + | ||
| 19 | + ```bash | ||
| 20 | + export ASCEND_SLOG_PRINT_TO_STDOUT=1 | ||
| 21 | + ``` | ||
| 22 | + | ||
| 23 | + For log-related information, refer to [Log Reference](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/maintenref/logreference/logreference_0001.html). For environment variable information, refer to [Environment Variable Reference](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/maintenref/envvar/envref_07_0001.html). | ||
| 24 | + | ||
| 25 | +* **aclnn exception error message acquisition** | ||
| 26 | + | ||
| 27 | + Obtain exception information during aclnn interface invocation through the aclGetRecentErrMsg interface (refer to [acl API (C)](https://www.hiascend.com/document/detail/en/canncommercial/latest/API/appdevgapi/aclcppdevg_03_0004.html)). The usage method is as follows: | ||
| 28 | + | ||
| 29 | + ```bash | ||
| 30 | + printf(aclGetRecentErrMsg()); | ||
| 31 | + ``` | ||
| 32 | + | ||
| 33 | + The printed error message example is as follows: | ||
| 34 | + | ||
| 35 | + ```bash | ||
| 36 | + [PID:646612] 2026-01-24-11:53:44.671.727 AclNN_Parameter_Error(EZ1001): Expected a proper Tensor but got null for argument addmmTensor.self. | ||
| 37 | + ``` | ||
| 38 | + | ||
| 39 | +### 2. Kernel Debugging | ||
| 40 | + | ||
| 41 | +Common debugging methods are as follows: | ||
| 42 | + | ||
| 43 | +* **printf** | ||
| 44 | + | ||
| 45 | + This interface supports printing Scalar-type data, such as integers, characters, and Boolean values. For detailed information, refer to [Ascend C API](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/API/ascendcopapi/atlasascendc_api_07_0003.html) in "Operator Debugging API > printf". | ||
| 46 | + | ||
| 47 | + ```c++ | ||
| 48 | + blockLength_ = tilingData->totalLength / AscendC::GetBlockNum(); | ||
| 49 | + tileNum_ = tilingData->tileNum; | ||
| 50 | + tileLength_ = blockLength_ / tileNum_ / BUFFER_NUM; | ||
| 51 | + // Print the current core computation Block length | ||
| 52 | + AscendC::PRINTF("Tiling blockLength is %llu\n", blockLength_); | ||
| 53 | + ``` | ||
| 54 | + | ||
| 55 | +* **DumpTensor** | ||
| 56 | + | ||
| 57 | + This interface supports dumping the content of a specified Tensor and also supports printing custom additional information, such as the current line number. For detailed information, refer to [Ascend C API](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/API/ascendcopapi/atlasascendc_api_07_0003.html) in "Operator Debugging API > DumpTensor". | ||
| 58 | + | ||
| 59 | + ```c++ | ||
| 60 | + AscendC::LocalTensor<T> zLocal = outputQueueZ.DeQue<T>(); | ||
| 61 | + // Print zLocal Tensor information | ||
| 62 | + DumpTensor(zLocal, 0, 128); | ||
| 63 | + AscendC::DataCopy(outputGMZ[progress * tileLength_], zLocal, tileLength_); | ||
| 64 | + ``` | ||
| 65 | + | ||
| 66 | +For troubleshooting in complex scenarios, such as operator hangs or GM/UB access out-of-bounds, you can use **step-by-step debugging**. For specific operations, refer to the [msDebug](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/devaids/optool/docs/en/quick_start/msdebug_quick_start.md) operator debugging tool. | ||
| 67 | + | ||
| 68 | +## Debugging and Troubleshooting (AI CPU Operators) | ||
| 69 | + | ||
| 70 | +If an operator execution failure or accuracy anomaly occurs during operator execution, you can print information at each stage, such as Kernel intermediate results, for problem analysis and troubleshooting. | ||
| 71 | + | ||
| 72 | +### 1. Host-Side Log Acquisition Method | ||
| 73 | + | ||
| 74 | + Refer to the AI Core operator [Host-Side Log Acquisition Method](#1-host-side-log-acquisition-method) | ||
| 75 | + | ||
| 76 | +### 2. Kernel Debugging | ||
| 77 | + | ||
| 78 | +Common debugging methods are as follows: | ||
| 79 | + | ||
| 80 | +* **KERNEL_LOG macro** | ||
| 81 | + | ||
| 82 | + You can print log information during operator execution through the following macros, including DEBUG, INFO, WARN, and ERROR level logs. | ||
| 83 | + | ||
| 84 | + ```Cpp | ||
| 85 | + KERNEL_LOG_DEBUG(fmt, ...) // The fmt parameter represents the format control string | ||
| 86 | + KERNEL_LOG_INFO(fmt, ...) | ||
| 87 | + KERNEL_LOG_WARN(fmt, ...) | ||
| 88 | + KERNEL_LOG_ERROR(fmt, ...) // ERROR level logs are printed by default | ||
| 89 | + ``` | ||
| 90 | + | ||
| 91 | + To print logs at non-ERROR levels, you need to configure the environment variable `ASCEND_GLOBAL_LOG_LEVEL` in advance. For specific usage, refer to [Environment Variable Reference](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/maintenref/envvar/envref_07_0001.html). | ||
| 92 | + | ||
| 93 | + The printing example is as follows: | ||
| 94 | + | ||
| 95 | + ```c++ | ||
| 96 | + Tensor* input0 = ctx.Input(kFirstInputIndex); | ||
| 97 | + Tensor* input1 = ctx.Input(kSecondInputIndex); | ||
| 98 | + Tensor* output = ctx.Output(0); | ||
| 99 | + | ||
| 100 | + if (input0 == nullptr || input1 == nullptr || output == nullptr) { | ||
| 101 | + // Print error information | ||
| 102 | + KERNEL_LOG_ERROR("Invalid argument"); | ||
| 103 | + return kParamInvalid; | ||
| 104 | + } | ||
| 105 | + | ||
| 106 | + int64_t num_elements = input0->NumElements(); | ||
| 107 | + // Print the number of input elements | ||
| 108 | + KERNEL_LOG_INFO("Num of elements is %ld", data_size); | ||
| 109 | + ``` | ||
| 110 | + | ||
| 111 | +## Performance Tuning | ||
| 112 | + | ||
| 113 | +### Method 1 (For Atlas A2/A3 Series Products) | ||
| 114 | + | ||
| 115 | +If execution accuracy degradation or abnormal memory usage occurs during operator execution, you can use the [msProf](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/devaids/optool/docs/en/quick_start/msopprof_quick_start.md) performance analysis tool to analyze the operator's performance metrics at each execution stage (such as throughput, memory usage, and latency), thereby identifying the root cause and performing targeted optimization. | ||
| 116 | + | ||
| 117 | +This chapter uses the [AddExample custom operator](../../../examples/add_example/) as an example to introduce the two commonly used methods in operator tuning: on-board performance collection and pipeline simulation. By collecting the on-board running pipeline metrics of the operator, you can analyze the operator's Bound scenario. Understanding the simulation pipeline diagram helps optimize the operator's internal pipeline. | ||
| 118 | + | ||
| 119 | +1. Prerequisites. | ||
| 120 | + | ||
| 121 | + After completing operator development and compilation, assuming the aclnn interface invocation method is used, the generated operator executable file (test_aclnn_add_example) is located in the `examples/add_example/examples/build/bin/` directory of this project. | ||
| 122 | + | ||
| 123 | +2. Collect performance data. | ||
| 124 | + | ||
| 125 | + When you need to collect the on-board running pipeline metrics of the operator, navigate to the directory where the operator executable file is located and execute the following command: | ||
| 126 | + | ||
| 127 | + ```bash | ||
| 128 | + msprof op ./test_aclnn_add_example | ||
| 129 | + ``` | ||
| 130 | + | ||
| 131 | + The collection results are in the `examples/add_example/examples/build/bin/OPPROF_*` directory of this project. After collection completes, the following information is printed: | ||
| 132 | + | ||
| 133 | + ``` text | ||
| 134 | + Op Name: AddExample_a1532827238e1555db7b997c7bce2928_high_performance_1 | ||
| 135 | + Op Type: vector | ||
| 136 | + Task Duration(us): 97.861954 | ||
| 137 | + Block Dim: 8 | ||
| 138 | + Mix Block Dim: | ||
| 139 | + Device Id: 0 | ||
| 140 | + Pid: 2776181 | ||
| 141 | + Current Freq: 1800 | ||
| 142 | + Rated Freq: 1800 | ||
| 143 | + ``` | ||
| 144 | + | ||
| 145 | + Task Duration is the current operator Kernel execution time, and Block Dim is the current operator execution core count. | ||
| 146 | + | ||
| 147 | + For detailed pipeline metrics of the operator, refer to the `ArithmeticUtilization` file under `OPPROF_*`, which contains the proportion of each pipeline. For specific descriptions, refer to the [msProf](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/devaids/optool/docs/en/quick_start/msopprof_quick_start.md) section "Performance Data Files > msprof op > ArithmeticUtilization (cube and vector type instruction latency and proportion)". | ||
| 148 | + | ||
| 149 | +3. Collect simulation pipeline diagrams. | ||
| 150 | + | ||
| 151 | + Before using the msProf tool for operator simulation tuning, execute the following command to configure the environment variable. | ||
| 152 | + | ||
| 153 | + ```bash | ||
| 154 | + export LD_LIBRARY_PATH=${INSTALL_DIR}/tools/simulator/Ascendxxxyy/lib:$LD_LIBRARY_PATH | ||
| 155 | + ``` | ||
| 156 | + | ||
| 157 | + Modify the above environment variable according to the actual CANN software package installation path and AI processor model. | ||
| 158 | + | ||
| 159 | + Then navigate to the directory where the operator executable file is located and execute the following command: | ||
| 160 | + | ||
| 161 | + ```bash | ||
| 162 | + msprof op simulator --output=$PWD/pipeline_auto --kernel-name"AddExample" ./test_aclnn_add_example | ||
| 163 | + ``` | ||
| 164 | + | ||
| 165 | + The collection results are in the `$PWD/pipeline_auto/OPPROF_**` directory of this project. | ||
| 166 | + The pipeline-related file path is `OPPROF**/simulator/visualize_data.bin`, which can be viewed using the [mindStudio Insight](https://www.hiascend.com/document/detail/en/mindstudio/latest/visualization_tool/MindStudioInsight/docs/en/user_guide/overview.md) tool. | ||
| 167 | + | ||
| 168 | +### Method 2 (For Ascend 950PR) | ||
| 169 | + | ||
| 170 | +If execution accuracy degradation or abnormal memory usage occurs during operator development, you can use the [CANN Simulator](./cann_simulator.md) simulation tool to analyze the operator's instruction pipeline situation, thereby identifying the root cause and performing targeted optimization. | ||
| 171 | + | ||
| 172 | +This chapter uses the [AddExample custom operator](../../../examples/add_example/) as an example to introduce the use of the simulation tool. It describes how to perform accuracy and performance tuning through the simulation tool. | ||
| 173 | + | ||
| 174 | +1. Prerequisites. | ||
| 175 | + | ||
| 176 | + After completing operator development and compilation, assuming the aclnn interface invocation method is used, the generated operator executable file (test_aclnn_add_example) is located in the `examples/add_example/examples/build/bin/` directory of this project. | ||
| 177 | + | ||
| 178 | +2. Execute the simulation command to generate simulation data. | ||
| 179 | + | ||
| 180 | + ```text | ||
| 181 | + cannsim record ./test_aclnn_add_example -s Ascend950 --gen-report | ||
| 182 | + ``` | ||
| 183 | + | ||
| 184 | + The simulation results are in the `examples/add_example/examples/build/bin/cannsim_*` directory of this project. The pipeline-related file is: | ||
| 185 | + | ||
| 186 | + ```text | ||
| 187 | + trace_core0.json | ||
| 188 | + ``` | ||
| 189 | + | ||
| 190 | +3. Enter "chrome://tracing" in the Chrome browser and drag the generated instruction pipeline diagram file (trace_core0.json) to the blank area to open it. For specific parameter descriptions, refer to the [Simulation Result Analysis](./cann_simulator.md#simulation-result-analysis-instructions) section in CANN Simulator. | ||
| @@ -0,0 +1,247 @@ | |||
| 1 | +# AI CPU Operator Development Guide | ||
| 2 | + | ||
| 3 | +## Overview | ||
| 4 | + | ||
| 5 | +> Note: | ||
| 6 | +> | ||
| 7 | +> 1. For basic concepts and AI CPU interfaces involved in operator development, refer to [TBE & AI CPU Operator Development](https://www.hiascend.com/document/detail/en/CANNCommunityEdition/latest/others/tbeaicpudevg/atlasopdev_10_0001.html) for detailed information. | ||
| 8 | +> 2. AI CPU operators are developed using the C++ language and run on the AI CPU hardware unit. | ||
| 9 | +> 3. build.sh: The commands involved in operator development can be viewed through `bash build.sh --help`. For function parameter descriptions, refer to [build Parameter Description](../context/build.md). | ||
| 10 | + | ||
| 11 | +This development guide uses the `AddExample` operator as an example to introduce the new operator development process and the deliverables involved. For complete sample code, visit the project `examples` directory. | ||
| 12 | + | ||
| 13 | +1. [Project Creation](#project-creation): Before developing an operator, complete the environment deployment and create the operator directory for subsequent operator compilation and deployment. | ||
| 14 | + | ||
| 15 | +2. [Operator Definition](#operator-definition): Determine the operator functionality and prototype definition. | ||
| 16 | + | ||
| 17 | +3. [Kernel Implementation](#kernel-implementation): Implement the Device-side operator kernel function. | ||
| 18 | + | ||
| 19 | +4. [aclnn Adaptation](#aclnn-adaptation): Custom operators are recommended to use the aclnn interface for invocation, which requires completing binary publishing in advance. **If you use graph mode to invoke the operator**, refer to the [Graph Mode Adaptation Guide](./graph_develop_guide.md). | ||
| 20 | + | ||
| 21 | +5. [Compilation and Deployment](#compilation-and-deployment): Complete the compilation and installation of the custom operator through the project compilation script. | ||
| 22 | + | ||
| 23 | +6. [Operator Verification](#operator-verification): Verify the custom operator functionality through common operator invocation methods. | ||
| 24 | + | ||
| 25 | +## Project Creation | ||
| 26 | + | ||
| 27 | +**1. Environment Deployment** | ||
| 28 | + | ||
| 29 | +Before developing an operator, complete the basic environment setup by following [Environment Deployment](../context/quick_install.md). | ||
| 30 | + | ||
| 31 | +**2. Directory Creation** | ||
| 32 | + | ||
| 33 | +Directory creation is an important step in operator development, providing a unified directory structure and file organization for subsequent code writing, compilation, and debugging. | ||
| 34 | + | ||
| 35 | +You can quickly create the operator directory through `build.sh`. Enter the project root directory and execute the following command: | ||
| 36 | + | ||
| 37 | +```bash | ||
| 38 | +# Create the specified operator directory, for example: bash build.sh --genop=activation/op_example | ||
| 39 | +# ${op_class} represents the operator type, such as the activation class. | ||
| 40 | +# ${op_name} represents the lowercase underscore form of the operator name. For example, the `AddExample` operator corresponds to add_example. New operators must not have the same name as existing operators. | ||
| 41 | +bash build.sh --genop_aicpu=${op_class}/${op_name} | ||
| 42 | +``` | ||
| 43 | + | ||
| 44 | +After the command executes successfully, the following message appears: | ||
| 45 | + | ||
| 46 | +```bash | ||
| 47 | +Create the AI CPU initial directory for ${op_name} under ${op_class} success | ||
| 48 | +``` | ||
| 49 | + | ||
| 50 | +After creation, the directory structure is as follows: | ||
| 51 | + | ||
| 52 | +```text | ||
| 53 | +${op_name} # Replace with the lowercase underscore form of the actual operator name | ||
| 54 | +├── examples # Operator invocation samples | ||
| 55 | +│ └── test_aclnn_${op_name}.cpp # Operator aclnn invocation sample | ||
| 56 | +├── op_host # Host-side implementation | ||
| 57 | +│ └── ${op_name}_infershape.cpp # InferShape implementation, implementing operator shape inference, inferring the output shape at runtime | ||
| 58 | +├── op_kernel_aicpu # Device-side Kernel implementation | ||
| 59 | +│ ├── ${op_name}_aicpu.cpp # Kernel entry file, containing the main function and scheduling logic | ||
| 60 | +│ ├── ${op_name}_aicpu.h # Kernel header file, including function declarations, structure definitions, and logic implementation | ||
| 61 | +│ └── ${op_name}.json # Operator information library, defining basic operator information such as name, input/output, and data types | ||
| 62 | +├── tests # UT implementation | ||
| 63 | +│ └── ut # Kernel/aclnn UT implementation | ||
| 64 | +└── CMakeLists.txt # Operator Cmakelist entry | ||
| 65 | +``` | ||
| 66 | + | ||
| 67 | +If `${op_class}` is a new operator category, you need to additionally add `${op_class}` to the `OP_CATEGORY_LIST` in `cmake/variables.cmake`; otherwise, normal compilation is not possible. | ||
| 68 | + | ||
| 69 | +## Operator Definition | ||
| 70 | + | ||
| 71 | +Operator definition requires two deliverables: `README.md` and `${op_name}.json` | ||
| 72 | + | ||
| 73 | +**Deliverable 1: README.md** | ||
| 74 | + | ||
| 75 | +Before developing an operator, determine the functionality and computation logic of the target operator. | ||
| 76 | + | ||
| 77 | +For an example of the custom `AddExample` operator, refer to [AddExample Operator Description](../../../examples/add_example_aicpu/README.md). | ||
| 78 | + | ||
| 79 | +**Deliverable 2: ${op_name}.json** | ||
| 80 | + | ||
| 81 | +Operator information library. | ||
| 82 | + | ||
| 83 | +For an example of the custom `AddExample` operator, refer to [AddExample Operator Information Library](../../../examples/add_example_aicpu/op_kernel_aicpu/add_example.json). | ||
| 84 | + | ||
| 85 | +## Kernel Implementation | ||
| 86 | + | ||
| 87 | +### Kernel Introduction | ||
| 88 | + | ||
| 89 | +Kernel is the core part of an operator executed on the NPU. The Kernel implementation includes the following steps: | ||
| 90 | + | ||
| 91 | +```mermaid | ||
| 92 | +graph LR | ||
| 93 | +H([Operator Class Declaration]) -->A([Compute Function Implementation]) | ||
| 94 | +A -->B([Register Operator]) | ||
| 95 | +``` | ||
| 96 | + | ||
| 97 | +### Code Implementation | ||
| 98 | + | ||
| 99 | +Kernel requires two deliverables: `${op_name}_aicpu.cpp` and `${op_name}_aicpu.h` | ||
| 100 | + | ||
| 101 | +**Deliverable 1: ${op_name}_aicpu.h** | ||
| 102 | + | ||
| 103 | +Operator class declaration | ||
| 104 | + | ||
| 105 | +The first step of Kernel implementation is to declare the operator class in the header file `op_kernel_aicpu/${op_name}_aicpu.h`. The operator class must inherit from the CpuKernel base class. | ||
| 106 | +For detailed implementation, refer to [add_example_aicpu.h](../../../examples/add_example_aicpu/op_kernel_aicpu/add_example_aicpu.h). | ||
| 107 | + | ||
| 108 | +```CPP | ||
| 109 | +// 1. Operator class declaration | ||
| 110 | +// Include the AI CPU base library header file | ||
| 111 | +#include "cpu_kernel.h" | ||
| 112 | +// Define the namespace aicpu (fixed, do not modify), and define the operator Compute implementation function | ||
| 113 | +namespace aicpu { | ||
| 114 | +// The operator class inherits from the CpuKernel base class | ||
| 115 | +class AddExampleCpuKernel : public CpuKernel { | ||
| 116 | + public: | ||
| 117 | + ~AddExampleCpuKernel() = default; | ||
| 118 | + // Declare the Compute function (needs to be overridden); the parameter CpuKernelContext is the context of CPUKernel, including operator input, output, and attribute information | ||
| 119 | + uint32_t Compute(CpuKernelContext &ctx) override; | ||
| 120 | +}; | ||
| 121 | +} // namespace aicpu | ||
| 122 | +``` | ||
| 123 | + | ||
| 124 | +**Deliverable 2: ${op_name}_aicpu.cpp** | ||
| 125 | + | ||
| 126 | +Compute function implementation and AI CPU operator registration | ||
| 127 | + | ||
| 128 | +Obtain the input/output Tensor information and perform validity checks, then implement the core computation logic (such as the addition operation), and set the computation result to the output Tensor. | ||
| 129 | + | ||
| 130 | +For detailed implementation, refer to [add_example_aicpu.cpp](../../../examples/add_example_aicpu/op_kernel_aicpu/add_example_aicpu.cpp). | ||
| 131 | + | ||
| 132 | +```C++ | ||
| 133 | +// 2. Compute function implementation | ||
| 134 | +#include "add_example_aicpu.h" | ||
| 135 | + | ||
| 136 | +namespace { | ||
| 137 | +// Operator name | ||
| 138 | +const char* const kAddExample = "AddExample"; | ||
| 139 | +const uint32_t kParamInvalid = 1; | ||
| 140 | +} // namespace | ||
| 141 | + | ||
| 142 | +// Define the namespace aicpu | ||
| 143 | +namespace aicpu { | ||
| 144 | +// Implement the Compute function of the custom operator class | ||
| 145 | +uint32_t AddExampleCpuKernel::Compute(CpuKernelContext& ctx) { | ||
| 146 | + // Obtain the input tensor from CpuKernelContext | ||
| 147 | + Tensor* input0 = ctx.Input(0); | ||
| 148 | + Tensor* input1 = ctx.Input(1); | ||
| 149 | + // Obtain the output tensor from CpuKernelContext | ||
| 150 | + Tensor* output = ctx.Output(0); | ||
| 151 | + | ||
| 152 | + // Perform basic validation on the tensor; check for null pointers | ||
| 153 | + if (input0 == nullptr || input1 == nullptr || output == nullptr) { | ||
| 154 | + return kParamInvalid; | ||
| 155 | + } | ||
| 156 | + | ||
| 157 | + // Obtain the data type of the input tensor | ||
| 158 | + auto data_type = static_cast<DataType>(input0->GetDataType()); | ||
| 159 | + // Obtain the data address of the input tensor, for example, the input data type is int32 | ||
| 160 | + auto input0_data = reinterpret_cast<int32_t*>(input0->GetData()); | ||
| 161 | + // Obtain the tensor shape | ||
| 162 | + auto input0_shape = input->GetTensorShape(); | ||
| 163 | + | ||
| 164 | + // Obtain the data address of the output tensor, for example, the output data type is int32 | ||
| 165 | + auto y = reinterpret_cast<int32_t*>(output->GetData()); | ||
| 166 | + | ||
| 167 | + // The AddCompute function executes the corresponding computation based on the input type. | ||
| 168 | + // Since C++ does not natively support half-precision floating-point types, you can use the third-party library Eigen (version 3.3.9 recommended) for representation. | ||
| 169 | + switch (data_type) { | ||
| 170 | + case DT_FLOAT: | ||
| 171 | + return AddCompute<float>(...); | ||
| 172 | + case DT_INT32: | ||
| 173 | + return AddCompute<int32>(...); | ||
| 174 | + .... | ||
| 175 | + default : return PARAM_INVALID; | ||
| 176 | + } | ||
| 177 | +} | ||
| 178 | + | ||
| 179 | +// 3. Register the operator Kernel implementation for the framework to obtain the Compute function of the operator Kernel. | ||
| 180 | +REGISTER_CPU_KERNEL(kAddExample, AddExampleCpuKernel); | ||
| 181 | +} // namespace aicpu | ||
| 182 | +``` | ||
| 183 | + | ||
| 184 | +## aclnn Adaptation | ||
| 185 | + | ||
| 186 | +After operator development and compilation are completed, the aclnn interface (a set of C-based APIs) is automatically generated. No additional configuration is required. You can directly invoke the aclnn interface in your application to call the operator. | ||
| 187 | + | ||
| 188 | +## Compilation and Deployment | ||
| 189 | + | ||
| 190 | +After operator development is completed, compile the operator project to generate a custom operator installation package *.run. The specific operations are as follows: | ||
| 191 | + | ||
| 192 | +1. **Preparation.** | ||
| 193 | + | ||
| 194 | + Complete the basic environment setup by following [Project Creation](#project-creation), and check whether the operator development deliverables are complete and in the corresponding operator category directory. | ||
| 195 | + | ||
| 196 | +2. **Compile the custom operator package.** | ||
| 197 | + | ||
| 198 | + Using the `AddExample` operator as an example, assuming the development deliverables are in the `examples` directory, the complete code is in the [add_example](../../../examples/add_example_aicpu) directory. | ||
| 199 | + | ||
| 200 | + ```bash | ||
| 201 | + # Compile the specified operator, for example: bash build.sh --pkg --ops=add_example | ||
| 202 | + bash build.sh --pkg --soc=${soc_version} --vendor_name=${vendor_name} --ops=${op_list} [--experimental] | ||
| 203 | + ``` | ||
| 204 | + | ||
| 205 | + - --soc: ${soc_version} represents the NPU model. For Atlas A2 series products, use "ascend910b" (default). For Atlas A3 series products, use "ascend910_93". For Ascend 950PR/Ascend 950DT products, use "ascend950". | ||
| 206 | + - --vendor_name (optional): ${vendor_name} represents the name of the custom operator package to build. The default name is custom. | ||
| 207 | + - --ops (optional): ${op_list} represents the operators to compile. If not specified, all operators are compiled by default. The format is "--ops=add_example". | ||
| 208 | + - --experimental (optional): If the operator being compiled is a contributed operator, configure --experimental. | ||
| 209 | + | ||
| 210 | + If the following message appears, the compilation is successful: | ||
| 211 | + | ||
| 212 | + ```bash | ||
| 213 | + Self-extractable archive "cann-ops-nn-${vendor_name}-linux.${arch}.run" successfully created. | ||
| 214 | + ``` | ||
| 215 | + | ||
| 216 | +3. **Install the custom operator package.** | ||
| 217 | + | ||
| 218 | + ```bash | ||
| 219 | + # Install the run package | ||
| 220 | + ./build_out/cann-ops-nn-${vendor_name}-linux.${arch}.run | ||
| 221 | + ``` | ||
| 222 | + | ||
| 223 | + The custom operator package is installed in the `${ASCEND_HOME_PATH}/opp/vendors` path. `${ASCEND_HOME_PATH}` represents the CANN software installation directory, which can be configured in the environment variable in advance. | ||
| 224 | + | ||
| 225 | +4. **(Optional) Uninstall the custom operator package.** | ||
| 226 | + | ||
| 227 | + After the custom operator package is installed, an `uninstall.sh` script is generated in the `${ASCEND_HOME_PATH}/opp/vendors/${vendor_name}_nn/scripts` directory. You can uninstall the custom operator package through this script. The command is as follows: | ||
| 228 | + | ||
| 229 | + ```bash | ||
| 230 | + bash ${ASCEND_HOME_PATH}/opp/vendors/${vendor_name}_nn/scripts/uninstall.sh | ||
| 231 | + ``` | ||
| 232 | + | ||
| 233 | +## Operator Verification | ||
| 234 | + | ||
| 235 | +Before verifying the operator, ensure that the environment variables are configured. The command is as follows: | ||
| 236 | + | ||
| 237 | +```bash | ||
| 238 | +export LD_LIBRARY_PATH=${ASCEND_HOME_PATH}/opp/vendors/${vendor_name}_nn/op_api/lib:${LD_LIBRARY_PATH} | ||
| 239 | +``` | ||
| 240 | + | ||
| 241 | +- **UT Verification** | ||
| 242 | + | ||
| 243 | + During operator development, you can quickly verify through UT verification (such as Kernel). | ||
| 244 | + | ||
| 245 | +- **aclnn Invocation Verification** | ||
| 246 | + | ||
| 247 | + After the developed operator is compiled and deployed, you can verify the functionality through the aclnn method. For the method, refer to [Operator Invocation Methods](../invocation/op_invocation.md). | ||
| @@ -0,0 +1,466 @@ | |||
| 1 | +# Cross-Platform Migration Guide for Operators | ||
| 2 | + | ||
| 3 | +This guide describes the key adaptation points and solutions for migrating operators across multiple platforms. Taking the migration of operators from the Atlas A2 series to the Ascend 950 series as an example, it compares hardware architecture differences and related adaptation points, and provides relevant operator adaptation samples. | ||
| 4 | + | ||
| 5 | +## I. Hardware Architecture and Specification Parameter Comparison | ||
| 6 | + | ||
| 7 | +### Atlas A2 Series Hardware Architecture | ||
| 8 | + | ||
| 9 | +<div align="center"> | ||
| 10 | + <img src="../figures/AtlasA2_hardware_architecture.png" width="900" alt="Atlas A2 Hardware Architecture" /> | ||
| 11 | +</div> | ||
| 12 | + | ||
| 13 | +### Ascend 950 Series Hardware Architecture | ||
| 14 | + | ||
| 15 | +<div align="center"> | ||
| 16 | + <img src="../figures/Ascend950_hardware_architecture.png" width="900" alt="Ascend 950 Hardware Architecture" /> | ||
| 17 | +</div> | ||
| 18 | + | ||
| 19 | +### Intergenerational Specification Parameter Comparison | ||
| 20 | + | ||
| 21 | +Multiple product models are typically divided based on different application scenarios, processes, or hardware configurations. Each model may have certain differences in performance, resource configuration, and other aspects. For ease of explanation and direct comparison, this section selects representative configurations as parameter display and difference analysis objects. For other related adjustments, refer to the actual manual or official release. | ||
| 22 | + | ||
| 23 | +<table> | ||
| 24 | + <tr> | ||
| 25 | + <th colspan="2" style="width: 25%;">Specification Item</th> | ||
| 26 | + <th style="width:37.5%;">Atlas A2</th> | ||
| 27 | + <th style="width:37.5%;">Ascend 950</th> | ||
| 28 | + </tr> | ||
| 29 | + <tr> | ||
| 30 | + <td rowspan="4">AICore</td> | ||
| 31 | + <td>Core Count</td> | ||
| 32 | + <td>24</td> | ||
| 33 | + <td>32</td> | ||
| 34 | + </tr> | ||
| 35 | + <tr> | ||
| 36 | + <td>Frequency</td> | ||
| 37 | + <td>1.8</td> | ||
| 38 | + <td>1.65</td> | ||
| 39 | + </tr> | ||
| 40 | + <tr> | ||
| 41 | + <td>Cube Computing Power</td> | ||
| 42 | + <td>353T/376T @BF16,FP16</td> | ||
| 43 | + <td>426T@BF16,FP16 757T@FP8,HIFP8,MXFP8,INT8 1514T@MXFP4</td> | ||
| 44 | + </tr> | ||
| 45 | + <tr> | ||
| 46 | + <td>Vector Computing Power (FP16)</td> | ||
| 47 | + <td>23.5T</td> | ||
| 48 | + <td>54T</td> | ||
| 49 | + </tr> | ||
| 50 | + <tr> | ||
| 51 | + <td rowspan="2">Memory</td> | ||
| 52 | + <td>Memory Capacity (GB)</td> | ||
| 53 | + <td>64</td> | ||
| 54 | + <td>128</td> | ||
| 55 | + </tr> | ||
| 56 | + <tr> | ||
| 57 | + <td>Memory Bandwidth</td> | ||
| 58 | + <td>1.6TB/s</td> | ||
| 59 | + <td>1.6TB/s</td> | ||
| 60 | + </tr> | ||
| 61 | +</table> | ||
| 62 | + | ||
| 63 | +## II. Adaptation Points Introduced by Hardware Capability Changes | ||
| 64 | + | ||
| 65 | +<table> | ||
| 66 | + <tr> | ||
| 67 | + <th style="width: 25%;">Hardware Unit</th> | ||
| 68 | + <th style="width:35%;">Hardware Capability Change</th> | ||
| 69 | + <th style="width:40%;">Typical Impact Scope</th> | ||
| 70 | + </tr> | ||
| 71 | + <tr> | ||
| 72 | + <td rowspan="5">Data Transfer Unit</td> | ||
| 73 | + <td>Removed the data path from L1 to GM</td> | ||
| 74 | + <td>Kernels that rely on L1 directly writing back to GM must be changed to use the L1 to UB to GM or L0C/FIXPIPE to GM path. Related DataCopy links, event synchronization, and buffer planning need to be adjusted.</td> | ||
| 75 | + </tr> | ||
| 76 | + <tr> | ||
| 77 | + <td>Removed the data paths from GM to L0A and L0B</td> | ||
| 78 | + <td>The direct GM to L0A/L0B connection is no longer available. Use GM to L1 to L0A/L0B instead. The L1 tiling strategy and MTE1/MTE2 pipelines need to be restructured.</td> | ||
| 79 | + </tr> | ||
| 80 | + <tr> | ||
| 81 | + <td>ND DMA flexible data transfer, supporting in-line ND to NZ conversion</td> | ||
| 82 | + <td>ND2NZ/DN2NZ can be used to complete format conversion during the MTE2 stage, reducing intermediate buffers and format conversion overhead. Pay attention to stride, alignment, and NZ shape mapping.</td> | ||
| 83 | + </tr> | ||
| 84 | + <tr> | ||
| 85 | + <td>Supports efficient Cube-to-Vector internal data paths: L1 to UB, L0C to UB, FIXP to UB</td> | ||
| 86 | + <td>Intermediate accumulation/activation/fusion (such as K-split accumulation and post-processing) can be performed on the UB side, reducing GM round trips. The corresponding synchronization and pipeline partitioning need to be adjusted.</td> | ||
| 87 | + </tr> | ||
| 88 | + <tr> | ||
| 89 | + <td>Introduced the collective communication accelerator CCU1.0</td> | ||
| 90 | + <td>For communication-computation fusion operators, adjust HcclServerType in Eager mode, and switch to the CCU series GE interfaces in Graph mode.</td> | ||
| 91 | + </tr> | ||
| 92 | + <tr> | ||
| 93 | + <td rowspan="3">Compute Unit</td> | ||
| 94 | + <td>Vector now supports the Regbase paradigm</td> | ||
| 95 | + <td>The memory access patterns, alignment methods, and register count assumptions that originally relied on Membase need to be re-examined. Templates and tiling may need to be updated to the Regbase version.</td> | ||
| 96 | + </tr> | ||
| 97 | + <tr> | ||
| 98 | + <td>Cube no longer supports int4_t</td> | ||
| 99 | + <td>All operators using int4_t need to switch to supported data types (such as int8) and update the quantization calculation logic.</td> | ||
| 100 | + </tr> | ||
| 101 | + <tr> | ||
| 102 | + <td>4:2 sparse matrix computation is not supported</td> | ||
| 103 | + <td>Kernels that originally relied on the 4:2 sparse feature for acceleration need to be changed to dense or other supported sparse strategies, and the performance expectation description needs to be updated.</td> | ||
| 104 | + </tr> | ||
| 105 | + <tr> | ||
| 106 | + <td rowspan="1">Storage Unit</td> | ||
| 107 | + <td>Local Buffer memory improvements: Cube L0C 256 KB, Vector UB 256 KB</td> | ||
| 108 | + <td>Larger L0C/UB allows for increasing the basic block size and double buffering capacity, reducing the number of K-split and block-split rounds. The L1/L0/UB ratio and tile size need to be re-evaluated.</td> | ||
| 109 | + </tr> | ||
| 110 | + <tr> | ||
| 111 | + <td rowspan="2">Other</td> | ||
| 112 | + <td>Performance optimization for multiple cores simultaneously accessing the same Global Memory address</td> | ||
| 113 | + <td>Templates related to matrix multiplication operators can be optimized.</td> | ||
| 114 | + </tr> | ||
| 115 | + <tr> | ||
| 116 | + <td>SIMT</td> | ||
| 117 | + <td>With the introduction of SIMT, thread-level parallelism can be used to handle branching and irregular computation, but it requires adaptation of thread partitioning, shared memory, and synchronization semantics. Some Vector implementations can be migrated to SIMT versions.</td> | ||
| 118 | + </tr> | ||
| 119 | +</table> | ||
| 120 | + | ||
| 121 | +## III. Recommended Migration Steps | ||
| 122 | + | ||
| 123 | +1. Confirm whether the compute units (Cube/Vector) involved in the operator and the supported data types of the corresponding units differ between platforms. | ||
| 124 | +2. Confirm whether the data transfer units involved (ND-to-NZ, GM-to-Lx, collective communication, and so on) differ between platforms. | ||
| 125 | +3. Modify item by item according to the hardware capability change points (Vector architecture, Cube supported data types, L1/L0/UB size, CCU communication, and so on). | ||
| 126 | +4. Refer to the operator migration samples to adjust or supplement the Atlas A2/Ascend 950 branching logic. | ||
| 127 | + | ||
| 128 | +## IV. Operator Migration Samples | ||
| 129 | + | ||
| 130 | +### Cube Matrix Computation Operators | ||
| 131 | + | ||
| 132 | +#### Global Memory Same-Address Access Conflict Optimization | ||
| 133 | + | ||
| 134 | +The Ascend 950 hardware introduces a new feature for parallel processing of same-address requests, eliminating the need to specifically avoid same-address access conflicts in various multi-core scenarios. During migration, the multi-core strategy designed for "offset-based conflict avoidance" on Atlas A2 can be simplified to a more regular sliding window template (such as row-group windowing with column-wise back-and-forth scanning), reducing invalid offsets and redundant address transformations. In practice, it is recommended to first retain the original tile size with the goal of functional equivalence, and then gradually relax the multi-core constraints. Observe key metrics such as MAC utilization, MTE2 utilization, and L2 hit rate based on profiling data to confirm whether the template adjustment brings stable benefits. | ||
| 135 | + | ||
| 136 | +<div align="center"> | ||
| 137 | + <img src="../figures/SWAT_sliding_window_template.png" width="900" alt="SWAT Sliding Window Template" /> | ||
| 138 | +</div> | ||
| 139 | + | ||
| 140 | +#### Tile Size Adjustment | ||
| 141 | + | ||
| 142 | +On Atlas A2, the L0C size is 128 KB, while on Ascend 950 it is increased to 256 KB. This means that a single instance can carry a larger accumulation result block. During migration, prioritize increasing the tile block partitioning granularity or the single-round processing depth in the K direction to reduce the number of block and K-split rounds, thereby reducing loop control and data transfer overhead. At the same time, rebalance the L1/L0/UB capacity budget to avoid pipeline breakpoints caused by L0C expansion squeezing A/B/scale buffering. | ||
| 143 | + | ||
| 144 | +### Vector Computation Operators | ||
| 145 | + | ||
| 146 | +#### SIMT | ||
| 147 | + | ||
| 148 | +The Ascend 950 series introduces a new SIMT unit. SIMT has significant advantages over SIMD in handling irregular discrete access, and is suitable for scenarios with discontinuous addresses, large variations in memory access span, and inconsistent branch paths (such as scatter/gather, index reordering, and sparse updates). | ||
| 149 | + | ||
| 150 | +During migration, it is recommended to prioritize identifying operator sub-processes that are "memory-access-dominated" and have "low vectorization efficiency." If the original SIMD implementation has a large number of mask branches, a high proportion of invalid lanes, or requires complex address assembly, that part can be rewritten to the SIMT path, which typically reduces control overhead and improves effective memory access throughput. | ||
| 151 | + | ||
| 152 | +In practice, focus on the following points: first, the thread task partitioning must match the data sparsity to avoid extremely unbalanced thread loads; second, reduce pipeline idle time caused by high-frequency random memory access, and try to complete index regularization and bucketing upstream; third, decouple boundary processing from the main path to avoid introducing too many branches in hot loops. After migration, it is recommended to compare the "pure SIMD implementation" and the "SIMD + SIMT hybrid implementation" and select the optimal strategy based on data distribution, rather than fixing a single path. | ||
| 153 | + | ||
| 154 | +**Taking the gather_v2 operator as an example: SIMD vs. SIMT implementation comparison** | ||
| 155 | + | ||
| 156 | +The gather_v2 operator performs gather based on the last axis after axis merging. Therefore, the template selection basis is: use the SIMT template when the last axis is less than or equal to 2048, and use the SIMD template when the last axis is greater than 2048. This is because when the last axis is small, multiple discontinuous small block addresses need to be accessed discretely, and SIMT is more efficient. The following compares the core differences between the two implementations: | ||
| 157 | + | ||
| 158 | +**1. Programming Model Differences** | ||
| 159 | + | ||
| 160 | +The SIMD implementation uses the traditional vectorized programming model, requiring explicit management of UB buffers and pipeline queues: | ||
| 161 | + | ||
| 162 | +```cpp | ||
| 163 | +// SIMD: Uses queue mechanism to manage data buffering | ||
| 164 | +TQueBind<QuePosition::VECIN, QuePosition::VECOUT, BUFFER_NUM> inQueue_; | ||
| 165 | +TBuf<QuePosition::VECCALC> indexBuf_; | ||
| 166 | + | ||
| 167 | +// SIMD: Row-by-row processing, explicit data transfer and synchronization | ||
| 168 | +for (int64_t j = 0; j < rows; j++) { | ||
| 169 | + INDICES_T index = GetIndex(yIdx, indiceEndIdx); // Scalar index read | ||
| 170 | + int64_t xIndex = index * tilingData_->innerSize; | ||
| 171 | + DataCopyPad(xLocal[j * colsAlign], xGm[offset], dataCoptExtParams, dataCopyPadExtParams); // Batch continuous data transfer in | ||
| 172 | +} | ||
| 173 | +inQueue_.EnQue<int8_t>(xLocal); // Enqueue for output | ||
| 174 | +``` | ||
| 175 | + | ||
| 176 | +The SIMT implementation uses a thread-level parallelism model, where each thread independently processes elements: | ||
| 177 | + | ||
| 178 | +```cpp | ||
| 179 | +// SIMT: Uses thread-level parallelism, no explicit buffer management required | ||
| 180 | +__simt_vf__ LAUNCH_BOUND(2048) void GatherSimt(...) { | ||
| 181 | + for (INDEX_SIZE_T index = Simt::GetThreadIdx(); | ||
| 182 | + index < currentCoreElements; | ||
| 183 | + index += Simt::GetThreadNum()) { // Thread jump-style parallelism | ||
| 184 | + // Each thread independently computes a single-point index and accesses memory | ||
| 185 | + INDEX_SIZE_T gatherI = Simt::UintDiv(yIndex, m0, shift0); | ||
| 186 | + INDICES_T indicesValue = indices[gatherI]; // Directly access GM based on the single-point index gatherI | ||
| 187 | + y[yIndex] = idxOutOfBound ? 0 : x[xIndex]; // Directly write back to GM | ||
| 188 | + } | ||
| 189 | +} | ||
| 190 | +``` | ||
| 191 | + | ||
| 192 | +**2. Memory Access Pattern Differences** | ||
| 193 | + | ||
| 194 | +| Feature | SIMD Implementation | SIMT Implementation | | ||
| 195 | +|------|----------|----------| | ||
| 196 | +| Data Access | Explicit data transfer to UB through DataCopyPad | Threads directly access GM through `__gm__` pointers | | ||
| 197 | +| Buffer Management | AllocTensor/EnQue/DeQue/FreeTensor required | No explicit buffer required; hardware manages automatically | | ||
| 198 | +| Synchronization Mechanism | Explicit event synchronization (HardEvent::MTE2_V, etc.) | Implicit synchronization between threads | | ||
| 199 | + | ||
| 200 | +**3. Applicable Scenario Differences** | ||
| 201 | + | ||
| 202 | +SIMD is suitable for scenarios with continuous access to large blocks of addresses, efficiently processing continuous data through vectorized instructions. | ||
| 203 | + | ||
| 204 | +SIMT is suitable for discrete memory access, with threads processing in parallel. | ||
| 205 | + | ||
| 206 | +#### Regbase | ||
| 207 | + | ||
| 208 | +The Ascend 950 series introduces the Regbase programming paradigm. Compared with the traditional Membase (Vector API) programming, Regbase is closer to the underlying hardware register operations and provides finer-grained vectorization control capabilities. | ||
| 209 | + | ||
| 210 | +**Features** | ||
| 211 | + | ||
| 212 | +- Uses the underlying APIs under the `AscendC::MicroAPI` namespace | ||
| 213 | +- Directly operates on registers `RegTensor<T>` instead of explicitly managing UB buffer queues | ||
| 214 | +- Implements flexible element-level mask control through `MaskReg` | ||
| 215 | + | ||
| 216 | +**Comparison with the Membase Programming Model** | ||
| 217 | + | ||
| 218 | +| Feature | Membase (Traditional Vector API) | Regbase (MicroAPI) | | ||
| 219 | +|------|---------------------------|---------------------| | ||
| 220 | +| Data Carrier | `LocalTensor<T>` + Queue mechanism | `RegTensor<T>` register | | ||
| 221 | +| Memory Management | Explicit Alloc/EnQue/DeQue/Free | Register auto-allocation | | ||
| 222 | +| Mask Control | Function parameter control | `MaskReg` register control | | ||
| 223 | +| Data Transfer | `DataCopy`/`DataCopyPad` | `MicroAPI::DataCopy` + distribution mode | | ||
| 224 | + | ||
| 225 | +**Code Examples** | ||
| 226 | + | ||
| 227 | +```cpp | ||
| 228 | +__simd_vf__ __aicore__ void GenIndexBuf(ubuf int32_t* helpAddr, int32_t colFactor) | ||
| 229 | +{ | ||
| 230 | + // Declare register tensors | ||
| 231 | + AscendC::MicroAPI::RegTensor<int32_t> v0; | ||
| 232 | + AscendC::MicroAPI::RegTensor<int32_t> v1; | ||
| 233 | + AscendC::MicroAPI::RegTensor<int32_t> vd1; | ||
| 234 | + | ||
| 235 | + // Create a full mask | ||
| 236 | + AscendC::MicroAPI::MaskReg preg = | ||
| 237 | + AscendC::MicroAPI::CreateMask<int32_t, AscendC::MicroAPI::MaskPattern::ALL>(); | ||
| 238 | + | ||
| 239 | + // Duplicate scalar to register | ||
| 240 | + AscendC::MicroAPI::Duplicate(v1, colFactor, preg); | ||
| 241 | + // Generate sequence [0, 1, 2, ...] | ||
| 242 | + AscendC::MicroAPI::Arange(v0, 0); | ||
| 243 | + // Vector operations | ||
| 244 | + AscendC::MicroAPI::Div(vd1, v0, v1, preg); | ||
| 245 | + AscendC::MicroAPI::Mul(vd2, vd1, v1, preg); | ||
| 246 | + AscendC::MicroAPI::Sub(vd3, v0, vd2, preg); | ||
| 247 | + // Write register data back to UB | ||
| 248 | + AscendC::MicroAPI::DataCopy(helpAddr, vd3, preg); | ||
| 249 | +} | ||
| 250 | +``` | ||
| 251 | + | ||
| 252 | +```cpp | ||
| 253 | +// Dynamic mask: Handle tail incomplete data | ||
| 254 | +__simd_vf__ __aicore__ void GatherProcess(ubuf int8_t* curYAddr, uint16_t repeatimes, uint16_t computeSize) | ||
| 255 | +{ | ||
| 256 | + MicroAPI::RegTensor<int8_t> vregTemp; | ||
| 257 | + MicroAPI::MaskReg preg; | ||
| 258 | + | ||
| 259 | + for (uint16_t r = 0; r < repeatTimes; r++) { | ||
| 260 | + // Update mask based on the number of remaining elements | ||
| 261 | + preg = MicroAPI::UpdateMask<int8_t>(sreg); | ||
| 262 | + // Create an address offset register | ||
| 263 | + MicroAPI::AddrReg offset = MicroAPI::CreateAddrReg<int8_t>(r, computeSize); | ||
| 264 | + MicroAPI::DataCopy(vregTemp, curXAddr, offset); | ||
| 265 | + // Store data with mask | ||
| 266 | + MicroAPI::DataCopy(curYAddr, vregTemp, offset, preg); | ||
| 267 | + } | ||
| 268 | +} | ||
| 269 | +``` | ||
| 270 | + | ||
| 271 | +```cpp | ||
| 272 | +// Data aggregation | ||
| 273 | +__VEC_SCOPE__ | ||
| 274 | +{ | ||
| 275 | + MicroAPI::RegTensor<uint32_t> indicesReg; | ||
| 276 | + MicroAPI::RegTensor<int32_t> vd0; | ||
| 277 | + | ||
| 278 | + for (uint16_t indices = 0; indices < indicesLoopNum; indices++) { | ||
| 279 | + // Load indices (E2B distribution mode: broadcast scalar to vector) | ||
| 280 | + MicroAPI::DataCopy<uint32_t, MicroAPI::LoadDist::DIST_E2B_B32>(indicesReg, indicesAddr); | ||
| 281 | + // Gather data aggregation based on indices | ||
| 282 | + MicroAPI::DataCopyGather(vd0, curXAddr, indicesReg, preg); | ||
| 283 | + // Data block copy output | ||
| 284 | + MicroAPI::DataCopy<int32_t, MicroAPI::DataCopyMode::DATA_BLOCK_COPY>( | ||
| 285 | + curYAddr, vd0, blockStride, preg); | ||
| 286 | + } | ||
| 287 | +} | ||
| 288 | +``` | ||
| 289 | + | ||
| 290 | +**Key Regbase API Descriptions** | ||
| 291 | + | ||
| 292 | +| API Category | API Name | Function Description | | ||
| 293 | +|---------|---------|----------| | ||
| 294 | +| Register Type | `RegTensor<T>` | Vector register tensor type | | ||
| 295 | +| Mask Type | `MaskReg` | Mask register type | | ||
| 296 | +| Mask Creation | `CreateMask<T, Pattern>()` | Create a mask (ALL/HALF and other patterns) | | ||
| 297 | +| Mask Update | `UpdateMask<T>(count)` | Dynamically update the mask based on the number of remaining elements | | ||
| 298 | +| Scalar Operation | `Duplicate(reg, val, mask)` | Duplicate a scalar value to all elements of a register | | ||
| 299 | +| Sequence Generation | `Arange(reg, start)` | Generate a continuous sequence | | ||
| 300 | +| Arithmetic Operations | `Add/Sub/Mul/Div(dst, src1, src2, mask)` | Vector arithmetic operations | | ||
| 301 | +| Scalar Operations | `Adds/Muls(dst, src, scalar, mask)` | Vector and scalar operations | | ||
| 302 | +| Type Conversion | `Cast<DT, ST>(dst, src, mask)` | Data type conversion | | ||
| 303 | +| Comparison Operations | `Compare<T, CMPMODE>(mask, src1, src2, pred)` | Vector comparison generates a mask | | ||
| 304 | +| Data Load | `DataCopy<T, LoadDist>(reg, addr)` | Load from UB to register | | ||
| 305 | +| Data Store | `DataCopy<T>(addr, reg, mask)` | Store from register to UB | | ||
| 306 | +| Gather | `DataCopyGather(dst, base, indices, mask)` | Collect data based on indices | | ||
| 307 | +| Address Offset | `CreateAddrReg<T>(loop, stride)` | Create a loop address offset register | | ||
| 308 | + | ||
| 309 | +**LoadDist (Distribution Mode) Descriptions** | ||
| 310 | + | ||
| 311 | +| Mode | Description | Typical Use | | ||
| 312 | +|------|------|----------| | ||
| 313 | +| `DIST_NORM` | Normal continuous load | Continuous data processing | | ||
| 314 | +| `DIST_UNPACK_B16` | 16-bit unpack load | FP16/BF16 to FP32 conversion | | ||
| 315 | +| `DIST_BRC_B32/B16` | Broadcast load | Scalar scale broadcast | | ||
| 316 | +| `DIST_E2B_B32` | Scalar to vector broadcast | Index value broadcast | | ||
| 317 | + | ||
| 318 | +**Migration Suggestions** | ||
| 319 | + | ||
| 320 | +1. Scenarios suitable for Regbase: Cases requiring fine-grained control of register allocation, complex mask logic, and Gather/Scatter memory access patterns. | ||
| 321 | +2. Scenarios to retain Membase: Simple continuous data transfer and computation, double-buffered pipelines. | ||
| 322 | +3. Hybrid use: Combine both paradigms within the same operator, using Regbase to handle core computation logic and Membase to manage data transfer. | ||
| 323 | + | ||
| 324 | +### Cube-Vector Fusion Operators | ||
| 325 | + | ||
| 326 | +#### MTE Data Transfer Path Changes | ||
| 327 | + | ||
| 328 | +The Ascend 950 new architecture introduces direct connection paths between UB-to-L1 and L0C-to-UB, enabling fast transfer of matrix computation data. This aims to simplify CV fusion operator development and improve performance. | ||
| 329 | +<div align="center"> | ||
| 330 | + <img src="../figures/Ascend950_CV_passthrough_link.png" width="700" alt="Ascend 950 New CV Passthrough Link" /> | ||
| 331 | +</div> | ||
| 332 | + | ||
| 333 | +**Matrix Move-In** | ||
| 334 | + | ||
| 335 | +Enable the UB-to-L1 (UB2L1) direct connection path. Through the DataCopy interface, the vector computation results of fusion operators can be directly moved into L1. | ||
| 336 | + | ||
| 337 | +**Matrix Move-Out** | ||
| 338 | + | ||
| 339 | +Enable the L0C-to-UB (L0C2UB) direct connection path. Through the DataCopy interface, the matrix computation results of fusion operators can be directly moved into UB for subsequent vector computation. | ||
| 340 | + | ||
| 341 | +For K-split or multi-stage fusion scenarios, the "L0C moved back to GM and then read back to UB" approach can be changed to "L0C directly to UB for accumulation/post-processing," reducing GM round-trip bandwidth pressure and latency. During migration, it is recommended to place intermediate result merging and activation/quantization pre-processing on the UB side, and explicitly sort out the event synchronization order of MTE1/MTE2/MTE3 and compute units to ensure continuous cross-unit pipelines, avoiding data visibility or synchronization timing issues introduced by the new paths. For the key enabling interface definitions, refer to: | ||
| 342 | + | ||
| 343 | +```cpp | ||
| 344 | +// 1. New: The move-in interface adds UB2L1 Nd2Nz move-in, supporting the form where both Src and Dst are LocalTensor | ||
| 345 | +template <typename T> | ||
| 346 | +__aicore__ inline void DataCopy(const LocalTensor<T>& dst, const LocalTensor<T>& src, const Nd2NzParams& intriParams); | ||
| 347 | + | ||
| 348 | +// 2. New: The move-out interface adds L0C2UB move-out, supporting direct move-out from L0C to UB, supporting the form where both Src and Dst are LocalTensor | ||
| 349 | +template <typename T, typename U, const FixpipeConfig& config = CFG_ROW_MAJOR> | ||
| 350 | +__aicore__ inline void Fixpipe(const LocalTensor<T>& dst, const LocalTensor<U>& src, const FixpipeParamsC310<config.format>& intriParams); | ||
| 351 | +template <CO2Layout format = CO2Layout::ROW_MAJOR> | ||
| 352 | +struct FixpipeParamsC310 { | ||
| 353 | + // ... | ||
| 354 | + uint8_t dualDstCtl = 0; | ||
| 355 | +}; | ||
| 356 | + | ||
| 357 | +// 3. Capability enhancement: The cross-core synchronization interface adds mode 3 | ||
| 358 | +template <uint8_t modeId, pipe_t pipe> | ||
| 359 | +__aicore__ inline void CrossCoreSetFlag(uint16_t flagId) | ||
| 360 | +template <uint8_t modeId = 0, pipe_t pipe = PIPE_S> | ||
| 361 | +__aicore__ inline void CrossCoreWaitFlag(uint16_t flagId) | ||
| 362 | + | ||
| 363 | +``` | ||
| 364 | + | ||
| 365 | +#### Cross-Core Synchronization Semaphore Matching | ||
| 366 | + | ||
| 367 | +`CrossCoreSetFlag` and `CrossCoreWaitFlag` are cross-core synchronization semaphore interfaces, widely used for data dependency and collaborative control between multiple cores. They essentially decouple and orderly advance data processing stages between different AICores in the form of "semaphores," and are commonly used in scenarios such as pipeline control, double-buffer switching, and cross-core collaboration. | ||
| 368 | + | ||
| 369 | +- `CrossCoreSetFlag`: After the current core (or thread) completes data processing for a certain stage, it actively sets the specified flag signal to inform the dependent party (generally another core or downstream pipeline stage) that this stage is complete and the subsequent process can continue. | ||
| 370 | +- `CrossCoreWaitFlag`: The current core (or thread) needs to wait for a certain flag signal to be set (that is, the dependent data or event is complete). After detecting the flag, it continues to execute downward. | ||
| 371 | + | ||
| 372 | +The essence of this semaphore mechanism is to ensure a consistent synchronization sequence between multiple threads or pipeline stages, preventing hardware exceptions such as data races or deadlocks caused by resources not being ready or dependencies not being completed. For detailed interface descriptions, refer to the official documentation: [CrossCoreSetFlag and CrossCoreWaitFlag Cross-Core Synchronization Interface Details](https://www.hiascend.com/document/detail/zh/CANNCommunityEdition/900beta1/API/ascendcopapi/atlasascendc_api_07_0273.html). | ||
| 373 | + | ||
| 374 | +On Ascend 950, the numbers of `CrossCoreWaitFlag` and `CrossCoreSetFlag` must strictly match, and it is recommended to design them in pairs within the same synchronization semantic domain, following the "produce first, then consume" order. On Atlas A2, if there are redundant `CrossCoreSetFlag` semaphores between operators, HWTS performs special handling to clear the counter. The Ascend 950 series, to reduce hardware overhead, no longer relies on this type of fallback mechanism, requiring that the synchronization semaphores within a single operator kernel match one-to-one. Otherwise, a deterministic hang will occur. | ||
| 375 | + | ||
| 376 | +During migration, focus on troubleshooting the following issues: first, abnormal branch early returns that cause only `Set` to be executed without the corresponding `Wait` (or vice versa); second, multi-stage pipelines reusing the same `flagId` but with overlapping lifecycles, causing "cross-stage crosstalk"; third, conditionally triggered synchronization within loops but with unaligned loop boundaries, resulting in inconsistent iteration counts. The above issues may be masked on Atlas A2 but directly exposed as blocking timeouts or deadlocks on Ascend 950. For operators with complex cross-core pipelines, first build a minimal dataset for single-stage verification, and then gradually add double buffering and multiple stages to reduce the complexity of locating synchronization issues. | ||
| 377 | + | ||
| 378 | +### Collective Communication Operators | ||
| 379 | + | ||
| 380 | +Ascend 950 introduces the collective communication accelerator CCU1.0, which reduces memory access requirements and scheduling latency. To effectively utilize this feature, the cross-chip communication method of operators is changed from AICPU on A2 to CCU communication. | ||
| 381 | + | ||
| 382 | +**Eager Mode** | ||
| 383 | + | ||
| 384 | +In the second-stage interface of the aclnn two-phase interface, specify the collective communication type for the operator executor aclOpExecutor. | ||
| 385 | + | ||
| 386 | +Taking the [MatmulAllReduce](https://gitcode.com/cann/ops-transformer/tree/master/mc2/matmul_all_reduce) operator migration as an example: | ||
| 387 | +Set the NnopbaseSetHcclServerType enum value. For A2, it is NNOPBASE_HCCL_SERVER_AICPU, and for 950, it is NNOPBASE_HCCL_SERVER_TYPE_CCU. | ||
| 388 | + | ||
| 389 | +```CPP | ||
| 390 | +// ... | ||
| 391 | +aclnnStatus aclnnMatmulAllReduce( | ||
| 392 | + void* workspace, uint64_t workspaceSize, aclOpExecutor* executor, const aclrtStream stream) | ||
| 393 | +{ | ||
| 394 | + // ... | ||
| 395 | + if (NnopbaseSetHcclServerType) { | ||
| 396 | + if (op::GetCurrentPlatformInfo().GetCurNpuArch() == NpuArch::DAV_3510) { | ||
| 397 | + NnopbaseSetHcclServerType(executor, NnopbaseHcclServerType::NNOPBASE_HCCL_SERVER_TYPE_CCU); | ||
| 398 | + } | ||
| 399 | + } | ||
| 400 | + // ... | ||
| 401 | + return ACLNN_SUCCESS; | ||
| 402 | +} | ||
| 403 | +``` | ||
| 404 | + | ||
| 405 | +**Graph Mode** | ||
| 406 | + | ||
| 407 | +1. In the CalcParamFunc callback interface used for resource computation and application, which involves auxiliary stream-related information, differentiate the collective communication type of the auxiliary stream for the GE context context. | ||
| 408 | +2. In the GenerateTask callback interface used for setting custom tasks and parameter customization on the main stream and auxiliary streams, differentiate between the two sets of GE KernelLaunch interfaces, and call the AICPU communication or CCU communication creation and customization processes respectively. | ||
| 409 | + | ||
| 410 | +For the static graph GE side, the task type for creating communication tasks is aicpu kfc server + kfc_stream for A2, and ccu server + ccu_stream for 950. The related code file is: [matmul_all_reduce_gen_task.cpp](https://gitcode.com/cann/ops-transformer/blob/master/mc2/matmul_all_reduce/op_graph/matmul_all_reduce_gen_task.cpp) | ||
| 411 | + | ||
| 412 | +```CPP | ||
| 413 | +// ... | ||
| 414 | +ge::Status MatmulAllReduceCalcParamFunc(gert::ExeResGenerationContext *context) | ||
| 415 | +{ | ||
| 416 | + if (Mc2GenTaskOpsUtils::IsTargetPlatformNpuArch(context->GetNodeName(), NPUARCH_A5)) { | ||
| 417 | + // 950 | ||
| 418 | + return Mc2GenTaskOpsUtils::CommonKFCMc2CalcParamFunc(context, "ccu server", "ccu_stream"); | ||
| 419 | + } | ||
| 420 | + // A2 | ||
| 421 | + return Mc2GenTaskOpsUtils::CommonKFCMc2CalcParamFunc(context, "aicpu kfc server", "kfc_stream"); | ||
| 422 | +} | ||
| 423 | +// ... | ||
| 424 | +``` | ||
| 425 | + | ||
| 426 | +The static graph GenTask invocation interfaces differ, and the processes are different. The related code file is: [matmul_all_reduce_gen_task.cpp](https://gitcode.com/cann/ops-transformer/blob/master/mc2/matmul_all_reduce/op_graph/matmul_all_reduce_gen_task.cpp) | ||
| 427 | + | ||
| 428 | +```CPP | ||
| 429 | +// ... | ||
| 430 | +// A2 | ||
| 431 | +ge::Status MatmulAllReduceGenTaskOpsUtils::MatmulAllReduceGenTaskCallback( | ||
| 432 | + const gert::ExeResGenerationContext *context, std::vector<std::vector<uint8_t>>& tasks) { | ||
| 433 | + // ... | ||
| 434 | + // aicpu task | ||
| 435 | + ge::KernelLaunchInfo aicpu_task = | ||
| 436 | + ge::KernelLaunchInfo::CreateAicpuKfcTask(context, SO_NAME.c_str(), KERNEL_NAME_V1.c_str()); | ||
| 437 | + // ... | ||
| 438 | +} | ||
| 439 | + | ||
| 440 | +// 950 | ||
| 441 | +ge::Status Mc2Arch35GenTaskOpsUtils::Mc2Arch35GenTaskCallBack(const gert::ExeResGenerationContext *context, std::vector<std::vector<uint8_t>> &tasks) { | ||
| 442 | + // ... | ||
| 443 | + // ccu task | ||
| 444 | + ge::KernelLaunchInfo ccuTask = ge::KernelLaunchInfo::CreateCcuTask(context, ccuGroups); | ||
| 445 | + // ... | ||
| 446 | +} | ||
| 447 | + | ||
| 448 | +ge::Status MatmulAllReduceGenTaskFunc(const gert::ExeResGenerationContext *context, std::vector<std::vector<uint8_t>> &tasks) | ||
| 449 | +{ | ||
| 450 | + if (Mc2GenTaskOpsUtils::IsTargetPlatformNpuArch(context->GetNodeName(), NPUARCH_A5)) { | ||
| 451 | + // 950 | ||
| 452 | + return Mc2Arch35GenTaskOpsUtils::Mc2Arch35GenTaskCallBack(context, tasks); | ||
| 453 | + } | ||
| 454 | + // A2 | ||
| 455 | + return MatmulAllReduceGenTaskOpsUtils::MatmulAllReduceGenTaskCallback(context, tasks); | ||
| 456 | +} | ||
| 457 | +// ... | ||
| 458 | +``` | ||
| 459 | + | ||
| 460 | +## V. Common Issues and Performance Tuning Suggestions (FAQ/Performance Tips) | ||
| 461 | + | ||
| 462 | +If the performance of an operator on Ascend 950 does not improve but declines, prioritize the following checks: | ||
| 463 | + | ||
| 464 | +1. Whether the Atlas A2 offset-based multi-core template is still being used. | ||
| 465 | +2. Whether CCU communication is not enabled and AICPU is still being used. | ||
| 466 | +3. Whether the tiling still uses the Atlas A2 L1/L0/UB partitioning strategy, resulting in the larger on-chip cache of Ascend 950 not being fully utilized. | ||
| @@ -0,0 +1,118 @@ | |||
| 1 | +# Graph Mode Adaptation Guide | ||
| 2 | + | ||
| 3 | +## Overview | ||
| 4 | + | ||
| 5 | +If a custom operator needs to run in graph mode, the overall process is consistent with the operator development guide ([AI Core Operator Development Guide](aicore_develop_guide.md)/[AI CPU Operator Development Guide](aicpu_develop_guide.md)). Note that **aclnn adaptation is not required**; only the following deliverable adaptations are needed. | ||
| 6 | + | ||
| 7 | +```text | ||
| 8 | +${op_name} # Replace with the lowercase underscore form of the actual operator name | ||
| 9 | +├── op_host # Host-side implementation | ||
| 10 | +│ └── ${op_name}_infershape.cpp # InferShape implementation, implementing operator shape inference, inferring the output shape at runtime | ||
| 11 | +├── op_graph # Graph fusion-related implementation | ||
| 12 | +│ ├── CMakeLists.txt # op_graph side cmakelist file | ||
| 13 | +│ ├── ${op_name}_graph_infer.cpp # InferDataType file, implementing operator type inference, inferring the output dataType at runtime | ||
| 14 | +└── └── ${op_name}_proto.h # Operator prototype definition, used for graph optimization and fusion phase operator identification | ||
| 15 | +``` | ||
| 16 | + | ||
| 17 | +This document uses the `AddExample` operator (assuming it is an AI Core operator) as an example to introduce the implementation of graph mode deliverables. The implementation for AI CPU operators entering the graph is basically similar. For complete code, refer to the `add_example` and `add_example_aicpu` directories under `examples`. | ||
| 18 | + | ||
| 19 | +## Shape and DataType Inference | ||
| 20 | + | ||
| 21 | +Graph mode requires two deliverables: `${op_name}_graph_infer.cpp` and `${op_name}_infershape.cpp` | ||
| 22 | + | ||
| 23 | +**Deliverable 1: ${op_name}_infershape.cpp** | ||
| 24 | + | ||
| 25 | +The InferShape function infers the output shape based on the input shape. | ||
| 26 | + | ||
| 27 | +The example is as follows. For the complete code of the `AddExample` operator, refer to [add_example_infershape.cpp](../../../examples/add_example/op_host/add_example_infershape.cpp) under `examples/add_example/op_host`. | ||
| 28 | + | ||
| 29 | +```C++ | ||
| 30 | +// The AddExample operator logic is adding two numbers, so the output shape is consistent with the input shape | ||
| 31 | +static ge::graphStatus InferShapeAddExample(gert::InferShapeContext* context) | ||
| 32 | +{ | ||
| 33 | + .... | ||
| 34 | + // Obtain the input shape | ||
| 35 | + const gert::Shape* xShape = context->GetInputShape(IDX_0); | ||
| 36 | + // Obtain the output shape | ||
| 37 | + gert::Shape* yShape = context->GetOutputShape(IDX_0); | ||
| 38 | + // Obtain the input DimNum | ||
| 39 | + auto xShapeSize = xShape->GetDimNum(); | ||
| 40 | + // Set the output DimNum | ||
| 41 | + yShape->SetDimNum(xShapeSize); | ||
| 42 | + // Set the input Dim values to the output one by one | ||
| 43 | + for (size_t i = 0; i < xShapeSize; i++) { | ||
| 44 | + int64_t dim = xShape->GetDim(i); | ||
| 45 | + yShape->SetDim(i, dim); | ||
| 46 | + } | ||
| 47 | + .... | ||
| 48 | +} | ||
| 49 | +// InferShape registration | ||
| 50 | +IMPL_OP_INFERSHAPE(AddExample).InferShape(InferShapeAddExample); | ||
| 51 | +``` | ||
| 52 | + | ||
| 53 | +**Deliverable 2: ${op_name}_graph_infer.cpp** | ||
| 54 | + | ||
| 55 | +The InferDataType function infers the output DataType based on the input DataType. The example is as follows. | ||
| 56 | + | ||
| 57 | +```C++ | ||
| 58 | +// The AddExample operator logic is adding two numbers, so the output dataType is consistent with the input dataType | ||
| 59 | +static ge::graphStatus InferDataTypeAddExample(gert::InferDataTypeContext* context) | ||
| 60 | +{ | ||
| 61 | + .... | ||
| 62 | + // Obtain the input dataType | ||
| 63 | + ge::DataType sizeDtype = context->GetInputDataType(IDX_0); | ||
| 64 | + // Set the input dataType to the output | ||
| 65 | + context->SetOutputDataType(IDX_0, sizeDtype); | ||
| 66 | + .... | ||
| 67 | +} | ||
| 68 | + | ||
| 69 | +// Register InferDataType | ||
| 70 | +IMPL_OP(AddExample).InferDataType(InferDataTypeAddExample); | ||
| 71 | +``` | ||
| 72 | + | ||
| 73 | +## Operator Prototype Configuration | ||
| 74 | + | ||
| 75 | +Graph mode invocation requires registering the operator prototype into [Graph Engine](https://www.hiascend.com/eng/cann/graph-engine) (abbreviated as GE) so that GE can identify the input, output, and attribute information of this type of operator. Registration is completed through the `REG_OP` interface. Developers need to define basic information such as the operator input, output tensor types, and quantities. | ||
| 76 | + | ||
| 77 | +Common tensor/attribute data type examples are as follows: | ||
| 78 | + | ||
| 79 | +|Tensor Type|Attribute Type|Example| | ||
| 80 | +|-----|------|-----| | ||
| 81 | +|int64|/|DT_INT64| | ||
| 82 | +|int32|/|DT_INT32| | ||
| 83 | +|int16|/|DT_INT16| | ||
| 84 | +|int8|/|DT_INT8| | ||
| 85 | +|double|/|DT_DOUBLE| | ||
| 86 | +|float32|/|DT_FLOAT| | ||
| 87 | +|float16|/|DT_FLOAT16| | ||
| 88 | +|bfloat16|/|DT_BF16| | ||
| 89 | +|complex128|/|DT_COMPLEX128| | ||
| 90 | +|complex64|/|DT_COMPLEX64| | ||
| 91 | +|complex32|/|DT_COMPLEX32| | ||
| 92 | +|/|int|Int| | ||
| 93 | +|/|bool|Bool| | ||
| 94 | +|/|string|String| | ||
| 95 | +|/|float|Float| | ||
| 96 | +|/|list|ListInt| | ||
| 97 | + | ||
| 98 | +Basic information is as follows: | ||
| 99 | + | ||
| 100 | +|Input/Output|Keyword|Example| | ||
| 101 | +|-----|------|-----| | ||
| 102 | +|Required input|INPUT|.INPUT(${name}, TensorType({input_dtype}))| | ||
| 103 | +|Optional input|OPTIONAL_INPUT|.OPTIONAL_INPUT(${name}, TensorType({optional_input_dtype}))| | ||
| 104 | +|Required attribute|REQUIRED_ATTR|.REQUIRED_ATTR(${name}, ${dtype})| | ||
| 105 | +|Optional attribute|ATTR|.ATTR(${name}, ${dtype}, ${default_value})| | ||
| 106 | +|Output|OUTPUT|.OUTPUT(${name}, TensorType({output_dtype}))| | ||
| 107 | + | ||
| 108 | +The sample code below shows how to register the `AddExample` operator: | ||
| 109 | + | ||
| 110 | +```CPP | ||
| 111 | +REG_OP(AddExample) | ||
| 112 | + .INPUT(x1, TensorType({DT_FLOAT})) | ||
| 113 | + .INPUT(x2, TensorType({DT_FLOAT})) | ||
| 114 | + .OUTPUT(y, TensorType({DT_FLOAT})) | ||
| 115 | + .OP_END_FACTORY_REG(AddExample) | ||
| 116 | +``` | ||
| 117 | + | ||
| 118 | +For complete code, refer to [add_example_proto.h](../../../examples/add_example/op_graph/add_example_proto.h) under the `examples/add_example/op_graph` directory. | ||
| @@ -0,0 +1,70 @@ | |||
| 1 | +# build Parameter Description | ||
| 2 | + | ||
| 3 | +## Introduction | ||
| 4 | + | ||
| 5 | +build.sh is the build script of this project, located in the project root directory by default. Its function is to automatically compile, link, and configure the source code, and finally generate executable files, library files, or other target files that can be installed or run directly. Specifically, the script configures different parameters to achieve multiple functions, including building multiple target libraries (such as libophost_nn.so), compiling operator packages, executing unit tests, etc. | ||
| 6 | + | ||
| 7 | +## Usage | ||
| 8 | + | ||
| 9 | +1. **Configure Environment Variables** | ||
| 10 | + | ||
| 11 | + Complete the basic environment setup by referring to [Environment Deployment](../context/quick_install.md). | ||
| 12 | + | ||
| 13 | + ```bash | ||
| 14 | + # Default path installation, taking root user as an example | ||
| 15 | + source /usr/local/Ascend/cann/set_env.sh | ||
| 16 | + ``` | ||
| 17 | + | ||
| 18 | +2. **Build Command Format** | ||
| 19 | + | ||
| 20 | + Taking the compile operator package command as an example, the format is as follows, where `--vendor_name` and `--ops` are optional in this scenario. | ||
| 21 | + | ||
| 22 | + ```bash | ||
| 23 | + bash build.sh --pkg --soc=${soc_version} [--vendor_name=${vendor_name}] [--ops=${op_list}] | ||
| 24 | + ``` | ||
| 25 | + | ||
| 26 | + For the meaning of all parameters, refer to the parameter description section below. Choose the appropriate parameters according to the actual situation. | ||
| 27 | + | ||
| 28 | +## Parameter Description | ||
| 29 | + | ||
| 30 | +build.sh supports multiple functions. You can view all function parameters through the following command. | ||
| 31 | + | ||
| 32 | +```bash | ||
| 33 | +bash build.sh --help | ||
| 34 | +``` | ||
| 35 | + | ||
| 36 | +| Parameter Name | Optional/Required | Parameter Description | | ||
| 37 | +|------------------|--------|-----------------------------------------------------------------------------| | ||
| 38 | +| -j${n} | Optional | Specifies the number of compilation threads. ${n} is the specific number of threads. The default value is 8 (such as -j8). If the number of threads exceeds the number of CPU cores, it will be automatically adjusted to the number of CPU cores. | | ||
| 39 | +| -v | Optional | View CMake compilation configuration information. | | ||
| 40 | +| -O${n} | Optional | Specifies the compilation optimization level. Supports O0/O1/O2/O3 (such as -O3). ${n} is the optimization level identifier. | | ||
| 41 | +| -u | Optional | Enables unit test (UT) compilation mode and compiles all UT targets. | | ||
| 42 | +| --help, -h | Optional | Prints script usage help information. | | ||
| 43 | +| --ops | Optional | Specifies the operators to be compiled, such as mat_mul_v3, mse_loss. Multiple operators are separated by English commas ",". Cannot be used with --ophost and --opapi at the same time. | | ||
| 44 | +| --soc | Optional | Specifies the NPU model. Only 1 NPU model is supported per compilation. | | ||
| 45 | +| --jit | Optional | In the static graph scenario, when compiling the `cann-${soc_name}-ops-nn_${cann_version}_linux-${arch}.run` package, you do not need to compile the operator binary files (the graph runtime will compile online). You can configure this option to improve compilation speed. | | ||
| 46 | +| --static | Optional | When configured, it means generating a static library file, including libcann_nn_static.a and aclnn interface header files. Combined with the --pkg parameter, it generates a static library compressed package.| | ||
| 47 | +| --vendor_name | Optional | Specifies the name of the custom operator package. The default value is custom. | | ||
| 48 | +| --build-type | Optional | Enables debug mode. Optional types: Release/Debug. The default is Release. When the value is Debug, it cannot be used with --mssanitizer, --oom, --dump_cce at the same time | | ||
| 49 | +| --debug | Optional | Enables debug mode. | | ||
| 50 | +| --cov | Optional | Reserved parameter, developers do not need to pay attention for now. | | ||
| 51 | +| --noexec | Optional | Only compiles the unit test binary file without automatically executing the compiled UT executable file. | | ||
| 52 | +| --opkernel | Optional | Compiles the binary kernel. | | ||
| 53 | +| --pkg | Optional | Generates the installation package. Cannot be used with -u (UT mode) or --ophost, --opapi at the same time. | | ||
| 54 | +| --asan | Optional | Enables host-side ASAN (AddressSanitizer) memory detection function. | | ||
| 55 | +| --valgrind | Optional | Reserved parameter, developers do not need to pay attention for now. | | ||
| 56 | +| --make_clean | Optional | Executes basic cleanup operations (cleans compilation products). The script exits after execution. | | ||
| 57 | +| --make_clean_all | Optional | Executes complete cleanup operations (deletes all compilation-related files). The script exits after execution. | | ||
| 58 | +| --ophost | Optional | Compiles the libophost_nn.so library. Cannot be used with --pkg, --ops at the same time. | | ||
| 59 | +| --opapi | Optional | Compiles the libopapi_nn.so library. Cannot be used with --pkg, --ops at the same time. | | ||
| 60 | +| --run_example | Optional | Compiles the sample of the specified operator and mode and executes the compiled executable file. Use --run_example --help to view the usage. | | ||
| 61 | +| --genop | Optional | Creates the AI Core custom operator initial directory. | | ||
| 62 | +| --genop_aicpu | Optional | Creates the AI CPU custom operator initial directory. | | ||
| 63 | +| --experimental | Optional | Compiles user operators in the experimental directory. | | ||
| 64 | +| --mssanitizer | Optional | Enables kernel-side mssanitizer memory detection function. | | ||
| 65 | +| --oom | Optional | Enables kernel-side oom memory detection function. | | ||
| 66 | +| --dump_cce | Optional | Enables kernel-side dump precompiled file function. | | ||
| 67 | +| --cann_3rd_lib_path| Optional | The directory where third-party libraries are stored in the offline compilation scenario. | | ||
| 68 | +| --simulator | Optional | Used in combination with --run_example to enable simulator mode to execute --run_example tasks. In simulator mode, the corresponding simulator library will be linked according to soc_version. | | ||
| 69 | +| --bisheng_flags | Optional | Specifies the BiSheng compiler compilation parameters. Multiple compilation parameters are separated by English commas ",". Cannot be used with --mssanitizer, --oom, --dump_cce at the same time. | | ||
| 70 | +| --kernel_template_input | Optional | Specifies the tilingKey template when compiling the kernel. Only one template can be specified. Used with --ops and only one operator can be specified. It will not compile the binary files of other operators that this operator depends on. | | ||