已合并
docs: improve AIV direct-drive sample documentation #38
Hyunbin创建于 7月20日
docs: improve AIV direct-drive sample documentation #38
已合并
Hyunbin创建于 7月20日
共 14 个文件变更+139-97
@@ -8,7 +8,7 @@
8 8 
9- 新增项目文档入口、快速开始、构建与测试、三方依赖与兼容性说明。9- 新增项目文档入口、快速开始、构建与测试、三方依赖与兼容性说明。
10- 新增AICore Hcomm点对点通信接口API参考文档,覆盖`Hcomm`、`Init`、`ReadNbi`、`WriteNbi`、`WriteWithNotifyNbi`、`AtomicFAA`、`AtomicCAS`、`Commit`、`Drain`。10- 新增AICore Hcomm点对点通信接口API参考文档,覆盖`Hcomm`、`Init`、`ReadNbi`、`WriteNbi`、`WriteWithNotifyNbi`、`AtomicFAA`、`AtomicCAS`、`Commit`、`Drain`。
11-- 新增AICore Hcomm `hcomm_write_read_nbi`样例入口,覆盖`WriteNbi`和`ReadNbi`两卡点对点通信流程。11+- 新增AIV直驱URMA Hcomm `hcomm_write_read_nbi`样例,覆盖`WriteNbi`和`ReadNbi`两卡点对点通信流程。
12- 新增Issue模板、贡献指南和资料贡献说明。12- 新增Issue模板、贡献指南和资料贡献说明。
13 13 
14### 说明14### 说明
@@ -20,13 +20,13 @@
20- 提供AICore侧Hcomm点对点通信接口,覆盖`Init`、`ReadNbi`、`WriteNbi`、`WriteWithNotifyNbi`、`AtomicFAA`、`AtomicCAS`、`Commit`、`Drain`。20- 提供AICore侧Hcomm点对点通信接口,覆盖`Init`、`ReadNbi`、`WriteNbi`、`WriteWithNotifyNbi`、`AtomicFAA`、`AtomicCAS`、`Commit`、`Drain`。
21- 提供AIV直驱Hcomm RoCE和UBC_CTP/URMA相关实现,主实现位于`src/aicore/hcomm/detail/`。21- 提供AIV直驱Hcomm RoCE和UBC_CTP/URMA相关实现,主实现位于`src/aicore/hcomm/detail/`。
22- 提供Hcomm UT工程,覆盖`ascend950pr_9599_AIV`的RoCE/URMA路径,以及`ascend910B1_AIC`基础接口用例。22- 提供Hcomm UT工程,覆盖`ascend950pr_9599_AIV`的RoCE/URMA路径,以及`ascend910B1_AIC`基础接口用例。
23-- 提供`hcomm_write_read_nbi`样例,演示AICore Kernel侧`WriteNbi`和`ReadNbi`点对点通信流程,并包含运行样例所需的Host侧资源准备流程。23+- 提供`hcomm_write_read_nbi`样例,演示AIV直驱URMA场景下`WriteNbi`和`ReadNbi`点对点通信流程,并包含运行样例所需的Host侧资源准备流程。
24 24 
25### 📖 资料文档25### 📖 资料文档
26 26 
27- 新增[快速开始](./docs/quick_start.md)、[构建与测试](./docs/guide/build_and_test.md)、[三方依赖与兼容性](./docs/guide/dependencies.md)说明。27- 新增[快速开始](./docs/quick_start.md)、[构建与测试](./docs/guide/build_and_test.md)、[三方依赖与兼容性](./docs/guide/dependencies.md)说明。
28- 新增[Hcomm使用说明](./docs/guide/hcomm_usage.md)和[API参考](./docs/api/README.md),覆盖当前公开的Hcomm接口。28- 新增[Hcomm使用说明](./docs/guide/hcomm_usage.md)和[API参考](./docs/api/README.md),覆盖当前公开的Hcomm接口。
29-- 新增[样例目录](./examples/README.md),提供Hcomm Kernel侧调用和端到端通信样例入口。29+- 新增[样例目录](./examples/README.md),提供Hcomm AIV直驱调用和端到端通信样例入口。
30 30 
31有关所有历史版本及更新的详细信息,请参阅[CHANGELOG.md](./CHANGELOG.md)。31有关所有历史版本及更新的详细信息,请参阅[CHANGELOG.md](./CHANGELOG.md)。
32 32 
@@ -42,10 +42,10 @@ asc-comm是面向昇腾AI处理器通信场景的开源仓,当前用于承载A
42| --- | --- |42| --- | --- |
43| AICore Hcomm公开接口 | 已提供Kernel侧`Init`、`ReadNbi`、`WriteNbi`、`WriteWithNotifyNbi`、`AtomicFAA`、`AtomicCAS`、`Commit`、`Drain`。 |43| AICore Hcomm公开接口 | 已提供Kernel侧`Init`、`ReadNbi`、`WriteNbi`、`WriteWithNotifyNbi`、`AtomicFAA`、`AtomicCAS`、`Commit`、`Drain`。 |
44| AIV直驱实现 | 已提供Hcomm RoCE和UBC_CTP/URMA相关实现,主实现位于`src/aicore/hcomm/detail/`。 |44| AIV直驱实现 | 已提供Hcomm RoCE和UBC_CTP/URMA相关实现,主实现位于`src/aicore/hcomm/detail/`。 |
45-| 样例配套流程 | `hcomm_write_read_nbi`包含运行样例所需的通信域创建、通信内存注册、P2P通道创建和远端内存获取流程。 |45+| AIV直驱样例配套流程 | `hcomm_write_read_nbi`包含AIV直驱URMA通信所需的通信域创建、通信内存注册、P2P通道创建和远端内存获取流程。 |
46| 协议能力 | `COMM_PROTOCOL_ROCE`支持读写、提交和等待;`COMM_PROTOCOL_UBC_CTP`支持读写、写通知、原子操作、提交和等待。 |46| 协议能力 | `COMM_PROTOCOL_ROCE`支持读写、提交和等待;`COMM_PROTOCOL_UBC_CTP`支持读写、写通知、原子操作、提交和等待。 |
47| UT验证 | UT覆盖`ascend950pr_9599_AIV`的RoCE/URMA路径,以及`ascend910B1_AIC`基础接口用例。 |47| UT验证 | UT覆盖`ascend950pr_9599_AIV`的RoCE/URMA路径,以及`ascend910B1_AIC`基础接口用例。 |
48-| 样例 | 提供`hcomm_write_read_nbi`样例,覆盖两卡`WriteNbi`/`ReadNbi`对称通信和结果校验流程。 |48+| AIV直驱样例 | 提供`hcomm_write_read_nbi`样例,覆盖两卡AIV直驱URMA `WriteNbi`/`ReadNbi`对称通信和结果校验流程。 |
49 49 
50### 如何使用Hcomm接口50### 如何使用Hcomm接口
51 51 
@@ -80,7 +80,7 @@ Hcomm Kernel侧使用时包含如下头文件:
80├── cmake # asc-comm CMake辅助模块80├── cmake # asc-comm CMake辅助模块
81├── docs # 项目文档介绍81├── docs # 项目文档介绍
82├── examples # asc-comm API样例目录82├── examples # asc-comm API样例目录
83-│ └── hcomm_write_read_nbi # Hcomm WriteNbi/ReadNbi两卡P2P通信样例83+│ └── hcomm_write_read_nbi # Hcomm AIV直驱URMA两卡P2P通信样例
84├── include # asc-comm API声明源代码84├── include # asc-comm API声明源代码
85│ ├── aicore/hcomm # AICore侧Hcomm公开接口85│ ├── aicore/hcomm # AICore侧Hcomm公开接口
86│ ├── ain # AIN相关API预留目录86│ ├── ain # AIN相关API预留目录
@@ -122,7 +122,7 @@ cmake --build build/ut-hcomm
122 122 
123更多环境准备、Docker、CANN包安装和UT依赖说明请参考[快速开始](./docs/quick_start.md)和[构建与测试](./docs/guide/build_and_test.md)。123更多环境准备、Docker、CANN包安装和UT依赖说明请参考[快速开始](./docs/quick_start.md)和[构建与测试](./docs/guide/build_and_test.md)。
124 124 
125-## 🧰clangd/IDE 支持125+## 🧰clangd/IDE支持
126 126 
127- 安装clangd,推荐使用15或以上版本。127- 安装clangd,推荐使用15或以上版本。
128- 配置本地IDE时,需要将CANN头文件目录和本仓`include/`目录加入索引路径。128- 配置本地IDE时,需要将CANN头文件目录和本仓`include/`目录加入索引路径。
@@ -163,7 +163,7 @@ cmake --build build/ut-hcomm
163 163 
164## 📌相关规划164## 📌相关规划
165 165 
166-- 持续补充AICore Hcomm端到端样例,覆盖更多协议路径和通信接口。166+- 持续补充AIV直驱Hcomm端到端样例,覆盖更多协议路径和通信接口。
167- 持续完善不同产品、协议路径下的构建验证和UT覆盖。167- 持续完善不同产品、协议路径下的构建验证和UT覆盖。
168- 持续补充API约束、使用说明和常见问题。168- 持续补充API约束、使用说明和常见问题。
169 169 
@@ -23,4 +23,4 @@ docs/
23| [API文档贡献指南](./api_contributing.md) | 新增或修改API文档时的结构、约束和检查要求。 |23| [API文档贡献指南](./api_contributing.md) | 新增或修改API文档时的结构、约束和检查要求。 |
24| [资料贡献指南](./doc_contributing.md) | README、docs、examples等资料文档的补充规范。 |24| [资料贡献指南](./doc_contributing.md) | README、docs、examples等资料文档的补充规范。 |
25| [贡献指南](../CONTRIBUTING.md) | Issue、开发、检查和PR提交流程。 |25| [贡献指南](../CONTRIBUTING.md) | Issue、开发、检查和PR提交流程。 |
26-| [样例](../examples/README.md) | asc-comm API使用样例入口,包含Hcomm WriteNbi/ReadNbi点对点通信样例。 |26+| [样例](../examples/README.md) | asc-comm API使用样例入口,包含AIV直驱URMA WriteNbi/ReadNbi点对点通信样例。 |
@@ -19,3 +19,8 @@
19```cpp19```cpp
20#include "hcomm/hcomm.h"20#include "hcomm/hcomm.h"
21```21```
22+ 
23+## 相关文档
24+ 
25+- [Hcomm使用说明](../guide/hcomm_usage.md)
26+- [AIV直驱URMA WriteNbi/ReadNbi样例](../../examples/hcomm_write_read_nbi/README.md)
@@ -61,3 +61,7 @@ class Hcomm;
61- `WriteWithNotifyNbi`仅支持`COMM_PROTOCOL_UBC_CTP`路径,`COMM_PROTOCOL_ROCE`路径会返回失败。61- `WriteWithNotifyNbi`仅支持`COMM_PROTOCOL_UBC_CTP`路径,`COMM_PROTOCOL_ROCE`路径会返回失败。
62- `AtomicFAA`和`AtomicCAS`仅支持`COMM_PROTOCOL_UBC_CTP`路径,数据类型仅支持`int32_t`、`uint32_t`、`int64_t`、`uint64_t`。62- `AtomicFAA`和`AtomicCAS`仅支持`COMM_PROTOCOL_UBC_CTP`路径,数据类型仅支持`int32_t`、`uint32_t`、`int64_t`、`uint64_t`。
63- 传入的`ChannelHandle`需要指向与协议匹配的通道实体。63- 传入的`ChannelHandle`需要指向与协议匹配的通道实体。
64+ 
65+## 相关样例
66+ 
67+[hcomm_write_read_nbi](../../../../examples/hcomm_write_read_nbi/README.md)演示两卡场景下,AIV Kernel通过`COMM_PROTOCOL_UBC_CTP`路径调用`WriteNbi`和`ReadNbi`。该样例不覆盖RoCE路径。
@@ -45,3 +45,7 @@ __aicore__ inline int32_t ReadNbi(
45 45 
46- 调用前通信通道需已完成初始化。46- 调用前通信通道需已完成初始化。
47- `COMM_PROTOCOL_UBC_CTP`路径下,`src`需要落在通道注册的远端buffer范围内,`dst`为本端目标地址。47- `COMM_PROTOCOL_UBC_CTP`路径下,`src`需要落在通道注册的远端buffer范围内,`dst`为本端目标地址。
48+ 
49+## 相关样例
50+ 
51+参考[hcomm_write_read_nbi](../../../../examples/hcomm_write_read_nbi/README.md)中的AIV直驱URMA两卡读写流程。
@@ -45,3 +45,7 @@ __aicore__ inline int32_t WriteNbi(
45 45 
46- 调用前通信通道需已完成初始化。46- 调用前通信通道需已完成初始化。
47- `COMM_PROTOCOL_UBC_CTP`路径下,`dst`需要落在通道注册的远端buffer范围内,`src`为本端源地址。47- `COMM_PROTOCOL_UBC_CTP`路径下,`dst`需要落在通道注册的远端buffer范围内,`src`为本端源地址。
48+ 
49+## 相关样例
50+ 
51+参考[hcomm_write_read_nbi](../../../../examples/hcomm_write_read_nbi/README.md)中的AIV直驱URMA两卡读写流程。
@@ -60,11 +60,15 @@ UT会优先查找系统GTest。若系统中没有GTest,可以通过`CANN_3RD_L
60cmake -S tests/ut -B build/ut-hcomm -DCANN_3RD_LIB_PATH=<path-to-third-party>60cmake -S tests/ut -B build/ut-hcomm -DCANN_3RD_LIB_PATH=<path-to-third-party>
61```61```
62 62 
63-`build.sh -t`当前未暴露`CANN_3RD_LIB_PATH`参数;需要指定离线GTest路径时,建议直接使用上述CMake命令构建UT。63+也可以通过构建脚本传入CANN third_party目录:
64+ 
65+```bash
66+bash build.sh -t --cann_3rd_lib_path=<path-to-third-party>
67+```
64 68 
65## 样例构建与运行69## 样例构建与运行
66 70 
67-`examples/hcomm_write_read_nbi`提供Hcomm `WriteNbi`和`ReadNbi`点对点通信样例。该样例使用独立CMake工程构建:71+`examples/hcomm_write_read_nbi`提供AIV直驱URMA `WriteNbi`和`ReadNbi`点对点通信样例。该样例使用独立CMake工程构建:
68 72 
69```bash73```bash
70source /usr/local/Ascend/cann/set_env.sh74source /usr/local/Ascend/cann/set_env.sh
@@ -9,7 +9,7 @@ asc-comm当前仓内构建主要用于环境检查、AICore Hcomm接口UT验证
9| 依赖 | 要求 | 说明 |9| 依赖 | 要求 | 说明 |
10| --- | --- | --- |10| --- | --- | --- |
11| CANN Toolkit | 与当前分支或Tag配套 | 执行`build.sh`前必须先`source ${install_path}/cann/set_env.sh`,脚本会检查`ASCEND_HOME_PATH`。 |11| CANN Toolkit | 与当前分支或Tag配套 | 执行`build.sh`前必须先`source ${install_path}/cann/set_env.sh`,脚本会检查`ASCEND_HOME_PATH`。 |
12-| CANN Runtime/HCCL/Hcomm | CANN 9.1.0或以上 | `hcomm_write_read_nbi`样例需要通信域创建、内存注册和P2P通道创建能力,并在链接阶段依赖`hcomm`库。 |12+| CANN Runtime/HCCL/Hcomm | CANN 9.1.0或以上 | `hcomm_write_read_nbi`样例需要通信域创建、内存注册和AIV P2P通道创建能力,并在链接阶段依赖Host侧`hcomm`库。 |
13| CMake | >= 3.16 | UT CMake入口为`tests/ut/CMakeLists.txt`。 |13| CMake | >= 3.16 | UT CMake入口为`tests/ut/CMakeLists.txt`。 |
14| C++ 编译器 | 支持C++17 | UT目标使用`CMAKE_CXX_STANDARD 17`,建议`gcc/g++ >= 7.3.0`且版本一致。 |14| C++ 编译器 | 支持C++17 | UT目标使用`CMAKE_CXX_STANDARD 17`,建议`gcc/g++ >= 7.3.0`且版本一致。 |
15| Python | Python 3 | UT可用于生成tiling头文件;OAT钩子要求Python 3.7+。源码和examples环境建议Python >= 3.9.0。 |15| Python | Python 3 | UT可用于生成tiling头文件;OAT钩子要求Python 3.7+。源码和examples环境建议Python >= 3.9.0。 |
@@ -27,7 +27,7 @@ asc-comm当前仓内构建主要用于环境检查、AICore Hcomm接口UT验证
27 27 
28## 样例运行依赖28## 样例运行依赖
29 29 
30-`examples/hcomm_write_read_nbi`样例支持Ascend 950PR/Ascend 950DT,运行时需要至少2张NPU。单卡环境可完成编译验证,但无法完成两卡点对点通信运行验证。30+`examples/hcomm_write_read_nbi`样例支持Ascend 950PR/Ascend 950DT,运行时需要至少2张NPU。单卡环境可完成编译验证,但无法完成两卡点对点通信运行验证。该样例固定使用`COMM_ENGINE_AIV`和`COMM_PROTOCOL_UBC_CTP`,不覆盖RoCE路径。
31 31 
32样例编译命令如下:32样例编译命令如下:
33 33 
@@ -57,7 +57,11 @@ cmake -S tests/ut -B build/ut-hcomm -DCANN_3RD_LIB_PATH=<third_party>
57cmake --build build/ut-hcomm57cmake --build build/ut-hcomm
58```58```
59 59 
60-`build.sh -t`会调用UT构建,但当前脚本未暴露`CANN_3RD_LIB_PATH`参数;需要指定离线GTest路径时,建议直接使用上面的CMake命令。60+`build.sh -t`会调用UT构建。需要指定离线GTest路径时,可以直接使用上面的CMake命令,也可以通过构建脚本传入:
61+ 
62+```bash
63+bash build.sh -t --cann_3rd_lib_path=<third_party>
64+```
61 65 
62## 集成依赖边界66## 集成依赖边界
63 67 
@@ -2,7 +2,7 @@
2 2 
3## 概述3## 概述
4 4 
5-Hcomm是asc-comm当前提供的AICore侧点对点通信接口。使用方通过`AscendC::Hcomm`模板选择通信协议,并通过`ChannelHandle`指定通信通道。当前主要覆盖RoCE和UBC_CTP/URMA两条路径。5+Hcomm是asc-comm当前提供的AICore侧点对点通信接口。使用方通过`AscendC::Hcomm`模板选择通信协议,并通过`ChannelHandle`指定通信通道。当前仓库重点承载AIV直驱实现,覆盖RoCE和UBC_CTP/URMA两条路径。
6 6 
7## 基本流程7## 基本流程
8 8 
@@ -67,7 +67,7 @@ ret = hcomm.Drain(channel);
67 67 
68## 样例68## 样例
69 69 
70-可参考[asc-comm样例](../../examples/README.md)中的`hcomm_write_read_nbi`了解Kernel侧接口调用方式和Host侧通信资源创建流程。70+可参考[hcomm_write_read_nbi](../../examples/hcomm_write_read_nbi/README.md)了解AIV Kernel侧接口调用方式和Host侧通信资源创建流程。该样例固定使用`COMM_ENGINE_AIV`和`COMM_PROTOCOL_UBC_CTP`,不覆盖RoCE路径。
71 71 
72该样例在两卡场景下对称执行`WriteNbi`和`ReadNbi`:72该样例在两卡场景下对称执行`WriteNbi`和`ReadNbi`:
73 73 
@@ -29,7 +29,7 @@
29>29>
30> - 为了保障开发体验环境的质量,推荐用户基于**容器化技术**完成**环境准备**。30> - 为了保障开发体验环境的质量,推荐用户基于**容器化技术**完成**环境准备**。
31> - 如不希望使用容器,也可在带NPU设备的主机上完成**环境准备**,请参考[CANN软件安装指南 - 在物理机上安装](https://www.hiascend.com/cann/download)。31> - 如不希望使用容器,也可在带NPU设备的主机上完成**环境准备**,请参考[CANN软件安装指南 - 在物理机上安装](https://www.hiascend.com/cann/download)。
32-> - 针对仅体验"编译本开源仓 + 编译验证样例"的用户,不要求主机带NPU设备,可跳过安装NPU驱动和固件,直接安装CANN包,请参考[下载安装CANN包](#cann-install)。`hcomm_write_read_nbi`样例运行需要至少2张NPU。32+> - 针对仅体验“编译本开源仓 + 编译验证样例”的用户,不要求主机带NPU设备,可跳过安装NPU驱动和固件,直接安装CANN包,请参考[下载安装CANN包](#cann-install)。`hcomm_write_read_nbi`样例运行需要至少2张NPU。
33 33 
34### 1️⃣ 云开发环境<a name="cloud-dev-env"></a>34### 1️⃣ 云开发环境<a name="cloud-dev-env"></a>
35 35 
@@ -42,7 +42,7 @@
42 42 
43 <p align="center"><img src="./figures/cloudIDE.png" alt="云平台" width="750px" height="90px"></p>43 <p align="center"><img src="./figures/cloudIDE.png" alt="云平台" width="750px" height="90px"></p>
44 44 
45-2. 根据页面提示创建NPU环境并配置规格,启动云开发环境后,单击"`连接 > WebIDE 或 Visual Studio Code`"进入一站式开发平台。开源项目的资源默认在`/mnt/workspace`目录下。45+2. 根据页面提示创建NPU环境并配置规格,启动云开发环境后,单击"`连接 > WebIDE或Visual Studio Code`"进入一站式开发平台。开源项目的资源默认在`/mnt/workspace`目录下。
46 46 
47 <p align="center"><img src="./figures/webIDE.png" alt="云平台" width="1000px" height="150px"></p>47 <p align="center"><img src="./figures/webIDE.png" alt="云平台" width="1000px" height="150px"></p>
48 48 
@@ -84,6 +84,7 @@
84 docker run --name <cann_container> \84 docker run --name <cann_container> \
85 --ipc=host --net=host --privileged \85 --ipc=host --net=host --privileged \
86 --device /dev/davinci0 \86 --device /dev/davinci0 \
87+ --device /dev/davinci1 \
87 --device /dev/davinci_manager \88 --device /dev/davinci_manager \
88 --device /dev/devmm_svm \89 --device /dev/devmm_svm \
89 --device /dev/hisi_hdc \90 --device /dev/hisi_hdc \
@@ -102,7 +103,7 @@
102 | `--ipc=host` | 与宿主机共享IPC命名空间,NPU进程间通信(共享内存、信号量)所需 | - |103 | `--ipc=host` | 与宿主机共享IPC命名空间,NPU进程间通信(共享内存、信号量)所需 | - |
103 | `--net=host` | 使用宿主机网络栈,避免容器网络转发带来的通信延迟 | - |104 | `--net=host` | 使用宿主机网络栈,避免容器网络转发带来的通信延迟 | - |
104 | `--privileged` | 赋予容器完整设备访问权限,NPU驱动正常工作所需 | - |105 | `--privileged` | 赋予容器完整设备访问权限,NPU驱动正常工作所需 | - |
105- | `--device /dev/davinci0` | 将宿主机的NPU设备卡映射到容器内,可指定映射多张NPU设备卡 | 必须根据实际情况调整:`davinci0`对应系统中的第0张NPU卡。请先在宿主机执行`npu-smi info`命令,根据输出显示的设备号(如`NPU 0`, `NPU 1`)来修改此编号 |106+ | `--device /dev/davinci<N>` | 将指定的NPU设备映射到容器内;如需使用多张设备,可多次使用`--device`参数 | 根据`npu-smi info`显示的设备号调整。运行`hcomm_write_read_nbi`样例时至少映射两张Ascend 950PR/Ascend 950DT设备。 |
106 | `--device /dev/davinci_manager` | 映射NPU设备管理接口 | - |107 | `--device /dev/davinci_manager` | 映射NPU设备管理接口 | - |
107 | `--device /dev/devmm_svm` | 映射设备内存管理接口 | - |108 | `--device /dev/devmm_svm` | 映射设备内存管理接口 | - |
108 | `--device /dev/hisi_hdc` | 映射主机与设备间的通信接口 | - |109 | `--device /dev/hisi_hdc` | 映射主机与设备间的通信接口 | - |
@@ -147,7 +148,7 @@ CANN包分为CANN toolkit包和CANN ops包。
147 ```148 ```
148 149 
149 > [!IMPORTANT] 安装说明150 > [!IMPORTANT] 安装说明
150- > [examples](../examples)中部分算子样例的编译运行依赖本包,若想完整体验样例编译运行流程,建议安装此包。151+ > 当前[Hcomm AIV直驱URMA样例](../examples/hcomm_write_read_nbi/README.md)不依赖ops包;仅在后续使用依赖算子包的功能时按需安装。
151 152 
152| 参数 | 说明 |153| 参数 | 说明 |
153| :--- | :--- |154| :--- | :--- |
@@ -234,6 +235,12 @@ UT依赖googletest。若系统中没有GTest,可以通过`CANN_3RD_LIB_PATH`
234bash build.sh -t235bash build.sh -t
235```236```
236 237 
238+需要指定CANN third_party目录时执行:
239+ 
240+```bash
241+bash build.sh -t --cann_3rd_lib_path=<path-to-third-party>
242+```
243+ 
237方式二:用户也可直接使用CMake命令指定离线GTest路径。244方式二:用户也可直接使用CMake命令指定离线GTest路径。
238 245 
239```bash246```bash
@@ -251,7 +258,7 @@ cmake --build build/ut-hcomm
251 258 
252### 🧩 样例验证<a name="sample-verify"></a>259### 🧩 样例验证<a name="sample-verify"></a>
253 260 
254-`examples/hcomm_write_read_nbi`提供Hcomm `WriteNbi`和`ReadNbi`点对点通信样例。样例支持Ascend 950PR/Ascend 950DT,要求CANN 9.1.0或以上版本。运行样例需要至少2张NPU;单卡环境仅支持编译验证。261+[hcomm_write_read_nbi](../examples/hcomm_write_read_nbi/README.md)提供AIV直驱URMA `WriteNbi`和`ReadNbi`点对点通信样例。样例支持Ascend 950PR/Ascend 950DT,要求CANN 9.1.0或以上版本。运行样例需要至少2张NPU;单卡环境仅支持编译验证。
255 262 
256进入样例目录后执行:263进入样例目录后执行:
257 264 
@@ -6,11 +6,11 @@
6 6 
7| 样例 | 说明 | 支持产品 |7| 样例 | 说明 | 支持产品 |
8| --- | --- | --- |8| --- | --- | --- |
9-| [hcomm_write_read_nbi](./hcomm_write_read_nbi/README.md) | 演示两卡场景下使用`Hcomm::WriteNbi`和`Hcomm::ReadNbi`完成点对点通信,并校验通信结果。 | Ascend 950PR/Ascend 950DT |9+| [hcomm_write_read_nbi](./hcomm_write_read_nbi/README.md) | 演示两卡场景下AIV Kernel通过URMA路径调用`Hcomm::WriteNbi`和`Hcomm::ReadNbi`,并校验通信结果。 | Ascend 950PR/Ascend 950DT |
10 10 
11## hcomm_write_read_nbi11## hcomm_write_read_nbi
12 12 
13-`hcomm_write_read_nbi`展示完整的Hcomm点对点通信流程,包括Host侧通信域创建、通信内存注册、P2P通道创建、Kernel侧`Init`、`WriteNbi`、`ReadNbi`和`Drain`调用。13+`hcomm_write_read_nbi`展示AIV直驱URMA点对点通信流程,包括Host侧通信域创建、通信内存注册、AIV P2P通道创建,以及Kernel侧`Init`、`WriteNbi`、`ReadNbi`和`Drain`调用。该样例不覆盖RoCE路径。
14 14 
15样例采用两卡对称执行方式:15样例采用两卡对称执行方式:
16 16 
@@ -47,4 +47,4 @@ make -j
47 47 
48- 样例支持Ascend 950PR/Ascend 950DT,CANN软件版本要求为9.1.0或以上。48- 样例支持Ascend 950PR/Ascend 950DT,CANN软件版本要求为9.1.0或以上。
49- 样例运行需要至少2张NPU;单卡环境仅支持编译验证。49- 样例运行需要至少2张NPU;单卡环境仅支持编译验证。
50-- 样例编译依赖CANN ASC CMake能力,并在链接阶段依赖`hcomm`库。50+- 样例编译依赖CANN ASC CMake能力,并在链接阶段依赖CANN `hcomm`库。
@@ -1,51 +1,54 @@
1-# Hcomm WriteNbi/ReadNbi 点对点通信样例 (AIV直驱URMA)1+# Hcomm AIV直驱URMA WriteNbi/ReadNbi点对点通信样例
2 2 
3## 概述3## 概述
4 4 
5-本样例展示如何在 Ascend C Kernel 中,基于 **AIV直驱URMA** 架构,使用 `Hcomm` 类的 **`WriteNbi`** 和 **`ReadNbi`** API 实现 NPU 间极低延迟的点对点(P2P)通信。样例中两张卡对称执行 `WriteNbi` + `ReadNbi`,通过地址偏移区分数据段,相互写入并读出数据,最后在 Host 侧校验结果一致性。5+本样例展示如何在Ascend C AIV Kernel中,基于**AIV直驱URMA**架构,使用`Hcomm`类的`WriteNbi`和`ReadNbi`接口实现NPU间低时延的点对点(P2P)通信。两张卡对称执行`WriteNbi`和`ReadNbi`,通过地址偏移区分数据段,最后在Host侧校验通信结果。
6 6 
7-## 支持的产品及 CANN 软件版本7+## 支持的产品及CANN软件版本
8 8 
9-| 产品 | CANN 软件版本 |9+| 产品 | CANN软件版本 |
10|------|-------------|10|------|-------------|
11| Ascend 950PR / Ascend 950DT | >= CANN 9.1.0 |11| Ascend 950PR / Ascend 950DT | >= CANN 9.1.0 |
12 12 
13## 目录结构介绍13## 目录结构介绍
14 14 
15```text15```text
16-├── hcomm_write_read_nbi16+├── hcomm_write_read_nbi
17-│ ├── CMakeLists.txt // 编译工程文件17+│ ├── CMakeLists.txt // 编译工程文件
18-│ ├── README.md // 样例说明文档18+│ ├── README.md // 样例说明文档
19-│ ├── hcomm_write_read_nbi.asc // Ascend C 样例实现(Kernel + Host)19+│ ├── README_en.md // 英文样例说明文档
20-│ ├── utils.cpp // 工具函数实现20+│ ├── hcomm_write_read_nbi.asc // Host侧资源准备与Kernel调用
21-│ └── utils.h // 工具函数声明21+│ ├── hcomm_write_read_nbi_kernel.cpp // AIV Kernel侧Hcomm调用
22+│ ├── hcomm_rw_def.h // Host与Kernel共享定义
23+│ ├── utils.cpp // 工具函数实现
24+│ └── utils.h // 工具函数声明
22```25```
23 26 
24## 样例描述27## 样例描述
25 28 
26### 样例功能29### 样例功能
27-本样例重点演示 **AIV 直驱 URMA** 场景下的 P2P 通信接口。与传统的集合通信(如 `AlltoAll`)不同,他们都是由 Host 发起通信,**但 AIV 直驱 URMA 的每次数据面的通信不再需要 Host 参与**。它们允许 NPU 硬件直接向目标 NPU 发起 URMA 操作,绕过传统软件协议栈,实现显存到显存的直接访问,具有极低的通信延迟。此特性非常适用于 MoE (Dispatch/Combine)、Pipeline 并行等对延迟敏感的分布式训练场景。30+本样例重点演示**AIV直驱URMA**场景下的P2P通信接口。Host侧只负责创建通信域、注册通信内存和建立通道;通信任务由AIV Kernel直接提交,数据面执行期间不需要Host逐次参与。该模式适用于MoE Dispatch/Combine、Pipeline并行等对通信时延敏感的场景。
28 31 
29| API | 数据流向 | 语义说明 (AIV直驱URMA) |32| API | 数据流向 | 语义说明 (AIV直驱URMA) |
30|------|---------|------|33|------|---------|------|
31-| `WriteNbi` | 本地 GM → 远端 GM | AIV 直驱 URMA 写接口:将本地 Global Memory 数据直接写入远端 NPU 指定地址,无需远端 CPU 介入。 |34+| `WriteNbi` | 本地GM → 远端GM | AIV直驱URMA写接口:将本地Global Memory数据直接写入远端NPU指定地址,无需远端CPU介入。 |
32-| `ReadNbi` | 远端 GM → 本地 GM | AIV 直驱 URMA 读接口:直接从远端 NPU 指定地址读取数据到本地 Global Memory,无需远端 CPU 介入。 |35+| `ReadNbi` | 远端GM → 本地GM | AIV直驱URMA读接口:直接从远端NPU指定地址读取数据到本地Global Memory,无需远端CPU介入。 |
33 36 
34### 样例实现37### 样例实现
35 38 
36-#### 1. Host 侧通信域准备39+#### 1. Host侧通信域准备
37-在 Ascend 950 系列上,通信域需以多进程方式创建(每个进程对应一个 rank)。关键步骤如下,重点体现 AIV 直驱模式的配置:40+在Ascend 950系列上,通信域需以多进程方式创建(每个进程对应一个rank)。关键步骤如下,重点体现AIV直驱模式的配置:
38 41 
39-1. **交换 RootInfo**:Rank 0 调用 `HcclGetRootInfo` 获取 root 信息,通过 TCP 发送给 Rank 1。42+1. **交换RootInfo**:Rank 0调用`HcclGetRootInfo`获取root信息,通过TCP发送给Rank 1。
40-2. **创建通信域**:各 Rank 调用 `HcclCommInitRootInfoConfig` 创建通信域。**注意:AIV 直驱模式下,无需配置 `hcclOpExpansionMode`。**43+2. **创建通信域**:各Rank调用`HcclCommInitRootInfoConfig`创建通信域。**注意:AIV直驱模式下,无需配置`hcclOpExpansionMode`。**
41-3. **注册通信内存**:调用 `HcclCommMemReg` 向通信域注册本卡的通信 buffer,Channel 创建时该内存信息会自动交换给对端。44+3. **注册通信内存**:调用`HcclCommMemReg`向通信域注册本卡的通信buffer,Channel创建时该内存信息会自动交换给对端。
42-4. **获取链路 Endpoint**:通过 `HcclRankGraphGetLayers` 和 `HcclRankGraphGetLinks` 获取本 Rank 到对端的物理链路 Endpoint 信息。45+4. **获取链路Endpoint**:通过`HcclRankGraphGetLayers`和`HcclRankGraphGetLinks`获取本Rank到对端的物理链路Endpoint信息。
43-5. **创建 P2P 通道 (AIV直驱)**:调用 `HcclChannelAcquire` 创建到对端的 P2P 通道。此处需明确指定引擎为 `COMM_ENGINE_AIV`,协议为 `COMM_PROTOCOL_UBC_CTP`(即 URMA 协议),并传入通知数量和待交换的内存句柄。46+5. **创建P2P通道(AIV直驱)**:调用`HcclChannelAcquire`创建到对端的P2P通道。此处需明确指定引擎为`COMM_ENGINE_AIV`、协议为`COMM_PROTOCOL_UBC_CTP`(URMA协议),并传入待交换的内存句柄。
44-6. **获取对端内存地址**:调用 `HcclChannelGetRemoteMems` 获取对端注册的内存地址,作为 Kernel 侧 `WriteNbi`/`ReadNbi` 的远端目标地址。47+6. **获取对端内存地址**:调用`HcclChannelGetRemoteMems`获取对端注册的内存地址,作为Kernel侧`WriteNbi`/`ReadNbi`的远端目标地址。
45-7. **下发 Context**:Host 预初始化 seg0(填充基于 rankId 的伪随机 pattern),将 `ChannelHandle` 和 buffer 地址封装到 `CommContext` 并下发到各卡 GM。48+7. **下发Context**:Host预初始化seg0(填充基于rankId的伪随机pattern),将`ChannelHandle`和buffer地址封装到`CommContext`并下发到各卡GM。
46 49 
47```cpp50```cpp
48-// 1. 各 rank 各自创建通信域 (AIV直驱模式无需配置 hcclOpExpansionMode)51+// 1. 各rank各自创建通信域 (AIV直驱模式无需配置hcclOpExpansionMode)
49HcclCommConfig config;52HcclCommConfig config;
50HcclCommConfigInit(&config);53HcclCommConfigInit(&config);
51config.hcclWorldRankID = rank;54config.hcclWorldRankID = rank;
@@ -54,63 +57,64 @@ HcclCommInitRootInfoConfig(nranks, &rootInfo, rank, &config, &comm);
54// 2. 注册通信内存57// 2. 注册通信内存
55HcclCommMemReg(comm, "shareBuf", &regMem, &memHandle);58HcclCommMemReg(comm, "shareBuf", &regMem, &memHandle);
56 59 
57-// 3. 获取链路 endpoint60+// 3. 获取链路endpoint
58HcclRankGraphGetLinks(comm, layerId, rank, peerRank, &links, &linkNum);61HcclRankGraphGetLinks(comm, layerId, rank, peerRank, &links, &linkNum);
59 62 
60-// 4. 创建 AIV 直驱 URMA P2P 通道63+// 4. 创建AIV直驱URMA P2P通道
61channelDesc.remoteRank = peerRank;64channelDesc.remoteRank = peerRank;
62-channelDesc.channelProtocol = COMM_PROTOCOL_UBC_CTP; // URMA 协议65+channelDesc.channelProtocol = COMM_PROTOCOL_UBC_CTP; // URMA协议
63channelDesc.memHandles = &memHandle;66channelDesc.memHandles = &memHandle;
64channelDesc.memHandleNum = 1;67channelDesc.memHandleNum = 1;
65-HcclChannelAcquire(comm, COMM_ENGINE_AIV, &channelDesc, 1, &channel); // 指定 COMM_ENGINE_AIV68+HcclChannelAcquire(comm, COMM_ENGINE_AIV, &channelDesc, 1, &channel); // 指定COMM_ENGINE_AIV
66 69 
67// 5. 获取对端内存地址70// 5. 获取对端内存地址
68HcclChannelGetRemoteMems(comm, channel, &memNum, &remoteMems, &memTags);71HcclChannelGetRemoteMems(comm, channel, &memNum, &remoteMems, &memTags);
69```72```
70 73 
71-#### 2. Kernel 侧执行74+#### 2. Kernel侧执行
72-通信流程分为三步:`Init()` → `WriteNbi()`/`ReadNbi()` → `Drain()`。两卡执行相同的 Kernel 逻辑,通过地址偏移区分三段数据:75+通信流程分为三步:`Init()` → `WriteNbi()`/`ReadNbi()` → `Drain()`。两卡执行相同的Kernel逻辑,通过地址偏移区分三段数据:
73-- **seg0** `[0, DATA_SIZE)`:本卡 pattern(Host 预初始化)76+- **seg0** `[0, DATA_SIZE)`:本卡pattern(Host预初始化)
74-- **seg1** `[DATA_SIZE, 2*DATA_SIZE)`:接收对端 `WriteNbi` 写入的数据77+- **seg1** `[DATA_SIZE, 2*DATA_SIZE)`:接收对端`WriteNbi`写入的数据
75-- **seg2** `[2*DATA_SIZE, 3*DATA_SIZE)`:接收本卡 `ReadNbi` 从对端 seg0 读回的数据78+- **seg2** `[2*DATA_SIZE, 3*DATA_SIZE)`:接收本卡`ReadNbi`从对端seg0读回的数据
76 79 
77-- **`Init`**:分配 UB 工作空间(>= 512 字节),用于存放 Hcomm 内部 WQE/CQE 等状态。80+- **`Init`**:分配UB工作空间(>= 512字节),用于存放Hcomm内部WQE/CQE等状态。
78-- **`WriteNbi` / `ReadNbi`**:调用 **AIV直驱URMA API** 将通信任务入队到 SQ。默认 `commit=true`,入队后立即触发 doorbell;如需批量优化,可改用 `commit=false` 多次入队后统一调用 `Commit()`。81+- **`WriteNbi` / `ReadNbi`**:调用**AIV直驱URMA API**将通信任务入队到SQ。默认`commit=true`,入队后立即触发doorbell;如需批量优化,可改用`commit=false`多次入队后统一调用`Commit()`。
79-- **`Drain`**:轮询 CQ 等待通信任务完成,返回后数据可见性才有保证。82+- **`Drain`**:轮询CQ等待通信任务完成,返回后数据可见性才有保证。
80 83 
81```cpp84```cpp
82-// 两卡对称执行 AIV 直驱 URMA 通信:85+// 两卡对称执行AIV直驱URMA通信:
83-// 1. 本卡 seg0 → 对端 seg1 (WriteNbi)86+// 1. 本卡seg0 → 对端seg1 (WriteNbi)
84hcomm_.WriteNbi(channel, remoteBuf + DATA_SIZE, localBuf, DATA_SIZE);87hcomm_.WriteNbi(channel, remoteBuf + DATA_SIZE, localBuf, DATA_SIZE);
85 88 
86-// 2. 对端 seg0 → 本卡 seg2 (ReadNbi)89+// 2. 对端seg0 → 本卡seg2 (ReadNbi)
87hcomm_.ReadNbi(channel, localBuf + 2 * DATA_SIZE, remoteBuf, DATA_SIZE);90hcomm_.ReadNbi(channel, localBuf + 2 * DATA_SIZE, remoteBuf, DATA_SIZE);
88 91 
89-// 3. 等待 AIV 硬件通信任务完成92+// 3. 等待AIV硬件通信任务完成
90hcomm_.Drain(channel);93hcomm_.Drain(channel);
91```94```
92 95 
93#### 3. 调用实现96#### 3. 调用实现
94-多进程对称执行:两卡同时 launch kernel,无需分阶段同步。Host 侧预初始化 seg0 后通过 `TcpBarrier` 确保两卡就绪,再各自调用 Kernel。使用内核调用符 `<<<>>>` 调用核函数。97+多进程对称执行:两卡同时launch kernel,无需分阶段同步。Host侧预初始化seg0后通过`TcpBarrier`确保两卡就绪,再各自调用Kernel。使用内核调用符`<<<>>>`调用核函数。
95 98 
96### 校验机制99### 校验机制
97-- Host 预初始化 seg0 时,通过两次线性同余伪随机数生成器(LCG,使用 Knuth 乘法哈希常数 `0x9E3779B9U` 等参数)生成用于通信校验的随机 pattern,确保数据来源可区分。100+- 每个rank在Host侧预初始化seg0时,通过线性同余生成器(LCG,使用Knuth乘法哈希常数`0x9E3779B9U`等参数)生成用于通信校验的随机pattern,确保不同rank的数据来源可区分。
Li-Jian
Li-JianLi-Jian7月20日

这些seg0/1/2 能否用显示的名字去定义。 比较难理解

likedislike
Hyunbin
Hyunbin
7月23日 评论:
98-- Kernel 执行完毕后,Host 侧通过 `aclrtMemcpy` 回读 `CommContext::testResult`。101+- Kernel执行完毕后,Host侧通过`aclrtMemcpy`回读`CommContext::testResult`。
99-- Host 侧进一步校验 seg1(对端 `WriteNbi` 写入)和 seg2(本卡 `ReadNbi` 读回)的数据是否与对端 pattern 完全一致。当校验结果码 testResult 为 0(即 seg1/seg2 数据与对端 pattern 完全一致)时,打印 test pass!。102+- Host侧进一步校验seg1(对端`WriteNbi`写入)和seg2(本卡`ReadNbi`读回)的数据是否与对端pattern完全一致。当校验结果码testResult为0(即seg1/seg2数据与对端pattern完全一致)时,打印test pass!。
100 103 
101## 编译与运行104## 编译与运行
102 105 
103在本样例根目录下执行如下步骤,编译并执行样例。106在本样例根目录下执行如下步骤,编译并执行样例。
104 107 
105### 1. 配置环境变量108### 1. 配置环境变量
106-请根据当前环境上 CANN 开发套件包的安装方式,配置环境变量:109+请根据当前环境上CANN开发套件包的安装方式,配置环境变量:
107```bash110```bash
108source ${install_path}/cann/set_env.sh111source ${install_path}/cann/set_env.sh
109```112```
110-> **说明:** `${install_path}` 为 CANN 包安装目录,未指定安装目录时默认安装至 `/usr/local/Ascend` 下。113+> **说明:** `${install_path}`为CANN包安装目录,未指定安装目录时默认安装至`/usr/local/Ascend`下。
111 114 
112-### 2. 编译工程115+### 2. 编译工程
113-在本样例目录下执行如下命令:116+ 
117+在本样例目录下执行如下命令:
114```bash118```bash
115mkdir -p build && cd build119mkdir -p build && cd build
116cmake -DCMAKE_ASC_ARCHITECTURES=dav-3510 ..120cmake -DCMAKE_ASC_ARCHITECTURES=dav-3510 ..
@@ -120,32 +124,32 @@ make -j
120### 3. 样例执行124### 3. 样例执行
121本样例支持两种执行方式:125本样例支持两种执行方式:
122 126 
123-**方式一:自动 Fork 多进程(推荐)**127+**方式一:自动Fork多进程(推荐)**
124-直接执行即可,主进程会自动 fork 两个子进程(两卡对称执行 WriteNbi + ReadNbi):128+直接执行即可,主进程会自动fork两个子进程(两卡对称执行WriteNbi + ReadNbi):
125```bash129```bash
126-./demo130+./demo
127```131```
128 132 
129**方式二:手动指定参数单进程运行**133**方式二:手动指定参数单进程运行**
130-可在两个不同的终端中分别手动启动 Rank 0 和 Rank 1:134+可在两个不同的终端中分别手动启动Rank 0和Rank 1:
131```bash135```bash
132-# 终端 1:启动 rank 0(绑定卡 0)136+# 终端1:启动rank 0(绑定卡0)
133-./demo 0 2 tcp://127.0.0.1:29621137+./demo 0 2 tcp://127.0.0.1:29621
134 138 
135-# 终端 2:启动 rank 1(绑定卡 1)139+# 终端2:启动rank 1(绑定卡1)
136-./demo 1 2 tcp://127.0.0.1:29621140+./demo 1 2 tcp://127.0.0.1:29621
137```141```
138 142 
139### 4. 编译选项说明143### 4. 编译选项说明
140| 选项 | 可选值 | 说明 |144| 选项 | 可选值 | 说明 |
141|------|--------|------|145|------|--------|------|
142-| `CMAKE_ASC_ARCHITECTURES` | `dav-3510`(默认) | NPU 架构:`dav-3510` 对应 Ascend 950PR / Ascend 950DT |146+| `CMAKE_ASC_ARCHITECTURES` | `dav-3510`(默认) | NPU架构:`dav-3510`对应Ascend 950PR / Ascend 950DT |
143 147 
144### 5. 执行结果148### 5. 执行结果
145-执行成功后,终端将输出如下信息,说明 AIV 直驱 URMA 通信成功(两卡对称写入读出、结果一致):149+执行成功后,终端将输出如下信息,说明AIV直驱URMA通信成功(两卡对称写入读出、结果一致):
146```text150```text
147rank 0 test pass!151rank 0 test pass!
148rank 1 test pass!152rank 1 test pass!
149test pass!153test pass!
150```154```
151-> **注意:** 单卡环境仅支持编译验证,实际运行本样例需至少配备 2 张 NPU。155+> **注意:** 单卡环境仅支持编译验证,实际运行本样例需至少配备2张NPU。
@@ -1,8 +1,8 @@
1-# Hcomm WriteNbi/ReadNbi Point-to-Point Communication Sample (AIV Direct-Drive URMA)1+# Hcomm AIV Direct-Drive URMA WriteNbi/ReadNbi Point-to-Point Communication Sample
2 2 
3## Overview3## Overview
4 4 
5-This sample demonstrates how to implement ultra-low latency point-to-point (P2P) communication between NPUs in an Ascend C Kernel using the **`WriteNbi`** and **`ReadNbi`** APIs of the `Hcomm` class, based on the **AIV Direct-Drive URMA** architecture. In this sample, two cards symmetrically execute `WriteNbi` + `ReadNbi`, differentiate data segments via address offsets, mutually write and read data, and finally validate the consistency of the results on the Host side.5+This sample demonstrates low-latency point-to-point (P2P) communication between NPUs from an Ascend C AIV Kernel. It uses the `WriteNbi` and `ReadNbi` APIs of `Hcomm` over the AIV direct-drive URMA path. Two devices execute the same write-and-read sequence, use address offsets to separate data segments, and validate the communication results on the Host.
6 6 
7## Supported Products and CANN Software Versions7## Supported Products and CANN Software Versions
8 8 
@@ -13,16 +13,21 @@ This sample demonstrates how to implement ultra-low latency point-to-point (P2P)
13## Directory Structure13## Directory Structure
14 14 
15```text15```text
16-├── hcomm_write_read_nbi16+├── hcomm_write_read_nbi
17-│ ├── CMakeLists.txt // CMake build file17+│ ├── CMakeLists.txt // CMake build file
18-│ ├── README.md // Sample documentation18+│ ├── README.md // Sample documentation
19-│ └── hcomm_write_read_nbi.asc // Ascend C sample implementation (Kernel + Host)19+│ ├── README_en.md // English sample documentation
20+│ ├── hcomm_write_read_nbi.asc // Host resource setup and Kernel invocation
21+│ ├── hcomm_write_read_nbi_kernel.cpp // Hcomm calls from the AIV Kernel
22+│ ├── hcomm_rw_def.h // Definitions shared by Host and Kernel
23+│ ├── utils.cpp // TCP helper implementation
24+│ └── utils.h // TCP helper declarations
20```25```
21 26 
22## Sample Description27## Sample Description
23 28 
24### Sample Functionality29### Sample Functionality
25-This sample focuses on demonstrating P2P communication interfaces under the **AIV Direct-Drive URMA** architecture. Unlike collective communication primitives (e.g., AlltoAll) that are initiated by the Host, **every data-plane communication in AIV Direct-Driven URMA no longer requires Host involvement**. They allow the NPU hardware to directly initiate URMA operations to the target NPU, bypassing traditional software protocol stacks and achieving direct Global Memory (GM) to GM access with ultra-low latency. This feature is highly suitable for latency-sensitive distributed training scenarios such as MoE (Dispatch/Combine) and Pipeline Parallelism.30+This sample focuses on P2P communication over the **AIV direct-drive URMA** path. The Host creates the communication domain, registers communication memory, and acquires the channel. The AIV Kernel then submits the communication operations directly, without requiring the Host to participate in each data-plane transfer. This mode is suitable for latency-sensitive workloads such as MoE Dispatch/Combine and pipeline parallelism.
26 31 
27| API | Data Flow | Semantic Description (AIV Direct-Drive URMA) |32| API | Data Flow | Semantic Description (AIV Direct-Drive URMA) |
28|-----|-----------|---------------------------------------------|33|-----|-----------|---------------------------------------------|
@@ -38,7 +43,7 @@ On the Ascend 950 series, the communication domain must be created in a multi-pr
382. **Create Communication Domain**: Each rank calls `HcclCommInitRootInfoConfig` to create the communication domain. **Note: In AIV direct-drive mode, there is no need to configure `hcclOpExpansionMode`.**432. **Create Communication Domain**: Each rank calls `HcclCommInitRootInfoConfig` to create the communication domain. **Note: In AIV direct-drive mode, there is no need to configure `hcclOpExpansionMode`.**
393. **Register Communication Memory**: Call `HcclCommMemReg` to register the local communication buffer with the communication domain. This memory information is automatically exchanged with the peer during channel creation.443. **Register Communication Memory**: Call `HcclCommMemReg` to register the local communication buffer with the communication domain. This memory information is automatically exchanged with the peer during channel creation.
404. **Obtain Link Endpoints**: Use `HcclRankGraphGetLayers` and `HcclRankGraphGetLinks` to obtain the physical link endpoint information from the local rank to the peer rank.454. **Obtain Link Endpoints**: Use `HcclRankGraphGetLayers` and `HcclRankGraphGetLinks` to obtain the physical link endpoint information from the local rank to the peer rank.
41-5. **Acquire P2P Channel (AIV Direct-Drive)**: Call `HcclChannelAcquire` to create the P2P channel to the peer. You must explicitly specify the engine as `COMM_ENGINE_AIV` and the protocol as `COMM_PROTOCOL_UBC_CTP` (i.e., the URMA protocol), along with the notification count and the memory handles to be exchanged.46+5. **Acquire P2P Channel (AIV Direct-Drive)**: Call `HcclChannelAcquire` to create the P2P channel to the peer. Specify `COMM_ENGINE_AIV` as the engine and `COMM_PROTOCOL_UBC_CTP` as the URMA protocol, and pass the memory handles to be exchanged.
426. **Obtain Remote Memory Address**: Call `HcclChannelGetRemoteMems` to retrieve the memory address registered by the peer, which serves as the remote target address for `WriteNbi`/`ReadNbi` in the Kernel.476. **Obtain Remote Memory Address**: Call `HcclChannelGetRemoteMems` to retrieve the memory address registered by the peer, which serves as the remote target address for `WriteNbi`/`ReadNbi` in the Kernel.
437. **Download Context**: The Host pre-initializes seg0 (filling it with a pseudo-random pattern based on `rankId`), encapsulates the `ChannelHandle` and buffer addresses into `CommContext`, and downloads it to the GM of each card.487. **Download Context**: The Host pre-initializes seg0 (filling it with a pseudo-random pattern based on `rankId`), encapsulates the `ChannelHandle` and buffer addresses into `CommContext`, and downloads it to the GM of each card.
44 49 
@@ -94,7 +99,7 @@ Multi-process symmetric execution: Both cards launch the kernel simultaneously w
94### Validation Mechanism99### Validation Mechanism
95- During Host pre-initialization of seg0, a pseudo-random pattern is generated using a Linear Congruential Generator (LCG, utilizing Knuth's multiplicative hash constant `0x9E3779B9U` and other parameters) to ensure the data source is distinguishable.100- During Host pre-initialization of seg0, a pseudo-random pattern is generated using a Linear Congruential Generator (LCG, utilizing Knuth's multiplicative hash constant `0x9E3779B9U` and other parameters) to ensure the data source is distinguishable.
96- After Kernel execution, the Host reads back `CommContext::testResult` via `aclrtMemcpy`.101- After Kernel execution, the Host reads back `CommContext::testResult` via `aclrtMemcpy`.
97-- The Host further validates whether the data in seg1 (written by peer's `WriteNbi`) and seg2 (read by local `ReadNbi`) perfectly matches the peer's pattern. when the result code testResult == 0 (i.e. seg1/seg2 match the peer pattern), it prints test pass!.102+- The Host then checks that seg1 (written by the peer's `WriteNbi`) and seg2 (read by the local `ReadNbi`) match the peer's pattern. When `testResult` is 0, it prints `test pass!`.
98 103 
99## Compilation and Execution104## Compilation and Execution
100 105 
@@ -107,8 +112,9 @@ source ${install_path}/cann/set_env.sh
107```112```
108> **Note:** `${install_path}` is the CANN package installation directory. If not specified, it defaults to `/usr/local/Ascend`.113> **Note:** `${install_path}` is the CANN package installation directory. If not specified, it defaults to `/usr/local/Ascend`.
109 114 
110-### 2. Build the Project115+### 2. Build the Project
111-Execute the following commands in the sample directory:116+ 
117+Execute the following commands in the sample directory:
112```bash118```bash
113mkdir -p build && cd build119mkdir -p build && cd build
114cmake -DCMAKE_ASC_ARCHITECTURES=dav-3510 ..120cmake -DCMAKE_ASC_ARCHITECTURES=dav-3510 ..
@@ -121,17 +127,17 @@ This sample supports two execution methods:
121**Method 1: Auto-Fork Multi-Process (Recommended)**127**Method 1: Auto-Fork Multi-Process (Recommended)**
122Execute directly. The main process will automatically fork two child processes (symmetrically executing WriteNbi + ReadNbi on two cards):128Execute directly. The main process will automatically fork two child processes (symmetrically executing WriteNbi + ReadNbi on two cards):
123```bash129```bash
124-./demo130+./demo
125```131```
126 132 
127**Method 2: Manual Single-Process Execution with Arguments**133**Method 2: Manual Single-Process Execution with Arguments**
128You can manually start Rank 0 and Rank 1 in two separate terminals:134You can manually start Rank 0 and Rank 1 in two separate terminals:
129```bash135```bash
130# Terminal 1: Start rank 0 (bound to device 0)136# Terminal 1: Start rank 0 (bound to device 0)
131-./demo 0 2 tcp://127.0.0.1:29621137+./demo 0 2 tcp://127.0.0.1:29621
132 138 
133# Terminal 2: Start rank 1 (bound to device 1)139# Terminal 2: Start rank 1 (bound to device 1)
134-./demo 1 2 tcp://127.0.0.1:29621140+./demo 1 2 tcp://127.0.0.1:29621
135```141```
136 142 
137### 4. Build Options Description143### 4. Build Options Description
@@ -146,4 +152,4 @@ rank 0 test pass!
146rank 1 test pass!152rank 1 test pass!
147test pass!153test pass!
148```154```
149-> **Note:** A single-card environment only supports compilation verification. Running this sample requires at least 2 NPUs.155+> **Note:** A single-card environment only supports compilation verification. Running this sample requires at least 2 NPUs.