已关闭
[Bug-Report|缺陷反馈]:simt_rma_ub2gm算子barrier_on_stream_kernel中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争 #565
h-glue创建于  20 天前关闭于  13 天前
h-glue
20 天前 创建

Thanks for sending an issue! Please fill in the following template to help quickly solve your problem.

Describe the current behavior / 问题描述 (Mandatory / 必填)

问题描述

使用 mssanitizer 检测 simt_rma_ub2gm 示例时,barrier_on_stream_kernel(集体化 barrier 设备核)报出 GM 地址上的 WAW hazard,指向团队 barrier 内部实现的多核并发写。

[mssanitizer] Start memcheck, racecheck, initcheck and synccheck sanitizer on kernel "barrier_on_stream_kernel(int, unsigned long)"
====== ERROR: Potential WAW hazard detected at GM in "barrier_on_stream_kernel(int, unsigned long)":
====== PIPE_S Write at WAW()+0x124000500200 in block 0 (aiv) on device 0 at pc current 0x5328 (serialNo:66)
====== #0 .../src/device/shmemi_device_common.hpp:69:25 // aclshmemi_store: *addr = val
====== #1 .../src/device/gm2gm/shmemi_device_cc.h:386:9 // aclshmemi_store(sync_counter, count)
====== #2 .../src/device/gm2gm/shmemi_device_cc.h:505:9 // aclshmemi_sync_npu_v3
====== #3 .../src/device/gm2gm/shmemi_device_cc_kernel.cpp:22:5 // aclshmemi_sync
====== PIPE_S Write at WAW()+0x124000500200 in block 1 (aiv) on device 0 at pc current 0x5328 (serialNo:106)
====== #0/#1/#2/#3 同上

地址 0x124000500200 为团队 barrier 的代际计数器 sync_counter(aclshmemi_get_team_sync_counter)。

根因

aclshmemi_sync_npu_v3(src/device/gm2gm/shmemi_device_cc.h 附近 L349-L391)中,L386 的计数器 store 被放在 if ASCEND_IS_AIV 内、但在 if (vec_id < k)(L373)之外:

if ASCEND_IS_AIV {
// Only the first k AIVs participate in remote polling.
if (vec_id < k) {
for (...) { /* sync_pool 自写/读远端 / }
}
aclshmemi_store((gm int32_t
)sync_counter, count); // L386: 每个 AIV 都执行
}
aclshmemi_sync_core();

  • aclshmemi_store 是普通非原子标量写(src/device/shmemi_device_common.hpp L66-L70:((gm T)addr) = val)。
  • 因此同一 device 上的每个 AIV 向量核都会并发地向 sync_counter 写入。device 0 上 aiv0(block0) 与 aiv1(block1) 的两次写入之间没有任何跨核 happens-before 排序边,属于非同步并发写数据竞争,且为非原子写。

为什么是 bug 而非工具误报

  1. 调用栈清晰定位到非原子 store,写入者为多个核,语义上确为并发写。
  2. 竞争地址是共享单个字的 sync_counter,非每 PE 独立槽位。
  3. 旁证:v4 分支(src/device/gm2gm/shmemi_device_cc.h L456-L459)用 if (vec_id == 0) 守卫 sync_counter 更新,只有 AIV0 写它,故 v4 无 WAW——v3 的放置属于实现疏漏。
  4. 目前"能跑"只是因为所有写者写入相同 count 值,是幂等覆盖,未破坏 barrier 语义,但不改变其数据竞争性质。

建议修复

对齐 v4,让 sync_counter 只在单个核(AIV0)上推进:

if ASCEND_IS_AIV {
if (vec_id < k) { /* ...pool 读写... / }
if (vec_id == 0) {
aclshmemi_store((gm int32_t
)sync_counter, count);
}
}

可选替代:对 sync_counter 使用原子/信号写(aclshmemi_signal_set 等)。

工具这边结合ai分析和日志指令打印得到非工具误报结论,是在release/v1.6.0分支跑的,未验证主线是否修复

Environment / 环境信息 (Mandatory / 必填)

环境信息

A5
shmem分支release/v1.6.0

Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)

bash scripts/build.sh -examples -enable_simt -soc_type Ascend950 -mssanitizer
time mssanitizer -t racecheck --log-level=error --full-backtrace=yes -- bash run.sh

Describe the expected behavior / 预期结果 (Mandatory / 必填)

修复后,用 mssanitizer 检测 simt_rma_ub2gm 时,barrier_on_stream_kernel 不再对 sync_counter 报出 WAW hazard,barrier 跨核/跨 PE 同步语义保持不变。

[mssanitizer] Start memcheck, racecheck, initcheck and synccheck sanitizer on kernel "barrier_on_stream_kernel(int, unsigned long)"
====== ERROR: Potential WAW hazard detected at GM in "barrier_on_stream_kernel(int, unsigned long)":
====== PIPE_S Write at WAW()+0x124000500200 in block 0 (aiv) on device 0 at pc current 0x5328 (serialNo:66)
====== #0 .../src/device/shmemi_device_common.hpp:69:25 // aclshmemi_store: *addr = val
====== #1 .../src/device/gm2gm/shmemi_device_cc.h:386:9 // aclshmemi_store(sync_counter, count)
====== #2 .../src/device/gm2gm/shmemi_device_cc.h:505:9 // aclshmemi_sync_npu_v3
====== #3 .../src/device/gm2gm/shmemi_device_cc_kernel.cpp:22:5 // aclshmemi_sync
====== PIPE_S Write at WAW()+0x124000500200 in block 1 (aiv) on device 0 at pc current 0x5328 (serialNo:106)
====== #0/#1/#2/#3 同上

Special notes for this issue/备注 (Optional / 选填)

likedislike
Hh-glue
20 天前 修改标题为 “【检测扫网】simt_rma_ub2gm算子aclshmem_sync_npu_v3 中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”,原标题为“aclshmem_sync_npu_v3 中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”
Hh-glue
20 天前 修改标题为 “【检测扫网】simt_rma_ub2gm算子barrier_on_stream_kernel中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”,原标题为“【检测扫网】simt_rma_ub2gm算子aclshmem_sync_npu_v3 中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”
Vector
Vector成员
20 天前 评论:

/assign @wanghan

likedislike
CANN-robot
CANN-robot成员
20 天前 评论:

Notice

This issue can not be assigned to wanghan. Please try to assign to the repository members.

likedislike
Vector
Vector成员
20 天前 评论:

/assign @mizuki_p

likedislike
CANN-robotCANN-robot成员
20 天前 将 mizuki_p 设为负责人
wanghan
wanghan成员
20 天前 评论:

感谢您发现的问题,正在确认中

likedislike
wanghanwanghan成员
18 天前 关联了pull request:修复aclshmemi_sync_npu_v3重复写的问题
Hh-glue
16 天前 修改标题为 “[工具检测]simt_rma_ub2gm算子barrier_on_stream_kernel中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”,原标题为“【检测扫网】simt_rma_ub2gm算子barrier_on_stream_kernel中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”
Hh-glue
16 天前 修改标题为 “[工具检测]:simt_rma_ub2gm算子barrier_on_stream_kernel中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”,原标题为“[工具检测]simt_rma_ub2gm算子barrier_on_stream_kernel中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”
wanghan
wanghan成员
16 天前 评论:

您好,这个问题已经得到了确认,并且提了PR进行修复,然而因为这个修改涉及到了一个大量使用的barrier函数,所以还需要花时间验证PR是否会影响性能

likedislike
wanghan
wanghan成员
13 天前 评论:

您好,目前这个写法是有意这样做的,为了让每个核都有相同的数据缓存而不需要调用刷新缓存的接口,因为刷新缓存的接口耗时较长,经过实测,改用缓存刷新+单卡写入的形式会造成该接口性能严重下降(8卡情况下,接口调用耗时1.32us->2.5us,不同的测试方法结果可能不一样,这个结果是直接对aclshmemi_sync_npu_v3调用的打点测试),所以决定仅在开启mssanitizer的情况下使用缓存刷新+单卡写入的形式以规避工具的检测

likedislike
wanghan
wanghan成员
13 天前 评论:

问题修复PR已合入

likedislike
wanghanwanghan成员
13 天前 issue状态由 进行中 改变为 已解决
wanghanwanghan成员
13 天前 关闭了 issue
CANN-robotCANN-robot成员
13 天前 添加了label:resolved
哈喽Kiter哈喽Kiter成员
7 天前 修改标题为 “[Bug-Report|缺陷反馈]:simt_rma_ub2gm算子barrier_on_stream_kernel中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”,原标题为“[工具检测]:simt_rma_ub2gm算子barrier_on_stream_kernel中多 AIV 对同一 sync_counter 并发非原子写导致的 WAW 数据竞争”