已关闭
[Bug-Report|缺陷反馈]: dequant_swiglu_quant arch35 kernel tmpBuffer stride不匹配导致VEC越界AIC Error #4746
caorenlei创建于  8月12日关闭于  8月17日
caorenlei成员
8月12日 创建

Describe the current behavior / 问题描述

dequant_swiglu_quant 算子 arch35 kernel 中,tmpBuffer 使用 xUbAlignB32_ 作为行 stride 分配空间(line 220),但在后续访问 tmpXPtr(指向 tmpBuffer)时,错误使用了 xTypeUbAlignB32_ 作为行 stride 计算地址。当输入数据类型为 bfloat16 时,xTypeUbAlignB32_ 与 xUbAlignB32_ 取值不同,导致地址计算偏移与实际 buffer 布局不匹配,引发 VEC(矢量计算单元)越界访问,产生 AIC Error。

Environment / 环境信息

  • 硬件:Ascend 950 (arch35)
  • 仓库:cann/ops-nn
  • 算子:dequant_swiglu_quant
  • 触发条件:输入数据类型为 bfloat16 时

Steps to reproduce the issue / 重现步骤

  1. 使用 dequant_swiglu_quant 算子,输入 x 数据类型设置为 bfloat16
  2. 执行算子计算
  3. 观察到 AIC Error,VEC 越界访问

Describe the expected behavior / 预期结果

tmpXPtr 的地址计算应使用与 tmpBuffer 分配时一致的 xUbAlignB32_ 作为行 stride,确保地址偏移在 buffer 范围内,不发生越界访问。

关键代码位置(quant/dequant_swiglu_quant/op_kernel/arch35/dequant_swiglu_quant.h):

  • Buffer 分配(line 220):pipe_->InitBuffer(tmpBuffer, tl_->UbFactorDimx * xUbAlignB32_ * sizeof(float));
  • 错误地址计算(line 510/665/682/749):tmpXPtr + i * xTypeUbAlignB32_ + ...(应使用 xUbAlignB32_)

修复后:上述 4 处地址计算均改为使用 xUbAlignB32_,与 buffer 分配 stride 保持一致。

Special notes for this issue/备注

关联 PR:https://gitcode.com/cann/ops-nn/pull/8561

likedislike
Ccaorenlei成员
8月12日 添加了label:bug-report
caorenlei成员
8月12日 评论:

/assign @caorenlei

likedislike
CANN-robotCANN-robot成员
8月12日 将 caorenlei 设为负责人
CANN-robotCANN-robot成员
8月17日 关闭了 issue
CANN-robotCANN-robot成员
8月17日 添加了label:resolved