已关闭
[Bug-Report|缺陷反馈]: assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误 #3378
gggxinmeng创建于  18 天前关闭于  18 天前
gggxinmeng成员
18 天前 创建

[Bug-Report|缺陷反馈]: assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误

Describe the current behavior / 问题描述 (Mandatory / 必填)

mix_1c2v_auto.json case 148(mix_1c2v_VCV_Serial_Parallel_float32_h64_nondiv_148)精度 FAILED(result[1]/[2] 误差 1e4~1e9 量级)。二分定位:7b0399991 PASS → 46c711463(feat(pass): Adaptation to new IR features,assemble 版本化)引入回归。

根因:版本化(多次写同一 destination 改为"独立 LogicalTensor 版本 + 共享 RawTensor")后,ReplaceTensor 的 InsertNeedCopy 插 copy 条件失效——分段 resultTile 不再有第二个不同 raw 的 consumer,sameAssembleOut 恒为 true,不再插 COPY_OUT/COPY_IN。BackwardAssemble 遂将 fillpad 分段 resultTile 的 raw 合并回父 tensor 的 [16,128] raw,codegen 按父 rawshape 生成共享布局(LocalLayout2Dim<16,128>,段B base+80)。TFillPad/TVecDup 属"写满整块"类算子,在共享布局下 pad/fill 越界,实测 fillpad 输出 3×128 有效数据只剩 1×64,其余被 pad 0 覆盖(输入 cat_s2 逐元素正确)。

Environment / 环境信息 (Mandatory / 必填)

  • 服务器/NPU 型号: Ascend950PR(A5)
  • PyPTO 版本/Commit: e292d21b9(好版本 7b0399991,引入回归 46c711463)
  • CANN 版本: 9.2.0
  • Python 版本: Python 3.10.20
  • 操作系统: Ubuntu 24.04.4 LTS
  • torch / torch_npu: 2.8.0+cpu / 2.8.0.post4

Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)

  1. 输入规格:test_cases/pass/mix_1c2v_auto.json,method=VCV_Serial_Parallel,batch=9, seq_len=35, hidden=64, tile_b=2, tile_s=8, dtype=float32, vector_tile_shapes_2d=[16,80](128 列拆 80+48 非整除,触发分段写回)
  2. 执行:
    python3 xrunfk.py -i=./test_cases/pass/mix_1c2v_auto.json -o=./output/pass_mix -s=148 -e=148 -d=1
    
  3. 受影响算子链:scope2 concat → view → fillpad → matmul → sub;scope3 sigmoid tiling 展开后的 VEC_DUP(fill 常量,同机制,表现为小值域指标超标)

Describe the expected behavior / 预期结果 (Mandatory / 必填)

三个输出精度全部 PASS(7b0399991 行为):fillpad 分段 dst 保持独立 raw(布局 <16,80>/<16,48> = 各自责任区),pad 不越界。

生成代码对比(fillpad 两段 dst):

// GOOD: 两段独立 buffer,布局=责任区
UBTileTensorFP32Dim2_5 ubTensor_67(UB_S88064_E93184_T,   Shape2Dim(16, 80));  // <16,80>
UBTileTensorFP32Dim2_6 ubTensor_69(UB_S101376_E104448_T, Shape2Dim(16, 48));  // <16,48>

// BAD: 共用同一 buffer,布局=父列宽 128,段B +80 偏移
UBTileTensorFP32Dim2_3 ubTensor_67(UB_S88064_E96256_T,          Shape2Dim(16, 80));  // <16,128>
UBTileTensorFP32Dim2_3 ubTensor_69((float*)UB_S88064_E96256_T + 80, Shape2Dim(16, 48));  // <16,128>

出错机制(pto-isa include/pto/npu/a5/TFillPad.hpp)——写入由编译期/运行期分工决定:

编译期(模板参数,来自 rawshape) 运行期(valid)
Cols、RowStride = Cols(RowMajor 下行距=列宽) srcValidRow / srcValidCol
决定写到哪:第 i 行 = base + i*RowStride,pad 终点 = 第 dstStride 列 决定搬多少:搬 srcValidRow × srcValidCol 个
unsigned padCols = dstStride - srcValidCol;  // 每行 pad 列数 = 有效列结束 → 布局行尾

padCols 隐含契约"布局列宽 == 本段责任区宽度"。共享布局(128)≠ 责任区(80/48)时契约破裂:

段A: padCols = 128-80 = 48 → 把 [80,128)(段B 责任区)当自己行尾 pad 填 0
段B: padCols = 128-48 = 80 → 行内 pad 起点 base+80+48 物理落在下一行行首,
                             跨行覆盖段A 已写入的数据(行距按 128 走)
行pad: 段B 声明区域终点超 buffer 尾 80 个元素,真实 UB 越界写

golden 逐算子对比(BAD,bs_valid=3 的迭代):cat_s2(fillpad 输入)3×128 逐元素正确;v3(fillpad 输出)只剩 1×64 有效,row1/2 整行及 row0 后 64 列被 pad 0 覆盖;c2 及以下全部受害。

图结构证据(compile_debug_mode=1 dump 的 Pass_27_ReplaceTensor 前后图):

GOOD Before: FILLPAD→resultTile ─┬→ ASSEMBLE→父tensor(直连)
                               └→ COPY_OUT→(读方向分支,第二个不同 raw 的 consumer)
GOOD After:  FILLPAD→resultTile ─┬→ COPY_OUT→COPY_IN→拷贝件→ASSEMBLE(pass 自动插 copy)
                               └→ COPY_OUT→(读方向)
BAD:         FILLPAD→resultTile → ASSEMBLE→父tensor(直连,不插 copy)

Special notes for this issue/备注 (Optional / 选填)

已验证修复(pr_pad_dts 6e94047de,多轮验证):replace_tensor.cpp 的 FindNeedToCopyAssemble 补充条件——producer 为 OP_FILLPAD/OP_VEC_DUP 且 assemble input shape != output shape(分段写回)时恢复插 copy,使 ASSEMBLE 消费拷贝件、分段计算 tile 保持独立 raw。此为恢复版本化前该 pass 的实际行为;elementwise/cube 类合并行为不变。修复后 mare/mere/rmse 全轮次 100% 通过(与 torch NPU 同水平)。

likedislike
Ggggxinmeng成员
18 天前 修改标题为 “[Bug-Report|缺陷反馈]: assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误”,原标题为“assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误”
Ggggxinmeng成员
18 天前 修改了issue 的描述
Ggggxinmeng成员
18 天前 issue状态由 待办的 改变为 已解决
Ggggxinmeng成员
18 天前 关闭了 issue