已关闭
[Bug-Report|缺陷反馈]: assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误 #3378
gggxinmeng创建于 18 天前关闭于 18 天前
Ggggxinmeng
18 天前 修改标题为 “[Bug-Report|缺陷反馈]: assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误”,原标题为“assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误”
18 天前 修改标题为 “[Bug-Report|缺陷反馈]: assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误”,原标题为“assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误”
18 天前 修改了issue 的描述
18 天前 issue状态由 待办的 改变为 已解决
18 天前 关闭了 issue
[Bug-Report|缺陷反馈]: assemble版本化后分段写回tile被合并为共享布局,fillpad/vecdup越界覆盖导致精度错误
Describe the current behavior / 问题描述 (Mandatory / 必填)
mix_1c2v_auto.json case 148(
mix_1c2v_VCV_Serial_Parallel_float32_h64_nondiv_148)精度 FAILED(result[1]/[2] 误差 1e4~1e9 量级)。二分定位:7b0399991PASS →46c711463(feat(pass): Adaptation to new IR features,assemble 版本化)引入回归。根因:版本化(多次写同一 destination 改为"独立 LogicalTensor 版本 + 共享 RawTensor")后,ReplaceTensor 的
InsertNeedCopy插 copy 条件失效——分段 resultTile 不再有第二个不同 raw 的 consumer,sameAssembleOut恒为 true,不再插 COPY_OUT/COPY_IN。BackwardAssemble遂将 fillpad 分段 resultTile 的 raw 合并回父 tensor 的 [16,128] raw,codegen 按父 rawshape 生成共享布局(LocalLayout2Dim<16,128>,段B base+80)。TFillPad/TVecDup 属"写满整块"类算子,在共享布局下 pad/fill 越界,实测 fillpad 输出 3×128 有效数据只剩 1×64,其余被 pad 0 覆盖(输入 cat_s2 逐元素正确)。Environment / 环境信息 (Mandatory / 必填)
Steps to reproduce the issue / 重现步骤 (Mandatory / 必填)
test_cases/pass/mix_1c2v_auto.json,method=VCV_Serial_Parallel,batch=9, seq_len=35, hidden=64, tile_b=2, tile_s=8, dtype=float32,vector_tile_shapes_2d=[16,80](128 列拆 80+48 非整除,触发分段写回)concat → view → fillpad → matmul → sub;scope3 sigmoid tiling 展开后的VEC_DUP(fill 常量,同机制,表现为小值域指标超标)Describe the expected behavior / 预期结果 (Mandatory / 必填)
三个输出精度全部 PASS(7b0399991 行为):fillpad 分段 dst 保持独立 raw(布局
<16,80>/<16,48>= 各自责任区),pad 不越界。Related log / screenshot / 日志 / 截图 (Mandatory / 必填)
生成代码对比(fillpad 两段 dst):
// GOOD: 两段独立 buffer,布局=责任区 UBTileTensorFP32Dim2_5 ubTensor_67(UB_S88064_E93184_T, Shape2Dim(16, 80)); // <16,80> UBTileTensorFP32Dim2_6 ubTensor_69(UB_S101376_E104448_T, Shape2Dim(16, 48)); // <16,48> // BAD: 共用同一 buffer,布局=父列宽 128,段B +80 偏移 UBTileTensorFP32Dim2_3 ubTensor_67(UB_S88064_E96256_T, Shape2Dim(16, 80)); // <16,128> UBTileTensorFP32Dim2_3 ubTensor_69((float*)UB_S88064_E96256_T + 80, Shape2Dim(16, 48)); // <16,128>出错机制(pto-isa
include/pto/npu/a5/TFillPad.hpp)——写入由编译期/运行期分工决定:Cols、RowStride = Cols(RowMajor 下行距=列宽)srcValidRow / srcValidColbase + i*RowStride,pad 终点 = 第 dstStride 列srcValidRow × srcValidCol个unsigned padCols = dstStride - srcValidCol; // 每行 pad 列数 = 有效列结束 → 布局行尾padCols隐含契约"布局列宽 == 本段责任区宽度"。共享布局(128)≠ 责任区(80/48)时契约破裂:golden 逐算子对比(BAD,bs_valid=3 的迭代):cat_s2(fillpad 输入)3×128 逐元素正确;v3(fillpad 输出)只剩 1×64 有效,row1/2 整行及 row0 后 64 列被 pad 0 覆盖;c2 及以下全部受害。
图结构证据(
compile_debug_mode=1dump 的 Pass_27_ReplaceTensor 前后图):Special notes for this issue/备注 (Optional / 选填)
已验证修复(pr_pad_dts
6e94047de,多轮验证):replace_tensor.cpp的FindNeedToCopyAssemble补充条件——producer 为OP_FILLPAD/OP_VEC_DUP且 assemble input shape != output shape(分段写回)时恢复插 copy,使 ASSEMBLE 消费拷贝件、分段计算 tile 保持独立 raw。此为恢复版本化前该 pass 的实际行为;elementwise/cube 类合并行为不变。修复后 mare/mere/rmse 全轮次 100% 通过(与 torch NPU 同水平)。