已开启
[bug][auto-vectorize-v2][multi-consumer] 开启多 consumer 门禁 CV Native FA 编译错误 #381
hujiajun创建于  7月31日
hujiajun成员
7月31日 创建

叠加 pr commit ee1db4dd

do not retry for auto-vec-v2
remove no use clean transform op
bugfix shape of deintlv intlv
add loop.wrap_producer_extract_slice
test for open mcf default *

问题是 ub overflow

likedislike
hujiajun成员
7月31日 评论:
// Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:25:44.061Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:25:44.061Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpoy18bk1i/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpoy18bk1i/kernel --enable-vf-merge-level=1
// [2026-07-31T14:25:44.061Z] RERUN
// [2026-07-31T14:25:51.784Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype4-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:25:51.784Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:25:51.784Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp9i5qdfr4/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp9i5qdfr4/kernel --enable-vf-merge-level=1
// [2026-07-31T14:25:51.784Z] RERUN
// [2026-07-31T14:26:01.172Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype4-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:26:01.172Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:26:01.173Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp_qnqchlz/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp_qnqchlz/kernel --enable-vf-merge-level=1
// [2026-07-31T14:26:01.173Z] RERUN
// [2026-07-31T14:26:08.893Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype4-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:26:08.893Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:26:08.893Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpcm8ob1m3/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpcm8ob1m3/kernel --enable-vf-merge-level=1
// [2026-07-31T14:26:08.893Z] RERUN
// [2026-07-31T14:26:12.150Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype4-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:26:12.150Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:26:12.150Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmplbzh40ta/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmplbzh40ta/kernel --enable-vf-merge-level=1
// [2026-07-31T14:26:12.150Z] FAILED

#loc = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)
#loc2 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":212:70)
#loc5 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":202:78)
#loc6 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":195:44)
#loc7 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:46)
#loc17 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":282:36)
#loc18 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":88:25)
#loc41 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":81:22)
#loc44 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":84:24)
#loc59 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:33)
#loc60 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:46)
#loc73 = loc(callsite(#loc6 at #loc7))
#loc78 = loc(callsite(#loc17 at #loc18))
#loc83 = loc(callsite(#loc41 at #loc5))
#loc86 = loc(callsite(#loc44 at #loc5))
#loc105 = loc(callsite(#loc59 at #loc2))
#loc106 = loc(callsite(#loc6 at #loc60))
#loc110 = loc(callsite(#loc44 at #loc2))
#loc121 = loc(callsite(#loc73 at #loc5))
#loc122 = loc(callsite(#loc78 at #loc2))
#loc124 = loc(callsite(#loc78 at #loc5))
#loc126 = loc(callsite(#loc106 at #loc2))
module attributes {hacc.target = #hacc.target<"Ascend950PR_9579">} {
  func.func @_attn_fwd(%arg0: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg1: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg2: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg3: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg4: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg5: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg6: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg7: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg8: memref<?xi32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg9: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg10: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg11: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg12: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg13: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg14: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg15: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)) attributes {SyncBlockLockArgIdx = 0 : i64, WorkspaceArgIdx = 1 : i64, global_kernel = "local", mix_mode = "mix", parallel_mode = "simd"} {
    %c8192 = arith.constant 8192 : index loc(#loc71)
    %c128 = arith.constant 128 : index loc(#loc71)
    %c1 = arith.constant 1 : index loc(#loc3)
    %c0 = arith.constant 0 : index loc(#loc4)
    %cst = arith.constant 1.000000e+00 : f32 loc(#loc72)
    %cst_0 = arith.constant 0xFF800000 : f32 loc(#loc121)
    %cst_1 = arith.constant -1.000000e+04 : f32 loc(#loc74)
    %cst_2 = arith.constant 5.000000e-01 : f32 loc(#loc75)
    %c0_i32 = arith.constant 0 : i32 loc(#loc71)
    %c1_i32 = arith.constant 1 : i32 loc(#loc76)
    %c64_i32 = arith.constant 64 : i32 loc(#loc11)
    %c8_i32 = arith.constant 8 : i32 loc(#loc12)
    %c8388608_i64 = arith.constant 8388608 : i64 loc(#loc13)
    %c1048576_i64 = arith.constant 1048576 : i64 loc(#loc14)
    %c128_i32 = arith.constant 128 : i32 loc(#loc77)
    %c8192_i32 = arith.constant 8192 : i32 loc(#loc16)
    %cst_3 = arith.constant 0.000000e+00 : f32 loc(#loc122)
    %0 = tensor.empty() : tensor<128x128xf32> loc(#loc79)
    %1 = linalg.fill ins(%cst_3 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
    %2 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc75)
    %3 = linalg.fill ins(%cst_1 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
    %4 = tensor.empty() : tensor<128xf32> loc(#loc121)
    %5 = linalg.fill ins(%cst_0 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc121)
    %6 = linalg.fill ins(%cst : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc72)
    %7 = arith.index_cast %arg13 : i32 to index loc(#loc20)
    %reinterpret_cast = memref.reinterpret_cast %arg8 to offset: [%7], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc21)
    %8 = memref.load %reinterpret_cast[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc22)
    %9 = arith.addi %7, %c1 : index loc(#loc3)
    %reinterpret_cast_4 = memref.reinterpret_cast %arg8 to offset: [%9], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc3)
    %10 = memref.load %reinterpret_cast_4[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc4)
    scf.for %arg16 = %8 to %10 step %c1_i32  : i32 {
      %11 = arith.divsi %arg16, %c64_i32 : i32 loc(#loc24)
      %12 = arith.remsi %arg16, %c64_i32 : i32 loc(#loc11)
      %13 = arith.divsi %11, %c8_i32 : i32 loc(#loc25)
      %14 = arith.remsi %11, %c8_i32 : i32 loc(#loc26)
      %15 = arith.muli %14, %arg9 : i32 loc(#loc27)
      %16 = arith.divsi %15, %c8_i32 : i32 loc(#loc12)
      %17 = arith.extsi %13 : i32 to i64 loc(#loc28)
      %18 = arith.muli %17, %c8388608_i64 : i64 loc(#loc13)
      %19 = arith.extsi %14 : i32 to i64 loc(#loc29)
      %20 = arith.muli %19, %c1048576_i64 : i64 loc(#loc30)
      %21 = arith.addi %18, %20 : i64 loc(#loc31)
      %22 = arith.extsi %16 : i32 to i64 loc(#loc32)
      %23 = arith.muli %22, %c1048576_i64 : i64 loc(#loc14)
      %24 = arith.addi %18, %23 : i64 loc(#loc33)
      %25 = arith.index_cast %21 : i64 to index loc(#loc31)
      %26 = arith.muli %12, %c128_i32 : i32 loc(#loc34)
      %27 = arith.maxsi %26, %c0_i32 : i32 loc(#loc35)
      %28 = arith.index_cast %27 : i32 to index loc(#loc35)
      %29 = arith.muli %28, %c128 : index loc(#loc35)
      %30 = arith.addi %29, %25 : index loc(#loc35)
      %reinterpret_cast_5 = memref.reinterpret_cast %arg2 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc35)
      %31 = arith.index_cast %24 : i64 to index loc(#loc33)
      %reinterpret_cast_6 = memref.reinterpret_cast %arg7 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc36)
      %alloc = memref.alloc() : memref<128x128xf16> loc(#loc37)
      memref.copy %reinterpret_cast_5, %alloc : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc37)
      %32 = bufferization.to_tensor %alloc restrict writable : memref<128x128xf16> loc(#loc37)
      %33:5 = scf.for %arg17 = %c0_i32 to %26 step %c128_i32 iter_args(%arg18 = %6, %arg19 = %1, %arg20 = %5, %arg21 = %c0_i32, %arg22 = %c0_i32) -> (tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32)  : i32 {
        %46 = arith.maxsi %arg21, %c0_i32 : i32 loc(#loc72)
        %47 = arith.index_cast %46 : i32 to index loc(#loc72)
        %48 = arith.muli %47, %c128 : index loc(#loc72)
        %49 = arith.addi %48, %31 : index loc(#loc72)
        %reinterpret_cast_8 = memref.reinterpret_cast %arg4 to offset: [%49], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc72)
        %50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc72)
        %51 = arith.index_cast %50 : i32 to index loc(#loc72)
        %52 = arith.muli %51, %c128 : index loc(#loc72)
        %53 = arith.addi %52, %31 : index loc(#loc72)
        %reinterpret_cast_9 = memref.reinterpret_cast %arg3 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc72)
        %alloc_10 = memref.alloc() : memref<128x128xf16> loc(#loc80)
        memref.copy %reinterpret_cast_9, %alloc_10 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc80)
        %54 = bufferization.to_tensor %alloc_10 restrict writable : memref<128x128xf16> loc(#loc80)
        %55 = tensor.empty() : tensor<128x128xf16> loc(#loc81)
        %transposed = linalg.transpose ins(%54 : tensor<128x128xf16>) outs(%55 : tensor<128x128xf16>) permutation = [1, 0]  loc(#loc81)
        %56 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xf16>, tensor<128x128xf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc82)
        %57 = arith.mulf %56, %2 : tensor<128x128xf32> loc(#loc83)
        %reduced = linalg.reduce ins(%57 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc41 at #loc5)), %init: f32 loc(callsite(#loc73 at #loc5))) {
            %72 = arith.maximumf %in, %init : f32 loc(#loc128)
            linalg.yield %72 : f32 loc(#loc121)
          } loc(#loc121)
        %58 = arith.maximumf %arg20, %reduced : tensor<128xf32> loc(#loc85)
        %broadcasted_11 = linalg.broadcast ins(%58 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc79)
        %59 = arith.subf %57, %broadcasted_11 : tensor<128x128xf32> loc(#loc79)
        %60 = math.exp %59 : tensor<128x128xf32> loc(#loc86)
        %61 = arith.truncf %60 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc87)
        %alloc_12 = memref.alloc() : memref<128x128xf16> loc(#loc88)
        memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc88)
        %62 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf16> loc(#loc88)
        %63 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc124)
        %reduced_13 = linalg.reduce ins(%60 : tensor<128x128xf32>) outs(%63 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc44 at #loc5)), %init: f32 loc(callsite(#loc78 at #loc5))) {
            %72 = arith.addf %in, %init : f32 loc(#loc129)
            linalg.yield %72 : f32 loc(#loc124)
          } loc(#loc124)
        %64 = arith.subf %arg20, %58 : tensor<128xf32> loc(#loc90)
        %65 = math.exp %64 : tensor<128xf32> loc(#loc91)
        %66 = arith.mulf %arg18, %65 : tensor<128xf32> loc(#loc92)
        %67 = arith.addf %66, %reduced_13 : tensor<128xf32> loc(#loc93)
        %broadcasted_14 = linalg.broadcast ins(%65 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc94)
        %68 = arith.mulf %arg19, %broadcasted_14 : tensor<128x128xf32> loc(#loc94)
        %69 = linalg.matmul {input_precision = "ieee"} ins(%61, %62 : tensor<128x128xf16>, tensor<128x128xf16>) outs(%68 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc95)
        %70 = arith.addi %arg21, %c128_i32 : i32 loc(#loc96)
        %71 = arith.addi %arg22, %c128_i32 : i32 loc(#loc97)
        scf.yield %67, %69, %58, %70, %71 : tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc98)
      } {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc72)
      %34 = arith.muli %12, %c128_i32 {tt.divisibility = dense<128> : tensor<1xi32>} : i32 loc(#loc99)
      %35 = arith.addi %12, %c1_i32 : i32 loc(#loc76)
      %36 = arith.muli %35, %c128_i32 : i32 loc(#loc100)
      %37:6 = scf.for %arg17 = %34 to %36 step %c128_i32 iter_args(%arg18 = %34, %arg19 = %33#0, %arg20 = %33#1, %arg21 = %33#2, %arg22 = %34, %arg23 = %34) -> (i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32)  : i32 {
        %46 = arith.maxsi %arg18, %c0_i32 : i32 loc(#loc71)
        %47 = arith.index_cast %46 : i32 to index loc(#loc71)
        %48 = arith.muli %28, %c8192 : index loc(#loc71)
        %49 = arith.addi %48, %47 : index loc(#loc71)
        %reinterpret_cast_8 = memref.reinterpret_cast %arg5 to offset: [%49], sizes: [128, 128], strides: [8192, 1] : memref<?xf32> to memref<128x128xf32, strided<[8192, 1], offset: ?>> loc(#loc71)
        %50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc71)
        %51 = arith.index_cast %50 : i32 to index loc(#loc71)
        %52 = arith.muli %51, %c128 : index loc(#loc71)
        %53 = arith.addi %52, %31 : index loc(#loc71)
        %reinterpret_cast_9 = memref.reinterpret_cast %arg4 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc71)
        %54 = arith.maxsi %arg23, %c0_i32 : i32 loc(#loc71)
        %55 = arith.index_cast %54 : i32 to index loc(#loc71)
        %56 = arith.muli %55, %c128 : index loc(#loc71)
        %57 = arith.addi %56, %31 : index loc(#loc71)
        %reinterpret_cast_10 = memref.reinterpret_cast %arg3 to offset: [%57], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc71)
        %alloc_11 = memref.alloc() : memref<128x128xf16> loc(#loc101)
        memref.copy %reinterpret_cast_10, %alloc_11 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc101)
        %58 = bufferization.to_tensor %alloc_11 restrict writable : memref<128x128xf16> loc(#loc101)
        %59 = tensor.empty() : tensor<128x128xf16> loc(#loc102)
        %transposed = linalg.transpose ins(%58 : tensor<128x128xf16>) outs(%59 : tensor<128x128xf16>) permutation = [1, 0]  loc(#loc102)
        %60 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xf16>, tensor<128x128xf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc103)
        %alloc_12 = memref.alloc() : memref<128x128xf32> loc(#loc104)
        memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf32, strided<[8192, 1], offset: ?>> to memref<128x128xf32> loc(#loc104)
        %61 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf32> loc(#loc104)
        %62 = arith.mulf %60, %2 : tensor<128x128xf32> loc(#loc75)
        %63 = arith.cmpf une, %61, %1 : tensor<128x128xf32> loc(#loc74)
        %64 = arith.select %63, %3, %1 : tensor<128x128xi1>, tensor<128x128xf32> loc(#loc74)
        %65 = arith.addf %62, %64 : tensor<128x128xf32> loc(#loc105)
        %reduced = linalg.reduce ins(%65 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc59 at #loc2)), %init: f32 loc(callsite(#loc106 at #loc2))) {
            %81 = arith.maximumf %in, %init : f32 loc(#loc130)
            linalg.yield %81 : f32 loc(#loc126)
          } loc(#loc126)
        %66 = arith.maximumf %arg21, %reduced : tensor<128xf32> loc(#loc107)
        %broadcasted_13 = linalg.broadcast ins(%66 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc108)
        %67 = arith.subf %65, %broadcasted_13 : tensor<128x128xf32> loc(#loc108)
        %68 = arith.addi %arg18, %c128_i32 : i32 loc(#loc109)
        %69 = math.exp %67 : tensor<128x128xf32> loc(#loc110)
        %70 = arith.truncf %69 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc111)
        %alloc_14 = memref.alloc() : memref<128x128xf16> loc(#loc112)
        memref.copy %reinterpret_cast_9, %alloc_14 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc112)
        %71 = bufferization.to_tensor %alloc_14 restrict writable : memref<128x128xf16> loc(#loc112)
        %72 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc122)
        %reduced_15 = linalg.reduce ins(%69 : tensor<128x128xf32>) outs(%72 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc44 at #loc2)), %init: f32 loc(callsite(#loc78 at #loc2))) {
            %81 = arith.addf %in, %init : f32 loc(#loc131)
            linalg.yield %81 : f32 loc(#loc122)
          } loc(#loc122)
        %73 = arith.subf %arg21, %66 : tensor<128xf32> loc(#loc113)
        %74 = math.exp %73 : tensor<128xf32> loc(#loc114)
        %75 = arith.mulf %arg19, %74 : tensor<128xf32> loc(#loc115)
        %76 = arith.addf %75, %reduced_15 : tensor<128xf32> loc(#loc116)
        %broadcasted_16 = linalg.broadcast ins(%74 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc117)
        %77 = arith.mulf %arg20, %broadcasted_16 : tensor<128x128xf32> loc(#loc117)
        %78 = linalg.matmul {input_precision = "ieee"} ins(%70, %71 : tensor<128x128xf16>, tensor<128x128xf16>) outs(%77 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc118)
        %79 = arith.addi %arg22, %c128_i32 : i32 loc(#loc119)
        %80 = arith.addi %arg23, %c128_i32 : i32 loc(#loc77)
        scf.yield %68, %76, %78, %66, %79, %80 : i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc120)
      } {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc71)
      %38 = math.log %37#1 : tensor<128xf32> loc(#loc64)
      %39 = arith.addf %37#3, %38 : tensor<128xf32> loc(#loc65)
      %broadcasted = linalg.broadcast ins(%37#1 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc66)
      %40 = arith.divf %37#2, %broadcasted : tensor<128x128xf32> loc(#loc66)
      %41 = arith.muli %11, %c8192_i32 : i32 loc(#loc16)
      %42 = arith.index_cast %41 : i32 to index loc(#loc16)
      %43 = arith.index_cast %26 : i32 to index loc(#loc34)
      %44 = arith.addi %42, %43 : index loc(#loc67)
      %reinterpret_cast_7 = memref.reinterpret_cast %arg6 to offset: [%44], sizes: [128], strides: [1] : memref<?xf32> to memref<128xf32, strided<[1], offset: ?>> loc(#loc67)
      bufferization.materialize_in_destination %39 in writable %reinterpret_cast_7 : (tensor<128xf32>, memref<128xf32, strided<[1], offset: ?>>) -> () loc(#loc68)
      %45 = arith.truncf %40 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc69)
      bufferization.materialize_in_destination %45 in writable %reinterpret_cast_6 : (tensor<128x128xf16>, memref<128x128xf16, strided<[128, 1], offset: ?>>) -> () loc(#loc70)
    } loc(#loc23)
    return loc(#loc)
  } loc(#loc)
} loc(#loc)
#loc1 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":63:33)
#loc3 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:53)
#loc4 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:28)
#loc8 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:56)
#loc9 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:22)
#loc10 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:47)
#loc11 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":139:33)
#loc12 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:39)
#loc13 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:40)
#loc14 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:72)
#loc15 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:46)
#loc16 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:35)
#loc19 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":83:22)
#loc20 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":123:24)
#loc21 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:49)
#loc22 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:30)
#loc23 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":137:51)
#loc24 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":138:35)
#loc25 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":140:31)
#loc26 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":141:30)
#loc27 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:23)
#loc28 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:28)
#loc29 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:61)
#loc30 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:73)
#loc31 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:52)
#loc32 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:60)
#loc33 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:52)
#loc34 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":152:34)
#loc35 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":154:12)
#loc36 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":180:12)
#loc37 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":193:20)
#loc38 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":66:20)
#loc39 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":69:27)
#loc40 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":70:23)
#loc42 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":168:27)
#loc43 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:35)
#loc45 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":85:22)
#loc46 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":86:20)
#loc47 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":271:15)
#loc48 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:34)
#loc49 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:28)
#loc50 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:20)
#loc51 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:28)
#loc52 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":94:20)
#loc53 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":87:28)
#loc54 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":97:46)
#loc55 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:8)
#loc56 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:27)
#loc57 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:52)
#loc58 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":75:27)
#loc61 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:35)
#loc62 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":78:18)
#loc63 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":79:54)
#loc64 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:27)
#loc65 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:15)
#loc66 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":216:20)
#loc67 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:43)
#loc68 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":219:25)
#loc69 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:37)
#loc70 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:30)
#loc71 = loc(callsite(#loc1 at #loc2))
#loc72 = loc(callsite(#loc1 at #loc5))
#loc74 = loc(callsite(#loc8 at #loc2))
#loc75 = loc(callsite(#loc9 at #loc2))
#loc76 = loc(callsite(#loc10 at #loc2))
#loc77 = loc(callsite(#loc15 at #loc2))
#loc79 = loc(callsite(#loc19 at #loc5))
#loc80 = loc(callsite(#loc38 at #loc5))
#loc81 = loc(callsite(#loc39 at #loc5))
#loc82 = loc(callsite(#loc40 at #loc5))
#loc84 = loc(callsite(#loc42 at #loc6))
#loc85 = loc(callsite(#loc43 at #loc5))
#loc87 = loc(callsite(#loc45 at #loc5))
#loc88 = loc(callsite(#loc46 at #loc5))
#loc89 = loc(callsite(#loc47 at #loc17))
#loc90 = loc(callsite(#loc48 at #loc5))
#loc91 = loc(callsite(#loc49 at #loc5))
#loc92 = loc(callsite(#loc50 at #loc5))
#loc93 = loc(callsite(#loc51 at #loc5))
#loc94 = loc(callsite(#loc52 at #loc5))
#loc95 = loc(callsite(#loc53 at #loc5))
#loc96 = loc(callsite(#loc54 at #loc5))
#loc97 = loc(callsite(#loc15 at #loc5))
#loc98 = loc(callsite(#loc55 at #loc5))
#loc99 = loc(callsite(#loc56 at #loc2))
#loc100 = loc(callsite(#loc57 at #loc2))
#loc101 = loc(callsite(#loc38 at #loc2))
#loc102 = loc(callsite(#loc39 at #loc2))
#loc103 = loc(callsite(#loc40 at #loc2))
#loc104 = loc(callsite(#loc58 at #loc2))
#loc107 = loc(callsite(#loc61 at #loc2))
#loc108 = loc(callsite(#loc62 at #loc2))
#loc109 = loc(callsite(#loc63 at #loc2))
#loc111 = loc(callsite(#loc45 at #loc2))
#loc112 = loc(callsite(#loc46 at #loc2))
#loc113 = loc(callsite(#loc48 at #loc2))
#loc114 = loc(callsite(#loc49 at #loc2))
#loc115 = loc(callsite(#loc50 at #loc2))
#loc116 = loc(callsite(#loc51 at #loc2))
#loc117 = loc(callsite(#loc52 at #loc2))
#loc118 = loc(callsite(#loc53 at #loc2))
#loc119 = loc(callsite(#loc54 at #loc2))
#loc120 = loc(callsite(#loc55 at #loc2))
#loc123 = loc(callsite(#loc84 at #loc7))
#loc125 = loc(callsite(#loc89 at #loc18))
#loc127 = loc(callsite(#loc84 at #loc60))
#loc128 = loc(callsite(#loc123 at #loc5))
#loc129 = loc(callsite(#loc125 at #loc5))
#loc130 = loc(callsite(#loc127 at #loc2))
#loc131 = loc(callsite(#loc125 at #loc2))
likedislike
hujiajun成员
7月31日 评论:
[2026-07-31T14:26:32.036Z] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:26:32.036Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:26:32.036Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpeebgzx7u/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpeebgzx7u/kernel --enable-vf-merge-level=1
[2026-07-31T14:26:32.036Z] RERUN
[2026-07-31T14:26:39.756Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype6-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:26:39.756Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:26:39.756Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp4prqx238/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp4prqx238/kernel --enable-vf-merge-level=1
[2026-07-31T14:26:39.756Z] RERUN
[2026-07-31T14:26:49.142Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype6-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:26:49.143Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:26:49.143Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp2qtn812j/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp2qtn812j/kernel --enable-vf-merge-level=1
[2026-07-31T14:26:49.143Z] RERUN
[2026-07-31T14:26:56.865Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype6-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:26:56.865Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:26:56.865Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmplgik8dih/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmplgik8dih/kernel --enable-vf-merge-level=1
[2026-07-31T14:26:56.865Z] RERUN
[2026-07-31T14:27:00.122Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype6-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:27:00.122Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:27:00.122Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp8hxi0lmy/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp8hxi0lmy/kernel --enable-vf-merge-level=1
[2026-07-31T14:27:00.122Z] FAILED
#loc = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)
#loc2 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":212:70)
#loc5 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":202:78)
#loc6 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":195:44)
#loc7 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:46)
#loc16 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":282:36)
#loc17 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":88:25)
#loc41 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":81:22)
#loc44 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":84:24)
#loc59 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:33)
#loc60 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:46)
#loc73 = loc(callsite(#loc6 at #loc7))
#loc78 = loc(callsite(#loc16 at #loc17))
#loc83 = loc(callsite(#loc41 at #loc5))
#loc86 = loc(callsite(#loc44 at #loc5))
#loc105 = loc(callsite(#loc59 at #loc2))
#loc106 = loc(callsite(#loc6 at #loc60))
#loc110 = loc(callsite(#loc44 at #loc2))
#loc121 = loc(callsite(#loc73 at #loc5))
#loc122 = loc(callsite(#loc78 at #loc2))
#loc124 = loc(callsite(#loc78 at #loc5))
#loc126 = loc(callsite(#loc106 at #loc2))
module attributes {hacc.target = #hacc.target<"Ascend950PR_9579">} {
  func.func @_attn_fwd(%arg0: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg1: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg2: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg3: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg4: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg5: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg6: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg7: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg8: memref<?xi32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg9: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg10: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg11: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg12: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg13: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg14: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg15: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)) attributes {SyncBlockLockArgIdx = 0 : i64, WorkspaceArgIdx = 1 : i64, global_kernel = "local", mix_mode = "mix", parallel_mode = "simd"} {
    %c1024 = arith.constant 1024 : index loc(#loc71)
    %c128 = arith.constant 128 : index loc(#loc71)
    %c1 = arith.constant 1 : index loc(#loc3)
    %c0 = arith.constant 0 : index loc(#loc4)
    %cst = arith.constant 1.000000e+00 : f32 loc(#loc72)
    %cst_0 = arith.constant 0xFF800000 : f32 loc(#loc121)
    %cst_1 = arith.constant -1.000000e+04 : f32 loc(#loc74)
    %cst_2 = arith.constant 5.000000e-01 : f32 loc(#loc75)
    %c1024_i32 = arith.constant 1024 : i32 loc(#loc10)
    %c0_i32 = arith.constant 0 : i32 loc(#loc71)
    %c1_i32 = arith.constant 1 : i32 loc(#loc76)
    %c8_i32 = arith.constant 8 : i32 loc(#loc12)
    %c1048576_i64 = arith.constant 1048576 : i64 loc(#loc13)
    %c131072_i64 = arith.constant 131072 : i64 loc(#loc14)
    %c128_i32 = arith.constant 128 : i32 loc(#loc77)
    %cst_3 = arith.constant 0.000000e+00 : f32 loc(#loc122)
    %0 = tensor.empty() : tensor<128x128xf32> loc(#loc79)
    %1 = linalg.fill ins(%cst_3 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
    %2 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc75)
    %3 = linalg.fill ins(%cst_1 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
    %4 = tensor.empty() : tensor<128xf32> loc(#loc121)
    %5 = linalg.fill ins(%cst_0 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc121)
    %6 = linalg.fill ins(%cst : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc72)
    %7 = arith.index_cast %arg13 : i32 to index loc(#loc19)
    %reinterpret_cast = memref.reinterpret_cast %arg8 to offset: [%7], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc20)
    %8 = memref.load %reinterpret_cast[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc21)
    %9 = arith.addi %7, %c1 : index loc(#loc3)
    %reinterpret_cast_4 = memref.reinterpret_cast %arg8 to offset: [%9], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc3)
    %10 = memref.load %reinterpret_cast_4[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc4)
    scf.for %arg16 = %8 to %10 step %c1_i32  : i32 {
      %11 = arith.divsi %arg16, %c8_i32 : i32 loc(#loc23)
      %12 = arith.remsi %arg16, %c8_i32 : i32 loc(#loc24)
      %13 = arith.divsi %11, %c8_i32 : i32 loc(#loc25)
      %14 = arith.remsi %11, %c8_i32 : i32 loc(#loc26)
      %15 = arith.muli %14, %arg9 : i32 loc(#loc27)
      %16 = arith.divsi %15, %c8_i32 : i32 loc(#loc12)
      %17 = arith.extsi %13 : i32 to i64 loc(#loc28)
      %18 = arith.muli %17, %c1048576_i64 : i64 loc(#loc13)
      %19 = arith.extsi %14 : i32 to i64 loc(#loc29)
      %20 = arith.muli %19, %c131072_i64 : i64 loc(#loc30)
      %21 = arith.addi %18, %20 : i64 loc(#loc31)
      %22 = arith.extsi %16 : i32 to i64 loc(#loc32)
      %23 = arith.muli %22, %c131072_i64 : i64 loc(#loc14)
      %24 = arith.addi %18, %23 : i64 loc(#loc33)
      %25 = arith.index_cast %21 : i64 to index loc(#loc31)
      %26 = arith.muli %12, %c128_i32 : i32 loc(#loc34)
      %27 = arith.maxsi %26, %c0_i32 : i32 loc(#loc35)
      %28 = arith.index_cast %27 : i32 to index loc(#loc35)
      %29 = arith.muli %28, %c128 : index loc(#loc35)
      %30 = arith.addi %29, %25 : index loc(#loc35)
      %reinterpret_cast_5 = memref.reinterpret_cast %arg2 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc35)
      %31 = arith.index_cast %24 : i64 to index loc(#loc33)
      %reinterpret_cast_6 = memref.reinterpret_cast %arg7 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc36)
      %alloc = memref.alloc() : memref<128x128xf16> loc(#loc37)
      memref.copy %reinterpret_cast_5, %alloc : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc37)
      %32 = bufferization.to_tensor %alloc restrict writable : memref<128x128xf16> loc(#loc37)
      %33:5 = scf.for %arg17 = %c0_i32 to %26 step %c128_i32 iter_args(%arg18 = %6, %arg19 = %1, %arg20 = %5, %arg21 = %c0_i32, %arg22 = %c0_i32) -> (tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32)  : i32 {
        %46 = arith.maxsi %arg21, %c0_i32 : i32 loc(#loc72)
        %47 = arith.index_cast %46 : i32 to index loc(#loc72)
        %48 = arith.muli %47, %c128 : index loc(#loc72)
        %49 = arith.addi %48, %31 : index loc(#loc72)
        %reinterpret_cast_8 = memref.reinterpret_cast %arg4 to offset: [%49], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc72)
        %50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc72)
        %51 = arith.index_cast %50 : i32 to index loc(#loc72)
        %52 = arith.muli %51, %c128 : index loc(#loc72)
        %53 = arith.addi %52, %31 : index loc(#loc72)
        %reinterpret_cast_9 = memref.reinterpret_cast %arg3 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc72)
        %alloc_10 = memref.alloc() : memref<128x128xf16> loc(#loc80)
        memref.copy %reinterpret_cast_9, %alloc_10 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc80)
        %54 = bufferization.to_tensor %alloc_10 restrict writable : memref<128x128xf16> loc(#loc80)
        %55 = tensor.empty() : tensor<128x128xf16> loc(#loc81)
        %transposed = linalg.transpose ins(%54 : tensor<128x128xf16>) outs(%55 : tensor<128x128xf16>) permutation = [1, 0]  loc(#loc81)
        %56 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xf16>, tensor<128x128xf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc82)
        %57 = arith.mulf %56, %2 : tensor<128x128xf32> loc(#loc83)
        %reduced = linalg.reduce ins(%57 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc41 at #loc5)), %init: f32 loc(callsite(#loc73 at #loc5))) {
            %72 = arith.maximumf %in, %init : f32 loc(#loc128)
            linalg.yield %72 : f32 loc(#loc121)
          } loc(#loc121)
        %58 = arith.maximumf %arg20, %reduced : tensor<128xf32> loc(#loc85)
        %broadcasted_11 = linalg.broadcast ins(%58 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc79)
        %59 = arith.subf %57, %broadcasted_11 : tensor<128x128xf32> loc(#loc79)
        %60 = math.exp %59 : tensor<128x128xf32> loc(#loc86)
        %61 = arith.truncf %60 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc87)
        %alloc_12 = memref.alloc() : memref<128x128xf16> loc(#loc88)
        memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc88)
        %62 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf16> loc(#loc88)
        %63 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc124)
        %reduced_13 = linalg.reduce ins(%60 : tensor<128x128xf32>) outs(%63 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc44 at #loc5)), %init: f32 loc(callsite(#loc78 at #loc5))) {
            %72 = arith.addf %in, %init : f32 loc(#loc129)
            linalg.yield %72 : f32 loc(#loc124)
          } loc(#loc124)
        %64 = arith.subf %arg20, %58 : tensor<128xf32> loc(#loc90)
        %65 = math.exp %64 : tensor<128xf32> loc(#loc91)
        %66 = arith.mulf %arg18, %65 : tensor<128xf32> loc(#loc92)
        %67 = arith.addf %66, %reduced_13 : tensor<128xf32> loc(#loc93)
        %broadcasted_14 = linalg.broadcast ins(%65 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc94)
        %68 = arith.mulf %arg19, %broadcasted_14 : tensor<128x128xf32> loc(#loc94)
        %69 = linalg.matmul {input_precision = "ieee"} ins(%61, %62 : tensor<128x128xf16>, tensor<128x128xf16>) outs(%68 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc95)
        %70 = arith.addi %arg21, %c128_i32 : i32 loc(#loc96)
        %71 = arith.addi %arg22, %c128_i32 : i32 loc(#loc97)
        scf.yield %67, %69, %58, %70, %71 : tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc98)
      } {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc72)
      %34 = arith.muli %12, %c128_i32 {tt.divisibility = dense<128> : tensor<1xi32>} : i32 loc(#loc99)
      %35 = arith.addi %12, %c1_i32 : i32 loc(#loc76)
      %36 = arith.muli %35, %c128_i32 : i32 loc(#loc100)
      %37:6 = scf.for %arg17 = %34 to %36 step %c128_i32 iter_args(%arg18 = %34, %arg19 = %33#0, %arg20 = %33#1, %arg21 = %33#2, %arg22 = %34, %arg23 = %34) -> (i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32)  : i32 {
        %46 = arith.maxsi %arg18, %c0_i32 : i32 loc(#loc71)
        %47 = arith.index_cast %46 : i32 to index loc(#loc71)
        %48 = arith.muli %28, %c1024 : index loc(#loc71)
        %49 = arith.addi %48, %47 : index loc(#loc71)
        %reinterpret_cast_8 = memref.reinterpret_cast %arg5 to offset: [%49], sizes: [128, 128], strides: [1024, 1] : memref<?xf32> to memref<128x128xf32, strided<[1024, 1], offset: ?>> loc(#loc71)
        %50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc71)
        %51 = arith.index_cast %50 : i32 to index loc(#loc71)
        %52 = arith.muli %51, %c128 : index loc(#loc71)
        %53 = arith.addi %52, %31 : index loc(#loc71)
        %reinterpret_cast_9 = memref.reinterpret_cast %arg4 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc71)
        %54 = arith.maxsi %arg23, %c0_i32 : i32 loc(#loc71)
        %55 = arith.index_cast %54 : i32 to index loc(#loc71)
        %56 = arith.muli %55, %c128 : index loc(#loc71)
        %57 = arith.addi %56, %31 : index loc(#loc71)
        %reinterpret_cast_10 = memref.reinterpret_cast %arg3 to offset: [%57], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc71)
        %alloc_11 = memref.alloc() : memref<128x128xf16> loc(#loc101)
        memref.copy %reinterpret_cast_10, %alloc_11 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc101)
        %58 = bufferization.to_tensor %alloc_11 restrict writable : memref<128x128xf16> loc(#loc101)
        %59 = tensor.empty() : tensor<128x128xf16> loc(#loc102)
        %transposed = linalg.transpose ins(%58 : tensor<128x128xf16>) outs(%59 : tensor<128x128xf16>) permutation = [1, 0]  loc(#loc102)
        %60 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xf16>, tensor<128x128xf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc103)
        %alloc_12 = memref.alloc() : memref<128x128xf32> loc(#loc104)
        memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf32, strided<[1024, 1], offset: ?>> to memref<128x128xf32> loc(#loc104)
        %61 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf32> loc(#loc104)
        %62 = arith.mulf %60, %2 : tensor<128x128xf32> loc(#loc75)
        %63 = arith.cmpf une, %61, %1 : tensor<128x128xf32> loc(#loc74)
        %64 = arith.select %63, %3, %1 : tensor<128x128xi1>, tensor<128x128xf32> loc(#loc74)
        %65 = arith.addf %62, %64 : tensor<128x128xf32> loc(#loc105)
        %reduced = linalg.reduce ins(%65 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc59 at #loc2)), %init: f32 loc(callsite(#loc106 at #loc2))) {
            %81 = arith.maximumf %in, %init : f32 loc(#loc130)
            linalg.yield %81 : f32 loc(#loc126)
          } loc(#loc126)
        %66 = arith.maximumf %arg21, %reduced : tensor<128xf32> loc(#loc107)
        %broadcasted_13 = linalg.broadcast ins(%66 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc108)
        %67 = arith.subf %65, %broadcasted_13 : tensor<128x128xf32> loc(#loc108)
        %68 = arith.addi %arg18, %c128_i32 : i32 loc(#loc109)
        %69 = math.exp %67 : tensor<128x128xf32> loc(#loc110)
        %70 = arith.truncf %69 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc111)
        %alloc_14 = memref.alloc() : memref<128x128xf16> loc(#loc112)
        memref.copy %reinterpret_cast_9, %alloc_14 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc112)
        %71 = bufferization.to_tensor %alloc_14 restrict writable : memref<128x128xf16> loc(#loc112)
        %72 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc122)
        %reduced_15 = linalg.reduce ins(%69 : tensor<128x128xf32>) outs(%72 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc44 at #loc2)), %init: f32 loc(callsite(#loc78 at #loc2))) {
            %81 = arith.addf %in, %init : f32 loc(#loc131)
            linalg.yield %81 : f32 loc(#loc122)
          } loc(#loc122)
        %73 = arith.subf %arg21, %66 : tensor<128xf32> loc(#loc113)
        %74 = math.exp %73 : tensor<128xf32> loc(#loc114)
        %75 = arith.mulf %arg19, %74 : tensor<128xf32> loc(#loc115)
        %76 = arith.addf %75, %reduced_15 : tensor<128xf32> loc(#loc116)
        %broadcasted_16 = linalg.broadcast ins(%74 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc117)
        %77 = arith.mulf %arg20, %broadcasted_16 : tensor<128x128xf32> loc(#loc117)
        %78 = linalg.matmul {input_precision = "ieee"} ins(%70, %71 : tensor<128x128xf16>, tensor<128x128xf16>) outs(%77 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc118)
        %79 = arith.addi %arg22, %c128_i32 : i32 loc(#loc119)
        %80 = arith.addi %arg23, %c128_i32 : i32 loc(#loc77)
        scf.yield %68, %76, %78, %66, %79, %80 : i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc120)
      } {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc71)
      %38 = math.log %37#1 : tensor<128xf32> loc(#loc64)
      %39 = arith.addf %37#3, %38 : tensor<128xf32> loc(#loc65)
      %broadcasted = linalg.broadcast ins(%37#1 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc66)
      %40 = arith.divf %37#2, %broadcasted : tensor<128x128xf32> loc(#loc66)
      %41 = arith.muli %11, %c1024_i32 : i32 loc(#loc10)
      %42 = arith.index_cast %41 : i32 to index loc(#loc10)
      %43 = arith.index_cast %26 : i32 to index loc(#loc34)
      %44 = arith.addi %42, %43 : index loc(#loc67)
      %reinterpret_cast_7 = memref.reinterpret_cast %arg6 to offset: [%44], sizes: [128], strides: [1] : memref<?xf32> to memref<128xf32, strided<[1], offset: ?>> loc(#loc67)
      bufferization.materialize_in_destination %39 in writable %reinterpret_cast_7 : (tensor<128xf32>, memref<128xf32, strided<[1], offset: ?>>) -> () loc(#loc68)
      %45 = arith.truncf %40 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc69)
      bufferization.materialize_in_destination %45 in writable %reinterpret_cast_6 : (tensor<128x128xf16>, memref<128x128xf16, strided<[128, 1], offset: ?>>) -> () loc(#loc70)
    } loc(#loc22)
    return loc(#loc)
  } loc(#loc)
} loc(#loc)
#loc1 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":63:33)
#loc3 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:53)
#loc4 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:28)
#loc8 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:56)
#loc9 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:22)
#loc10 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:35)
#loc11 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:47)
#loc12 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:39)
#loc13 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:40)
#loc14 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:72)
#loc15 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:46)
#loc18 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":83:22)
#loc19 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":123:24)
#loc20 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:49)
#loc21 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:30)
#loc22 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":137:51)
#loc23 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":138:35)
#loc24 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":139:33)
#loc25 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":140:31)
#loc26 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":141:30)
#loc27 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:23)
#loc28 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:28)
#loc29 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:61)
#loc30 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:73)
#loc31 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:52)
#loc32 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:60)
#loc33 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:52)
#loc34 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":152:34)
#loc35 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":154:12)
#loc36 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":180:12)
#loc37 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":193:20)
#loc38 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":66:20)
#loc39 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":69:27)
#loc40 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":70:23)
#loc42 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":168:27)
#loc43 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:35)
#loc45 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":85:22)
#loc46 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":86:20)
#loc47 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":271:15)
#loc48 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:34)
#loc49 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:28)
#loc50 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:20)
#loc51 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:28)
#loc52 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":94:20)
#loc53 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":87:28)
#loc54 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":97:46)
#loc55 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:8)
#loc56 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:27)
#loc57 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:52)
#loc58 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":75:27)
#loc61 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:35)
#loc62 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":78:18)
#loc63 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":79:54)
#loc64 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:27)
#loc65 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:15)
#loc66 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":216:20)
#loc67 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:43)
#loc68 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":219:25)
#loc69 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:37)
#loc70 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:30)
#loc71 = loc(callsite(#loc1 at #loc2))
#loc72 = loc(callsite(#loc1 at #loc5))
#loc74 = loc(callsite(#loc8 at #loc2))
#loc75 = loc(callsite(#loc9 at #loc2))
#loc76 = loc(callsite(#loc11 at #loc2))
#loc77 = loc(callsite(#loc15 at #loc2))
#loc79 = loc(callsite(#loc18 at #loc5))
#loc80 = loc(callsite(#loc38 at #loc5))
#loc81 = loc(callsite(#loc39 at #loc5))
#loc82 = loc(callsite(#loc40 at #loc5))
#loc84 = loc(callsite(#loc42 at #loc6))
#loc85 = loc(callsite(#loc43 at #loc5))
#loc87 = loc(callsite(#loc45 at #loc5))
#loc88 = loc(callsite(#loc46 at #loc5))
#loc89 = loc(callsite(#loc47 at #loc16))
#loc90 = loc(callsite(#loc48 at #loc5))
#loc91 = loc(callsite(#loc49 at #loc5))
#loc92 = loc(callsite(#loc50 at #loc5))
#loc93 = loc(callsite(#loc51 at #loc5))
#loc94 = loc(callsite(#loc52 at #loc5))
#loc95 = loc(callsite(#loc53 at #loc5))
#loc96 = loc(callsite(#loc54 at #loc5))
#loc97 = loc(callsite(#loc15 at #loc5))
#loc98 = loc(callsite(#loc55 at #loc5))
#loc99 = loc(callsite(#loc56 at #loc2))
#loc100 = loc(callsite(#loc57 at #loc2))
#loc101 = loc(callsite(#loc38 at #loc2))
#loc102 = loc(callsite(#loc39 at #loc2))
#loc103 = loc(callsite(#loc40 at #loc2))
#loc104 = loc(callsite(#loc58 at #loc2))
#loc107 = loc(callsite(#loc61 at #loc2))
#loc108 = loc(callsite(#loc62 at #loc2))
#loc109 = loc(callsite(#loc63 at #loc2))
#loc111 = loc(callsite(#loc45 at #loc2))
#loc112 = loc(callsite(#loc46 at #loc2))
#loc113 = loc(callsite(#loc48 at #loc2))
#loc114 = loc(callsite(#loc49 at #loc2))
#loc115 = loc(callsite(#loc50 at #loc2))
#loc116 = loc(callsite(#loc51 at #loc2))
#loc117 = loc(callsite(#loc52 at #loc2))
#loc118 = loc(callsite(#loc53 at #loc2))
#loc119 = loc(callsite(#loc54 at #loc2))
#loc120 = loc(callsite(#loc55 at #loc2))
#loc123 = loc(callsite(#loc84 at #loc7))
#loc125 = loc(callsite(#loc89 at #loc17))
#loc127 = loc(callsite(#loc84 at #loc60))
#loc128 = loc(callsite(#loc123 at #loc5))
#loc129 = loc(callsite(#loc125 at #loc5))
#loc130 = loc(callsite(#loc127 at #loc2))
#loc131 = loc(callsite(#loc125 at #loc2))
likedislike
hujiajun成员
7月31日 评论:
// Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:31.306Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:31.307Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpaeq3vsbv/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpaeq3vsbv/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:31.307Z] RERUN
// [2026-07-31T14:28:39.030Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype12-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:39.030Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:39.030Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpsjg92szj/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpsjg92szj/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:39.030Z] RERUN
// [2026-07-31T14:28:48.421Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype12-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:48.421Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:48.421Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmptyz1abkc/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmptyz1abkc/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:48.421Z] RERUN
// [2026-07-31T14:28:56.144Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype12-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:56.144Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:56.144Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp0rg0uh8p/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp0rg0uh8p/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:56.144Z] RERUN
// [2026-07-31T14:28:59.399Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype12-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:59.399Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:59.400Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpft63b7xw/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpft63b7xw/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:59.400Z] FAILED

#loc = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)
#loc2 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":212:70)
#loc5 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":202:78)
#loc6 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":195:44)
#loc7 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:46)
#loc17 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":282:36)
#loc18 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":88:25)
#loc41 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":81:22)
#loc44 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":84:24)
#loc59 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:33)
#loc60 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:46)
#loc73 = loc(callsite(#loc6 at #loc7))
#loc78 = loc(callsite(#loc17 at #loc18))
#loc83 = loc(callsite(#loc41 at #loc5))
#loc86 = loc(callsite(#loc44 at #loc5))
#loc105 = loc(callsite(#loc59 at #loc2))
#loc106 = loc(callsite(#loc6 at #loc60))
#loc110 = loc(callsite(#loc44 at #loc2))
#loc121 = loc(callsite(#loc73 at #loc5))
#loc122 = loc(callsite(#loc78 at #loc2))
#loc124 = loc(callsite(#loc78 at #loc5))
#loc126 = loc(callsite(#loc106 at #loc2))
module attributes {hacc.target = #hacc.target<"Ascend950PR_9579">} {
  func.func @_attn_fwd(%arg0: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg1: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg2: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg3: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg4: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg5: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg6: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg7: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg8: memref<?xi32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg9: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg10: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg11: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg12: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg13: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg14: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg15: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)) attributes {SyncBlockLockArgIdx = 0 : i64, WorkspaceArgIdx = 1 : i64, global_kernel = "local", mix_mode = "mix", parallel_mode = "simd"} {
    %c8192 = arith.constant 8192 : index loc(#loc71)
    %c128 = arith.constant 128 : index loc(#loc71)
    %c1 = arith.constant 1 : index loc(#loc3)
    %c0 = arith.constant 0 : index loc(#loc4)
    %cst = arith.constant 1.000000e+00 : f32 loc(#loc72)
    %cst_0 = arith.constant 0xFF800000 : f32 loc(#loc121)
    %cst_1 = arith.constant -1.000000e+04 : f32 loc(#loc74)
    %cst_2 = arith.constant 5.000000e-01 : f32 loc(#loc75)
    %c0_i32 = arith.constant 0 : i32 loc(#loc71)
    %c1_i32 = arith.constant 1 : i32 loc(#loc76)
    %c64_i32 = arith.constant 64 : i32 loc(#loc11)
    %c8_i32 = arith.constant 8 : i32 loc(#loc12)
    %c8388608_i64 = arith.constant 8388608 : i64 loc(#loc13)
    %c1048576_i64 = arith.constant 1048576 : i64 loc(#loc14)
    %c128_i32 = arith.constant 128 : i32 loc(#loc77)
    %c8192_i32 = arith.constant 8192 : i32 loc(#loc16)
    %cst_3 = arith.constant 0.000000e+00 : f32 loc(#loc122)
    %0 = tensor.empty() : tensor<128x128xf32> loc(#loc79)
    %1 = linalg.fill ins(%cst_3 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
    %2 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc75)
    %3 = linalg.fill ins(%cst_1 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
    %4 = tensor.empty() : tensor<128xf32> loc(#loc121)
    %5 = linalg.fill ins(%cst_0 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc121)
    %6 = linalg.fill ins(%cst : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc72)
    %7 = arith.index_cast %arg13 : i32 to index loc(#loc20)
    %reinterpret_cast = memref.reinterpret_cast %arg8 to offset: [%7], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc21)
    %8 = memref.load %reinterpret_cast[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc22)
    %9 = arith.addi %7, %c1 : index loc(#loc3)
    %reinterpret_cast_4 = memref.reinterpret_cast %arg8 to offset: [%9], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc3)
    %10 = memref.load %reinterpret_cast_4[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc4)
    scf.for %arg16 = %8 to %10 step %c1_i32  : i32 {
      %11 = arith.divsi %arg16, %c64_i32 : i32 loc(#loc24)
      %12 = arith.remsi %arg16, %c64_i32 : i32 loc(#loc11)
      %13 = arith.divsi %11, %c8_i32 : i32 loc(#loc25)
      %14 = arith.remsi %11, %c8_i32 : i32 loc(#loc26)
      %15 = arith.muli %14, %arg9 : i32 loc(#loc27)
      %16 = arith.divsi %15, %c8_i32 : i32 loc(#loc12)
      %17 = arith.extsi %13 : i32 to i64 loc(#loc28)
      %18 = arith.muli %17, %c8388608_i64 : i64 loc(#loc13)
      %19 = arith.extsi %14 : i32 to i64 loc(#loc29)
      %20 = arith.muli %19, %c1048576_i64 : i64 loc(#loc30)
      %21 = arith.addi %18, %20 : i64 loc(#loc31)
      %22 = arith.extsi %16 : i32 to i64 loc(#loc32)
      %23 = arith.muli %22, %c1048576_i64 : i64 loc(#loc14)
      %24 = arith.addi %18, %23 : i64 loc(#loc33)
      %25 = arith.index_cast %21 : i64 to index loc(#loc31)
      %26 = arith.muli %12, %c128_i32 : i32 loc(#loc34)
      %27 = arith.maxsi %26, %c0_i32 : i32 loc(#loc35)
      %28 = arith.index_cast %27 : i32 to index loc(#loc35)
      %29 = arith.muli %28, %c128 : index loc(#loc35)
      %30 = arith.addi %29, %25 : index loc(#loc35)
      %reinterpret_cast_5 = memref.reinterpret_cast %arg2 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc35)
      %31 = arith.index_cast %24 : i64 to index loc(#loc33)
      %reinterpret_cast_6 = memref.reinterpret_cast %arg7 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc36)
      %alloc = memref.alloc() : memref<128x128xbf16> loc(#loc37)
      memref.copy %reinterpret_cast_5, %alloc : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc37)
      %32 = bufferization.to_tensor %alloc restrict writable : memref<128x128xbf16> loc(#loc37)
      %33:5 = scf.for %arg17 = %c0_i32 to %26 step %c128_i32 iter_args(%arg18 = %6, %arg19 = %1, %arg20 = %5, %arg21 = %c0_i32, %arg22 = %c0_i32) -> (tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32)  : i32 {
        %46 = arith.maxsi %arg21, %c0_i32 : i32 loc(#loc72)
        %47 = arith.index_cast %46 : i32 to index loc(#loc72)
        %48 = arith.muli %47, %c128 : index loc(#loc72)
        %49 = arith.addi %48, %31 : index loc(#loc72)
        %reinterpret_cast_8 = memref.reinterpret_cast %arg4 to offset: [%49], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc72)
        %50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc72)
        %51 = arith.index_cast %50 : i32 to index loc(#loc72)
        %52 = arith.muli %51, %c128 : index loc(#loc72)
        %53 = arith.addi %52, %31 : index loc(#loc72)
        %reinterpret_cast_9 = memref.reinterpret_cast %arg3 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc72)
        %alloc_10 = memref.alloc() : memref<128x128xbf16> loc(#loc80)
        memref.copy %reinterpret_cast_9, %alloc_10 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc80)
        %54 = bufferization.to_tensor %alloc_10 restrict writable : memref<128x128xbf16> loc(#loc80)
        %55 = tensor.empty() : tensor<128x128xbf16> loc(#loc81)
        %transposed = linalg.transpose ins(%54 : tensor<128x128xbf16>) outs(%55 : tensor<128x128xbf16>) permutation = [1, 0]  loc(#loc81)
        %56 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc82)
        %57 = arith.mulf %56, %2 : tensor<128x128xf32> loc(#loc83)
        %reduced = linalg.reduce ins(%57 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc41 at #loc5)), %init: f32 loc(callsite(#loc73 at #loc5))) {
            %72 = arith.maximumf %in, %init : f32 loc(#loc128)
            linalg.yield %72 : f32 loc(#loc121)
          } loc(#loc121)
        %58 = arith.maximumf %arg20, %reduced : tensor<128xf32> loc(#loc85)
        %broadcasted_11 = linalg.broadcast ins(%58 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc79)
        %59 = arith.subf %57, %broadcasted_11 : tensor<128x128xf32> loc(#loc79)
        %60 = math.exp %59 : tensor<128x128xf32> loc(#loc86)
        %61 = arith.truncf %60 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc87)
        %alloc_12 = memref.alloc() : memref<128x128xbf16> loc(#loc88)
        memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc88)
        %62 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xbf16> loc(#loc88)
        %63 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc124)
        %reduced_13 = linalg.reduce ins(%60 : tensor<128x128xf32>) outs(%63 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc44 at #loc5)), %init: f32 loc(callsite(#loc78 at #loc5))) {
            %72 = arith.addf %in, %init : f32 loc(#loc129)
            linalg.yield %72 : f32 loc(#loc124)
          } loc(#loc124)
        %64 = arith.subf %arg20, %58 : tensor<128xf32> loc(#loc90)
        %65 = math.exp %64 : tensor<128xf32> loc(#loc91)
        %66 = arith.mulf %arg18, %65 : tensor<128xf32> loc(#loc92)
        %67 = arith.addf %66, %reduced_13 : tensor<128xf32> loc(#loc93)
        %broadcasted_14 = linalg.broadcast ins(%65 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc94)
        %68 = arith.mulf %arg19, %broadcasted_14 : tensor<128x128xf32> loc(#loc94)
        %69 = linalg.matmul {input_precision = "ieee"} ins(%61, %62 : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%68 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc95)
        %70 = arith.addi %arg21, %c128_i32 : i32 loc(#loc96)
        %71 = arith.addi %arg22, %c128_i32 : i32 loc(#loc97)
        scf.yield %67, %69, %58, %70, %71 : tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc98)
      } {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc72)
      %34 = arith.muli %12, %c128_i32 {tt.divisibility = dense<128> : tensor<1xi32>} : i32 loc(#loc99)
      %35 = arith.addi %12, %c1_i32 : i32 loc(#loc76)
      %36 = arith.muli %35, %c128_i32 : i32 loc(#loc100)
      %37:6 = scf.for %arg17 = %34 to %36 step %c128_i32 iter_args(%arg18 = %34, %arg19 = %33#0, %arg20 = %33#1, %arg21 = %33#2, %arg22 = %34, %arg23 = %34) -> (i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32)  : i32 {
        %46 = arith.maxsi %arg18, %c0_i32 : i32 loc(#loc71)
        %47 = arith.index_cast %46 : i32 to index loc(#loc71)
        %48 = arith.muli %28, %c8192 : index loc(#loc71)
        %49 = arith.addi %48, %47 : index loc(#loc71)
        %reinterpret_cast_8 = memref.reinterpret_cast %arg5 to offset: [%49], sizes: [128, 128], strides: [8192, 1] : memref<?xf32> to memref<128x128xf32, strided<[8192, 1], offset: ?>> loc(#loc71)
        %50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc71)
        %51 = arith.index_cast %50 : i32 to index loc(#loc71)
        %52 = arith.muli %51, %c128 : index loc(#loc71)
        %53 = arith.addi %52, %31 : index loc(#loc71)
        %reinterpret_cast_9 = memref.reinterpret_cast %arg4 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc71)
        %54 = arith.maxsi %arg23, %c0_i32 : i32 loc(#loc71)
        %55 = arith.index_cast %54 : i32 to index loc(#loc71)
        %56 = arith.muli %55, %c128 : index loc(#loc71)
        %57 = arith.addi %56, %31 : index loc(#loc71)
        %reinterpret_cast_10 = memref.reinterpret_cast %arg3 to offset: [%57], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc71)
        %alloc_11 = memref.alloc() : memref<128x128xbf16> loc(#loc101)
        memref.copy %reinterpret_cast_10, %alloc_11 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc101)
        %58 = bufferization.to_tensor %alloc_11 restrict writable : memref<128x128xbf16> loc(#loc101)
        %59 = tensor.empty() : tensor<128x128xbf16> loc(#loc102)
        %transposed = linalg.transpose ins(%58 : tensor<128x128xbf16>) outs(%59 : tensor<128x128xbf16>) permutation = [1, 0]  loc(#loc102)
        %60 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc103)
        %alloc_12 = memref.alloc() : memref<128x128xf32> loc(#loc104)
        memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf32, strided<[8192, 1], offset: ?>> to memref<128x128xf32> loc(#loc104)
        %61 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf32> loc(#loc104)
        %62 = arith.mulf %60, %2 : tensor<128x128xf32> loc(#loc75)
        %63 = arith.cmpf une, %61, %1 : tensor<128x128xf32> loc(#loc74)
        %64 = arith.select %63, %3, %1 : tensor<128x128xi1>, tensor<128x128xf32> loc(#loc74)
        %65 = arith.addf %62, %64 : tensor<128x128xf32> loc(#loc105)
        %reduced = linalg.reduce ins(%65 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc59 at #loc2)), %init: f32 loc(callsite(#loc106 at #loc2))) {
            %81 = arith.maximumf %in, %init : f32 loc(#loc130)
            linalg.yield %81 : f32 loc(#loc126)
          } loc(#loc126)
        %66 = arith.maximumf %arg21, %reduced : tensor<128xf32> loc(#loc107)
        %broadcasted_13 = linalg.broadcast ins(%66 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc108)
        %67 = arith.subf %65, %broadcasted_13 : tensor<128x128xf32> loc(#loc108)
        %68 = arith.addi %arg18, %c128_i32 : i32 loc(#loc109)
        %69 = math.exp %67 : tensor<128x128xf32> loc(#loc110)
        %70 = arith.truncf %69 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc111)
        %alloc_14 = memref.alloc() : memref<128x128xbf16> loc(#loc112)
        memref.copy %reinterpret_cast_9, %alloc_14 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc112)
        %71 = bufferization.to_tensor %alloc_14 restrict writable : memref<128x128xbf16> loc(#loc112)
        %72 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc122)
        %reduced_15 = linalg.reduce ins(%69 : tensor<128x128xf32>) outs(%72 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc44 at #loc2)), %init: f32 loc(callsite(#loc78 at #loc2))) {
            %81 = arith.addf %in, %init : f32 loc(#loc131)
            linalg.yield %81 : f32 loc(#loc122)
          } loc(#loc122)
        %73 = arith.subf %arg21, %66 : tensor<128xf32> loc(#loc113)
        %74 = math.exp %73 : tensor<128xf32> loc(#loc114)
        %75 = arith.mulf %arg19, %74 : tensor<128xf32> loc(#loc115)
        %76 = arith.addf %75, %reduced_15 : tensor<128xf32> loc(#loc116)
        %broadcasted_16 = linalg.broadcast ins(%74 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc117)
        %77 = arith.mulf %arg20, %broadcasted_16 : tensor<128x128xf32> loc(#loc117)
        %78 = linalg.matmul {input_precision = "ieee"} ins(%70, %71 : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%77 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc118)
        %79 = arith.addi %arg22, %c128_i32 : i32 loc(#loc119)
        %80 = arith.addi %arg23, %c128_i32 : i32 loc(#loc77)
        scf.yield %68, %76, %78, %66, %79, %80 : i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc120)
      } {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc71)
      %38 = math.log %37#1 : tensor<128xf32> loc(#loc64)
      %39 = arith.addf %37#3, %38 : tensor<128xf32> loc(#loc65)
      %broadcasted = linalg.broadcast ins(%37#1 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc66)
      %40 = arith.divf %37#2, %broadcasted : tensor<128x128xf32> loc(#loc66)
      %41 = arith.muli %11, %c8192_i32 : i32 loc(#loc16)
      %42 = arith.index_cast %41 : i32 to index loc(#loc16)
      %43 = arith.index_cast %26 : i32 to index loc(#loc34)
      %44 = arith.addi %42, %43 : index loc(#loc67)
      %reinterpret_cast_7 = memref.reinterpret_cast %arg6 to offset: [%44], sizes: [128], strides: [1] : memref<?xf32> to memref<128xf32, strided<[1], offset: ?>> loc(#loc67)
      bufferization.materialize_in_destination %39 in writable %reinterpret_cast_7 : (tensor<128xf32>, memref<128xf32, strided<[1], offset: ?>>) -> () loc(#loc68)
      %45 = arith.truncf %40 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc69)
      bufferization.materialize_in_destination %45 in writable %reinterpret_cast_6 : (tensor<128x128xbf16>, memref<128x128xbf16, strided<[128, 1], offset: ?>>) -> () loc(#loc70)
    } loc(#loc23)
    return loc(#loc)
  } loc(#loc)
} loc(#loc)
#loc1 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":63:33)
#loc3 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:53)
#loc4 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:28)
#loc8 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:56)
#loc9 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:22)
#loc10 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:47)
#loc11 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":139:33)
#loc12 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:39)
#loc13 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:40)
#loc14 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:72)
#loc15 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:46)
#loc16 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:35)
#loc19 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":83:22)
#loc20 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":123:24)
#loc21 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:49)
#loc22 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:30)
#loc23 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":137:51)
#loc24 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":138:35)
#loc25 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":140:31)
#loc26 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":141:30)
#loc27 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:23)
#loc28 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:28)
#loc29 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:61)
#loc30 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:73)
#loc31 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:52)
#loc32 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:60)
#loc33 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:52)
#loc34 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":152:34)
#loc35 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":154:12)
#loc36 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":180:12)
#loc37 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":193:20)
#loc38 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":66:20)
#loc39 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":69:27)
#loc40 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":70:23)
#loc42 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":168:27)
#loc43 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:35)
#loc45 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":85:22)
#loc46 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":86:20)
#loc47 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":271:15)
#loc48 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:34)
#loc49 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:28)
#loc50 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:20)
#loc51 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:28)
#loc52 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":94:20)
#loc53 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":87:28)
#loc54 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":97:46)
#loc55 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:8)
#loc56 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:27)
#loc57 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:52)
#loc58 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":75:27)
#loc61 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:35)
#loc62 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":78:18)
#loc63 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":79:54)
#loc64 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:27)
#loc65 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:15)
#loc66 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":216:20)
#loc67 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:43)
#loc68 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":219:25)
#loc69 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:37)
#loc70 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:30)
#loc71 = loc(callsite(#loc1 at #loc2))
#loc72 = loc(callsite(#loc1 at #loc5))
#loc74 = loc(callsite(#loc8 at #loc2))
#loc75 = loc(callsite(#loc9 at #loc2))
#loc76 = loc(callsite(#loc10 at #loc2))
#loc77 = loc(callsite(#loc15 at #loc2))
#loc79 = loc(callsite(#loc19 at #loc5))
#loc80 = loc(callsite(#loc38 at #loc5))
#loc81 = loc(callsite(#loc39 at #loc5))
#loc82 = loc(callsite(#loc40 at #loc5))
#loc84 = loc(callsite(#loc42 at #loc6))
#loc85 = loc(callsite(#loc43 at #loc5))
#loc87 = loc(callsite(#loc45 at #loc5))
#loc88 = loc(callsite(#loc46 at #loc5))
#loc89 = loc(callsite(#loc47 at #loc17))
#loc90 = loc(callsite(#loc48 at #loc5))
#loc91 = loc(callsite(#loc49 at #loc5))
#loc92 = loc(callsite(#loc50 at #loc5))
#loc93 = loc(callsite(#loc51 at #loc5))
#loc94 = loc(callsite(#loc52 at #loc5))
#loc95 = loc(callsite(#loc53 at #loc5))
#loc96 = loc(callsite(#loc54 at #loc5))
#loc97 = loc(callsite(#loc15 at #loc5))
#loc98 = loc(callsite(#loc55 at #loc5))
#loc99 = loc(callsite(#loc56 at #loc2))
#loc100 = loc(callsite(#loc57 at #loc2))
#loc101 = loc(callsite(#loc38 at #loc2))
#loc102 = loc(callsite(#loc39 at #loc2))
#loc103 = loc(callsite(#loc40 at #loc2))
#loc104 = loc(callsite(#loc58 at #loc2))
#loc107 = loc(callsite(#loc61 at #loc2))
#loc108 = loc(callsite(#loc62 at #loc2))
#loc109 = loc(callsite(#loc63 at #loc2))
#loc111 = loc(callsite(#loc45 at #loc2))
#loc112 = loc(callsite(#loc46 at #loc2))
#loc113 = loc(callsite(#loc48 at #loc2))
#loc114 = loc(callsite(#loc49 at #loc2))
#loc115 = loc(callsite(#loc50 at #loc2))
#loc116 = loc(callsite(#loc51 at #loc2))
#loc117 = loc(callsite(#loc52 at #loc2))
#loc118 = loc(callsite(#loc53 at #loc2))
#loc119 = loc(callsite(#loc54 at #loc2))
#loc120 = loc(callsite(#loc55 at #loc2))
#loc123 = loc(callsite(#loc84 at #loc7))
#loc125 = loc(callsite(#loc89 at #loc18))
#loc127 = loc(callsite(#loc84 at #loc60))
#loc128 = loc(callsite(#loc123 at #loc5))
#loc129 = loc(callsite(#loc125 at #loc5))
#loc130 = loc(callsite(#loc127 at #loc2))
#loc131 = loc(callsite(#loc125 at #loc2))
likedislike
hujiajun成员
7月31日 评论:
// [2026-07-31T14:29:18.826Z] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:18.826Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:18.826Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpg2gwakv8/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpg2gwakv8/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:18.826Z] RERUN
// [2026-07-31T14:29:26.544Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype14-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:26.544Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:26.544Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpe5zws9b9/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpe5zws9b9/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:26.544Z] RERUN
// [2026-07-31T14:29:35.968Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype14-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:35.968Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:35.968Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmphdpje0xa/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmphdpje0xa/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:35.968Z] RERUN
// [2026-07-31T14:29:43.690Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype14-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:43.690Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:43.690Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp9dcmatiq/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp9dcmatiq/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:43.690Z] RERUN
// [2026-07-31T14:29:46.205Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype14-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:46.206Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:46.206Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpgsxkhje7/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpgsxkhje7/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:46.206Z] FAILED

#loc = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)
#loc2 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":212:70)
#loc5 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":202:78)
#loc6 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":195:44)
#loc7 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:46)
#loc16 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":282:36)
#loc17 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":88:25)
#loc41 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":81:22)
#loc44 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":84:24)
#loc59 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:33)
#loc60 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:46)
#loc73 = loc(callsite(#loc6 at #loc7))
#loc78 = loc(callsite(#loc16 at #loc17))
#loc83 = loc(callsite(#loc41 at #loc5))
#loc86 = loc(callsite(#loc44 at #loc5))
#loc105 = loc(callsite(#loc59 at #loc2))
#loc106 = loc(callsite(#loc6 at #loc60))
#loc110 = loc(callsite(#loc44 at #loc2))
#loc121 = loc(callsite(#loc73 at #loc5))
#loc122 = loc(callsite(#loc78 at #loc2))
#loc124 = loc(callsite(#loc78 at #loc5))
#loc126 = loc(callsite(#loc106 at #loc2))
module attributes {hacc.target = #hacc.target<"Ascend950PR_9579">} {
  func.func @_attn_fwd(%arg0: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg1: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg2: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg3: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg4: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg5: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg6: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg7: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg8: memref<?xi32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg9: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg10: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg11: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg12: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg13: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg14: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg15: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)) attributes {SyncBlockLockArgIdx = 0 : i64, WorkspaceArgIdx = 1 : i64, global_kernel = "local", mix_mode = "mix", parallel_mode = "simd"} {
    %c1024 = arith.constant 1024 : index loc(#loc71)
    %c128 = arith.constant 128 : index loc(#loc71)
    %c1 = arith.constant 1 : index loc(#loc3)
    %c0 = arith.constant 0 : index loc(#loc4)
    %cst = arith.constant 1.000000e+00 : f32 loc(#loc72)
    %cst_0 = arith.constant 0xFF800000 : f32 loc(#loc121)
    %cst_1 = arith.constant -1.000000e+04 : f32 loc(#loc74)
    %cst_2 = arith.constant 5.000000e-01 : f32 loc(#loc75)
    %c1024_i32 = arith.constant 1024 : i32 loc(#loc10)
    %c0_i32 = arith.constant 0 : i32 loc(#loc71)
    %c1_i32 = arith.constant 1 : i32 loc(#loc76)
    %c8_i32 = arith.constant 8 : i32 loc(#loc12)
    %c1048576_i64 = arith.constant 1048576 : i64 loc(#loc13)
    %c131072_i64 = arith.constant 131072 : i64 loc(#loc14)
    %c128_i32 = arith.constant 128 : i32 loc(#loc77)
    %cst_3 = arith.constant 0.000000e+00 : f32 loc(#loc122)
    %0 = tensor.empty() : tensor<128x128xf32> loc(#loc79)
    %1 = linalg.fill ins(%cst_3 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
    %2 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc75)
    %3 = linalg.fill ins(%cst_1 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
    %4 = tensor.empty() : tensor<128xf32> loc(#loc121)
    %5 = linalg.fill ins(%cst_0 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc121)
    %6 = linalg.fill ins(%cst : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc72)
    %7 = arith.index_cast %arg13 : i32 to index loc(#loc19)
    %reinterpret_cast = memref.reinterpret_cast %arg8 to offset: [%7], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc20)
    %8 = memref.load %reinterpret_cast[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc21)
    %9 = arith.addi %7, %c1 : index loc(#loc3)
    %reinterpret_cast_4 = memref.reinterpret_cast %arg8 to offset: [%9], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc3)
    %10 = memref.load %reinterpret_cast_4[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc4)
    scf.for %arg16 = %8 to %10 step %c1_i32  : i32 {
      %11 = arith.divsi %arg16, %c8_i32 : i32 loc(#loc23)
      %12 = arith.remsi %arg16, %c8_i32 : i32 loc(#loc24)
      %13 = arith.divsi %11, %c8_i32 : i32 loc(#loc25)
      %14 = arith.remsi %11, %c8_i32 : i32 loc(#loc26)
      %15 = arith.muli %14, %arg9 : i32 loc(#loc27)
      %16 = arith.divsi %15, %c8_i32 : i32 loc(#loc12)
      %17 = arith.extsi %13 : i32 to i64 loc(#loc28)
      %18 = arith.muli %17, %c1048576_i64 : i64 loc(#loc13)
      %19 = arith.extsi %14 : i32 to i64 loc(#loc29)
      %20 = arith.muli %19, %c131072_i64 : i64 loc(#loc30)
      %21 = arith.addi %18, %20 : i64 loc(#loc31)
      %22 = arith.extsi %16 : i32 to i64 loc(#loc32)
      %23 = arith.muli %22, %c131072_i64 : i64 loc(#loc14)
      %24 = arith.addi %18, %23 : i64 loc(#loc33)
      %25 = arith.index_cast %21 : i64 to index loc(#loc31)
      %26 = arith.muli %12, %c128_i32 : i32 loc(#loc34)
      %27 = arith.maxsi %26, %c0_i32 : i32 loc(#loc35)
      %28 = arith.index_cast %27 : i32 to index loc(#loc35)
      %29 = arith.muli %28, %c128 : index loc(#loc35)
      %30 = arith.addi %29, %25 : index loc(#loc35)
      %reinterpret_cast_5 = memref.reinterpret_cast %arg2 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc35)
      %31 = arith.index_cast %24 : i64 to index loc(#loc33)
      %reinterpret_cast_6 = memref.reinterpret_cast %arg7 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc36)
      %alloc = memref.alloc() : memref<128x128xbf16> loc(#loc37)
      memref.copy %reinterpret_cast_5, %alloc : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc37)
      %32 = bufferization.to_tensor %alloc restrict writable : memref<128x128xbf16> loc(#loc37)
      %33:5 = scf.for %arg17 = %c0_i32 to %26 step %c128_i32 iter_args(%arg18 = %6, %arg19 = %1, %arg20 = %5, %arg21 = %c0_i32, %arg22 = %c0_i32) -> (tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32)  : i32 {
        %46 = arith.maxsi %arg21, %c0_i32 : i32 loc(#loc72)
        %47 = arith.index_cast %46 : i32 to index loc(#loc72)
        %48 = arith.muli %47, %c128 : index loc(#loc72)
        %49 = arith.addi %48, %31 : index loc(#loc72)
        %reinterpret_cast_8 = memref.reinterpret_cast %arg4 to offset: [%49], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc72)
        %50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc72)
        %51 = arith.index_cast %50 : i32 to index loc(#loc72)
        %52 = arith.muli %51, %c128 : index loc(#loc72)
        %53 = arith.addi %52, %31 : index loc(#loc72)
        %reinterpret_cast_9 = memref.reinterpret_cast %arg3 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc72)
        %alloc_10 = memref.alloc() : memref<128x128xbf16> loc(#loc80)
        memref.copy %reinterpret_cast_9, %alloc_10 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc80)
        %54 = bufferization.to_tensor %alloc_10 restrict writable : memref<128x128xbf16> loc(#loc80)
        %55 = tensor.empty() : tensor<128x128xbf16> loc(#loc81)
        %transposed = linalg.transpose ins(%54 : tensor<128x128xbf16>) outs(%55 : tensor<128x128xbf16>) permutation = [1, 0]  loc(#loc81)
        %56 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc82)
        %57 = arith.mulf %56, %2 : tensor<128x128xf32> loc(#loc83)
        %reduced = linalg.reduce ins(%57 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc41 at #loc5)), %init: f32 loc(callsite(#loc73 at #loc5))) {
            %72 = arith.maximumf %in, %init : f32 loc(#loc128)
            linalg.yield %72 : f32 loc(#loc121)
          } loc(#loc121)
        %58 = arith.maximumf %arg20, %reduced : tensor<128xf32> loc(#loc85)
        %broadcasted_11 = linalg.broadcast ins(%58 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc79)
        %59 = arith.subf %57, %broadcasted_11 : tensor<128x128xf32> loc(#loc79)
        %60 = math.exp %59 : tensor<128x128xf32> loc(#loc86)
        %61 = arith.truncf %60 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc87)
        %alloc_12 = memref.alloc() : memref<128x128xbf16> loc(#loc88)
        memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc88)
        %62 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xbf16> loc(#loc88)
        %63 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc124)
        %reduced_13 = linalg.reduce ins(%60 : tensor<128x128xf32>) outs(%63 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc44 at #loc5)), %init: f32 loc(callsite(#loc78 at #loc5))) {
            %72 = arith.addf %in, %init : f32 loc(#loc129)
            linalg.yield %72 : f32 loc(#loc124)
          } loc(#loc124)
        %64 = arith.subf %arg20, %58 : tensor<128xf32> loc(#loc90)
        %65 = math.exp %64 : tensor<128xf32> loc(#loc91)
        %66 = arith.mulf %arg18, %65 : tensor<128xf32> loc(#loc92)
        %67 = arith.addf %66, %reduced_13 : tensor<128xf32> loc(#loc93)
        %broadcasted_14 = linalg.broadcast ins(%65 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc94)
        %68 = arith.mulf %arg19, %broadcasted_14 : tensor<128x128xf32> loc(#loc94)
        %69 = linalg.matmul {input_precision = "ieee"} ins(%61, %62 : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%68 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc95)
        %70 = arith.addi %arg21, %c128_i32 : i32 loc(#loc96)
        %71 = arith.addi %arg22, %c128_i32 : i32 loc(#loc97)
        scf.yield %67, %69, %58, %70, %71 : tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc98)
      } {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc72)
      %34 = arith.muli %12, %c128_i32 {tt.divisibility = dense<128> : tensor<1xi32>} : i32 loc(#loc99)
      %35 = arith.addi %12, %c1_i32 : i32 loc(#loc76)
      %36 = arith.muli %35, %c128_i32 : i32 loc(#loc100)
      %37:6 = scf.for %arg17 = %34 to %36 step %c128_i32 iter_args(%arg18 = %34, %arg19 = %33#0, %arg20 = %33#1, %arg21 = %33#2, %arg22 = %34, %arg23 = %34) -> (i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32)  : i32 {
        %46 = arith.maxsi %arg18, %c0_i32 : i32 loc(#loc71)
        %47 = arith.index_cast %46 : i32 to index loc(#loc71)
        %48 = arith.muli %28, %c1024 : index loc(#loc71)
        %49 = arith.addi %48, %47 : index loc(#loc71)
        %reinterpret_cast_8 = memref.reinterpret_cast %arg5 to offset: [%49], sizes: [128, 128], strides: [1024, 1] : memref<?xf32> to memref<128x128xf32, strided<[1024, 1], offset: ?>> loc(#loc71)
        %50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc71)
        %51 = arith.index_cast %50 : i32 to index loc(#loc71)
        %52 = arith.muli %51, %c128 : index loc(#loc71)
        %53 = arith.addi %52, %31 : index loc(#loc71)
        %reinterpret_cast_9 = memref.reinterpret_cast %arg4 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc71)
        %54 = arith.maxsi %arg23, %c0_i32 : i32 loc(#loc71)
        %55 = arith.index_cast %54 : i32 to index loc(#loc71)
        %56 = arith.muli %55, %c128 : index loc(#loc71)
        %57 = arith.addi %56, %31 : index loc(#loc71)
        %reinterpret_cast_10 = memref.reinterpret_cast %arg3 to offset: [%57], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc71)
        %alloc_11 = memref.alloc() : memref<128x128xbf16> loc(#loc101)
        memref.copy %reinterpret_cast_10, %alloc_11 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc101)
        %58 = bufferization.to_tensor %alloc_11 restrict writable : memref<128x128xbf16> loc(#loc101)
        %59 = tensor.empty() : tensor<128x128xbf16> loc(#loc102)
        %transposed = linalg.transpose ins(%58 : tensor<128x128xbf16>) outs(%59 : tensor<128x128xbf16>) permutation = [1, 0]  loc(#loc102)
        %60 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc103)
        %alloc_12 = memref.alloc() : memref<128x128xf32> loc(#loc104)
        memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf32, strided<[1024, 1], offset: ?>> to memref<128x128xf32> loc(#loc104)
        %61 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf32> loc(#loc104)
        %62 = arith.mulf %60, %2 : tensor<128x128xf32> loc(#loc75)
        %63 = arith.cmpf une, %61, %1 : tensor<128x128xf32> loc(#loc74)
        %64 = arith.select %63, %3, %1 : tensor<128x128xi1>, tensor<128x128xf32> loc(#loc74)
        %65 = arith.addf %62, %64 : tensor<128x128xf32> loc(#loc105)
        %reduced = linalg.reduce ins(%65 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc59 at #loc2)), %init: f32 loc(callsite(#loc106 at #loc2))) {
            %81 = arith.maximumf %in, %init : f32 loc(#loc130)
            linalg.yield %81 : f32 loc(#loc126)
          } loc(#loc126)
        %66 = arith.maximumf %arg21, %reduced : tensor<128xf32> loc(#loc107)
        %broadcasted_13 = linalg.broadcast ins(%66 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc108)
        %67 = arith.subf %65, %broadcasted_13 : tensor<128x128xf32> loc(#loc108)
        %68 = arith.addi %arg18, %c128_i32 : i32 loc(#loc109)
        %69 = math.exp %67 : tensor<128x128xf32> loc(#loc110)
        %70 = arith.truncf %69 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc111)
        %alloc_14 = memref.alloc() : memref<128x128xbf16> loc(#loc112)
        memref.copy %reinterpret_cast_9, %alloc_14 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc112)
        %71 = bufferization.to_tensor %alloc_14 restrict writable : memref<128x128xbf16> loc(#loc112)
        %72 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc122)
        %reduced_15 = linalg.reduce ins(%69 : tensor<128x128xf32>) outs(%72 : tensor<128xf32>) dimensions = [1] 
          (%in: f32 loc(callsite(#loc44 at #loc2)), %init: f32 loc(callsite(#loc78 at #loc2))) {
            %81 = arith.addf %in, %init : f32 loc(#loc131)
            linalg.yield %81 : f32 loc(#loc122)
          } loc(#loc122)
        %73 = arith.subf %arg21, %66 : tensor<128xf32> loc(#loc113)
        %74 = math.exp %73 : tensor<128xf32> loc(#loc114)
        %75 = arith.mulf %arg19, %74 : tensor<128xf32> loc(#loc115)
        %76 = arith.addf %75, %reduced_15 : tensor<128xf32> loc(#loc116)
        %broadcasted_16 = linalg.broadcast ins(%74 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc117)
        %77 = arith.mulf %arg20, %broadcasted_16 : tensor<128x128xf32> loc(#loc117)
        %78 = linalg.matmul {input_precision = "ieee"} ins(%70, %71 : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%77 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc118)
        %79 = arith.addi %arg22, %c128_i32 : i32 loc(#loc119)
        %80 = arith.addi %arg23, %c128_i32 : i32 loc(#loc77)
        scf.yield %68, %76, %78, %66, %79, %80 : i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc120)
      } {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc71)
      %38 = math.log %37#1 : tensor<128xf32> loc(#loc64)
      %39 = arith.addf %37#3, %38 : tensor<128xf32> loc(#loc65)
      %broadcasted = linalg.broadcast ins(%37#1 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1]  loc(#loc66)
      %40 = arith.divf %37#2, %broadcasted : tensor<128x128xf32> loc(#loc66)
      %41 = arith.muli %11, %c1024_i32 : i32 loc(#loc10)
      %42 = arith.index_cast %41 : i32 to index loc(#loc10)
      %43 = arith.index_cast %26 : i32 to index loc(#loc34)
      %44 = arith.addi %42, %43 : index loc(#loc67)
      %reinterpret_cast_7 = memref.reinterpret_cast %arg6 to offset: [%44], sizes: [128], strides: [1] : memref<?xf32> to memref<128xf32, strided<[1], offset: ?>> loc(#loc67)
      bufferization.materialize_in_destination %39 in writable %reinterpret_cast_7 : (tensor<128xf32>, memref<128xf32, strided<[1], offset: ?>>) -> () loc(#loc68)
      %45 = arith.truncf %40 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc69)
      bufferization.materialize_in_destination %45 in writable %reinterpret_cast_6 : (tensor<128x128xbf16>, memref<128x128xbf16, strided<[128, 1], offset: ?>>) -> () loc(#loc70)
    } loc(#loc22)
    return loc(#loc)
  } loc(#loc)
} loc(#loc)
#loc1 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":63:33)
#loc3 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:53)
#loc4 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:28)
#loc8 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:56)
#loc9 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:22)
#loc10 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:35)
#loc11 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:47)
#loc12 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:39)
#loc13 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:40)
#loc14 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:72)
#loc15 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:46)
#loc18 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":83:22)
#loc19 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":123:24)
#loc20 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:49)
#loc21 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:30)
#loc22 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":137:51)
#loc23 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":138:35)
#loc24 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":139:33)
#loc25 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":140:31)
#loc26 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":141:30)
#loc27 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:23)
#loc28 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:28)
#loc29 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:61)
#loc30 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:73)
#loc31 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:52)
#loc32 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:60)
#loc33 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:52)
#loc34 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":152:34)
#loc35 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":154:12)
#loc36 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":180:12)
#loc37 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":193:20)
#loc38 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":66:20)
#loc39 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":69:27)
#loc40 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":70:23)
#loc42 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":168:27)
#loc43 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:35)
#loc45 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":85:22)
#loc46 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":86:20)
#loc47 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":271:15)
#loc48 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:34)
#loc49 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:28)
#loc50 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:20)
#loc51 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:28)
#loc52 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":94:20)
#loc53 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":87:28)
#loc54 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":97:46)
#loc55 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:8)
#loc56 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:27)
#loc57 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:52)
#loc58 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":75:27)
#loc61 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:35)
#loc62 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":78:18)
#loc63 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":79:54)
#loc64 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:27)
#loc65 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:15)
#loc66 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":216:20)
#loc67 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:43)
#loc68 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":219:25)
#loc69 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:37)
#loc70 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:30)
#loc71 = loc(callsite(#loc1 at #loc2))
#loc72 = loc(callsite(#loc1 at #loc5))
#loc74 = loc(callsite(#loc8 at #loc2))
#loc75 = loc(callsite(#loc9 at #loc2))
#loc76 = loc(callsite(#loc11 at #loc2))
#loc77 = loc(callsite(#loc15 at #loc2))
#loc79 = loc(callsite(#loc18 at #loc5))
#loc80 = loc(callsite(#loc38 at #loc5))
#loc81 = loc(callsite(#loc39 at #loc5))
#loc82 = loc(callsite(#loc40 at #loc5))
#loc84 = loc(callsite(#loc42 at #loc6))
#loc85 = loc(callsite(#loc43 at #loc5))
#loc87 = loc(callsite(#loc45 at #loc5))
#loc88 = loc(callsite(#loc46 at #loc5))
#loc89 = loc(callsite(#loc47 at #loc16))
#loc90 = loc(callsite(#loc48 at #loc5))
#loc91 = loc(callsite(#loc49 at #loc5))
#loc92 = loc(callsite(#loc50 at #loc5))
#loc93 = loc(callsite(#loc51 at #loc5))
#loc94 = loc(callsite(#loc52 at #loc5))
#loc95 = loc(callsite(#loc53 at #loc5))
#loc96 = loc(callsite(#loc54 at #loc5))
#loc97 = loc(callsite(#loc15 at #loc5))
#loc98 = loc(callsite(#loc55 at #loc5))
#loc99 = loc(callsite(#loc56 at #loc2))
#loc100 = loc(callsite(#loc57 at #loc2))
#loc101 = loc(callsite(#loc38 at #loc2))
#loc102 = loc(callsite(#loc39 at #loc2))
#loc103 = loc(callsite(#loc40 at #loc2))
#loc104 = loc(callsite(#loc58 at #loc2))
#loc107 = loc(callsite(#loc61 at #loc2))
#loc108 = loc(callsite(#loc62 at #loc2))
#loc109 = loc(callsite(#loc63 at #loc2))
#loc111 = loc(callsite(#loc45 at #loc2))
#loc112 = loc(callsite(#loc46 at #loc2))
#loc113 = loc(callsite(#loc48 at #loc2))
#loc114 = loc(callsite(#loc49 at #loc2))
#loc115 = loc(callsite(#loc50 at #loc2))
#loc116 = loc(callsite(#loc51 at #loc2))
#loc117 = loc(callsite(#loc52 at #loc2))
#loc118 = loc(callsite(#loc53 at #loc2))
#loc119 = loc(callsite(#loc54 at #loc2))
#loc120 = loc(callsite(#loc55 at #loc2))
#loc123 = loc(callsite(#loc84 at #loc7))
#loc125 = loc(callsite(#loc89 at #loc17))
#loc127 = loc(callsite(#loc84 at #loc60))
#loc128 = loc(callsite(#loc123 at #loc5))
#loc129 = loc(callsite(#loc125 at #loc5))
#loc130 = loc(callsite(#loc127 at #loc2))
#loc131 = loc(callsite(#loc125 at #loc2))
likedislike
SSL25成员
8月1日 关联了看板:AscendNPU IR项目
SL25成员
8月1日 评论:

/label add triaged

likedislike
ascend-robotascend-robot成员
8月1日 添加了label:triaged
Hhujiajun成员
8月2日 修改了issue 的描述
Hhujiajun成员
8月2日 修改标题为 “[bug][auto-vectorize-v2][multi-consumer] 开启多 consumer 门禁 CV Native FA 编译错误”,原标题为“开启多 consumer 门禁 CV Native FA 编译错误”
ascend-robotascend-robot成员
8月2日 添加了label:bug
Hhujiajun成员
8月2日 修改了issue 的描述
Hhujiajun成员
8月6日 关联了pull request:feat: promote Hu_JJN to be reviewr of ascendnpu-ir repo