已开启
[bug][auto-vectorize-v2][multi-consumer] 开启多 consumer 门禁 CV Native FA 编译错误 #381
hujiajun创建于 7月31日
hujiajun
7月31日 评论:
7月31日 评论:
// Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:25:44.061Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:25:44.061Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpoy18bk1i/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpoy18bk1i/kernel --enable-vf-merge-level=1
// [2026-07-31T14:25:44.061Z] RERUN
// [2026-07-31T14:25:51.784Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype4-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:25:51.784Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:25:51.784Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp9i5qdfr4/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp9i5qdfr4/kernel --enable-vf-merge-level=1
// [2026-07-31T14:25:51.784Z] RERUN
// [2026-07-31T14:26:01.172Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype4-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:26:01.172Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:26:01.173Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp_qnqchlz/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp_qnqchlz/kernel --enable-vf-merge-level=1
// [2026-07-31T14:26:01.173Z] RERUN
// [2026-07-31T14:26:08.893Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype4-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:26:08.893Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:26:08.893Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpcm8ob1m3/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpcm8ob1m3/kernel --enable-vf-merge-level=1
// [2026-07-31T14:26:08.893Z] RERUN
// [2026-07-31T14:26:12.150Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype4-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/BnCfqAhYPDyUVqsglHTZmhEyCf_quyZ3a0UrunUwOiw
// [2026-07-31T14:26:12.150Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:26:12.150Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmplbzh40ta/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmplbzh40ta/kernel --enable-vf-merge-level=1
// [2026-07-31T14:26:12.150Z] FAILED
#loc = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)
#loc2 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":212:70)
#loc5 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":202:78)
#loc6 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":195:44)
#loc7 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:46)
#loc17 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":282:36)
#loc18 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":88:25)
#loc41 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":81:22)
#loc44 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":84:24)
#loc59 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:33)
#loc60 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:46)
#loc73 = loc(callsite(#loc6 at #loc7))
#loc78 = loc(callsite(#loc17 at #loc18))
#loc83 = loc(callsite(#loc41 at #loc5))
#loc86 = loc(callsite(#loc44 at #loc5))
#loc105 = loc(callsite(#loc59 at #loc2))
#loc106 = loc(callsite(#loc6 at #loc60))
#loc110 = loc(callsite(#loc44 at #loc2))
#loc121 = loc(callsite(#loc73 at #loc5))
#loc122 = loc(callsite(#loc78 at #loc2))
#loc124 = loc(callsite(#loc78 at #loc5))
#loc126 = loc(callsite(#loc106 at #loc2))
module attributes {hacc.target = #hacc.target<"Ascend950PR_9579">} {
func.func @_attn_fwd(%arg0: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg1: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg2: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg3: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg4: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg5: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg6: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg7: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg8: memref<?xi32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg9: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg10: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg11: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg12: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg13: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg14: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg15: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)) attributes {SyncBlockLockArgIdx = 0 : i64, WorkspaceArgIdx = 1 : i64, global_kernel = "local", mix_mode = "mix", parallel_mode = "simd"} {
%c8192 = arith.constant 8192 : index loc(#loc71)
%c128 = arith.constant 128 : index loc(#loc71)
%c1 = arith.constant 1 : index loc(#loc3)
%c0 = arith.constant 0 : index loc(#loc4)
%cst = arith.constant 1.000000e+00 : f32 loc(#loc72)
%cst_0 = arith.constant 0xFF800000 : f32 loc(#loc121)
%cst_1 = arith.constant -1.000000e+04 : f32 loc(#loc74)
%cst_2 = arith.constant 5.000000e-01 : f32 loc(#loc75)
%c0_i32 = arith.constant 0 : i32 loc(#loc71)
%c1_i32 = arith.constant 1 : i32 loc(#loc76)
%c64_i32 = arith.constant 64 : i32 loc(#loc11)
%c8_i32 = arith.constant 8 : i32 loc(#loc12)
%c8388608_i64 = arith.constant 8388608 : i64 loc(#loc13)
%c1048576_i64 = arith.constant 1048576 : i64 loc(#loc14)
%c128_i32 = arith.constant 128 : i32 loc(#loc77)
%c8192_i32 = arith.constant 8192 : i32 loc(#loc16)
%cst_3 = arith.constant 0.000000e+00 : f32 loc(#loc122)
%0 = tensor.empty() : tensor<128x128xf32> loc(#loc79)
%1 = linalg.fill ins(%cst_3 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
%2 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc75)
%3 = linalg.fill ins(%cst_1 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
%4 = tensor.empty() : tensor<128xf32> loc(#loc121)
%5 = linalg.fill ins(%cst_0 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc121)
%6 = linalg.fill ins(%cst : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc72)
%7 = arith.index_cast %arg13 : i32 to index loc(#loc20)
%reinterpret_cast = memref.reinterpret_cast %arg8 to offset: [%7], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc21)
%8 = memref.load %reinterpret_cast[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc22)
%9 = arith.addi %7, %c1 : index loc(#loc3)
%reinterpret_cast_4 = memref.reinterpret_cast %arg8 to offset: [%9], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc3)
%10 = memref.load %reinterpret_cast_4[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc4)
scf.for %arg16 = %8 to %10 step %c1_i32 : i32 {
%11 = arith.divsi %arg16, %c64_i32 : i32 loc(#loc24)
%12 = arith.remsi %arg16, %c64_i32 : i32 loc(#loc11)
%13 = arith.divsi %11, %c8_i32 : i32 loc(#loc25)
%14 = arith.remsi %11, %c8_i32 : i32 loc(#loc26)
%15 = arith.muli %14, %arg9 : i32 loc(#loc27)
%16 = arith.divsi %15, %c8_i32 : i32 loc(#loc12)
%17 = arith.extsi %13 : i32 to i64 loc(#loc28)
%18 = arith.muli %17, %c8388608_i64 : i64 loc(#loc13)
%19 = arith.extsi %14 : i32 to i64 loc(#loc29)
%20 = arith.muli %19, %c1048576_i64 : i64 loc(#loc30)
%21 = arith.addi %18, %20 : i64 loc(#loc31)
%22 = arith.extsi %16 : i32 to i64 loc(#loc32)
%23 = arith.muli %22, %c1048576_i64 : i64 loc(#loc14)
%24 = arith.addi %18, %23 : i64 loc(#loc33)
%25 = arith.index_cast %21 : i64 to index loc(#loc31)
%26 = arith.muli %12, %c128_i32 : i32 loc(#loc34)
%27 = arith.maxsi %26, %c0_i32 : i32 loc(#loc35)
%28 = arith.index_cast %27 : i32 to index loc(#loc35)
%29 = arith.muli %28, %c128 : index loc(#loc35)
%30 = arith.addi %29, %25 : index loc(#loc35)
%reinterpret_cast_5 = memref.reinterpret_cast %arg2 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc35)
%31 = arith.index_cast %24 : i64 to index loc(#loc33)
%reinterpret_cast_6 = memref.reinterpret_cast %arg7 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc36)
%alloc = memref.alloc() : memref<128x128xf16> loc(#loc37)
memref.copy %reinterpret_cast_5, %alloc : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc37)
%32 = bufferization.to_tensor %alloc restrict writable : memref<128x128xf16> loc(#loc37)
%33:5 = scf.for %arg17 = %c0_i32 to %26 step %c128_i32 iter_args(%arg18 = %6, %arg19 = %1, %arg20 = %5, %arg21 = %c0_i32, %arg22 = %c0_i32) -> (tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32) : i32 {
%46 = arith.maxsi %arg21, %c0_i32 : i32 loc(#loc72)
%47 = arith.index_cast %46 : i32 to index loc(#loc72)
%48 = arith.muli %47, %c128 : index loc(#loc72)
%49 = arith.addi %48, %31 : index loc(#loc72)
%reinterpret_cast_8 = memref.reinterpret_cast %arg4 to offset: [%49], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc72)
%50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc72)
%51 = arith.index_cast %50 : i32 to index loc(#loc72)
%52 = arith.muli %51, %c128 : index loc(#loc72)
%53 = arith.addi %52, %31 : index loc(#loc72)
%reinterpret_cast_9 = memref.reinterpret_cast %arg3 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc72)
%alloc_10 = memref.alloc() : memref<128x128xf16> loc(#loc80)
memref.copy %reinterpret_cast_9, %alloc_10 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc80)
%54 = bufferization.to_tensor %alloc_10 restrict writable : memref<128x128xf16> loc(#loc80)
%55 = tensor.empty() : tensor<128x128xf16> loc(#loc81)
%transposed = linalg.transpose ins(%54 : tensor<128x128xf16>) outs(%55 : tensor<128x128xf16>) permutation = [1, 0] loc(#loc81)
%56 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xf16>, tensor<128x128xf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc82)
%57 = arith.mulf %56, %2 : tensor<128x128xf32> loc(#loc83)
%reduced = linalg.reduce ins(%57 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc41 at #loc5)), %init: f32 loc(callsite(#loc73 at #loc5))) {
%72 = arith.maximumf %in, %init : f32 loc(#loc128)
linalg.yield %72 : f32 loc(#loc121)
} loc(#loc121)
%58 = arith.maximumf %arg20, %reduced : tensor<128xf32> loc(#loc85)
%broadcasted_11 = linalg.broadcast ins(%58 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc79)
%59 = arith.subf %57, %broadcasted_11 : tensor<128x128xf32> loc(#loc79)
%60 = math.exp %59 : tensor<128x128xf32> loc(#loc86)
%61 = arith.truncf %60 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc87)
%alloc_12 = memref.alloc() : memref<128x128xf16> loc(#loc88)
memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc88)
%62 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf16> loc(#loc88)
%63 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc124)
%reduced_13 = linalg.reduce ins(%60 : tensor<128x128xf32>) outs(%63 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc44 at #loc5)), %init: f32 loc(callsite(#loc78 at #loc5))) {
%72 = arith.addf %in, %init : f32 loc(#loc129)
linalg.yield %72 : f32 loc(#loc124)
} loc(#loc124)
%64 = arith.subf %arg20, %58 : tensor<128xf32> loc(#loc90)
%65 = math.exp %64 : tensor<128xf32> loc(#loc91)
%66 = arith.mulf %arg18, %65 : tensor<128xf32> loc(#loc92)
%67 = arith.addf %66, %reduced_13 : tensor<128xf32> loc(#loc93)
%broadcasted_14 = linalg.broadcast ins(%65 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc94)
%68 = arith.mulf %arg19, %broadcasted_14 : tensor<128x128xf32> loc(#loc94)
%69 = linalg.matmul {input_precision = "ieee"} ins(%61, %62 : tensor<128x128xf16>, tensor<128x128xf16>) outs(%68 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc95)
%70 = arith.addi %arg21, %c128_i32 : i32 loc(#loc96)
%71 = arith.addi %arg22, %c128_i32 : i32 loc(#loc97)
scf.yield %67, %69, %58, %70, %71 : tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc98)
} {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc72)
%34 = arith.muli %12, %c128_i32 {tt.divisibility = dense<128> : tensor<1xi32>} : i32 loc(#loc99)
%35 = arith.addi %12, %c1_i32 : i32 loc(#loc76)
%36 = arith.muli %35, %c128_i32 : i32 loc(#loc100)
%37:6 = scf.for %arg17 = %34 to %36 step %c128_i32 iter_args(%arg18 = %34, %arg19 = %33#0, %arg20 = %33#1, %arg21 = %33#2, %arg22 = %34, %arg23 = %34) -> (i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32) : i32 {
%46 = arith.maxsi %arg18, %c0_i32 : i32 loc(#loc71)
%47 = arith.index_cast %46 : i32 to index loc(#loc71)
%48 = arith.muli %28, %c8192 : index loc(#loc71)
%49 = arith.addi %48, %47 : index loc(#loc71)
%reinterpret_cast_8 = memref.reinterpret_cast %arg5 to offset: [%49], sizes: [128, 128], strides: [8192, 1] : memref<?xf32> to memref<128x128xf32, strided<[8192, 1], offset: ?>> loc(#loc71)
%50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc71)
%51 = arith.index_cast %50 : i32 to index loc(#loc71)
%52 = arith.muli %51, %c128 : index loc(#loc71)
%53 = arith.addi %52, %31 : index loc(#loc71)
%reinterpret_cast_9 = memref.reinterpret_cast %arg4 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc71)
%54 = arith.maxsi %arg23, %c0_i32 : i32 loc(#loc71)
%55 = arith.index_cast %54 : i32 to index loc(#loc71)
%56 = arith.muli %55, %c128 : index loc(#loc71)
%57 = arith.addi %56, %31 : index loc(#loc71)
%reinterpret_cast_10 = memref.reinterpret_cast %arg3 to offset: [%57], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc71)
%alloc_11 = memref.alloc() : memref<128x128xf16> loc(#loc101)
memref.copy %reinterpret_cast_10, %alloc_11 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc101)
%58 = bufferization.to_tensor %alloc_11 restrict writable : memref<128x128xf16> loc(#loc101)
%59 = tensor.empty() : tensor<128x128xf16> loc(#loc102)
%transposed = linalg.transpose ins(%58 : tensor<128x128xf16>) outs(%59 : tensor<128x128xf16>) permutation = [1, 0] loc(#loc102)
%60 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xf16>, tensor<128x128xf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc103)
%alloc_12 = memref.alloc() : memref<128x128xf32> loc(#loc104)
memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf32, strided<[8192, 1], offset: ?>> to memref<128x128xf32> loc(#loc104)
%61 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf32> loc(#loc104)
%62 = arith.mulf %60, %2 : tensor<128x128xf32> loc(#loc75)
%63 = arith.cmpf une, %61, %1 : tensor<128x128xf32> loc(#loc74)
%64 = arith.select %63, %3, %1 : tensor<128x128xi1>, tensor<128x128xf32> loc(#loc74)
%65 = arith.addf %62, %64 : tensor<128x128xf32> loc(#loc105)
%reduced = linalg.reduce ins(%65 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc59 at #loc2)), %init: f32 loc(callsite(#loc106 at #loc2))) {
%81 = arith.maximumf %in, %init : f32 loc(#loc130)
linalg.yield %81 : f32 loc(#loc126)
} loc(#loc126)
%66 = arith.maximumf %arg21, %reduced : tensor<128xf32> loc(#loc107)
%broadcasted_13 = linalg.broadcast ins(%66 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc108)
%67 = arith.subf %65, %broadcasted_13 : tensor<128x128xf32> loc(#loc108)
%68 = arith.addi %arg18, %c128_i32 : i32 loc(#loc109)
%69 = math.exp %67 : tensor<128x128xf32> loc(#loc110)
%70 = arith.truncf %69 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc111)
%alloc_14 = memref.alloc() : memref<128x128xf16> loc(#loc112)
memref.copy %reinterpret_cast_9, %alloc_14 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc112)
%71 = bufferization.to_tensor %alloc_14 restrict writable : memref<128x128xf16> loc(#loc112)
%72 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc122)
%reduced_15 = linalg.reduce ins(%69 : tensor<128x128xf32>) outs(%72 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc44 at #loc2)), %init: f32 loc(callsite(#loc78 at #loc2))) {
%81 = arith.addf %in, %init : f32 loc(#loc131)
linalg.yield %81 : f32 loc(#loc122)
} loc(#loc122)
%73 = arith.subf %arg21, %66 : tensor<128xf32> loc(#loc113)
%74 = math.exp %73 : tensor<128xf32> loc(#loc114)
%75 = arith.mulf %arg19, %74 : tensor<128xf32> loc(#loc115)
%76 = arith.addf %75, %reduced_15 : tensor<128xf32> loc(#loc116)
%broadcasted_16 = linalg.broadcast ins(%74 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc117)
%77 = arith.mulf %arg20, %broadcasted_16 : tensor<128x128xf32> loc(#loc117)
%78 = linalg.matmul {input_precision = "ieee"} ins(%70, %71 : tensor<128x128xf16>, tensor<128x128xf16>) outs(%77 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc118)
%79 = arith.addi %arg22, %c128_i32 : i32 loc(#loc119)
%80 = arith.addi %arg23, %c128_i32 : i32 loc(#loc77)
scf.yield %68, %76, %78, %66, %79, %80 : i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc120)
} {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc71)
%38 = math.log %37#1 : tensor<128xf32> loc(#loc64)
%39 = arith.addf %37#3, %38 : tensor<128xf32> loc(#loc65)
%broadcasted = linalg.broadcast ins(%37#1 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc66)
%40 = arith.divf %37#2, %broadcasted : tensor<128x128xf32> loc(#loc66)
%41 = arith.muli %11, %c8192_i32 : i32 loc(#loc16)
%42 = arith.index_cast %41 : i32 to index loc(#loc16)
%43 = arith.index_cast %26 : i32 to index loc(#loc34)
%44 = arith.addi %42, %43 : index loc(#loc67)
%reinterpret_cast_7 = memref.reinterpret_cast %arg6 to offset: [%44], sizes: [128], strides: [1] : memref<?xf32> to memref<128xf32, strided<[1], offset: ?>> loc(#loc67)
bufferization.materialize_in_destination %39 in writable %reinterpret_cast_7 : (tensor<128xf32>, memref<128xf32, strided<[1], offset: ?>>) -> () loc(#loc68)
%45 = arith.truncf %40 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc69)
bufferization.materialize_in_destination %45 in writable %reinterpret_cast_6 : (tensor<128x128xf16>, memref<128x128xf16, strided<[128, 1], offset: ?>>) -> () loc(#loc70)
} loc(#loc23)
return loc(#loc)
} loc(#loc)
} loc(#loc)
#loc1 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":63:33)
#loc3 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:53)
#loc4 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:28)
#loc8 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:56)
#loc9 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:22)
#loc10 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:47)
#loc11 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":139:33)
#loc12 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:39)
#loc13 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:40)
#loc14 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:72)
#loc15 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:46)
#loc16 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:35)
#loc19 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":83:22)
#loc20 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":123:24)
#loc21 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:49)
#loc22 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:30)
#loc23 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":137:51)
#loc24 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":138:35)
#loc25 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":140:31)
#loc26 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":141:30)
#loc27 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:23)
#loc28 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:28)
#loc29 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:61)
#loc30 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:73)
#loc31 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:52)
#loc32 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:60)
#loc33 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:52)
#loc34 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":152:34)
#loc35 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":154:12)
#loc36 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":180:12)
#loc37 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":193:20)
#loc38 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":66:20)
#loc39 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":69:27)
#loc40 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":70:23)
#loc42 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":168:27)
#loc43 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:35)
#loc45 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":85:22)
#loc46 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":86:20)
#loc47 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":271:15)
#loc48 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:34)
#loc49 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:28)
#loc50 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:20)
#loc51 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:28)
#loc52 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":94:20)
#loc53 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":87:28)
#loc54 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":97:46)
#loc55 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:8)
#loc56 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:27)
#loc57 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:52)
#loc58 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":75:27)
#loc61 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:35)
#loc62 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":78:18)
#loc63 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":79:54)
#loc64 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:27)
#loc65 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:15)
#loc66 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":216:20)
#loc67 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:43)
#loc68 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":219:25)
#loc69 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:37)
#loc70 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:30)
#loc71 = loc(callsite(#loc1 at #loc2))
#loc72 = loc(callsite(#loc1 at #loc5))
#loc74 = loc(callsite(#loc8 at #loc2))
#loc75 = loc(callsite(#loc9 at #loc2))
#loc76 = loc(callsite(#loc10 at #loc2))
#loc77 = loc(callsite(#loc15 at #loc2))
#loc79 = loc(callsite(#loc19 at #loc5))
#loc80 = loc(callsite(#loc38 at #loc5))
#loc81 = loc(callsite(#loc39 at #loc5))
#loc82 = loc(callsite(#loc40 at #loc5))
#loc84 = loc(callsite(#loc42 at #loc6))
#loc85 = loc(callsite(#loc43 at #loc5))
#loc87 = loc(callsite(#loc45 at #loc5))
#loc88 = loc(callsite(#loc46 at #loc5))
#loc89 = loc(callsite(#loc47 at #loc17))
#loc90 = loc(callsite(#loc48 at #loc5))
#loc91 = loc(callsite(#loc49 at #loc5))
#loc92 = loc(callsite(#loc50 at #loc5))
#loc93 = loc(callsite(#loc51 at #loc5))
#loc94 = loc(callsite(#loc52 at #loc5))
#loc95 = loc(callsite(#loc53 at #loc5))
#loc96 = loc(callsite(#loc54 at #loc5))
#loc97 = loc(callsite(#loc15 at #loc5))
#loc98 = loc(callsite(#loc55 at #loc5))
#loc99 = loc(callsite(#loc56 at #loc2))
#loc100 = loc(callsite(#loc57 at #loc2))
#loc101 = loc(callsite(#loc38 at #loc2))
#loc102 = loc(callsite(#loc39 at #loc2))
#loc103 = loc(callsite(#loc40 at #loc2))
#loc104 = loc(callsite(#loc58 at #loc2))
#loc107 = loc(callsite(#loc61 at #loc2))
#loc108 = loc(callsite(#loc62 at #loc2))
#loc109 = loc(callsite(#loc63 at #loc2))
#loc111 = loc(callsite(#loc45 at #loc2))
#loc112 = loc(callsite(#loc46 at #loc2))
#loc113 = loc(callsite(#loc48 at #loc2))
#loc114 = loc(callsite(#loc49 at #loc2))
#loc115 = loc(callsite(#loc50 at #loc2))
#loc116 = loc(callsite(#loc51 at #loc2))
#loc117 = loc(callsite(#loc52 at #loc2))
#loc118 = loc(callsite(#loc53 at #loc2))
#loc119 = loc(callsite(#loc54 at #loc2))
#loc120 = loc(callsite(#loc55 at #loc2))
#loc123 = loc(callsite(#loc84 at #loc7))
#loc125 = loc(callsite(#loc89 at #loc18))
#loc127 = loc(callsite(#loc84 at #loc60))
#loc128 = loc(callsite(#loc123 at #loc5))
#loc129 = loc(callsite(#loc125 at #loc5))
#loc130 = loc(callsite(#loc127 at #loc2))
#loc131 = loc(callsite(#loc125 at #loc2))


hujiajun
7月31日 评论:
7月31日 评论:
[2026-07-31T14:26:32.036Z] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:26:32.036Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:26:32.036Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpeebgzx7u/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpeebgzx7u/kernel --enable-vf-merge-level=1
[2026-07-31T14:26:32.036Z] RERUN
[2026-07-31T14:26:39.756Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype6-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:26:39.756Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:26:39.756Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp4prqx238/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp4prqx238/kernel --enable-vf-merge-level=1
[2026-07-31T14:26:39.756Z] RERUN
[2026-07-31T14:26:49.142Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype6-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:26:49.143Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:26:49.143Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp2qtn812j/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp2qtn812j/kernel --enable-vf-merge-level=1
[2026-07-31T14:26:49.143Z] RERUN
[2026-07-31T14:26:56.865Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype6-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:26:56.865Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:26:56.865Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmplgik8dih/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmplgik8dih/kernel --enable-vf-merge-level=1
[2026-07-31T14:26:56.865Z] RERUN
[2026-07-31T14:27:00.122Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype6-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/5XVQxRyAGDk3gf66Xd-L88jBPwwGIjZ2v3bNIEeRE0Q
[2026-07-31T14:27:00.122Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
[2026-07-31T14:27:00.122Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp8hxi0lmy/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp8hxi0lmy/kernel --enable-vf-merge-level=1
[2026-07-31T14:27:00.122Z] FAILED
#loc = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)
#loc2 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":212:70)
#loc5 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":202:78)
#loc6 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":195:44)
#loc7 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:46)
#loc16 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":282:36)
#loc17 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":88:25)
#loc41 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":81:22)
#loc44 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":84:24)
#loc59 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:33)
#loc60 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:46)
#loc73 = loc(callsite(#loc6 at #loc7))
#loc78 = loc(callsite(#loc16 at #loc17))
#loc83 = loc(callsite(#loc41 at #loc5))
#loc86 = loc(callsite(#loc44 at #loc5))
#loc105 = loc(callsite(#loc59 at #loc2))
#loc106 = loc(callsite(#loc6 at #loc60))
#loc110 = loc(callsite(#loc44 at #loc2))
#loc121 = loc(callsite(#loc73 at #loc5))
#loc122 = loc(callsite(#loc78 at #loc2))
#loc124 = loc(callsite(#loc78 at #loc5))
#loc126 = loc(callsite(#loc106 at #loc2))
module attributes {hacc.target = #hacc.target<"Ascend950PR_9579">} {
func.func @_attn_fwd(%arg0: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg1: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg2: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg3: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg4: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg5: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg6: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg7: memref<?xf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg8: memref<?xi32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg9: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg10: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg11: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg12: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg13: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg14: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg15: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)) attributes {SyncBlockLockArgIdx = 0 : i64, WorkspaceArgIdx = 1 : i64, global_kernel = "local", mix_mode = "mix", parallel_mode = "simd"} {
%c1024 = arith.constant 1024 : index loc(#loc71)
%c128 = arith.constant 128 : index loc(#loc71)
%c1 = arith.constant 1 : index loc(#loc3)
%c0 = arith.constant 0 : index loc(#loc4)
%cst = arith.constant 1.000000e+00 : f32 loc(#loc72)
%cst_0 = arith.constant 0xFF800000 : f32 loc(#loc121)
%cst_1 = arith.constant -1.000000e+04 : f32 loc(#loc74)
%cst_2 = arith.constant 5.000000e-01 : f32 loc(#loc75)
%c1024_i32 = arith.constant 1024 : i32 loc(#loc10)
%c0_i32 = arith.constant 0 : i32 loc(#loc71)
%c1_i32 = arith.constant 1 : i32 loc(#loc76)
%c8_i32 = arith.constant 8 : i32 loc(#loc12)
%c1048576_i64 = arith.constant 1048576 : i64 loc(#loc13)
%c131072_i64 = arith.constant 131072 : i64 loc(#loc14)
%c128_i32 = arith.constant 128 : i32 loc(#loc77)
%cst_3 = arith.constant 0.000000e+00 : f32 loc(#loc122)
%0 = tensor.empty() : tensor<128x128xf32> loc(#loc79)
%1 = linalg.fill ins(%cst_3 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
%2 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc75)
%3 = linalg.fill ins(%cst_1 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
%4 = tensor.empty() : tensor<128xf32> loc(#loc121)
%5 = linalg.fill ins(%cst_0 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc121)
%6 = linalg.fill ins(%cst : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc72)
%7 = arith.index_cast %arg13 : i32 to index loc(#loc19)
%reinterpret_cast = memref.reinterpret_cast %arg8 to offset: [%7], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc20)
%8 = memref.load %reinterpret_cast[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc21)
%9 = arith.addi %7, %c1 : index loc(#loc3)
%reinterpret_cast_4 = memref.reinterpret_cast %arg8 to offset: [%9], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc3)
%10 = memref.load %reinterpret_cast_4[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc4)
scf.for %arg16 = %8 to %10 step %c1_i32 : i32 {
%11 = arith.divsi %arg16, %c8_i32 : i32 loc(#loc23)
%12 = arith.remsi %arg16, %c8_i32 : i32 loc(#loc24)
%13 = arith.divsi %11, %c8_i32 : i32 loc(#loc25)
%14 = arith.remsi %11, %c8_i32 : i32 loc(#loc26)
%15 = arith.muli %14, %arg9 : i32 loc(#loc27)
%16 = arith.divsi %15, %c8_i32 : i32 loc(#loc12)
%17 = arith.extsi %13 : i32 to i64 loc(#loc28)
%18 = arith.muli %17, %c1048576_i64 : i64 loc(#loc13)
%19 = arith.extsi %14 : i32 to i64 loc(#loc29)
%20 = arith.muli %19, %c131072_i64 : i64 loc(#loc30)
%21 = arith.addi %18, %20 : i64 loc(#loc31)
%22 = arith.extsi %16 : i32 to i64 loc(#loc32)
%23 = arith.muli %22, %c131072_i64 : i64 loc(#loc14)
%24 = arith.addi %18, %23 : i64 loc(#loc33)
%25 = arith.index_cast %21 : i64 to index loc(#loc31)
%26 = arith.muli %12, %c128_i32 : i32 loc(#loc34)
%27 = arith.maxsi %26, %c0_i32 : i32 loc(#loc35)
%28 = arith.index_cast %27 : i32 to index loc(#loc35)
%29 = arith.muli %28, %c128 : index loc(#loc35)
%30 = arith.addi %29, %25 : index loc(#loc35)
%reinterpret_cast_5 = memref.reinterpret_cast %arg2 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc35)
%31 = arith.index_cast %24 : i64 to index loc(#loc33)
%reinterpret_cast_6 = memref.reinterpret_cast %arg7 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc36)
%alloc = memref.alloc() : memref<128x128xf16> loc(#loc37)
memref.copy %reinterpret_cast_5, %alloc : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc37)
%32 = bufferization.to_tensor %alloc restrict writable : memref<128x128xf16> loc(#loc37)
%33:5 = scf.for %arg17 = %c0_i32 to %26 step %c128_i32 iter_args(%arg18 = %6, %arg19 = %1, %arg20 = %5, %arg21 = %c0_i32, %arg22 = %c0_i32) -> (tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32) : i32 {
%46 = arith.maxsi %arg21, %c0_i32 : i32 loc(#loc72)
%47 = arith.index_cast %46 : i32 to index loc(#loc72)
%48 = arith.muli %47, %c128 : index loc(#loc72)
%49 = arith.addi %48, %31 : index loc(#loc72)
%reinterpret_cast_8 = memref.reinterpret_cast %arg4 to offset: [%49], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc72)
%50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc72)
%51 = arith.index_cast %50 : i32 to index loc(#loc72)
%52 = arith.muli %51, %c128 : index loc(#loc72)
%53 = arith.addi %52, %31 : index loc(#loc72)
%reinterpret_cast_9 = memref.reinterpret_cast %arg3 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc72)
%alloc_10 = memref.alloc() : memref<128x128xf16> loc(#loc80)
memref.copy %reinterpret_cast_9, %alloc_10 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc80)
%54 = bufferization.to_tensor %alloc_10 restrict writable : memref<128x128xf16> loc(#loc80)
%55 = tensor.empty() : tensor<128x128xf16> loc(#loc81)
%transposed = linalg.transpose ins(%54 : tensor<128x128xf16>) outs(%55 : tensor<128x128xf16>) permutation = [1, 0] loc(#loc81)
%56 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xf16>, tensor<128x128xf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc82)
%57 = arith.mulf %56, %2 : tensor<128x128xf32> loc(#loc83)
%reduced = linalg.reduce ins(%57 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc41 at #loc5)), %init: f32 loc(callsite(#loc73 at #loc5))) {
%72 = arith.maximumf %in, %init : f32 loc(#loc128)
linalg.yield %72 : f32 loc(#loc121)
} loc(#loc121)
%58 = arith.maximumf %arg20, %reduced : tensor<128xf32> loc(#loc85)
%broadcasted_11 = linalg.broadcast ins(%58 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc79)
%59 = arith.subf %57, %broadcasted_11 : tensor<128x128xf32> loc(#loc79)
%60 = math.exp %59 : tensor<128x128xf32> loc(#loc86)
%61 = arith.truncf %60 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc87)
%alloc_12 = memref.alloc() : memref<128x128xf16> loc(#loc88)
memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc88)
%62 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf16> loc(#loc88)
%63 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc124)
%reduced_13 = linalg.reduce ins(%60 : tensor<128x128xf32>) outs(%63 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc44 at #loc5)), %init: f32 loc(callsite(#loc78 at #loc5))) {
%72 = arith.addf %in, %init : f32 loc(#loc129)
linalg.yield %72 : f32 loc(#loc124)
} loc(#loc124)
%64 = arith.subf %arg20, %58 : tensor<128xf32> loc(#loc90)
%65 = math.exp %64 : tensor<128xf32> loc(#loc91)
%66 = arith.mulf %arg18, %65 : tensor<128xf32> loc(#loc92)
%67 = arith.addf %66, %reduced_13 : tensor<128xf32> loc(#loc93)
%broadcasted_14 = linalg.broadcast ins(%65 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc94)
%68 = arith.mulf %arg19, %broadcasted_14 : tensor<128x128xf32> loc(#loc94)
%69 = linalg.matmul {input_precision = "ieee"} ins(%61, %62 : tensor<128x128xf16>, tensor<128x128xf16>) outs(%68 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc95)
%70 = arith.addi %arg21, %c128_i32 : i32 loc(#loc96)
%71 = arith.addi %arg22, %c128_i32 : i32 loc(#loc97)
scf.yield %67, %69, %58, %70, %71 : tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc98)
} {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc72)
%34 = arith.muli %12, %c128_i32 {tt.divisibility = dense<128> : tensor<1xi32>} : i32 loc(#loc99)
%35 = arith.addi %12, %c1_i32 : i32 loc(#loc76)
%36 = arith.muli %35, %c128_i32 : i32 loc(#loc100)
%37:6 = scf.for %arg17 = %34 to %36 step %c128_i32 iter_args(%arg18 = %34, %arg19 = %33#0, %arg20 = %33#1, %arg21 = %33#2, %arg22 = %34, %arg23 = %34) -> (i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32) : i32 {
%46 = arith.maxsi %arg18, %c0_i32 : i32 loc(#loc71)
%47 = arith.index_cast %46 : i32 to index loc(#loc71)
%48 = arith.muli %28, %c1024 : index loc(#loc71)
%49 = arith.addi %48, %47 : index loc(#loc71)
%reinterpret_cast_8 = memref.reinterpret_cast %arg5 to offset: [%49], sizes: [128, 128], strides: [1024, 1] : memref<?xf32> to memref<128x128xf32, strided<[1024, 1], offset: ?>> loc(#loc71)
%50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc71)
%51 = arith.index_cast %50 : i32 to index loc(#loc71)
%52 = arith.muli %51, %c128 : index loc(#loc71)
%53 = arith.addi %52, %31 : index loc(#loc71)
%reinterpret_cast_9 = memref.reinterpret_cast %arg4 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc71)
%54 = arith.maxsi %arg23, %c0_i32 : i32 loc(#loc71)
%55 = arith.index_cast %54 : i32 to index loc(#loc71)
%56 = arith.muli %55, %c128 : index loc(#loc71)
%57 = arith.addi %56, %31 : index loc(#loc71)
%reinterpret_cast_10 = memref.reinterpret_cast %arg3 to offset: [%57], sizes: [128, 128], strides: [128, 1] : memref<?xf16> to memref<128x128xf16, strided<[128, 1], offset: ?>> loc(#loc71)
%alloc_11 = memref.alloc() : memref<128x128xf16> loc(#loc101)
memref.copy %reinterpret_cast_10, %alloc_11 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc101)
%58 = bufferization.to_tensor %alloc_11 restrict writable : memref<128x128xf16> loc(#loc101)
%59 = tensor.empty() : tensor<128x128xf16> loc(#loc102)
%transposed = linalg.transpose ins(%58 : tensor<128x128xf16>) outs(%59 : tensor<128x128xf16>) permutation = [1, 0] loc(#loc102)
%60 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xf16>, tensor<128x128xf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc103)
%alloc_12 = memref.alloc() : memref<128x128xf32> loc(#loc104)
memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf32, strided<[1024, 1], offset: ?>> to memref<128x128xf32> loc(#loc104)
%61 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf32> loc(#loc104)
%62 = arith.mulf %60, %2 : tensor<128x128xf32> loc(#loc75)
%63 = arith.cmpf une, %61, %1 : tensor<128x128xf32> loc(#loc74)
%64 = arith.select %63, %3, %1 : tensor<128x128xi1>, tensor<128x128xf32> loc(#loc74)
%65 = arith.addf %62, %64 : tensor<128x128xf32> loc(#loc105)
%reduced = linalg.reduce ins(%65 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc59 at #loc2)), %init: f32 loc(callsite(#loc106 at #loc2))) {
%81 = arith.maximumf %in, %init : f32 loc(#loc130)
linalg.yield %81 : f32 loc(#loc126)
} loc(#loc126)
%66 = arith.maximumf %arg21, %reduced : tensor<128xf32> loc(#loc107)
%broadcasted_13 = linalg.broadcast ins(%66 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc108)
%67 = arith.subf %65, %broadcasted_13 : tensor<128x128xf32> loc(#loc108)
%68 = arith.addi %arg18, %c128_i32 : i32 loc(#loc109)
%69 = math.exp %67 : tensor<128x128xf32> loc(#loc110)
%70 = arith.truncf %69 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc111)
%alloc_14 = memref.alloc() : memref<128x128xf16> loc(#loc112)
memref.copy %reinterpret_cast_9, %alloc_14 : memref<128x128xf16, strided<[128, 1], offset: ?>> to memref<128x128xf16> loc(#loc112)
%71 = bufferization.to_tensor %alloc_14 restrict writable : memref<128x128xf16> loc(#loc112)
%72 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc122)
%reduced_15 = linalg.reduce ins(%69 : tensor<128x128xf32>) outs(%72 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc44 at #loc2)), %init: f32 loc(callsite(#loc78 at #loc2))) {
%81 = arith.addf %in, %init : f32 loc(#loc131)
linalg.yield %81 : f32 loc(#loc122)
} loc(#loc122)
%73 = arith.subf %arg21, %66 : tensor<128xf32> loc(#loc113)
%74 = math.exp %73 : tensor<128xf32> loc(#loc114)
%75 = arith.mulf %arg19, %74 : tensor<128xf32> loc(#loc115)
%76 = arith.addf %75, %reduced_15 : tensor<128xf32> loc(#loc116)
%broadcasted_16 = linalg.broadcast ins(%74 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc117)
%77 = arith.mulf %arg20, %broadcasted_16 : tensor<128x128xf32> loc(#loc117)
%78 = linalg.matmul {input_precision = "ieee"} ins(%70, %71 : tensor<128x128xf16>, tensor<128x128xf16>) outs(%77 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc118)
%79 = arith.addi %arg22, %c128_i32 : i32 loc(#loc119)
%80 = arith.addi %arg23, %c128_i32 : i32 loc(#loc77)
scf.yield %68, %76, %78, %66, %79, %80 : i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc120)
} {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc71)
%38 = math.log %37#1 : tensor<128xf32> loc(#loc64)
%39 = arith.addf %37#3, %38 : tensor<128xf32> loc(#loc65)
%broadcasted = linalg.broadcast ins(%37#1 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc66)
%40 = arith.divf %37#2, %broadcasted : tensor<128x128xf32> loc(#loc66)
%41 = arith.muli %11, %c1024_i32 : i32 loc(#loc10)
%42 = arith.index_cast %41 : i32 to index loc(#loc10)
%43 = arith.index_cast %26 : i32 to index loc(#loc34)
%44 = arith.addi %42, %43 : index loc(#loc67)
%reinterpret_cast_7 = memref.reinterpret_cast %arg6 to offset: [%44], sizes: [128], strides: [1] : memref<?xf32> to memref<128xf32, strided<[1], offset: ?>> loc(#loc67)
bufferization.materialize_in_destination %39 in writable %reinterpret_cast_7 : (tensor<128xf32>, memref<128xf32, strided<[1], offset: ?>>) -> () loc(#loc68)
%45 = arith.truncf %40 : tensor<128x128xf32> to tensor<128x128xf16> loc(#loc69)
bufferization.materialize_in_destination %45 in writable %reinterpret_cast_6 : (tensor<128x128xf16>, memref<128x128xf16, strided<[128, 1], offset: ?>>) -> () loc(#loc70)
} loc(#loc22)
return loc(#loc)
} loc(#loc)
} loc(#loc)
#loc1 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":63:33)
#loc3 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:53)
#loc4 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:28)
#loc8 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:56)
#loc9 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:22)
#loc10 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:35)
#loc11 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:47)
#loc12 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:39)
#loc13 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:40)
#loc14 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:72)
#loc15 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:46)
#loc18 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":83:22)
#loc19 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":123:24)
#loc20 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:49)
#loc21 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:30)
#loc22 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":137:51)
#loc23 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":138:35)
#loc24 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":139:33)
#loc25 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":140:31)
#loc26 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":141:30)
#loc27 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:23)
#loc28 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:28)
#loc29 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:61)
#loc30 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:73)
#loc31 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:52)
#loc32 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:60)
#loc33 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:52)
#loc34 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":152:34)
#loc35 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":154:12)
#loc36 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":180:12)
#loc37 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":193:20)
#loc38 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":66:20)
#loc39 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":69:27)
#loc40 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":70:23)
#loc42 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":168:27)
#loc43 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:35)
#loc45 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":85:22)
#loc46 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":86:20)
#loc47 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":271:15)
#loc48 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:34)
#loc49 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:28)
#loc50 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:20)
#loc51 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:28)
#loc52 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":94:20)
#loc53 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":87:28)
#loc54 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":97:46)
#loc55 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:8)
#loc56 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:27)
#loc57 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:52)
#loc58 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":75:27)
#loc61 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:35)
#loc62 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":78:18)
#loc63 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":79:54)
#loc64 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:27)
#loc65 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:15)
#loc66 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":216:20)
#loc67 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:43)
#loc68 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":219:25)
#loc69 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:37)
#loc70 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:30)
#loc71 = loc(callsite(#loc1 at #loc2))
#loc72 = loc(callsite(#loc1 at #loc5))
#loc74 = loc(callsite(#loc8 at #loc2))
#loc75 = loc(callsite(#loc9 at #loc2))
#loc76 = loc(callsite(#loc11 at #loc2))
#loc77 = loc(callsite(#loc15 at #loc2))
#loc79 = loc(callsite(#loc18 at #loc5))
#loc80 = loc(callsite(#loc38 at #loc5))
#loc81 = loc(callsite(#loc39 at #loc5))
#loc82 = loc(callsite(#loc40 at #loc5))
#loc84 = loc(callsite(#loc42 at #loc6))
#loc85 = loc(callsite(#loc43 at #loc5))
#loc87 = loc(callsite(#loc45 at #loc5))
#loc88 = loc(callsite(#loc46 at #loc5))
#loc89 = loc(callsite(#loc47 at #loc16))
#loc90 = loc(callsite(#loc48 at #loc5))
#loc91 = loc(callsite(#loc49 at #loc5))
#loc92 = loc(callsite(#loc50 at #loc5))
#loc93 = loc(callsite(#loc51 at #loc5))
#loc94 = loc(callsite(#loc52 at #loc5))
#loc95 = loc(callsite(#loc53 at #loc5))
#loc96 = loc(callsite(#loc54 at #loc5))
#loc97 = loc(callsite(#loc15 at #loc5))
#loc98 = loc(callsite(#loc55 at #loc5))
#loc99 = loc(callsite(#loc56 at #loc2))
#loc100 = loc(callsite(#loc57 at #loc2))
#loc101 = loc(callsite(#loc38 at #loc2))
#loc102 = loc(callsite(#loc39 at #loc2))
#loc103 = loc(callsite(#loc40 at #loc2))
#loc104 = loc(callsite(#loc58 at #loc2))
#loc107 = loc(callsite(#loc61 at #loc2))
#loc108 = loc(callsite(#loc62 at #loc2))
#loc109 = loc(callsite(#loc63 at #loc2))
#loc111 = loc(callsite(#loc45 at #loc2))
#loc112 = loc(callsite(#loc46 at #loc2))
#loc113 = loc(callsite(#loc48 at #loc2))
#loc114 = loc(callsite(#loc49 at #loc2))
#loc115 = loc(callsite(#loc50 at #loc2))
#loc116 = loc(callsite(#loc51 at #loc2))
#loc117 = loc(callsite(#loc52 at #loc2))
#loc118 = loc(callsite(#loc53 at #loc2))
#loc119 = loc(callsite(#loc54 at #loc2))
#loc120 = loc(callsite(#loc55 at #loc2))
#loc123 = loc(callsite(#loc84 at #loc7))
#loc125 = loc(callsite(#loc89 at #loc17))
#loc127 = loc(callsite(#loc84 at #loc60))
#loc128 = loc(callsite(#loc123 at #loc5))
#loc129 = loc(callsite(#loc125 at #loc5))
#loc130 = loc(callsite(#loc127 at #loc2))
#loc131 = loc(callsite(#loc125 at #loc2))


hujiajun
7月31日 评论:
7月31日 评论:
// Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:31.306Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:31.307Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpaeq3vsbv/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpaeq3vsbv/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:31.307Z] RERUN
// [2026-07-31T14:28:39.030Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype12-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:39.030Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:39.030Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpsjg92szj/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpsjg92szj/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:39.030Z] RERUN
// [2026-07-31T14:28:48.421Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype12-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:48.421Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:48.421Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmptyz1abkc/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmptyz1abkc/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:48.421Z] RERUN
// [2026-07-31T14:28:56.144Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype12-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:56.144Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:56.144Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp0rg0uh8p/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp0rg0uh8p/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:56.144Z] RERUN
// [2026-07-31T14:28:59.399Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-8192-128-True-dtype12-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/w8JWiyRW-PyeHY0T0w0dKv93p6R-nTK6O_BbfUq5Yis
// [2026-07-31T14:28:59.399Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:28:59.400Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpft63b7xw/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpft63b7xw/kernel --enable-vf-merge-level=1
// [2026-07-31T14:28:59.400Z] FAILED
#loc = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)
#loc2 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":212:70)
#loc5 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":202:78)
#loc6 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":195:44)
#loc7 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:46)
#loc17 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":282:36)
#loc18 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":88:25)
#loc41 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":81:22)
#loc44 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":84:24)
#loc59 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:33)
#loc60 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:46)
#loc73 = loc(callsite(#loc6 at #loc7))
#loc78 = loc(callsite(#loc17 at #loc18))
#loc83 = loc(callsite(#loc41 at #loc5))
#loc86 = loc(callsite(#loc44 at #loc5))
#loc105 = loc(callsite(#loc59 at #loc2))
#loc106 = loc(callsite(#loc6 at #loc60))
#loc110 = loc(callsite(#loc44 at #loc2))
#loc121 = loc(callsite(#loc73 at #loc5))
#loc122 = loc(callsite(#loc78 at #loc2))
#loc124 = loc(callsite(#loc78 at #loc5))
#loc126 = loc(callsite(#loc106 at #loc2))
module attributes {hacc.target = #hacc.target<"Ascend950PR_9579">} {
func.func @_attn_fwd(%arg0: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg1: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg2: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg3: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg4: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg5: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg6: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg7: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg8: memref<?xi32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg9: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg10: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg11: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg12: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg13: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg14: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg15: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)) attributes {SyncBlockLockArgIdx = 0 : i64, WorkspaceArgIdx = 1 : i64, global_kernel = "local", mix_mode = "mix", parallel_mode = "simd"} {
%c8192 = arith.constant 8192 : index loc(#loc71)
%c128 = arith.constant 128 : index loc(#loc71)
%c1 = arith.constant 1 : index loc(#loc3)
%c0 = arith.constant 0 : index loc(#loc4)
%cst = arith.constant 1.000000e+00 : f32 loc(#loc72)
%cst_0 = arith.constant 0xFF800000 : f32 loc(#loc121)
%cst_1 = arith.constant -1.000000e+04 : f32 loc(#loc74)
%cst_2 = arith.constant 5.000000e-01 : f32 loc(#loc75)
%c0_i32 = arith.constant 0 : i32 loc(#loc71)
%c1_i32 = arith.constant 1 : i32 loc(#loc76)
%c64_i32 = arith.constant 64 : i32 loc(#loc11)
%c8_i32 = arith.constant 8 : i32 loc(#loc12)
%c8388608_i64 = arith.constant 8388608 : i64 loc(#loc13)
%c1048576_i64 = arith.constant 1048576 : i64 loc(#loc14)
%c128_i32 = arith.constant 128 : i32 loc(#loc77)
%c8192_i32 = arith.constant 8192 : i32 loc(#loc16)
%cst_3 = arith.constant 0.000000e+00 : f32 loc(#loc122)
%0 = tensor.empty() : tensor<128x128xf32> loc(#loc79)
%1 = linalg.fill ins(%cst_3 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
%2 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc75)
%3 = linalg.fill ins(%cst_1 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
%4 = tensor.empty() : tensor<128xf32> loc(#loc121)
%5 = linalg.fill ins(%cst_0 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc121)
%6 = linalg.fill ins(%cst : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc72)
%7 = arith.index_cast %arg13 : i32 to index loc(#loc20)
%reinterpret_cast = memref.reinterpret_cast %arg8 to offset: [%7], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc21)
%8 = memref.load %reinterpret_cast[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc22)
%9 = arith.addi %7, %c1 : index loc(#loc3)
%reinterpret_cast_4 = memref.reinterpret_cast %arg8 to offset: [%9], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc3)
%10 = memref.load %reinterpret_cast_4[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc4)
scf.for %arg16 = %8 to %10 step %c1_i32 : i32 {
%11 = arith.divsi %arg16, %c64_i32 : i32 loc(#loc24)
%12 = arith.remsi %arg16, %c64_i32 : i32 loc(#loc11)
%13 = arith.divsi %11, %c8_i32 : i32 loc(#loc25)
%14 = arith.remsi %11, %c8_i32 : i32 loc(#loc26)
%15 = arith.muli %14, %arg9 : i32 loc(#loc27)
%16 = arith.divsi %15, %c8_i32 : i32 loc(#loc12)
%17 = arith.extsi %13 : i32 to i64 loc(#loc28)
%18 = arith.muli %17, %c8388608_i64 : i64 loc(#loc13)
%19 = arith.extsi %14 : i32 to i64 loc(#loc29)
%20 = arith.muli %19, %c1048576_i64 : i64 loc(#loc30)
%21 = arith.addi %18, %20 : i64 loc(#loc31)
%22 = arith.extsi %16 : i32 to i64 loc(#loc32)
%23 = arith.muli %22, %c1048576_i64 : i64 loc(#loc14)
%24 = arith.addi %18, %23 : i64 loc(#loc33)
%25 = arith.index_cast %21 : i64 to index loc(#loc31)
%26 = arith.muli %12, %c128_i32 : i32 loc(#loc34)
%27 = arith.maxsi %26, %c0_i32 : i32 loc(#loc35)
%28 = arith.index_cast %27 : i32 to index loc(#loc35)
%29 = arith.muli %28, %c128 : index loc(#loc35)
%30 = arith.addi %29, %25 : index loc(#loc35)
%reinterpret_cast_5 = memref.reinterpret_cast %arg2 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc35)
%31 = arith.index_cast %24 : i64 to index loc(#loc33)
%reinterpret_cast_6 = memref.reinterpret_cast %arg7 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc36)
%alloc = memref.alloc() : memref<128x128xbf16> loc(#loc37)
memref.copy %reinterpret_cast_5, %alloc : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc37)
%32 = bufferization.to_tensor %alloc restrict writable : memref<128x128xbf16> loc(#loc37)
%33:5 = scf.for %arg17 = %c0_i32 to %26 step %c128_i32 iter_args(%arg18 = %6, %arg19 = %1, %arg20 = %5, %arg21 = %c0_i32, %arg22 = %c0_i32) -> (tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32) : i32 {
%46 = arith.maxsi %arg21, %c0_i32 : i32 loc(#loc72)
%47 = arith.index_cast %46 : i32 to index loc(#loc72)
%48 = arith.muli %47, %c128 : index loc(#loc72)
%49 = arith.addi %48, %31 : index loc(#loc72)
%reinterpret_cast_8 = memref.reinterpret_cast %arg4 to offset: [%49], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc72)
%50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc72)
%51 = arith.index_cast %50 : i32 to index loc(#loc72)
%52 = arith.muli %51, %c128 : index loc(#loc72)
%53 = arith.addi %52, %31 : index loc(#loc72)
%reinterpret_cast_9 = memref.reinterpret_cast %arg3 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc72)
%alloc_10 = memref.alloc() : memref<128x128xbf16> loc(#loc80)
memref.copy %reinterpret_cast_9, %alloc_10 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc80)
%54 = bufferization.to_tensor %alloc_10 restrict writable : memref<128x128xbf16> loc(#loc80)
%55 = tensor.empty() : tensor<128x128xbf16> loc(#loc81)
%transposed = linalg.transpose ins(%54 : tensor<128x128xbf16>) outs(%55 : tensor<128x128xbf16>) permutation = [1, 0] loc(#loc81)
%56 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc82)
%57 = arith.mulf %56, %2 : tensor<128x128xf32> loc(#loc83)
%reduced = linalg.reduce ins(%57 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc41 at #loc5)), %init: f32 loc(callsite(#loc73 at #loc5))) {
%72 = arith.maximumf %in, %init : f32 loc(#loc128)
linalg.yield %72 : f32 loc(#loc121)
} loc(#loc121)
%58 = arith.maximumf %arg20, %reduced : tensor<128xf32> loc(#loc85)
%broadcasted_11 = linalg.broadcast ins(%58 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc79)
%59 = arith.subf %57, %broadcasted_11 : tensor<128x128xf32> loc(#loc79)
%60 = math.exp %59 : tensor<128x128xf32> loc(#loc86)
%61 = arith.truncf %60 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc87)
%alloc_12 = memref.alloc() : memref<128x128xbf16> loc(#loc88)
memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc88)
%62 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xbf16> loc(#loc88)
%63 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc124)
%reduced_13 = linalg.reduce ins(%60 : tensor<128x128xf32>) outs(%63 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc44 at #loc5)), %init: f32 loc(callsite(#loc78 at #loc5))) {
%72 = arith.addf %in, %init : f32 loc(#loc129)
linalg.yield %72 : f32 loc(#loc124)
} loc(#loc124)
%64 = arith.subf %arg20, %58 : tensor<128xf32> loc(#loc90)
%65 = math.exp %64 : tensor<128xf32> loc(#loc91)
%66 = arith.mulf %arg18, %65 : tensor<128xf32> loc(#loc92)
%67 = arith.addf %66, %reduced_13 : tensor<128xf32> loc(#loc93)
%broadcasted_14 = linalg.broadcast ins(%65 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc94)
%68 = arith.mulf %arg19, %broadcasted_14 : tensor<128x128xf32> loc(#loc94)
%69 = linalg.matmul {input_precision = "ieee"} ins(%61, %62 : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%68 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc95)
%70 = arith.addi %arg21, %c128_i32 : i32 loc(#loc96)
%71 = arith.addi %arg22, %c128_i32 : i32 loc(#loc97)
scf.yield %67, %69, %58, %70, %71 : tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc98)
} {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc72)
%34 = arith.muli %12, %c128_i32 {tt.divisibility = dense<128> : tensor<1xi32>} : i32 loc(#loc99)
%35 = arith.addi %12, %c1_i32 : i32 loc(#loc76)
%36 = arith.muli %35, %c128_i32 : i32 loc(#loc100)
%37:6 = scf.for %arg17 = %34 to %36 step %c128_i32 iter_args(%arg18 = %34, %arg19 = %33#0, %arg20 = %33#1, %arg21 = %33#2, %arg22 = %34, %arg23 = %34) -> (i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32) : i32 {
%46 = arith.maxsi %arg18, %c0_i32 : i32 loc(#loc71)
%47 = arith.index_cast %46 : i32 to index loc(#loc71)
%48 = arith.muli %28, %c8192 : index loc(#loc71)
%49 = arith.addi %48, %47 : index loc(#loc71)
%reinterpret_cast_8 = memref.reinterpret_cast %arg5 to offset: [%49], sizes: [128, 128], strides: [8192, 1] : memref<?xf32> to memref<128x128xf32, strided<[8192, 1], offset: ?>> loc(#loc71)
%50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc71)
%51 = arith.index_cast %50 : i32 to index loc(#loc71)
%52 = arith.muli %51, %c128 : index loc(#loc71)
%53 = arith.addi %52, %31 : index loc(#loc71)
%reinterpret_cast_9 = memref.reinterpret_cast %arg4 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc71)
%54 = arith.maxsi %arg23, %c0_i32 : i32 loc(#loc71)
%55 = arith.index_cast %54 : i32 to index loc(#loc71)
%56 = arith.muli %55, %c128 : index loc(#loc71)
%57 = arith.addi %56, %31 : index loc(#loc71)
%reinterpret_cast_10 = memref.reinterpret_cast %arg3 to offset: [%57], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc71)
%alloc_11 = memref.alloc() : memref<128x128xbf16> loc(#loc101)
memref.copy %reinterpret_cast_10, %alloc_11 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc101)
%58 = bufferization.to_tensor %alloc_11 restrict writable : memref<128x128xbf16> loc(#loc101)
%59 = tensor.empty() : tensor<128x128xbf16> loc(#loc102)
%transposed = linalg.transpose ins(%58 : tensor<128x128xbf16>) outs(%59 : tensor<128x128xbf16>) permutation = [1, 0] loc(#loc102)
%60 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc103)
%alloc_12 = memref.alloc() : memref<128x128xf32> loc(#loc104)
memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf32, strided<[8192, 1], offset: ?>> to memref<128x128xf32> loc(#loc104)
%61 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf32> loc(#loc104)
%62 = arith.mulf %60, %2 : tensor<128x128xf32> loc(#loc75)
%63 = arith.cmpf une, %61, %1 : tensor<128x128xf32> loc(#loc74)
%64 = arith.select %63, %3, %1 : tensor<128x128xi1>, tensor<128x128xf32> loc(#loc74)
%65 = arith.addf %62, %64 : tensor<128x128xf32> loc(#loc105)
%reduced = linalg.reduce ins(%65 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc59 at #loc2)), %init: f32 loc(callsite(#loc106 at #loc2))) {
%81 = arith.maximumf %in, %init : f32 loc(#loc130)
linalg.yield %81 : f32 loc(#loc126)
} loc(#loc126)
%66 = arith.maximumf %arg21, %reduced : tensor<128xf32> loc(#loc107)
%broadcasted_13 = linalg.broadcast ins(%66 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc108)
%67 = arith.subf %65, %broadcasted_13 : tensor<128x128xf32> loc(#loc108)
%68 = arith.addi %arg18, %c128_i32 : i32 loc(#loc109)
%69 = math.exp %67 : tensor<128x128xf32> loc(#loc110)
%70 = arith.truncf %69 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc111)
%alloc_14 = memref.alloc() : memref<128x128xbf16> loc(#loc112)
memref.copy %reinterpret_cast_9, %alloc_14 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc112)
%71 = bufferization.to_tensor %alloc_14 restrict writable : memref<128x128xbf16> loc(#loc112)
%72 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc122)
%reduced_15 = linalg.reduce ins(%69 : tensor<128x128xf32>) outs(%72 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc44 at #loc2)), %init: f32 loc(callsite(#loc78 at #loc2))) {
%81 = arith.addf %in, %init : f32 loc(#loc131)
linalg.yield %81 : f32 loc(#loc122)
} loc(#loc122)
%73 = arith.subf %arg21, %66 : tensor<128xf32> loc(#loc113)
%74 = math.exp %73 : tensor<128xf32> loc(#loc114)
%75 = arith.mulf %arg19, %74 : tensor<128xf32> loc(#loc115)
%76 = arith.addf %75, %reduced_15 : tensor<128xf32> loc(#loc116)
%broadcasted_16 = linalg.broadcast ins(%74 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc117)
%77 = arith.mulf %arg20, %broadcasted_16 : tensor<128x128xf32> loc(#loc117)
%78 = linalg.matmul {input_precision = "ieee"} ins(%70, %71 : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%77 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc118)
%79 = arith.addi %arg22, %c128_i32 : i32 loc(#loc119)
%80 = arith.addi %arg23, %c128_i32 : i32 loc(#loc77)
scf.yield %68, %76, %78, %66, %79, %80 : i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc120)
} {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc71)
%38 = math.log %37#1 : tensor<128xf32> loc(#loc64)
%39 = arith.addf %37#3, %38 : tensor<128xf32> loc(#loc65)
%broadcasted = linalg.broadcast ins(%37#1 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc66)
%40 = arith.divf %37#2, %broadcasted : tensor<128x128xf32> loc(#loc66)
%41 = arith.muli %11, %c8192_i32 : i32 loc(#loc16)
%42 = arith.index_cast %41 : i32 to index loc(#loc16)
%43 = arith.index_cast %26 : i32 to index loc(#loc34)
%44 = arith.addi %42, %43 : index loc(#loc67)
%reinterpret_cast_7 = memref.reinterpret_cast %arg6 to offset: [%44], sizes: [128], strides: [1] : memref<?xf32> to memref<128xf32, strided<[1], offset: ?>> loc(#loc67)
bufferization.materialize_in_destination %39 in writable %reinterpret_cast_7 : (tensor<128xf32>, memref<128xf32, strided<[1], offset: ?>>) -> () loc(#loc68)
%45 = arith.truncf %40 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc69)
bufferization.materialize_in_destination %45 in writable %reinterpret_cast_6 : (tensor<128x128xbf16>, memref<128x128xbf16, strided<[128, 1], offset: ?>>) -> () loc(#loc70)
} loc(#loc23)
return loc(#loc)
} loc(#loc)
} loc(#loc)
#loc1 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":63:33)
#loc3 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:53)
#loc4 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:28)
#loc8 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:56)
#loc9 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:22)
#loc10 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:47)
#loc11 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":139:33)
#loc12 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:39)
#loc13 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:40)
#loc14 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:72)
#loc15 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:46)
#loc16 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:35)
#loc19 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":83:22)
#loc20 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":123:24)
#loc21 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:49)
#loc22 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:30)
#loc23 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":137:51)
#loc24 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":138:35)
#loc25 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":140:31)
#loc26 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":141:30)
#loc27 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:23)
#loc28 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:28)
#loc29 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:61)
#loc30 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:73)
#loc31 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:52)
#loc32 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:60)
#loc33 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:52)
#loc34 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":152:34)
#loc35 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":154:12)
#loc36 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":180:12)
#loc37 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":193:20)
#loc38 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":66:20)
#loc39 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":69:27)
#loc40 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":70:23)
#loc42 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":168:27)
#loc43 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:35)
#loc45 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":85:22)
#loc46 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":86:20)
#loc47 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":271:15)
#loc48 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:34)
#loc49 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:28)
#loc50 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:20)
#loc51 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:28)
#loc52 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":94:20)
#loc53 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":87:28)
#loc54 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":97:46)
#loc55 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:8)
#loc56 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:27)
#loc57 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:52)
#loc58 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":75:27)
#loc61 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:35)
#loc62 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":78:18)
#loc63 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":79:54)
#loc64 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:27)
#loc65 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:15)
#loc66 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":216:20)
#loc67 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:43)
#loc68 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":219:25)
#loc69 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:37)
#loc70 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:30)
#loc71 = loc(callsite(#loc1 at #loc2))
#loc72 = loc(callsite(#loc1 at #loc5))
#loc74 = loc(callsite(#loc8 at #loc2))
#loc75 = loc(callsite(#loc9 at #loc2))
#loc76 = loc(callsite(#loc10 at #loc2))
#loc77 = loc(callsite(#loc15 at #loc2))
#loc79 = loc(callsite(#loc19 at #loc5))
#loc80 = loc(callsite(#loc38 at #loc5))
#loc81 = loc(callsite(#loc39 at #loc5))
#loc82 = loc(callsite(#loc40 at #loc5))
#loc84 = loc(callsite(#loc42 at #loc6))
#loc85 = loc(callsite(#loc43 at #loc5))
#loc87 = loc(callsite(#loc45 at #loc5))
#loc88 = loc(callsite(#loc46 at #loc5))
#loc89 = loc(callsite(#loc47 at #loc17))
#loc90 = loc(callsite(#loc48 at #loc5))
#loc91 = loc(callsite(#loc49 at #loc5))
#loc92 = loc(callsite(#loc50 at #loc5))
#loc93 = loc(callsite(#loc51 at #loc5))
#loc94 = loc(callsite(#loc52 at #loc5))
#loc95 = loc(callsite(#loc53 at #loc5))
#loc96 = loc(callsite(#loc54 at #loc5))
#loc97 = loc(callsite(#loc15 at #loc5))
#loc98 = loc(callsite(#loc55 at #loc5))
#loc99 = loc(callsite(#loc56 at #loc2))
#loc100 = loc(callsite(#loc57 at #loc2))
#loc101 = loc(callsite(#loc38 at #loc2))
#loc102 = loc(callsite(#loc39 at #loc2))
#loc103 = loc(callsite(#loc40 at #loc2))
#loc104 = loc(callsite(#loc58 at #loc2))
#loc107 = loc(callsite(#loc61 at #loc2))
#loc108 = loc(callsite(#loc62 at #loc2))
#loc109 = loc(callsite(#loc63 at #loc2))
#loc111 = loc(callsite(#loc45 at #loc2))
#loc112 = loc(callsite(#loc46 at #loc2))
#loc113 = loc(callsite(#loc48 at #loc2))
#loc114 = loc(callsite(#loc49 at #loc2))
#loc115 = loc(callsite(#loc50 at #loc2))
#loc116 = loc(callsite(#loc51 at #loc2))
#loc117 = loc(callsite(#loc52 at #loc2))
#loc118 = loc(callsite(#loc53 at #loc2))
#loc119 = loc(callsite(#loc54 at #loc2))
#loc120 = loc(callsite(#loc55 at #loc2))
#loc123 = loc(callsite(#loc84 at #loc7))
#loc125 = loc(callsite(#loc89 at #loc18))
#loc127 = loc(callsite(#loc84 at #loc60))
#loc128 = loc(callsite(#loc123 at #loc5))
#loc129 = loc(callsite(#loc125 at #loc5))
#loc130 = loc(callsite(#loc127 at #loc2))
#loc131 = loc(callsite(#loc125 at #loc2))


hujiajun
7月31日 评论:
7月31日 评论:
// [2026-07-31T14:29:18.826Z] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:18.826Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:18.826Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpg2gwakv8/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpg2gwakv8/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:18.826Z] RERUN
// [2026-07-31T14:29:26.544Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype14-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:26.544Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:26.544Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpe5zws9b9/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpe5zws9b9/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:26.544Z] RERUN
// [2026-07-31T14:29:35.968Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype14-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:35.968Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:35.968Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmphdpje0xa/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmphdpje0xa/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:35.968Z] RERUN
// [2026-07-31T14:29:43.690Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype14-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:43.690Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:43.690Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmp9dcmatiq/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmp9dcmatiq/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:43.690Z] RERUN
// [2026-07-31T14:29:46.205Z] triton-ops-ascend-npu-ir-ci/native/fa/test_fa_fwd.py::test_op[128-8-8-1024-128-True-dtype14-128-128] Dumping intermediate results to /home/jenkins/.triton/dump/iIJPQ90w4g8LT6w3nloF1DdxDDAGYk8lZx_DuEMN-z4
// [2026-07-31T14:29:46.206Z] SSBUFFER return code=2, will fallback to enable_dynamic_cv_pipeline=False
// [2026-07-31T14:29:46.206Z] [DEBUG] cmd_list: /home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/bishengir-toolkit/tools/bishengir/bin/bishengir-compile /tmp/tmpgsxkhje7/kernel.ttadapter.mlir --target=Ascend950PR_9579 --enable-auto-multi-buffer=False --enable-auto-bind-sub-block=True --disable-ffts --enable-hivm-graph-sync-solver=True --set-workspace-multibuffer=2 --limit-auto-multi-buffer-of-local-buffer=no-limit --enable-mixed-cv=True --enable-flatten=False --enable-auto-blockify-loop --enable-hfusion-compile=true --enable-triton-kernel-compile=true --append-bisheng-options=-cce-link-aicore-ll-module /home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/backends/ascend/lib/libdevice.10.bc --bishengir-print-ir-after=hivm-graph-sync-solver -o /tmp/tmpgsxkhje7/kernel --enable-vf-merge-level=1
// [2026-07-31T14:29:46.206Z] FAILED
#loc = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)
#loc2 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":212:70)
#loc5 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":202:78)
#loc6 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":195:44)
#loc7 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:46)
#loc16 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":282:36)
#loc17 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":88:25)
#loc41 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":81:22)
#loc44 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":84:24)
#loc59 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:33)
#loc60 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:46)
#loc73 = loc(callsite(#loc6 at #loc7))
#loc78 = loc(callsite(#loc16 at #loc17))
#loc83 = loc(callsite(#loc41 at #loc5))
#loc86 = loc(callsite(#loc44 at #loc5))
#loc105 = loc(callsite(#loc59 at #loc2))
#loc106 = loc(callsite(#loc6 at #loc60))
#loc110 = loc(callsite(#loc44 at #loc2))
#loc121 = loc(callsite(#loc73 at #loc5))
#loc122 = loc(callsite(#loc78 at #loc2))
#loc124 = loc(callsite(#loc78 at #loc5))
#loc126 = loc(callsite(#loc106 at #loc2))
module attributes {hacc.target = #hacc.target<"Ascend950PR_9579">} {
func.func @_attn_fwd(%arg0: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg1: memref<?xi8> loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg2: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg3: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg4: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg5: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg6: memref<?xf32> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg7: memref<?xbf16> {tt.divisibility = 16 : i32, tt.tensor_kind = 1 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg8: memref<?xi32> {tt.divisibility = 16 : i32, tt.tensor_kind = 0 : i32} loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg9: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg10: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg11: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg12: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg13: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg14: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0), %arg15: i32 loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":105:0)) attributes {SyncBlockLockArgIdx = 0 : i64, WorkspaceArgIdx = 1 : i64, global_kernel = "local", mix_mode = "mix", parallel_mode = "simd"} {
%c1024 = arith.constant 1024 : index loc(#loc71)
%c128 = arith.constant 128 : index loc(#loc71)
%c1 = arith.constant 1 : index loc(#loc3)
%c0 = arith.constant 0 : index loc(#loc4)
%cst = arith.constant 1.000000e+00 : f32 loc(#loc72)
%cst_0 = arith.constant 0xFF800000 : f32 loc(#loc121)
%cst_1 = arith.constant -1.000000e+04 : f32 loc(#loc74)
%cst_2 = arith.constant 5.000000e-01 : f32 loc(#loc75)
%c1024_i32 = arith.constant 1024 : i32 loc(#loc10)
%c0_i32 = arith.constant 0 : i32 loc(#loc71)
%c1_i32 = arith.constant 1 : i32 loc(#loc76)
%c8_i32 = arith.constant 8 : i32 loc(#loc12)
%c1048576_i64 = arith.constant 1048576 : i64 loc(#loc13)
%c131072_i64 = arith.constant 131072 : i64 loc(#loc14)
%c128_i32 = arith.constant 128 : i32 loc(#loc77)
%cst_3 = arith.constant 0.000000e+00 : f32 loc(#loc122)
%0 = tensor.empty() : tensor<128x128xf32> loc(#loc79)
%1 = linalg.fill ins(%cst_3 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
%2 = linalg.fill ins(%cst_2 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc75)
%3 = linalg.fill ins(%cst_1 : f32) outs(%0 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc74)
%4 = tensor.empty() : tensor<128xf32> loc(#loc121)
%5 = linalg.fill ins(%cst_0 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc121)
%6 = linalg.fill ins(%cst : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc72)
%7 = arith.index_cast %arg13 : i32 to index loc(#loc19)
%reinterpret_cast = memref.reinterpret_cast %arg8 to offset: [%7], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc20)
%8 = memref.load %reinterpret_cast[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc21)
%9 = arith.addi %7, %c1 : index loc(#loc3)
%reinterpret_cast_4 = memref.reinterpret_cast %arg8 to offset: [%9], sizes: [1], strides: [1] : memref<?xi32> to memref<1xi32, strided<[1], offset: ?>> loc(#loc3)
%10 = memref.load %reinterpret_cast_4[%c0] : memref<1xi32, strided<[1], offset: ?>> loc(#loc4)
scf.for %arg16 = %8 to %10 step %c1_i32 : i32 {
%11 = arith.divsi %arg16, %c8_i32 : i32 loc(#loc23)
%12 = arith.remsi %arg16, %c8_i32 : i32 loc(#loc24)
%13 = arith.divsi %11, %c8_i32 : i32 loc(#loc25)
%14 = arith.remsi %11, %c8_i32 : i32 loc(#loc26)
%15 = arith.muli %14, %arg9 : i32 loc(#loc27)
%16 = arith.divsi %15, %c8_i32 : i32 loc(#loc12)
%17 = arith.extsi %13 : i32 to i64 loc(#loc28)
%18 = arith.muli %17, %c1048576_i64 : i64 loc(#loc13)
%19 = arith.extsi %14 : i32 to i64 loc(#loc29)
%20 = arith.muli %19, %c131072_i64 : i64 loc(#loc30)
%21 = arith.addi %18, %20 : i64 loc(#loc31)
%22 = arith.extsi %16 : i32 to i64 loc(#loc32)
%23 = arith.muli %22, %c131072_i64 : i64 loc(#loc14)
%24 = arith.addi %18, %23 : i64 loc(#loc33)
%25 = arith.index_cast %21 : i64 to index loc(#loc31)
%26 = arith.muli %12, %c128_i32 : i32 loc(#loc34)
%27 = arith.maxsi %26, %c0_i32 : i32 loc(#loc35)
%28 = arith.index_cast %27 : i32 to index loc(#loc35)
%29 = arith.muli %28, %c128 : index loc(#loc35)
%30 = arith.addi %29, %25 : index loc(#loc35)
%reinterpret_cast_5 = memref.reinterpret_cast %arg2 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc35)
%31 = arith.index_cast %24 : i64 to index loc(#loc33)
%reinterpret_cast_6 = memref.reinterpret_cast %arg7 to offset: [%30], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc36)
%alloc = memref.alloc() : memref<128x128xbf16> loc(#loc37)
memref.copy %reinterpret_cast_5, %alloc : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc37)
%32 = bufferization.to_tensor %alloc restrict writable : memref<128x128xbf16> loc(#loc37)
%33:5 = scf.for %arg17 = %c0_i32 to %26 step %c128_i32 iter_args(%arg18 = %6, %arg19 = %1, %arg20 = %5, %arg21 = %c0_i32, %arg22 = %c0_i32) -> (tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32) : i32 {
%46 = arith.maxsi %arg21, %c0_i32 : i32 loc(#loc72)
%47 = arith.index_cast %46 : i32 to index loc(#loc72)
%48 = arith.muli %47, %c128 : index loc(#loc72)
%49 = arith.addi %48, %31 : index loc(#loc72)
%reinterpret_cast_8 = memref.reinterpret_cast %arg4 to offset: [%49], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc72)
%50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc72)
%51 = arith.index_cast %50 : i32 to index loc(#loc72)
%52 = arith.muli %51, %c128 : index loc(#loc72)
%53 = arith.addi %52, %31 : index loc(#loc72)
%reinterpret_cast_9 = memref.reinterpret_cast %arg3 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc72)
%alloc_10 = memref.alloc() : memref<128x128xbf16> loc(#loc80)
memref.copy %reinterpret_cast_9, %alloc_10 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc80)
%54 = bufferization.to_tensor %alloc_10 restrict writable : memref<128x128xbf16> loc(#loc80)
%55 = tensor.empty() : tensor<128x128xbf16> loc(#loc81)
%transposed = linalg.transpose ins(%54 : tensor<128x128xbf16>) outs(%55 : tensor<128x128xbf16>) permutation = [1, 0] loc(#loc81)
%56 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc82)
%57 = arith.mulf %56, %2 : tensor<128x128xf32> loc(#loc83)
%reduced = linalg.reduce ins(%57 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc41 at #loc5)), %init: f32 loc(callsite(#loc73 at #loc5))) {
%72 = arith.maximumf %in, %init : f32 loc(#loc128)
linalg.yield %72 : f32 loc(#loc121)
} loc(#loc121)
%58 = arith.maximumf %arg20, %reduced : tensor<128xf32> loc(#loc85)
%broadcasted_11 = linalg.broadcast ins(%58 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc79)
%59 = arith.subf %57, %broadcasted_11 : tensor<128x128xf32> loc(#loc79)
%60 = math.exp %59 : tensor<128x128xf32> loc(#loc86)
%61 = arith.truncf %60 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc87)
%alloc_12 = memref.alloc() : memref<128x128xbf16> loc(#loc88)
memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc88)
%62 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xbf16> loc(#loc88)
%63 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc124)
%reduced_13 = linalg.reduce ins(%60 : tensor<128x128xf32>) outs(%63 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc44 at #loc5)), %init: f32 loc(callsite(#loc78 at #loc5))) {
%72 = arith.addf %in, %init : f32 loc(#loc129)
linalg.yield %72 : f32 loc(#loc124)
} loc(#loc124)
%64 = arith.subf %arg20, %58 : tensor<128xf32> loc(#loc90)
%65 = math.exp %64 : tensor<128xf32> loc(#loc91)
%66 = arith.mulf %arg18, %65 : tensor<128xf32> loc(#loc92)
%67 = arith.addf %66, %reduced_13 : tensor<128xf32> loc(#loc93)
%broadcasted_14 = linalg.broadcast ins(%65 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc94)
%68 = arith.mulf %arg19, %broadcasted_14 : tensor<128x128xf32> loc(#loc94)
%69 = linalg.matmul {input_precision = "ieee"} ins(%61, %62 : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%68 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc95)
%70 = arith.addi %arg21, %c128_i32 : i32 loc(#loc96)
%71 = arith.addi %arg22, %c128_i32 : i32 loc(#loc97)
scf.yield %67, %69, %58, %70, %71 : tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc98)
} {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc72)
%34 = arith.muli %12, %c128_i32 {tt.divisibility = dense<128> : tensor<1xi32>} : i32 loc(#loc99)
%35 = arith.addi %12, %c1_i32 : i32 loc(#loc76)
%36 = arith.muli %35, %c128_i32 : i32 loc(#loc100)
%37:6 = scf.for %arg17 = %34 to %36 step %c128_i32 iter_args(%arg18 = %34, %arg19 = %33#0, %arg20 = %33#1, %arg21 = %33#2, %arg22 = %34, %arg23 = %34) -> (i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32) : i32 {
%46 = arith.maxsi %arg18, %c0_i32 : i32 loc(#loc71)
%47 = arith.index_cast %46 : i32 to index loc(#loc71)
%48 = arith.muli %28, %c1024 : index loc(#loc71)
%49 = arith.addi %48, %47 : index loc(#loc71)
%reinterpret_cast_8 = memref.reinterpret_cast %arg5 to offset: [%49], sizes: [128, 128], strides: [1024, 1] : memref<?xf32> to memref<128x128xf32, strided<[1024, 1], offset: ?>> loc(#loc71)
%50 = arith.maxsi %arg22, %c0_i32 : i32 loc(#loc71)
%51 = arith.index_cast %50 : i32 to index loc(#loc71)
%52 = arith.muli %51, %c128 : index loc(#loc71)
%53 = arith.addi %52, %31 : index loc(#loc71)
%reinterpret_cast_9 = memref.reinterpret_cast %arg4 to offset: [%53], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc71)
%54 = arith.maxsi %arg23, %c0_i32 : i32 loc(#loc71)
%55 = arith.index_cast %54 : i32 to index loc(#loc71)
%56 = arith.muli %55, %c128 : index loc(#loc71)
%57 = arith.addi %56, %31 : index loc(#loc71)
%reinterpret_cast_10 = memref.reinterpret_cast %arg3 to offset: [%57], sizes: [128, 128], strides: [128, 1] : memref<?xbf16> to memref<128x128xbf16, strided<[128, 1], offset: ?>> loc(#loc71)
%alloc_11 = memref.alloc() : memref<128x128xbf16> loc(#loc101)
memref.copy %reinterpret_cast_10, %alloc_11 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc101)
%58 = bufferization.to_tensor %alloc_11 restrict writable : memref<128x128xbf16> loc(#loc101)
%59 = tensor.empty() : tensor<128x128xbf16> loc(#loc102)
%transposed = linalg.transpose ins(%58 : tensor<128x128xbf16>) outs(%59 : tensor<128x128xbf16>) permutation = [1, 0] loc(#loc102)
%60 = linalg.matmul {input_precision = "ieee"} ins(%32, %transposed : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%1 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc103)
%alloc_12 = memref.alloc() : memref<128x128xf32> loc(#loc104)
memref.copy %reinterpret_cast_8, %alloc_12 : memref<128x128xf32, strided<[1024, 1], offset: ?>> to memref<128x128xf32> loc(#loc104)
%61 = bufferization.to_tensor %alloc_12 restrict writable : memref<128x128xf32> loc(#loc104)
%62 = arith.mulf %60, %2 : tensor<128x128xf32> loc(#loc75)
%63 = arith.cmpf une, %61, %1 : tensor<128x128xf32> loc(#loc74)
%64 = arith.select %63, %3, %1 : tensor<128x128xi1>, tensor<128x128xf32> loc(#loc74)
%65 = arith.addf %62, %64 : tensor<128x128xf32> loc(#loc105)
%reduced = linalg.reduce ins(%65 : tensor<128x128xf32>) outs(%5 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc59 at #loc2)), %init: f32 loc(callsite(#loc106 at #loc2))) {
%81 = arith.maximumf %in, %init : f32 loc(#loc130)
linalg.yield %81 : f32 loc(#loc126)
} loc(#loc126)
%66 = arith.maximumf %arg21, %reduced : tensor<128xf32> loc(#loc107)
%broadcasted_13 = linalg.broadcast ins(%66 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc108)
%67 = arith.subf %65, %broadcasted_13 : tensor<128x128xf32> loc(#loc108)
%68 = arith.addi %arg18, %c128_i32 : i32 loc(#loc109)
%69 = math.exp %67 : tensor<128x128xf32> loc(#loc110)
%70 = arith.truncf %69 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc111)
%alloc_14 = memref.alloc() : memref<128x128xbf16> loc(#loc112)
memref.copy %reinterpret_cast_9, %alloc_14 : memref<128x128xbf16, strided<[128, 1], offset: ?>> to memref<128x128xbf16> loc(#loc112)
%71 = bufferization.to_tensor %alloc_14 restrict writable : memref<128x128xbf16> loc(#loc112)
%72 = linalg.fill ins(%cst_3 : f32) outs(%4 : tensor<128xf32>) -> tensor<128xf32> loc(#loc122)
%reduced_15 = linalg.reduce ins(%69 : tensor<128x128xf32>) outs(%72 : tensor<128xf32>) dimensions = [1]
(%in: f32 loc(callsite(#loc44 at #loc2)), %init: f32 loc(callsite(#loc78 at #loc2))) {
%81 = arith.addf %in, %init : f32 loc(#loc131)
linalg.yield %81 : f32 loc(#loc122)
} loc(#loc122)
%73 = arith.subf %arg21, %66 : tensor<128xf32> loc(#loc113)
%74 = math.exp %73 : tensor<128xf32> loc(#loc114)
%75 = arith.mulf %arg19, %74 : tensor<128xf32> loc(#loc115)
%76 = arith.addf %75, %reduced_15 : tensor<128xf32> loc(#loc116)
%broadcasted_16 = linalg.broadcast ins(%74 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc117)
%77 = arith.mulf %arg20, %broadcasted_16 : tensor<128x128xf32> loc(#loc117)
%78 = linalg.matmul {input_precision = "ieee"} ins(%70, %71 : tensor<128x128xbf16>, tensor<128x128xbf16>) outs(%77 : tensor<128x128xf32>) -> tensor<128x128xf32> loc(#loc118)
%79 = arith.addi %arg22, %c128_i32 : i32 loc(#loc119)
%80 = arith.addi %arg23, %c128_i32 : i32 loc(#loc77)
scf.yield %68, %76, %78, %66, %79, %80 : i32, tensor<128xf32>, tensor<128x128xf32>, tensor<128xf32>, i32, i32 loc(#loc120)
} {tt.divisibility_arg1 = dense<128> : tensor<1xi32>} loc(#loc71)
%38 = math.log %37#1 : tensor<128xf32> loc(#loc64)
%39 = arith.addf %37#3, %38 : tensor<128xf32> loc(#loc65)
%broadcasted = linalg.broadcast ins(%37#1 : tensor<128xf32>) outs(%0 : tensor<128x128xf32>) dimensions = [1] loc(#loc66)
%40 = arith.divf %37#2, %broadcasted : tensor<128x128xf32> loc(#loc66)
%41 = arith.muli %11, %c1024_i32 : i32 loc(#loc10)
%42 = arith.index_cast %41 : i32 to index loc(#loc10)
%43 = arith.index_cast %26 : i32 to index loc(#loc34)
%44 = arith.addi %42, %43 : index loc(#loc67)
%reinterpret_cast_7 = memref.reinterpret_cast %arg6 to offset: [%44], sizes: [128], strides: [1] : memref<?xf32> to memref<128xf32, strided<[1], offset: ?>> loc(#loc67)
bufferization.materialize_in_destination %39 in writable %reinterpret_cast_7 : (tensor<128xf32>, memref<128xf32, strided<[1], offset: ?>>) -> () loc(#loc68)
%45 = arith.truncf %40 : tensor<128x128xf32> to tensor<128x128xbf16> loc(#loc69)
bufferization.materialize_in_destination %45 in writable %reinterpret_cast_6 : (tensor<128x128xbf16>, memref<128x128xbf16, strided<[128, 1], offset: ?>>) -> () loc(#loc70)
} loc(#loc22)
return loc(#loc)
} loc(#loc)
} loc(#loc)
#loc1 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":63:33)
#loc3 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:53)
#loc4 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":132:28)
#loc8 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:56)
#loc9 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":76:22)
#loc10 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:35)
#loc11 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:47)
#loc12 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:39)
#loc13 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:40)
#loc14 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:72)
#loc15 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:46)
#loc18 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":83:22)
#loc19 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":123:24)
#loc20 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:49)
#loc21 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":131:30)
#loc22 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":137:51)
#loc23 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":138:35)
#loc24 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":139:33)
#loc25 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":140:31)
#loc26 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":141:30)
#loc27 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":142:23)
#loc28 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:28)
#loc29 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:61)
#loc30 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:73)
#loc31 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":144:52)
#loc32 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:60)
#loc33 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":145:52)
#loc34 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":152:34)
#loc35 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":154:12)
#loc36 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":180:12)
#loc37 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":193:20)
#loc38 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":66:20)
#loc39 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":69:27)
#loc40 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":70:23)
#loc42 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":168:27)
#loc43 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":82:35)
#loc45 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":85:22)
#loc46 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":86:20)
#loc47 = loc("/home/jenkins/miniconda3/envs/CI_B020/lib/python3.11/site-packages/triton/language/standard.py":271:15)
#loc48 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:34)
#loc49 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":90:28)
#loc50 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:20)
#loc51 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":92:28)
#loc52 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":94:20)
#loc53 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":87:28)
#loc54 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":97:46)
#loc55 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":98:8)
#loc56 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:27)
#loc57 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":47:52)
#loc58 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":75:27)
#loc61 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":77:35)
#loc62 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":78:18)
#loc63 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":79:54)
#loc64 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:27)
#loc65 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":215:15)
#loc66 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":216:20)
#loc67 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":217:43)
#loc68 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":219:25)
#loc69 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:37)
#loc70 = loc("/home/jenkins/10.50.84.213/workspace/AscendNPU-IR-CI/BiShengMAX-RegBase-Quick-Pipeline@2/AscendNPU-IR_trustbuild-CI/triton-ops-ascend-npu-ir-ci/native/fa/test_fa_forward_modify_core_exp.py":220:30)
#loc71 = loc(callsite(#loc1 at #loc2))
#loc72 = loc(callsite(#loc1 at #loc5))
#loc74 = loc(callsite(#loc8 at #loc2))
#loc75 = loc(callsite(#loc9 at #loc2))
#loc76 = loc(callsite(#loc11 at #loc2))
#loc77 = loc(callsite(#loc15 at #loc2))
#loc79 = loc(callsite(#loc18 at #loc5))
#loc80 = loc(callsite(#loc38 at #loc5))
#loc81 = loc(callsite(#loc39 at #loc5))
#loc82 = loc(callsite(#loc40 at #loc5))
#loc84 = loc(callsite(#loc42 at #loc6))
#loc85 = loc(callsite(#loc43 at #loc5))
#loc87 = loc(callsite(#loc45 at #loc5))
#loc88 = loc(callsite(#loc46 at #loc5))
#loc89 = loc(callsite(#loc47 at #loc16))
#loc90 = loc(callsite(#loc48 at #loc5))
#loc91 = loc(callsite(#loc49 at #loc5))
#loc92 = loc(callsite(#loc50 at #loc5))
#loc93 = loc(callsite(#loc51 at #loc5))
#loc94 = loc(callsite(#loc52 at #loc5))
#loc95 = loc(callsite(#loc53 at #loc5))
#loc96 = loc(callsite(#loc54 at #loc5))
#loc97 = loc(callsite(#loc15 at #loc5))
#loc98 = loc(callsite(#loc55 at #loc5))
#loc99 = loc(callsite(#loc56 at #loc2))
#loc100 = loc(callsite(#loc57 at #loc2))
#loc101 = loc(callsite(#loc38 at #loc2))
#loc102 = loc(callsite(#loc39 at #loc2))
#loc103 = loc(callsite(#loc40 at #loc2))
#loc104 = loc(callsite(#loc58 at #loc2))
#loc107 = loc(callsite(#loc61 at #loc2))
#loc108 = loc(callsite(#loc62 at #loc2))
#loc109 = loc(callsite(#loc63 at #loc2))
#loc111 = loc(callsite(#loc45 at #loc2))
#loc112 = loc(callsite(#loc46 at #loc2))
#loc113 = loc(callsite(#loc48 at #loc2))
#loc114 = loc(callsite(#loc49 at #loc2))
#loc115 = loc(callsite(#loc50 at #loc2))
#loc116 = loc(callsite(#loc51 at #loc2))
#loc117 = loc(callsite(#loc52 at #loc2))
#loc118 = loc(callsite(#loc53 at #loc2))
#loc119 = loc(callsite(#loc54 at #loc2))
#loc120 = loc(callsite(#loc55 at #loc2))
#loc123 = loc(callsite(#loc84 at #loc7))
#loc125 = loc(callsite(#loc89 at #loc17))
#loc127 = loc(callsite(#loc84 at #loc60))
#loc128 = loc(callsite(#loc123 at #loc5))
#loc129 = loc(callsite(#loc125 at #loc5))
#loc130 = loc(callsite(#loc127 at #loc2))
#loc131 = loc(callsite(#loc125 at #loc2))


8月1日 添加了label:triaged
8月2日 添加了label:bug
叠加 pr commit ee1db4dd
do not retry for auto-vec-v2
remove no use clean transform op
bugfix shape of deintlv intlv
add loop.wrap_producer_extract_slice
test for open mcf default *
问题是 ub overflow