std.runtime.blackBox(@Bench 框架对其返回值自动调用)在 -O2 下无法阻止 LICM 将循环不变的读取提升到测量循环之外。
std.runtime.blackBox
@Bench
-O2
根因在 LLVM 侧:CJRuntimeLowering 在 CJ 流水线最开头(PassBuilder 中最早一批 pass)就把 llvm.cj.blackhole intrinsic 降级为 CJ_LLVM_BlackHole 函数调用,并附加 ReadOnly 函数属性。readonly 调用不写内存,LICM 可合法地把循环不变 load 提升到该调用之前。实测后果:基准测试被计时的内层循环里,被测的数组/字符串读取每批只执行一次,循环内只剩一次结果 store——测量失效(实测数字 ≈ 框架空载开销,与真实访问成本相差数倍)。
CJRuntimeLowering
llvm.cj.blackhole
CJ_LLVM_BlackHole
ReadOnly
readonly
对 micro-benchmark 全目录扫描(212 个 @Bench 文件 / 1457 个测量闭包),确认 11 个读取类用例受害:BenchmarkArrayBracketsGet(6 个)、BenchmarkArrayBracketsRangeGet(1 个)、BenchmarkArrayListBrackets String 组(4 个)。
BenchmarkArrayBracketsGet
BenchmarkArrayBracketsRangeGet
BenchmarkArrayListBrackets
blackBox 应同时构成对 DSE 和 LICM 的屏障:@Bench 测量循环内、被 blackBox 消费(含框架对返回值的自动消费)的循环不变读取,必须在每次迭代执行,不得被提升到循环 preheader。
blackBox
import std.unittest.* import std.unittest.testmacro.* @Test public class TestRepro { static var arr = Array<Int64>(1024, {i => i}) @Bench func get_val(): Int64 { return arr[512] // 批内不变下标,框架自动 blackBox 返回值 } }
cjc -O2 --int-overflow=wrapping --no-sub-pkg --test repro.cj --save-temps=.,llvm-dis 优化后 bitcode:数组元素 load 出现在测量循环的 preheader(每批一次),而不是内层循环体内(每次迭代一次)。
cjc -O2 --int-overflow=wrapping --no-sub-pkg --test repro.cj --save-temps=.
llvm-dis
对照:llvm.cj.blackhole 保持 intrinsic 形态(默认 may-read/may-write 内存效果、不带 readonly)时 LICM 无法跨越,load 留在循环内。问题只在于降级过早 + 附加了 ReadOnly。
由关联 PR 修复:llvm.cj.blackhole 不再在 CJRuntimeLowering 中提前降级,以 intrinsic 形态(默认内存效果)活过整个优化流水线,由 CJRewriteStatepoint 在所有优化结束后将其零开销消除(有用结果替换为参数、无用调用删除),新/旧两套 pass manager 路径均已覆盖。
CJRewriteStatepoint
Cangjie Compiler: 0.0.1 (cjnative) Target: x86_64-unknown-linux-gnu
发生了什么问题? | Describe the issue that occurred.
std.runtime.blackBox(@Bench框架对其返回值自动调用)在-O2下无法阻止 LICM 将循环不变的读取提升到测量循环之外。根因在 LLVM 侧:
CJRuntimeLowering在 CJ 流水线最开头(PassBuilder 中最早一批 pass)就把llvm.cj.blackholeintrinsic 降级为CJ_LLVM_BlackHole函数调用,并附加ReadOnly函数属性。readonly调用不写内存,LICM 可合法地把循环不变 load 提升到该调用之前。实测后果:基准测试被计时的内层循环里,被测的数组/字符串读取每批只执行一次,循环内只剩一次结果 store——测量失效(实测数字 ≈ 框架空载开销,与真实访问成本相差数倍)。对 micro-benchmark 全目录扫描(212 个 @Bench 文件 / 1457 个测量闭包),确认 11 个读取类用例受害:
BenchmarkArrayBracketsGet(6 个)、BenchmarkArrayBracketsRangeGet(1 个)、BenchmarkArrayListBracketsString 组(4 个)。期望行为是什么? | What is the expected behavior?
blackBox应同时构成对 DSE 和 LICM 的屏障:@Bench测量循环内、被blackBox消费(含框架对返回值的自动消费)的循环不变读取,必须在每次迭代执行,不得被提升到循环 preheader。如何复现该缺陷? | How to reproduce the bug.
import std.unittest.* import std.unittest.testmacro.* @Test public class TestRepro { static var arr = Array<Int64>(1024, {i => i}) @Bench func get_val(): Int64 { return arr[512] // 批内不变下标,框架自动 blackBox 返回值 } }cjc -O2 --int-overflow=wrapping --no-sub-pkg --test repro.cj --save-temps=.,llvm-dis优化后 bitcode:数组元素 load 出现在测量循环的 preheader(每批一次),而不是内层循环体内(每次迭代一次)。对照:
llvm.cj.blackhole保持 intrinsic 形态(默认 may-read/may-write 内存效果、不带readonly)时 LICM 无法跨越,load 留在循环内。问题只在于降级过早 + 附加了ReadOnly。其他补充信息 | Additional information
由关联 PR 修复:
llvm.cj.blackhole不再在CJRuntimeLowering中提前降级,以 intrinsic 形态(默认内存效果)活过整个优化流水线,由CJRewriteStatepoint在所有优化结束后将其零开销消除(有用结果替换为参数、无用调用删除),新/旧两套 pass manager 路径均已覆盖。cjc版本信息 | cjc version information
Cangjie Compiler: 0.0.1 (cjnative)
Target: x86_64-unknown-linux-gnu
分支版本信息 | Branch version information