| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
[libc][gpu] Disable loop unrolling in the throughput benchmark loop (#153971) This patch makes GPU throughput benchmark results more comparable across targets by disabling loop unrolling in the benchmark loop. Motivation: * PTX (post-LTO) evidence on NVPTX: for libc sin, the generated PTX shows the throughput loop unrolled 8x at N=128 (one iteration advances the input pointer by 64 bytes = 8 doubles), interleaving eight independent chains before the back-edge. This hides latency and significantly reduces cycles/call as the batch size N grows. * Observed scaling (NVPTX measurements): with unrolling enabled, sin dropped from ~3,100 cycles/call at N=1 to ~360 at N=128. After enforcing #pragma clang loop unroll(disable), results stabilized (e.g., from ~3100 cycles/call at N=1 to ~2700 at N=128). * libdevice contrast: the libdevice sin path did not exhibit a similar drop in our measurements, and the PTX appears as compact internal calls rather than a long FMA chain, leaving less ILP for the outer loop to extract. What this change does: * Applies #pragma clang loop unroll(disable) to the GPU throughput() loop in both NVPTX and AMDGPU backends. Leaving unrolling entirely to the optimizer makes apples-to-apples comparisons uneven (e.g., libc vs. vendor). Disabling unrolling yields fairer, more consistent numbers. | 11 个月前 | |
[libc][gpu] Disable loop unrolling in the throughput benchmark loop (#153971) This patch makes GPU throughput benchmark results more comparable across targets by disabling loop unrolling in the benchmark loop. Motivation: * PTX (post-LTO) evidence on NVPTX: for libc sin, the generated PTX shows the throughput loop unrolled 8x at N=128 (one iteration advances the input pointer by 64 bytes = 8 doubles), interleaving eight independent chains before the back-edge. This hides latency and significantly reduces cycles/call as the batch size N grows. * Observed scaling (NVPTX measurements): with unrolling enabled, sin dropped from ~3,100 cycles/call at N=1 to ~360 at N=128. After enforcing #pragma clang loop unroll(disable), results stabilized (e.g., from ~3100 cycles/call at N=1 to ~2700 at N=128). * libdevice contrast: the libdevice sin path did not exhibit a similar drop in our measurements, and the PTX appears as compact internal calls rather than a long FMA chain, leaving less ILP for the outer loop to extract. What this change does: * Applies #pragma clang loop unroll(disable) to the GPU throughput() loop in both NVPTX and AMDGPU backends. Leaving unrolling entirely to the optimizer makes apples-to-apples comparisons uneven (e.g., libc vs. vendor). Disabling unrolling yields fairer, more consistent numbers. | 11 个月前 | |
[libc] Add AMDGPU Timing to CMake (#99603) libc/benchmarks/gpu/timing/CMakeLists.txt did not correctly build amdgpu utils. This PR fixes that issue by adding amdgpu to the loop that adds the correct sub directories. | 2 年前 | |
[libc] Add Timing Utils for AMDGPU (#96828) PR for adding AMDGPU timing utils for benchmarking. I was not able to test this code since I do not have an AMD GPU, but I was able to successfully compile this code using -DRUNTIMES_amdgcn-amd-amdhsa_LIBC_GPU_TEST_ARCHITECTURE=gfx90a -DRUNTIMES_amdgcn-amd-amdhsa_LIBC_GPU_LOADER_EXECUTABLE=echo -DRUNTIMES_amdgcn_amd-amdhsa_LIBC_GPU_TARGET_ARCHITECTURE=gfx90a to force the code to compile without having an AMD gpu on my machine. @jhuber6 | 2 年前 |