| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
[libc][gpu] Add exp/log benchmarks and flexible input generation (#155727) This patch adds GPU benchmarks for the exp ( exp, expf, expf16) and log (log, logf, logf16) families of math functions. Adding these benchmarks revealed a key limitation in the existing framework: the input generation mechanism was hardcoded to a single strategy that sampled numbers with a uniform distribution of their unbiased exponents. While this strategy is effective for values spanning multiple orders of magnitude, it is not suitable for linear ranges. The previous framework lacked the flexibility to support this. ### Summary of Changes **1. Framework Refactoring for Flexible Input Sampling:** The GPU benchmark framework was refactored to support multiple, pluggable input sampling strategies. * **Random.h:** A new header was created to house the RandomGenerator and the new distribution classes. * **Distribution Classes:** Two sampling strategies were implemented: * UniformExponent: Formalizes the previous logic of sampling numbers with a uniform distribution of their unbiased exponents. It can now also be configured to produce only positive values, which is essential for functions like log. * UniformLinear: A new strategy that samples numbers from a uniform distribution over a linear interval [min, max). * **MathPerf Update:** The MathPerf class was updated with a generic run_throughput method that is templated on a distribution object. This makes the framework extensible to future sampling strategies. **2. New Benchmarks for exp and log:** Using the newly refactored framework, benchmarks were added for exp, expf, expf16, log, logf, and logf16. The test intervals were carefully chosen to measure the performance of distinct behavioral regions of each function. | 11 个月前 | |
[libc][gpu] Disable loop unrolling in the throughput benchmark loop (#153971) This patch makes GPU throughput benchmark results more comparable across targets by disabling loop unrolling in the benchmark loop. Motivation: * PTX (post-LTO) evidence on NVPTX: for libc sin, the generated PTX shows the throughput loop unrolled 8x at N=128 (one iteration advances the input pointer by 64 bytes = 8 doubles), interleaving eight independent chains before the back-edge. This hides latency and significantly reduces cycles/call as the batch size N grows. * Observed scaling (NVPTX measurements): with unrolling enabled, sin dropped from ~3,100 cycles/call at N=1 to ~360 at N=128. After enforcing #pragma clang loop unroll(disable), results stabilized (e.g., from ~3100 cycles/call at N=1 to ~2700 at N=128). * libdevice contrast: the libdevice sin path did not exhibit a similar drop in our measurements, and the PTX appears as compact internal calls rather than a long FMA chain, leaving less ILP for the outer loop to extract. What this change does: * Applies #pragma clang loop unroll(disable) to the GPU throughput() loop in both NVPTX and AMDGPU backends. Leaving unrolling entirely to the optimizer makes apples-to-apples comparisons uneven (e.g., libc vs. vendor). Disabling unrolling yields fairer, more consistent numbers. | 11 个月前 | |
[libc][gpu] Add exp/log benchmarks and flexible input generation (#155727) This patch adds GPU benchmarks for the exp ( exp, expf, expf16) and log (log, logf, logf16) families of math functions. Adding these benchmarks revealed a key limitation in the existing framework: the input generation mechanism was hardcoded to a single strategy that sampled numbers with a uniform distribution of their unbiased exponents. While this strategy is effective for values spanning multiple orders of magnitude, it is not suitable for linear ranges. The previous framework lacked the flexibility to support this. ### Summary of Changes **1. Framework Refactoring for Flexible Input Sampling:** The GPU benchmark framework was refactored to support multiple, pluggable input sampling strategies. * **Random.h:** A new header was created to house the RandomGenerator and the new distribution classes. * **Distribution Classes:** Two sampling strategies were implemented: * UniformExponent: Formalizes the previous logic of sampling numbers with a uniform distribution of their unbiased exponents. It can now also be configured to produce only positive values, which is essential for functions like log. * UniformLinear: A new strategy that samples numbers from a uniform distribution over a linear interval [min, max). * **MathPerf Update:** The MathPerf class was updated with a generic run_throughput method that is templated on a distribution object. This makes the framework extensible to future sampling strategies. **2. New Benchmarks for exp and log:** Using the newly refactored framework, benchmarks were added for exp, expf, expf16, log, logf, and logf16. The test intervals were carefully chosen to measure the performance of distinct behavioral regions of each function. | 11 个月前 | |
[libc] Polish GPU benchmarking (#153900) This patch provides cleanups and improvements for the GPU benchmarking infrastructure. The key changes are: - Fix benchmark convergence bug: Round up the scaled iteration count (ceil) to ensure it grows properly. The previous truncation logic causes the iteration count to get stuck. - Resolve remaining compiler warning. - Remove unused BenchmarkLogger files: This is dead code that added maintenance and cognitive overhead without providing functionality. - Improve build hygiene: Clean up headers and CMake dependencies to strictly follow the 'include what you use' (IWYU) principle. | 11 个月前 | |
[libc][gpu] Add exp/log benchmarks and flexible input generation (#155727) This patch adds GPU benchmarks for the exp ( exp, expf, expf16) and log (log, logf, logf16) families of math functions. Adding these benchmarks revealed a key limitation in the existing framework: the input generation mechanism was hardcoded to a single strategy that sampled numbers with a uniform distribution of their unbiased exponents. While this strategy is effective for values spanning multiple orders of magnitude, it is not suitable for linear ranges. The previous framework lacked the flexibility to support this. ### Summary of Changes **1. Framework Refactoring for Flexible Input Sampling:** The GPU benchmark framework was refactored to support multiple, pluggable input sampling strategies. * **Random.h:** A new header was created to house the RandomGenerator and the new distribution classes. * **Distribution Classes:** Two sampling strategies were implemented: * UniformExponent: Formalizes the previous logic of sampling numbers with a uniform distribution of their unbiased exponents. It can now also be configured to produce only positive values, which is essential for functions like log. * UniformLinear: A new strategy that samples numbers from a uniform distribution over a linear interval [min, max). * **MathPerf Update:** The MathPerf class was updated with a generic run_throughput method that is templated on a distribution object. This makes the framework extensible to future sampling strategies. **2. New Benchmarks for exp and log:** Using the newly refactored framework, benchmarks were added for exp, expf, expf16, log, logf, and logf16. The test intervals were carefully chosen to measure the performance of distinct behavioral regions of each function. | 11 个月前 | |
| 1 年前 | ||
[libc][gpu] Add exp/log benchmarks and flexible input generation (#155727) This patch adds GPU benchmarks for the exp ( exp, expf, expf16) and log (log, logf, logf16) families of math functions. Adding these benchmarks revealed a key limitation in the existing framework: the input generation mechanism was hardcoded to a single strategy that sampled numbers with a uniform distribution of their unbiased exponents. While this strategy is effective for values spanning multiple orders of magnitude, it is not suitable for linear ranges. The previous framework lacked the flexibility to support this. ### Summary of Changes **1. Framework Refactoring for Flexible Input Sampling:** The GPU benchmark framework was refactored to support multiple, pluggable input sampling strategies. * **Random.h:** A new header was created to house the RandomGenerator and the new distribution classes. * **Distribution Classes:** Two sampling strategies were implemented: * UniformExponent: Formalizes the previous logic of sampling numbers with a uniform distribution of their unbiased exponents. It can now also be configured to produce only positive values, which is essential for functions like log. * UniformLinear: A new strategy that samples numbers from a uniform distribution over a linear interval [min, max). * **MathPerf Update:** The MathPerf class was updated with a generic run_throughput method that is templated on a distribution object. This makes the framework extensible to future sampling strategies. **2. New Benchmarks for exp and log:** Using the newly refactored framework, benchmarks were added for exp, expf, expf16, log, logf, and logf16. The test intervals were carefully chosen to measure the performance of distinct behavioral regions of each function. | 11 个月前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 11 个月前 | ||
| 11 个月前 | ||
| 11 个月前 | ||
| 11 个月前 | ||
| 11 个月前 | ||
| 1 年前 | ||
| 11 个月前 |