| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
[MLIR][ExecutionEngine] don't leak -Wweak-vtables (#164498) I'm not 100% what this is used for in this lib but the compile flag leaks out and prevents (in certain compile scenarios) linking mlir_c_runner_utils. | 10 个月前 | |
[mlir][arith] Add support for sitofp, uitofp to ArithToAPFloat (#169284) Add support for arith.sitofp and arith.uitofp. | 9 个月前 | |
[MLIR][AArch64] Change some tests to ensure SVE vector length is the same throughout the function (#147506) This change only applies to functions the can be reasonably expected to use SVE registers. Modifying vector length in the middle of a function might cause incorrect stack deallocation if there are callee-saved SVE registers or incorrect access to SVE stack slots. Addresses (non-issue) https://github.com/llvm/llvm-project/issues/143670 | 1 年前 | |
Reland [mlir] Workaround for export lib generation on Windows for mlir_arm_sme_abi_stubs #73147 (#73238) https://github.com/llvm/llvm-project/pull/73147 Fixed the visibility macro | 2 年前 | |
Rename llvm::ThreadPool -> llvm::DefaultThreadPool (NFC) (#83702) The base class llvm::ThreadPoolInterface will be renamed llvm::ThreadPool in a subsequent commit. This is a breaking change: clients who use to create a ThreadPool must now create a DefaultThreadPool instead. | 2 年前 | |
[mlir] Remove filtering of deprecated rocm-agent-enumerator value gfx000 (#166634) Getting a gfx000 result from the rocm-agent-enumerator command was deprecated beginning with the release of ROCm 7, but the MLIR build system still filters it from results when looking for ROCm agents. This PR removes that filtering. There are a few other uses of "gfx000" in MLIR source, but those are used as default options for running some passes, and, to my understanding, have a semantically different meaning to the dummy result returned from rocm-agent-enumerator and don't need to be changed. | 10 个月前 | |
[mlir] Make the print function in CRunnerUtil platform agnostic (#86767) The platform running on Apple Silicon does not seem to support the negative nan. It causes the test failure where we explicitly specify the negative nan bit pattern and check the output printed by the CRunnerUtil function. We can make the print function in the utility platform agnostic by using the standard library functions (i.e. std::isnan and std::signbit) so that we can run the test across platforms that do not support the negative bit pattern. I have added two test cases that would fail in the Apple Silicon platform without print function changes. $ uname -a Darwin Kernel Version 23.3.0: Wed Dec 20 21:30:44 PST 2023; root:xnu-10002.81.5~7/RELEASE_ARM64_T6000 arm64 See: https://discourse.llvm.org/t/test-failure-of-sparse-sign-test-in-apple-silicon/77876/3 | 2 年前 | |
[mlir][gpu] Change GPU modules to globals (#135478) Load/unload GPU modules in global ctors/dtors instead of each time when launching a kernel. Loading GPU modules is a heavy-weight operation and synchronizes the GPU context. Now that the modules are loaded ahead of time, asynchronously launched kernels can run concurrently, see https://discourse.llvm.org/t/how-to-lower-the-combination-of-async-gpu-ops-in-gpu-dialect. The implementations of embedBinary() and launchKernel() use slightly different mechanics at the moment but I prefer to not change the latter more than necessary as part of this PR. I will prepare a follow-up NFC for launchKernel() to align them again. | 1 年前 | |
[MLIR][ExecutionEngine] don't dump decls (#164478) Currently ExecutionEngine tries to dump all functions declared in the module, even those which are "external" (i.e., linked/loaded at runtime). E.g. mlir func.func private @printF32(f32) func.func @supported_arg_types(%arg0: i32, %arg1: f32) { call @printF32(%arg1) : (f32) -> () return } fails with Could not compile printF32: Symbols not found: [ __mlir_printF32 ] Program aborted due to an unhandled Error: Symbols not found: [ __mlir_printF32 ] even though printF32 can be provided at final build time (i.e., when the object file is linked to some executable or shlib). E.g, if our own libmlir_c_runner_utils is linked. So just skip functions which have no bodies during dump (i.e., are decls without defns). | 10 个月前 | |
[mlir][sparse] Fix the calling convention of __truncsfbf2 on windows x64 It also wants us to return the value in XMM0. | 2 年前 | |
[MLIR] Apply clang-tidy fixes for misc-use-internal-linkage in JitRunner.cpp (NFC) | 11 个月前 | |
[MLIR] Apply clang-tidy fixes for misc-use-internal-linkage in LevelZeroRuntimeWrappers.cpp (NFC) | 11 个月前 | |
[PassBuilder] Support O0 in default pipelines The default and pre-link pipeline builders currently require you to call a separate method for optimization level O0, even though they have perfectly well-defined O0 optimization pipelines. Accept O0 optimization level and call buildO0DefaultPipeline() internally, so all consumers don't need to repeat this. Differential Revision: https://reviews.llvm.org/D146200 | 3 年前 | |
[MLIR][AMDGPU] Add ability to do 16-bit Memset with HIP APIs (#108587) CC: @krzysz00 @manupak | 1 年前 | |
[mlir][verifyMemref] Fix bug and support more types for verifyMemref (#77682) 1. Fix a bug in verifyMemref to pass in data instead of baseptr, which didn't verify data correctly. 2. Add == for f16 and bf16. 3. Add a comprehensive test of verifyMemref for all supported types. | 2 年前 | |
[mlir][sparse] provide an AoS "view" into sparse runtime support lib (#87116) Note that even though the sparse runtime support lib always uses SoA storage for COO storage (and provides correct codegen by means of views into this storage), in some rare cases we need the true physical SoA storage as a coordinate buffer. This PR provides that functionality by means of a (costly) coordinate buffer call. Since this is currently only used for testing/debugging by means of the sparse_tensor.print method, this solution is acceptable. If we ever want a performing version of this, we should truly support AoS storage of COO in addition to the SoA used right now. | 2 年前 | |
[mlir] Remove the mlir-spirv-cpu-runner (move to mlir-cpu-runner) (#114563) This commit builds on and completes the work done in 9f6c632ecda08bfff76b798c46d5d7cfde57b5e9 to eliminate the need for a separate mlir-spirv-cpu-runner binary. Since the MLIR processing is already done outside this runner, the only real difference between it and the mlir-cpu-runner is the final linking step between the nested LLVM IR modules. By moving this step into mlir-cpu-runner behind a new command-line flag ( --link-nested-modules), this commit is able to completely remove the runner component of the mlir-spirv-cpu-runner. The runtime libraries and the tests are moved and renamed to fit into the Execution Engine and Integration tests, following the model of the similar migration done for the CUDA Runner in D97463. | 1 年前 | |
[mlir] SYCL runtime wrapper: add memcpy support. (#141647) | 1 年前 | |
[MLIR] Apply clang-tidy fixes for misc-use-internal-linkage in VulkanRuntime.cpp (NFC) | 1 年前 | |
[mlir] Remove mlir-vulkan-runner and GPUToVulkan conversion passes (#123750) This follows up on 733be4ed7dcf976719f424c0cb81b77a14f91f5a, which made mlir-vulkan-runner and its associated passes redundant, and completes the main goal of #73457. The mlir-vulkan-runner tests become part of the integration test suite, and the Vulkan runner runtime components become part of ExecutionEngine, just as was done when removing other target-specific runners. | 1 年前 | |
[mlir] Rename mlir-cpu-runner to mlir-runner (#123776) With the removal of mlir-vulkan-runner (as part of #73457) in e7e3c45bc70904e24e2b3221ac8521e67eb84668, mlir-cpu-runner is now the only runner for all CPU and GPU targets, and the "cpu" name has been misleading for some time already. This commit renames it to mlir-runner. | 1 年前 |
| 文件 | 最后提交记录 | 最后更新时间 |
|---|---|---|
| 10 个月前 | ||
| 9 个月前 | ||
| 1 年前 | ||
| 2 年前 | ||
| 2 年前 | ||
| 10 个月前 | ||
| 2 年前 | ||
| 1 年前 | ||
| 10 个月前 | ||
| 2 年前 | ||
| 11 个月前 | ||
| 11 个月前 | ||
| 3 年前 | ||
| 1 年前 | ||
| 2 年前 | ||
| 2 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 | ||
| 1 年前 |