TensorRT-LLM:基于 TensorRT 的 LLM 推理优化库项目

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

分支43Tags94

项目介绍

TensorRT-LLM为用户提供了简便易用的Python API,以定义大型语言模型(LLMs)并构建包含最先进优化技术的TensorRT引擎,这些优化技术能够确保在NVIDIA GPU上高效地进行推理。此外,TensorRT-LLM还包括用于创建Python和C++运行时的组件,以执行这些TensorRT引擎。【此简介由AI生成】

定制我的领域
11914.29 K2.64 K访问 GitHub