vllm:基于 PyTorch 生态的 LLM 推理与服务库项目

A high-throughput and memory-efficient inference and serving engine for LLMs

Branch176Tags99

Introduction

A high-throughput and memory-efficient inference and serving engine for LLMs

Customize your domain