A high-throughput and memory-efficient inference and serving engine for LLMs
面向大规模语言模型的高通量、内存高效的推理与服务引擎【此简介由AI生成】