Running large language models on a single GPU for throughput-oriented scenarios.
在面向吞吐量的场景下,利用单个GPU运行大型语言模型。【此简介由AI生成】