An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs
一个用于在现代消费级 GPU 上本地运行大语言模型的优化量化与推理库【此简介由AI生成】