LLMLingua:基于语言模型的提示词压缩工具项目

[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.

分支6Tags8

项目介绍

为加快大型语言模型(LLMs)的推理速度并提升其对关键信息的感知能力,对提示(prompt)和键-值缓存(KV-Cache)进行压缩,该方法在保证性能损失最小的情况下,实现了高达20倍的压缩率。【此简介由AI生成】

定制我的领域
376.45 K401访问 GitHub