exllama:A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.

A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.

分支2Tags0

项目介绍

一种更为内存高效的HF变换器中Llama实现的改写,适用于量化权重使用。【此简介由AI生成】

定制我的领域
342.93 K222访问 GitHub