A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.
一种更为内存高效的HF变换器中Llama实现的改写,适用于量化权重使用。【此简介由AI生成】