distributed-llama:基于分布式技术的LLM推理加速项目

Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.

分支56Tags64

项目介绍

“Tensor并行 suffice。在家中利用任意设备,在AI集群上运行大规模语言模型。分配工作负载,分割内存使用,提升推理速度。”【此简介由AI生成】

定制我的领域
513.01 K242访问 GitHub