Hugging Face Blog·· 2025-01-23AI 评分42
用 KVPress 掌握 LLM 长上下文:NVIDIA 的 KV Cache 压缩工具包
Mastering Long Contexts in LLMs with KVPress
AI 导读
NVIDIA 推出 Python 工具包 KVPress,通过一系列压缩算法降低长上下文 LLM 的 KV Cache 内存占用。以 Llama 3-70B 在 bfloat16 下处理 1M token 为例,KV Cache 需 327.6GB,占约 470GB 总内存的 70%。
来源:Hugging Face Blog · huggingface.co