跳到正文
原文
Hugging Face Blog·· 2024-11-20AI 评分36

Hugging Face 详解 LayerSkip 自推测解码:用 transformers 加速文本生成

Faster Text Generation with Self-Speculative Decoding

AI 导读

Hugging Face 博客介绍 LayerSkip 提出的自推测解码,用同一 LLM 的早期层生成 draft token、深层网络负责验证,无需额外小模型即可加速文本生成并降低显存占用。

来源:Hugging Face Blog · huggingface.co