Hugging Face Blog·· 2024-10-29精选AI 评分62
Intel Labs 与 Hugging Face 提出 Universal Assisted Generation:任意辅助模型加速解码
Universal Assisted Generation: Faster Decoding with Any Assistant Model
AI 导读
Intel Labs 与 Hugging Face 提出 Universal Assisted Generation(UAG),让辅助生成不再要求目标模型与辅助模型共享同一 tokenizer,可跨模型家族配对,对任意 decoder 或 MoE 模型实现 1.5x-2.0x 推理加速。
推荐理由
Intel Labs 与 Hugging Face 提出跨 tokenizer 的辅助生成方法,读者可了解其 2-way tokenizer 翻译思路与实测加速幅度。
来源:Hugging Face Blog · huggingface.co