跳到正文
原文
Hugging Face Blog·· 2026-08-25AI 评分52

Multiverse Computing 提出 Quantization-Aware Healing:4-bit 压缩模型在 7/9 基准上超过其全精度版本

Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

AI 导读

Multiverse Computing 发布论文提出 Quantization-Aware Healing(QAH),在 GPT-OSS 120B 压缩到 60B 参数并量化为 MXFP4 后,直接从压缩前的原始模型蒸馏,而非从恢复出的 bfloat16 检查点蒸馏。

来源:Hugging Face Blog · huggingface.co