MoBiE: Efficient Inference of Mixture of Binary Experts under Post-Training Quantization
AI 摘要
MoBiE是一种面向MoE-LLM的二值化框架,旨在提高推理效率并缓解量化引起的路由偏差。
主要贡献
- 使用联合SVD分解减少专家间的冗余
- 结合全局损失梯度和局部Hessian度量来增强权重重要性估计
- 引入输入零空间引导的误差约束来缓解路由失真
方法论
提出MoBiE框架,通过联合SVD、梯度增强的Hessian度量和误差约束来优化MoE-LLM的二值化过程。
原文摘要
Mixture-of-Experts (MoE) based large language models (LLMs) offer strong performance but suffer from high memory and computation costs. Weight binarization provides extreme efficiency, yet existing binary methods designed for dense LLMs struggle with MoE-specific issues, including cross-expert redundancy, task-agnostic importance estimation, and quantization-induced routing shifts. To this end, we propose MoBiE, the first binarization framework tailored for MoE-based LLMs. MoBiE is built on three core innovations: 1. using joint SVD decomposition to reduce cross-expert redundancy; 2. integrating global loss gradients into local Hessian metrics to enhance weight importance estimation; 3. introducing an error constraint guided by the input null space to mitigate routing distortion. Notably, MoBiE achieves these optimizations while incurring no additional storage overhead, striking a balance between efficiency and model performance. Extensive experiments demonstrate that MoBiE consistently outperforms state-of-the-art binary methods across multiple MoE-based LLMs and benchmarks. For example, on Qwen3-30B-A3B, MoBiE reduces perplexity by 52.2$\%$, improves average zero-shot performance by 43.4$\%$, achieves over 2 $\times$ inference speedup, and further shortens quantization time. The code is available at https://github.com/Kishon-zzx/MoBiE.