LLM Memory & RAG 相关度: 8/10

MUXQ: Mixed-to-Uniform Precision MatriX Quantization via Low-Rank Outlier Decomposition

Seoungsub Lee, In Seo Kim, Seon Wook Kim
arXiv: 2604.04701v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

MUXQ通过低秩分解优化激活量化,降低大模型在NPU上的推理成本,保持精度。

主要贡献

  • 提出了MUXQ混合精度量化方法,解决激活异常值问题
  • 引入辅助矩阵重分布异常值,降低量化损失
  • 在GPT-2模型上验证了MUXQ的有效性,实现低精度推理

方法论

通过检测激活中的异常通道,引入辅助矩阵重新分配幅度,从而实现低精度量化并保持硬件友好型计算结构。

原文摘要

Large language models (LLMs) have achieved outstanding performance across a wide range of natural language processing tasks, but their enormous parameter counts impose ubstantial memory and computational overheads. This challenge is particularly critical in NPU-based on-device environments, where FP16/FP32 computation is inefficient and integer (INT) quantization is therefore essential. However, existing methods, including ZeroQuant, LLM.int8(), and SmoothQuant, do not fully address input-activation outliers and the associated hardware inefficiencies. To overcome these limitations, we propose MUXQ (Mixed-to-Uniform Quantization). MUXQ detects outlier channels in input activations and introduces a small auxiliary matrix that redistributes outlier magnitudes across channels, thereby alleviating the outlier problem. This enables even activation outliers to be quantized at low-precision INT levels while preserving a hardware-friendly computation structure. Experiments on GPT-2 models at three scales (0.1B, 0.3B, and 0.7B parameters) using the WikiText-2 dataset show that MUXQ consistently achieves lower perplexity than naive quantization. In particular, under per-tensor quantization, MUXQ quantizes both activations and weights to INT8 while maintaining accuracy close to that of FP16. With only modest computational overhead, MUXQ enables stable low-precision inference and can be readily combined with other quantization techniques. These results suggest that MUXQ provides a promising direction for efficient and accurate LLM inference on edge devices.

标签

LLM 量化 低精度推理 NPU 模型优化

arXiv 分类

cs.LG cs.AI