LLM Reasoning 相关度: 7/10

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate

Jaemin Kim, Sungkyun Kim, Junyeol Lee, Jiwon Seo
arXiv: 2604.13806v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

DASH-Q提出了一种基于对角Hessian近似的超低比特后训练量化方法,提升LLM的量化精度。

主要贡献

  • 提出DASH-Q量化框架
  • 使用对角Hessian近似,降低噪声影响
  • 在超低比特量化下表现优于现有方法

方法论

采用对角Hessian近似和迭代加权最小二乘法,过滤采样噪声,保留重要特征信息,实现鲁棒的量化。

原文摘要

Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging. Post-Training Quantization (PTQ) reduces memory footprint without retraining by leveraging a small calibration set. Recent Hessian-based PTQ methods compensate quantization error via cross-channel dependencies, but such approaches degrade at low bit-widths due to noisy curvature estimates from limited calibration data. We propose DASH-Q, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares. By discarding noise-prone dependencies, DASH-Q filters sampling noise while prioritizing the preservation of salient feature power. We outperform other PTQ baselines in ultra low-bit regime, improving zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines across five baseline LLM models, while showing robust and stable performance with very small calibration data.

标签

量化 后训练量化 低比特量化 Hessian

arXiv 分类

cs.LG