Agent Tuning & Optimization 相关度: 7/10

Towards Faster Language Model Inference Using Mixture-of-Experts Flow Matching

Aihua Li
arXiv: 2604.15009v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

提出了混合专家流匹配(MoE-FM)框架,加速非自回归语言模型推理。

主要贡献

  • 提出了MoE-FM框架,解决流匹配在语言建模中的限制
  • 开发了基于MoE-FM的非自回归语言模型YAN
  • YAN在多个任务上达到与自回归模型相当的生成质量,且推理速度大幅提升

方法论

利用混合专家模型分解复杂潜在空间,构建局部专业化的向量场,并使用Transformer和Mamba架构实现YAN。

原文摘要

Flow matching retains the generation quality of diffusion models while enabling substantially faster inference, making it a compelling paradigm for generative modeling. However, when applied to language modeling, it exhibits fundamental limitations in representing complex latent distributions with irregular geometries, such as anisotropy and multimodality. To address these challenges, we propose a mixture-of-experts flow matching (MoE-FM) framework, which captures complex global transport geometries in latent space by decomposing them into locally specialized vector fields. Building on MoE-FM, we develop a non-autoregressive (NAR) language modeling approach, named YAN, instantiated with both Transformer and Mamba architectures. Across multiple downstream tasks, YAN achieves generation quality on par with both autoregressive (AR) and diffusion-based NAR language models, while requiring as few as three sampling steps. This yields a $40\times$ speedup over AR baselines and up to a $10^3\times$ speedup over diffusion language models, demonstrating substantial efficiency advantages for language modeling.

标签

语言模型 流匹配 非自回归 推理加速

arXiv 分类

cs.AI cs.LG