Multimodal Learning 相关度: 9/10

Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation

Quoc-Huy Trinh, Mustapha Abdullahi, Bo Zhao, Debesh Jha
arXiv: 2604.04579v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

Firebolt-VL提出了一种高效的视觉语言模型,通过LFM解码器和Token-Grid关联模块提升性能。

主要贡献

  • 提出了基于Liquid Foundation Model (LFM) 的视觉语言模型Firebolt-VL。
  • 引入Token-Grid Correlation Module增强视觉定位能力。
  • 实现了高效的线性时间推理和精确的细粒度理解。

方法论

使用LFM解码器替代Transformer,并设计Token-Grid相关模块,通过FiLM调节状态空间模型,选择性强调相关视觉区域。

原文摘要

Recent advances in multimodal large language models (MLLMs) have enabled impressive progress in vision-language understanding, yet their high computational cost limits deployment in resource-constrained scenarios such as personal assistants, document understanding, and smart cameras. Most existing methods rely on Transformer-based cross-attention, whose quadratic complexity hinders efficiency. Moreover, small vision-language models often struggle to precisely capture fine-grained, task-relevant visual regions, leading to degraded performance on fine-grained reasoning tasks that limit their effectiveness in the real world. To address these issues, we introduce Firebolt-VL, an efficient vision-language model that replaces the Transformer-based decoder with a Liquid Foundation Model (LFM) decoder. To further enhance visual grounding, we propose a Token-Grid Correlation Module, which computes lightweight correlations between text tokens and image patches and modulates via the state-space model with FiLM conditioning. This enables the model to selectively emphasize visual regions relevant to the textual prompt while maintaining linear-time inference. Experimental results across multiple benchmarks demonstrate that Firebolt-VL achieves accurate, fine-grained understanding with significantly improved efficiency. Our model and code are available at: https://fireboltvl.github.io

标签

视觉语言理解 多模态学习 模型效率 细粒度推理

arXiv 分类

cs.CV