Multimodal Learning 相关度: 9/10

Rethinking Model Efficiency: Multi-Agent Inference with Large Models

Sixun Dong, Juhua Hu, Steven Li, Wei Wen, Qi Qian
arXiv: 2604.04929v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

通过多智能体推理,复用小模型的推理token,提升大模型在视觉语言任务中的效率。

主要贡献

  • 分析了VLM中不同组件的延迟瓶颈
  • 提出了多智能体推理框架
  • 通过复用小模型推理token提升大模型效率

方法论

通过模拟数据分析和真实benchmark测试,提出并验证了一种利用小模型推理token辅助大模型的多智能体推理方法。

原文摘要

Most vision-language models (VLMs) apply a large language model (LLM) as the decoder, where the response tokens are generated sequentially through autoregression. Therefore, the number of output tokens can be the bottleneck of the end-to-end latency. However, different models may require vastly different numbers of output tokens to achieve comparable performance. In this work, we conduct a comprehensive analysis of the latency across different components of VLMs on simulated data. The experiment shows that a large model with fewer output tokens can be more efficient than a small model with a long output sequence. The empirical study on diverse real-world benchmarks confirms the observation that a large model can achieve better or comparable performance as a small model with significantly fewer output tokens. To leverage the efficiency of large models, we propose a multi-agent inference framework that keeps large models with short responses but transfers the key reasoning tokens from the small model when necessary. The comparison on benchmark tasks demonstrates that by reusing the reasoning tokens from small models, it can help approach the performance of a large model with its own reasoning, which confirms the effectiveness of our proposal.

标签

多智能体 视觉语言模型 模型效率 推理

arXiv 分类

cs.CV