Multimodal Learning 相关度: 9/10

Multimodal Latent Reasoning via Predictive Embeddings

Ashutosh Adhikari, Mirella Lapata
arXiv: 2604.08065v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

Pearl通过预测嵌入对齐,在隐空间学习工具使用,提升多模态推理,避免显式工具调用。

主要贡献

  • 提出Pearl框架,用于隐空间多模态推理
  • 使用预测嵌入学习,避免重建误差
  • 提升视觉语言模型在感知任务上的表现

方法论

Pearl框架通过JEPA启发的预测嵌入对齐,在隐空间学习专家工具使用轨迹,实现多步工具推理。

原文摘要

Tool-augmented multimodal reasoning enables visual language models (VLMs) to improve perception by interacting with external tools (e.g., cropping, depth estimation). However, such approaches incur substantial inference overhead, require specialized supervision, and are prone to erroneous tool calls. We propose Pearl (Predictive Embedding Alignment for Reasoning in Latent space), a JEPA-inspired framework that learns from expert tool-use trajectories entirely in the latent space, eliminating the need for explicit tool invocation at inference time. Unlike reconstruction-based latent reasoning methods, which autoregressively generate latent tokens and suffer from training-inference mismatch and limited support for multi-step tool use, Pearl directly learns predictive embeddings from multimodal trajectories while preserving the standard vision-language generation pipeline: it is model-agnostic, simple to train, and naturally supports trajectories with multiple tool calls. Experiments across multiple perception benchmarks show that Pearl matches or outperforms standard supervised fine-tuning and reconstruction-based latent reasoning approaches. Furthermore, we provide empirical evidence that reconstruction-based methods primarily learn embeddings rather than image edits in latent space, motivating predictive embedding learning as a more principled alternative.

标签

多模态学习 视觉语言模型 工具使用 隐空间推理

arXiv 分类

cs.LG