FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
AI 摘要
FlowInOne提出了一种统一的多模态生成框架,将所有输入转化为视觉提示,实现图像输入输出的流程匹配。
主要贡献
- 提出FlowInOne框架,将多模态生成统一为视觉流程
- 引入VisPrompt-5M数据集,包含5百万视觉提示对
- 提出VP-Bench基准测试,评估生成效果
方法论
将文本、布局、指令等多种模态输入转化为视觉提示,使用单一的Flow Matching模型实现图像到图像的转换。
原文摘要
Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We challenge this paradigm by asking whether all modalities, including textual descriptions, spatial layouts, and editing instructions, can be unified into a single visual representation. We present FlowInOne, a framework that reformulates multimodal generation as a purely visual flow, converting all inputs into visual prompts and enabling a clean image-in, image-out pipeline governed by a single flow matching model. This vision-centric formulation naturally eliminates cross-modal alignment bottlenecks, noise scheduling, and task-specific architectural branches, unifying text-to-image generation, layout-guided editing, and visual instruction following under one coherent paradigm. To support this, we introduce VisPrompt-5M, a large-scale dataset of 5 million visual prompt pairs spanning diverse tasks including physics-aware force dynamics and trajectory prediction, alongside VP-Bench, a rigorously curated benchmark assessing instruction faithfulness, spatial precision, visual realism, and content consistency. Extensive experiments demonstrate that FlowInOne achieves state-of-the-art performance across all unified generation tasks, surpassing both open-source models and competitive commercial systems, establishing a new foundation for fully vision-centric generative modeling where perception and creation coexist within a single continuous visual space.