Multimodal Learning 相关度: 9/10

EEG2Vision: A Multimodal EEG-Based Framework for 2D Visual Reconstruction in Cognitive Neuroscience

Emanuele Balloni, Emanuele Frontoni, Chiara Matti, Marina Paolanti, Roberto Pierdicca, Emiliano Santarnecchi
arXiv: 2604.08063v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

EEG2Vision通过多模态方法,提升低分辨率脑电信号重建视觉图像的质量。

主要贡献

  • 提出一个端到端的EEG-to-image框架EEG2Vision
  • 使用提示引导的后重建增强机制提升图像质量
  • 系统评估了不同分辨率脑电信号的重建性能

方法论

利用脑电信号条件扩散重建,结合多模态大语言模型提取语义信息,通过图像到图像的扩散模型优化图像。

原文摘要

Reconstructing visual stimuli from non-invasive electroencephalography (EEG) remains challenging due to its low spatial resolution and high noise, particularly under realistic low-density electrode configurations. To address this, we present EEG2Vision, a modular, end-to-end EEG-to-image framework that systematically evaluates reconstruction performance across different EEG resolutions (128, 64, 32, and 24 channels) and enhances visual quality through a prompt-guided post-reconstruction boosting mechanism. Starting from EEG-conditioned diffusion reconstruction, the boosting stage uses a multimodal large language model to extract semantic descriptions and leverages image-to-image diffusion to refine geometry and perceptual coherence while preserving EEG-grounded structure. Our experiments show that semantic decoding accuracy degrades significantly with channel reduction (e.g., 50-way Top-1 Acc from 89% to 38%), while reconstruction quality slight decreases (e.g., FID from 76.77 to 80.51). The proposed boosting consistently improves perceptual metrics across all configurations, achieving up to 9.71% IS gains in low-channel settings. A user study confirms the clear perceptual preference for boosted reconstructions. The proposed approach significantly boosts the feasibility of real-time brain-2-image applications using low-resolution EEG devices, potentially unlocking this type of applications outside laboratory settings.

标签

脑电信号 图像重建 多模态学习 扩散模型

arXiv 分类

cs.CV