Multimodal Learning 相关度: 8/10

Your Pre-trained Diffusion Model Secretly Knows Restoration

Sudarshan Rajagopalan, Vishal M. Patel
arXiv: 2604.04924v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

发现预训练扩散模型内在的图像修复能力,通过学习Prompt嵌入来实现高效修复。

主要贡献

  • 揭示预训练扩散模型隐藏的修复能力
  • 提出基于prompt学习的轻量级修复方法
  • 提出diffusion bridge解决训练和推理的错位问题

方法论

通过学习文本编码器输出的Prompt嵌入,并使用Diffusion Bridge对齐训练和推理过程,从而激活预训练扩散模型的修复能力。

原文摘要

Pre-trained diffusion models have enabled significant advancements in All-in-One Restoration (AiOR), offering improved perceptual quality and generalization. However, diffusion-based restoration methods primarily rely on fine-tuning or Control-Net style modules to leverage the pre-trained diffusion model's priors for AiOR. In this work, we show that these pre-trained diffusion models inherently possess restoration behavior, which can be unlocked by directly learning prompt embeddings at the output of the text encoder. Interestingly, this behavior is largely inaccessible through text prompts and text-token embedding optimization. Furthermore, we observe that naive prompt learning is unstable because the forward noising process using degraded images is misaligned with the reverse sampling trajectory. To resolve this, we train prompts within a diffusion bridge formulation that aligns training and inference dynamics, enforcing a coherent denoising path from noisy degraded states to clean images. Building on these insights, we introduce our lightweight learned prompts on the pre-trained WAN video model and FLUX image models, converting them into high-performing restoration models. Extensive experiments demonstrate that our approach achieves competitive performance and generalization across diverse degradations, while avoiding fine-tuning and restoration-specific control modules.

标签

扩散模型 图像修复 Prompt学习 Diffusion Bridge 预训练模型

arXiv 分类

cs.CV cs.AI