Multimodal Learning 相关度: 9/10

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details

Dewei Zhou, You Li, Zongxin Yang, Yi Yang
arXiv: 2604.06870v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

RefineAnything提出了一种基于扩散模型的区域特异性图像修复方法,专注于细节恢复和背景保持。

主要贡献

  • 提出了区域特异性图像修复问题设定
  • 提出了Focus-and-Refine策略
  • 提出了Boundary Consistency Loss
  • 构建了Refine-30K数据集和RefineEval基准

方法论

基于扩散模型,采用Focus-and-Refine策略,通过裁剪和缩放将分辨率预算重新分配给目标区域,并使用边界一致性损失。

原文摘要

We introduce region-specific image refinement as a dedicated problem setting: given an input image and a user-specified region (e.g., a scribble mask or a bounding box), the goal is to restore fine-grained details while keeping all non-edited pixels strictly unchanged. Despite rapid progress in image generation, modern models still frequently suffer from local detail collapse (e.g., distorted text, logos, and thin structures). Existing instruction-driven editing models emphasize coarse-grained semantic edits and often either overlook subtle local defects or inadvertently change the background, especially when the region of interest occupies only a small portion of a fixed-resolution input. We present RefineAnything, a multimodal diffusion-based refinement model that supports both reference-based and reference-free refinement. Building on a counter-intuitive observation that crop-and-resize can substantially improve local reconstruction under a fixed VAE input resolution, we propose Focus-and-Refine, a region-focused refinement-and-paste-back strategy that improves refinement effectiveness and efficiency by reallocating the resolution budget to the target region, while a blended-mask paste-back guarantees strict background preservation. We further introduce a boundary-aware Boundary Consistency Loss to reduce seam artifacts and improve paste-back naturalness. To support this new setting, we construct Refine-30K (20K reference-based and 10K reference-free samples) and introduce RefineEval, a benchmark that evaluates both edited-region fidelity and background consistency. On RefineEval, RefineAnything achieves strong improvements over competitive baselines and near-perfect background preservation, establishing a practical solution for high-precision local refinement. Project Page: https://limuloo.github.io/RefineAnything/.

标签

图像修复 扩散模型 多模态 区域特异性

arXiv 分类

cs.CV