Multimodal Learning 相关度: 8/10

HI-MoE: Hierarchical Instance-Conditioned Mixture-of-Experts for Object Detection

Vadim Vashkelis, Natalia Trukhina
arXiv: 2604.04908v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

HI-MoE提出了一种分层实例条件混合专家模型用于目标检测,尤其提升了小目标检测性能。

主要贡献

  • 提出分层实例条件混合专家模型HI-MoE
  • 设计了两阶段路由机制:场景路由和实例路由
  • 在COCO和LVIS数据集上验证了HI-MoE的有效性

方法论

采用DETR架构,引入两阶段路由机制,首先进行场景级别的专家选择,然后进行实例级别的专家分配,实现稀疏计算。

原文摘要

Mixture-of-Experts (MoE) architectures enable conditional computation by activating only a subset of model parameters for each input. Although sparse routing has been highly effective in language models and has also shown promise in vision, most vision MoE methods operate at the image or patch level. This granularity is poorly aligned with object detection, where the fundamental unit of reasoning is an object query corresponding to a candidate instance. We propose Hierarchical Instance-Conditioned Mixture-of-Experts (HI-MoE), a DETR-style detection architecture that performs routing in two stages: a lightweight scene router first selects a scene-consistent expert subset, and an instance router then assigns each object query to a small number of experts within that subset. This design aims to preserve sparse computation while better matching the heterogeneous, instance-centric structure of detection. In the current draft, experiments are concentrated on COCO with preliminary specialization analysis on LVIS. Under these settings, HI-MoE improves over a dense DINO baseline and over simpler token-level or instance-only routing variants, with especially strong gains on small objects. We also provide an initial visualization of expert specialization patterns. We present the method, ablations, and current limitations in a form intended to support further experimental validation.

标签

目标检测 混合专家模型 DETR 稀疏计算

arXiv 分类

cs.LG