Multimodal Learning 相关度: 8/10

MONETA: Multimodal Industry Classification through Geographic Information with Multi Agent Systems

Arda Yüksel, Gabriel Thiem, Susanne Walter, Patrick Felka, Gabriela Alves Werb, Ivan Habernal
arXiv: 2604.07956v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出了MONETA多模态行业分类基准数据集,利用多模态大语言模型进行行业分类。

主要贡献

  • 构建了包含文本和地理空间信息的多模态行业分类基准数据集MONETA
  • 提出了基于多模态大语言模型的行业分类方法
  • 验证了多轮设计、上下文丰富和分类解释的有效性

方法论

利用文本和地理空间信息,结合多模态大语言模型进行行业分类,并通过多轮交互和上下文增强提高分类准确率。

原文摘要

Industry classification schemes are integral parts of public and corporate databases as they classify businesses based on economic activity. Due to the size of the company registers, manual annotation is costly, and fine-tuning models with every update in industry classification schemes requires significant data collection. We replicate the manual expert verification by using existing or easily retrievable multimodal resources for industry classification. We present MONETA, the first multimodal industry classification benchmark with text (Website, Wikipedia, Wikidata) and geospatial sources (OpenStreetMap and satellite imagery). Our dataset enlists 1,000 businesses in Europe with 20 economic activity labels according to EU guidelines (NACE). Our training-free baseline reaches 62.10% and 74.10% with open and closed-source Multimodal Large Language Models (MLLM). We observe an increase of up to 22.80% with the combination of multi-turn design, context enrichment, and classification explanations. We will release our dataset and the enhanced guidelines.

标签

多模态学习 行业分类 地理信息 大语言模型

arXiv 分类

cs.AI