Multimodal Learning 相关度: 9/10

AtlasOCR: Building the First Open-Source Darija OCR Model with Vision Language Models

Imane Momayiz, Soufiane Ait Elaouad, Abdeljalil Elmajjodi, Haitame Bouanane
arXiv: 2604.08070v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

AtlasOCR是首个开源Darija OCR模型,通过微调VLM实现,在Darija和标准阿拉伯语OCR任务上表现出色。

主要贡献

  • 构建首个开源Darija OCR模型
  • 创建Darija-specific数据集
  • 采用高效的VLM微调策略

方法论

利用OCRSmith合成数据并收集真实数据,使用QLoRA和Unsloth高效微调Qwen2.5-VL 3B,并进行超参数优化。

原文摘要

Darija, the Moroccan Arabic dialect, is rich in visual content yet lacks specialized Optical Character Recognition (OCR) tools. This paper introduces AtlasOCR, the first open-source Darija OCR model built by fine-tuning a 3B parameter Vision Language Model (VLM). We detail our comprehensive approach, from curating a unique Darija-specific dataset leveraging both synthetic generation with our OCRSmith library and carefully sourced real-world data, to implementing efficient fine-tuning strategies. We utilize QLoRA and Unsloth for parameter-efficient training of Qwen2.5-VL 3B and present comprehensive ablation studies optimizing key hyperparameters. Our evaluation on the newly curated AtlasOCRBench and the established KITAB-Bench demonstrates state-of-the-art performance, challenging larger models and highlighting AtlasOCR's robustness and generalization capabilities for both Darija and standard Arabic OCR tasks.

标签

OCR Darija VLM Fine-tuning

arXiv 分类

cs.CV cs.AI