Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy
AI 摘要
Revise框架通过模拟OCR错误进行纠正,提升文档结构化表示和下游任务性能。
主要贡献
- 提出了Revise框架,系统性纠正OCR错误
- 构建了OCR错误分层分类体系
- 设计了逼真的OCR错误合成数据生成策略
方法论
构建分层OCR错误分类体系,利用合成数据训练纠错模型,在字符、单词和结构层面进行纠正。
原文摘要
Recent advances in Large Language Models (LLMs) have significantly improved the field of Document AI, demonstrating remarkable performance on document understanding tasks such as question answering. However, existing approaches primarily focus on solving specific tasks, lacking the capability to structurally organize and manage document information. To address this limitation, we propose Revise, a framework that systematically corrects errors introduced by OCR at the character, word, and structural levels. Specifically, Revise employs a comprehensive hierarchical taxonomy of common OCR errors and a synthetic data generation strategy that realistically simulates such errors to train an effective correction model. Experimental results demonstrate that Revise effectively corrects OCR outputs, enabling more structured representation and systematic management of document contents. Consequently, our method significantly enhances downstream performance in document retrieval and question answering tasks, highlighting the potential to overcome the structural management limitations of existing Document AI frameworks.