LLM Memory & RAG 相关度: 7/10

CLEAR: Cross-Lingual Enhancement in Alignment via Reverse-training

Seungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang, Dongsuk Oh, Heuiseok Lim
arXiv: 2604.05821v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

CLEAR通过反向训练增强跨语言对齐,提升跨语言检索,尤其在低资源语言上效果显著。

主要贡献

  • 提出CLEAR损失函数,利用反向训练提升跨语言对齐
  • 在跨语言检索任务中取得了显著提升,尤其在低资源语言上
  • 实验证明CLEAR在多语言训练中也具有潜力

方法论

利用英语作为桥梁,通过反向训练加强目标语言和英语之间的对齐,优化跨语言检索。

原文摘要

Existing multilingual embedding models often encounter challenges in cross-lingual scenarios due to imbalanced linguistic resources and less consideration of cross-lingual alignment during training. Although standardized contrastive learning approaches for cross-lingual adaptation are widely adopted, they may struggle to capture fundamental alignment between languages and degrade performance in well-aligned languages such as English. To address these challenges, we propose Cross-Lingual Enhancement in Retrieval via Reverse-training (CLEAR), a novel loss function utilizing a reverse training scheme to improve retrieval performance across diverse cross-lingual retrieval scenarios. CLEAR leverages an English passage as a bridge to strengthen alignments between the target language and English, ensuring robust performance in the cross-lingual retrieval task. Our extensive experiments demonstrate that CLEAR achieves notable improvements in cross-lingual scenarios, with gains up to 15%, particularly in low-resource languages, while minimizing performance degradation in English. Furthermore, our findings highlight that CLEAR offers promising effectiveness even in multilingual training, suggesting its potential for broad application and scalability. We release the code at https://github.com/dltmddbs100/CLEAR.

标签

跨语言学习 检索 对比学习 低资源语言

arXiv 分类

cs.CL cs.IR