THIVLVC: Retrieval Augmented Dependency Parsing for Latin
AI 摘要
THIVLVC利用检索增强依赖分析提升拉丁语依存句法分析性能。
主要贡献
- 提出基于检索增强的拉丁语依存句法分析系统THIVLVC
- 利用CIRCSE treebank进行检索
- 在Seneca诗歌数据集上取得显著提升
方法论
两阶段方法:检索相似句子,提示大型语言模型优化UDPipe的baseline parse结果。
原文摘要
We describe THIVLVC, a two-stage system for the EvaLatin 2026 Dependency Parsing task. Given a Latin sentence, we retrieve structurally similar entries from the CIRCSE treebank using sentence length and POS n-gram similarity, then prompt a large language model to refine the baseline parse from UDPipe using the retrieved examples and UD annotation guidelines. We submit two configurations: one without retrieval and one with retrieval (RAG). On poetry (Seneca), THIVLVC improves CLAS by +17 points over the UDPipe baseline; on prose (Thomas Aquinas), the gain is +1.5 CLAS. A double-blind error analysis of 300 divergences between our system and the gold standard reveals that, among unanimous annotator decisions, 53.3% favour THIVLVC, showing annotation inconsistencies both within and across treebanks.