LLM Memory & RAG 相关度: 7/10

JUÁ - A Benchmark for Information Retrieval in Brazilian Legal Text Collections

Jayr Pereira, Leandro Fernandes, Erick de Brito, Roberto Lotufo, Luiz Bonifacio
arXiv: 2604.06098v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

JUÁ是一个巴西法律信息检索的公共基准,旨在支持可复现和可比较的评估。

主要贡献

  • 构建了巴西法律信息检索的公共基准JUÁ
  • 评估了词汇、密集和BM25重排序的管道
  • 提供了跨多个巴西法律领域的评估框架

方法论

构建包含多种法律文档类型的数据集,并使用词汇、密集和BM25方法进行检索评估,最后公开数据集和评估结果。

原文摘要

Legal information retrieval in Portuguese remains difficult to evaluate systematically because available datasets differ widely in document type, query style, and relevance definition. We present \textsc{JUÁ}, a public benchmark for Brazilian legal retrieval designed to support more reproducible and comparable evaluation across heterogeneous legal collections. More broadly, \textsc{JUÁ} is intended not only as a benchmark, but as a continuous evaluation infrastructure for Brazilian legal IR, combining shared protocols, common ranking metrics, fixed splits when applicable, and a public leaderboard. The benchmark covers jurisprudence retrieval as well as broader legislative, regulatory, and question-driven legal search. We evaluate lexical, dense, and BM25-based reranking pipelines, including a domain-adapted Qwen embedding model fine-tuned on \textsc{JUÁ}-aligned supervision. Results show that the benchmark is sufficiently heterogeneous to distinguish retrieval paradigms and reveal substantial cross-dataset trade-offs. Domain adaptation yields its clearest gains on the supervision-aligned \textsc{JUÁ-Juris} subset, while BM25 remains highly competitive on other collections, especially in settings with strong lexical and institutional phrasing cues. Overall, \textsc{JUÁ} provides a practical evaluation framework for studying legal retrieval across multiple Brazilian legal domains under a common benchmark design.

标签

信息检索 法律领域 基准测试 巴西

arXiv 分类

cs.IR cs.CL