LLM Reasoning 相关度: 8/10

BenGER: A Collaborative Web Platform for End-to-End Benchmarking of German Legal Tasks

Sebastian Nagl, Matthias Grabmair
arXiv: 2604.13583v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

BenGER是一个端到端评估德国法律LLM的开源Web平台,集成了任务创建、标注、模型运行和评估。

主要贡献

  • 创建BenGER平台,集成法律任务评估流程
  • 支持多组织协作和权限管理
  • 提供多种评估指标

方法论

BenGER平台通过网页界面整合任务设计、专家标注、LLM运行和评估,支持多种指标和协作管理。

原文摘要

Evaluating large language models (LLMs) for legal reasoning requires workflows that span task design, expert annotation, model execution, and metric-based evaluation. In practice, these steps are split across platforms and scripts, limiting transparency, reproducibility, and participation by non-technical legal experts. We present the BenGER (Benchmark for German Law) framework, an open-source web platform that integrates task creation, collaborative annotation, configurable LLM runs, and evaluation with lexical, semantic, factual, and judge-based metrics. BenGER supports multi-organization projects with tenant isolation and role-based access control, and can optionally provide formative, reference-grounded feedback to annotators. We will demonstrate a live deployment showing end-to-end benchmark creation and analysis.

标签

LLM 法律 Benchmark 评估平台

arXiv 分类

cs.CL cs.AI