LLM Reasoning 相关度: 7/10

DeepTest Tool Competition 2026: Benchmarking an LLM-Based Automotive Assistant

Lev Sorokin, Ivan Vasilev, Samuele Pasini
arXiv: 2604.12615v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

ICSE 2026 DeepTest研讨会举办了首届LLM测试竞赛,评估LLM汽车助手的信息检索应用。

主要贡献

  • 评估了多个LLM测试工具在汽车手册信息检索任务上的性能
  • 提出了评估LLM系统在安全警告方面的能力的方法
  • 提供了LLM测试竞赛的实验方法和结果

方法论

使用竞赛的方式,评估不同工具发现LLM汽车助手未能提及手册警告的能力,并评估其效率和测试多样性。

原文摘要

This report summarizes the results of the first edition of the Large Language Model (LLM) Testing competition, held as part of the DeepTest workshop at ICSE 2026. Four tools competed in benchmarking an LLM-based car manual information retrieval application, with the objective of identifying user inputs for which the system fails to appropriately mention warnings contained in the manual. The testing solutions were evaluated based on their effectiveness in exposing failures and the diversity of the discovered failure-revealing tests. We report on the experimental methodology, the competitors, and the results.

标签

LLM Testing Benchmarking Automotive Assistant Information Retrieval

arXiv 分类

cs.AI