LLM Reasoning 相关度: 8/10

GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding

Weicai Long, Yusen Hou, Junning Feng, Houcheng Su, Shuo Yang, Donglin Xie, Yanlin Zhang
arXiv: 2604.05774v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

GenomeQA基准测试通用LLM在基因组序列理解任务上的表现,揭示其优势与局限。

主要贡献

  • 构建GenomeQA基准,包含多种基因组推断任务。
  • 评估了多个通用LLM在基因组序列理解任务上的性能。
  • 分析了LLM在不同任务上的表现差异,揭示其对序列信号的利用能力。

方法论

构建包含5200个样本的GenomeQA基准,评估LLM在六个基因组推断任务家族上的表现,并与随机基线进行比较。

原文摘要

Large Language Models (LLMs) are increasingly adopted as conversational assistants in genomics, where they are mainly used to reason over biological knowledge, annotations, and analysis outputs through natural language interfaces. However, existing benchmarks either focus on specialized DNA models trained for sequence prediction or evaluate biological knowledge using text-only questions, leaving the behavior of general-purpose LLMs when directly exposed to raw genome sequences underexplored. We introduce GenomeQA, a benchmark designed to provide a controlled evaluation setting for general-purpose LLMs on sequence-based genome inference tasks. GenomeQA comprises 5,200 samples drawn from multiple biological databases, with sequence lengths ranging from 6 to 1,000 base pairs (bp), spanning six task families: Enhancer and Promoter Identification, Splice Site Identification, Taxonomic Classification, Histone Mark Prediction, Transcription Factor Binding Site Prediction, and TF Motif Prediction. Across six frontier LLMs, we find that models consistently outperform random baselines and can exploit local sequence signals such as GC content and short motifs, while performance degrades on tasks that require more indirect or multi-step inference over sequence patterns. GenomeQA establishes a diagnostic benchmark for studying and improving the use of general-purpose LLMs on raw genomic sequences.

标签

LLM 基因组学 基准测试 序列理解 生物信息学

arXiv 分类

q-bio.GN cs.CL