IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages
AI 摘要
IndicDB是针对印度语言多语言Text-to-SQL的基准测试,揭示了LLM的“Indic Gap”。
主要贡献
- 构建了包含20个数据库和237个表的多语言Text-to-SQL基准IndicDB。
- 提出了一个迭代三代理框架(Architect, Auditor, Refiner)用于生成高质量的数据库模式。
- 评估了多个SOTA模型在IndicDB上的跨语言语义解析性能,揭示了英语到印度语言的性能下降(Indic Gap)。
方法论
使用三代理框架转换开放数据为关系数据库,并生成包含七种语言的Text-to-SQL任务,评估模型性能。
原文摘要
While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and simplified schemas, leaving a gap in real-world, non-Western applications. We present IndicDB, a multilingual Text-to-SQL benchmark for evaluating cross-lingual semantic parsing across diverse Indic languages. The relational schemas are sourced from open-data platforms, including the National Data and Analytics Platform (NDAP) and the India Data Portal (IDP), ensuring realistic administrative data complexity. IndicDB comprises 20 databases across 237 tables. To convert denormalized government data into rich relational structures, we employ an iterative three-agent framework (Architect, Auditor, Refiner) to ensure structural rigor and high relational density (11.85 tables per database; join depths up to six). Our pipeline is value-aware, difficulty-calibrated, and join-enforced, generating 15,617 tasks across English, Hindi, and five Indic languages. We evaluate cross-lingual semantic parsing performance of state-of-the-art models (DeepSeek v3.2, MiniMax 2.7, LLaMA 3.3, Qwen3) across seven linguistic variants. Results show a 9.00% performance drop from English to Indic languages, revealing an "Indic Gap" driven by harder schema linking, increased structural ambiguity, and limited external knowledge. IndicDB serves as a rigorous benchmark for multilingual Text-to-SQL. Code and data: https://anonymous.4open.science/r/multilingualText2Sql-Indic--DDCC/