GLeMM: A large-scale multilingual dataset for morphological research
AI 摘要
GLeMM是一个大规模多语种词法数据集,旨在支持词法研究的数据驱动方法和实验验证。
主要贡献
- 构建了一个大规模的多语种衍生词资源GLeMM
- 实现了跨七种欧洲语言(德语、英语、西班牙语、法语、意大利语、波兰语、俄语)的统一设计
- 自动标注了词法特征,并编码了语义描述
方法论
利用Wiktionary数据,通过自动化流程构建数据集,并自动标注词法和语义信息。
原文摘要
In derivational morphology, what mechanisms govern the variation in form-meaning relations between words? The answers to this type of questions are typically based on intuition and on observations drawn from limited data, even when a wide range of languages is considered. Many of these studies are difficult to replicate and generalize. To address this issue, we present GLeMM, a new derivational resource designed for experimentation and data-driven description in morphology. GLeMM is characterized by (i) its large size, (ii) its extensive coverage (currently amounting to seven European languages, i.e., German, English, Spanish, French, Italian, Polish, Russian, (iii) its fully automated design, identical across all languages, (iv) the automatic annotation of morphological features on each entry, as well as (v) the encoding of semantic descriptions for a significant subset of these entries. It enables researchers to address difficult questions, such as the role of form and meaning in word-formation, and to develop and experimentally test computational methods that identify the structures of derivational morphology. The article describes how GLeMM is created using Wiktionary articles and presents various case studies illustrating possible applications of the resource.