LLM Reasoning 相关度: 8/10

A Triadic Suffix Tokenization Scheme for Numerical Reasoning

Olga Chetverina
arXiv: 2604.11582v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出一种三位后缀分词(TST)方案,旨在提高LLM的数值推理能力,解决数字分词不一致问题。

主要贡献

  • 提出Triadic Suffix Tokenization (TST)分词方案
  • 设计两种TST的实现变体:词汇表方法和后缀标记方法
  • TST方案具有可扩展性,可适应任意精度和范围

方法论

提出一种确定性的分词方案,将数字分成三位一组,并使用显式量级标记进行注释,以此解决数字的位置和十进制结构问题。

原文摘要

Standard subword tokenization methods fragment numbers inconsistently, causing large language models (LLMs) to lose positional and decimal structure - a primary driver of errors in arithmetic and scientific reasoning. We introduce Triadic Suffix Tokenization (TST), a deterministic scheme that partitions digits into three-digit triads and annotates each triad with an explicit magnitude marker. Critically, the scheme defines a fixed, one-to-one mapping between suffixes and orders of magnitude for the integer part (thousands, millions, billions, etc.) and a parallel system of replicated markers for fractional depth (tenths, thousandths, millionths, etc.). Unlike approaches that rely on positional inference, this method provides a consistent gradient signal, which should ensure stable convergence. Two implementation variants are proposed: (1) a vocabulary-based approach that adds at most 10,000 fixed tokens to an existing vocabulary, covering 33 orders of magnitude ($10^{-15}$ to $10^{18}$); and (2) a suffix-marker approach that uses a small set of special tokens to denote magnitude dynamically. Both variants preserve exact digits while making order-of-magnitude relationships transparent at the token level. The framework is inherently scalable, allowing for linear vocabulary expansion to accommodate arbitrary precision and range. TST is architecture-agnostic and can be integrated as a drop-in preprocessing step. Experimental validation is deferred to future work.

标签

分词 数值推理 LLM tokenization

arXiv 分类

cs.CL cs.AI cs.LG