When Do We Need LLMs? A Diagnostic for Language-Driven Bandits
AI 摘要
论文研究了文本和数值结合的CMAB问题,提出了LLMP-UCB算法,并探讨了何时需要使用LLM。
主要贡献
- 提出LLMP-UCB算法,利用LLM进行不确定性估计
- 发现基于文本嵌入的轻量级数值bandit可媲美LLM
- 提出基于embedding的几何诊断方法,指导LLM使用
方法论
提出LLMP-UCB算法,并与基于文本嵌入的bandit进行实验对比,提出了基于embedding的几何诊断方法。
原文摘要
We study Contextual Multi-Armed Bandits (CMABs) for non-episodic sequential decision making problems where the context includes both textual and numerical information (e.g., recommendation systems, dynamic portfolio adjustments, offer selection; all frequent problems in finance). While Large Language Models (LLMs) are increasingly applied to these settings, utilizing LLMs for reasoning at every decision step is computationally expensive and uncertainty estimates are difficult to obtain. To address this, we introduce LLMP-UCB, a bandit algorithm that derives uncertainty estimates from LLMs via repeated inference. However, our experiments demonstrate that lightweight numerical bandits operating on text embeddings (dense or Matryoshka) match or exceed the accuracy of LLM-based solutions at a fraction of their cost. We further show that embedding dimensionality is a practical lever on the exploration-exploitation balance, enabling cost--performance tradeoffs without prompt complexity. Finally, to guide practitioners, we propose a geometric diagnostic based on the arms' embedding to decide when to use LLM-driven reasoning versus a lightweight numerical bandit. Our results provide a principled deployment framework for cost-effective, uncertainty-aware decision systems with broad applicability across AI use cases in financial services.