AI Agents 相关度: 9/10

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill

Xuanbo Su, Wenhao Hu, Le Zhan, Yanqi Yang, Leo Huang
arXiv: 2604.07054v1 发布: 2026-04-08 更新: 2026-04-08

AI 摘要

提出了SalesLLM,一个用于评估LLM销售技能的双语基准,并训练了CustomerLM以提高模拟逼真度。

主要贡献

  • 构建了一个双语(中/英)销售对话基准SalesLLM,包含真实场景和可控难度。
  • 提出了一个全自动评估流程,结合了LLM评估器和微调BERT分类器。
  • 训练了一个用户模型CustomerLM,减少了角色反转,提升了模拟真实度。

方法论

构建SalesLLM基准,训练CustomerLM,并设计全自动评估流程,包含LLM评估器和BERT分类器。

原文摘要

Sales dialogues require multi-turn, goal-directed persuasion under asymmetric incentives, which makes them a challenging setting for large language models (LLMs). Yet existing dialogue benchmarks rarely measure deal progression and outcomes. We introduce SalesLLM, a bilingual (ZH/EN) benchmark derived from realistic applications covering Financial Services and Consumer Goods, built from 30,074 scripted configurations and 1,805 curated multi-turn scenarios with controllable difficulty and personas. We propose a fully automatic evaluation pipeline that combines (i) an LLM-based rater for sales-process progress, and (ii) fine-tuned BERT classifiers for end-of-dialogue buying intent. To improve simulation fidelity, we train a user model, CustomerLM, with SFT and DPO on 8,000 crowdworker-involved sales conversations, reducing role inversion from 17.44% (GPT-4o) to 8.8%. SalesLLM scores correlate strongly with expert human ratings (Pearson r=0.98). Experiments across 15 mainstream LLMs reveal substantial variability: top-performance LLMs are competitive with human-level performance while the less capable ones are worse than human. SalesLLM serves as a scalable benchmark for developing and evaluating outcome-oriented sales agents.

标签

LLM benchmark sales dialogue evaluation

arXiv 分类

cs.CL