Agent Tuning & Optimization 相关度: 8/10

DDO-RM for LLM Preference Optimization: A Minimal Held-Out Benchmark against DPO

Tiantian Zhang, Jierui Zuo, Wenping Wang
arXiv: 2604.11119v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

论文提出DDO-RM优化策略,在LLM偏好优化任务中优于DPO,但结果初步。

主要贡献

  • 提出DDO-RM算法
  • 对比DDO-RM和DPO在LLM偏好优化上的性能
  • 初步验证DDO-RM在pairwise偏好学习上的有效性

方法论

DDO-RM将prompt视为决策问题,构建响应策略分布,利用reward model引导目标分布。

原文摘要

This paper reorganizes the current manuscript around the DPO versus DDO-RM preference-optimization project and focuses on two parts: the algorithmic view and the preliminary held-out benchmark. The benchmark asks a narrow question: even in a minimal pairwise chosen-versus-rejected setting, can a reward-guided decision-distribution update outperform a direct pairwise objective? We compare Direct Preference Optimization (DPO) against DDO-RM on EleutherAI/pythia-410m using HuggingFaceH4/ultrafeedback\_binarized, evaluate on the held-out test\_prefs split, and report results for seeds 42, 13, and 3407. Algorithmically, DDO-RM treats each prompt as a finite decision problem over candidate responses. Instead of optimizing only a binary chosen-rejected relation, it forms a policy distribution over candidates, centers reward-model scores under that distribution, and distills a reward-guided target distribution back into the policy. In the current public benchmark, DDO-RM improves mean pair accuracy from 0.5238 to 0.5602, AUC from 0.5315 to 0.5382, and mean margin from 0.1377 to 0.5353 relative to DPO. These are encouraging but still preliminary results: the study covers one model family, one dataset, one held-out evaluation split, and three seeds.

标签

LLM 偏好优化 DPO DDO-RM 奖励模型

arXiv 分类

stat.ML cs.LG