Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
偏好树优化:通过前瞻模拟增强面向目标的对话
机构 * Reichman University(赖希曼大学) ; School of Computer Science(计算机科学学院) ; School of Communications(传播学院)
AI总结 本研究提出偏好树优化(PTO)框架,结合带前瞻的偏好树与直接偏好优化(DPO),在动机访谈领域通过虚拟患者等模拟对话生成偏好数据,训练的对话智能体在关键指标上优于基线。
Comments 13 pages, 4 figures. Accepted at an ICLR 2025 workshop