Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
比较语言模型和人类的探索-利用策略:来自标准多臂老虎机实验的见解
机构 * Department of Mechanical & Industrial Engineering, University of Toronto(机械与工业工程系,多伦多大学) ; The Rotman School of Management, University of Toronto(罗特曼管理学院,多伦多大学) ; Department of Psychiatry, University of Toronto(精神病学系,多伦多大学)
AI总结 本文通过多臂老虎机实验比较语言模型、人类和算法的探索-利用策略,发现启用思考使语言模型行为更接近人类,但在复杂环境中表现受限。