Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
通过逐步偏好调优的多模态智能体迭代工具使用探索
机构 * Beijing Key Laboratory of Intelligent Information Technology, School of Computer Science & Technology, Beijing Institute of Technology(北京智能信息科技重点实验室,计算机科学与技术学院,北京理工大学) ; State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室,BIGAI) ; State Key Laboratory of General Artificial Intelligence, Peking University(通用人工智能国家重点实验室,北京大学) ; Harbin Institute of Technology(哈尔滨工业大学) ; Guangdong Laboratory of Machine Perception and Intelligent Computing, Shenzhen MSU-BIT University(广东机器感知与智能计算实验室,深圳MSU-BIT大学) ; Department of Automation, Tsinghua University(自动化系,清华大学)
专题命中 VLM训练与架构 :vision language model(abstract);分类 cs.CV
AI总结 提出SPORT方法,通过任务合成、步骤采样、步骤验证和偏好调优的迭代循环,使多模态智能体无需预收集数据即可自主探索和优化工具使用策略,在GTA和GAIA基准上分别提升6.41%和3.64%。
Comments 24 pages