Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles
Maestro:通过强化学习协调分层模型-技能集合
Jinyang Wu, Guocheng Zhai, Ruihan Jin, Yuhao Shen, Zhengxi Lu, Fan Zhang, Haoran Luo, Zheng Lian, Zhengqi Wen, Jianhua Tao
机构
*
Tsinghua University(清华大学)
;
Zhejiang University(浙江大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Nanyang Technological University(南洋理工大学)
;
Tongji University(同济大学)
机构
*
Department of Data Science, City University of Hong Kong(香港城市大学数据科学系)
;
Hong Kong Institute of AI for Science, City University of Hong Kong(香港城市大学人工智能科学研究院)
;
Li Auto Inc.
;
Beihang University(北航大学)
CommentsMajor revision: substantially reorganized the manuscript and added a theoretical explanation section. The replacement is intended for the same arXiv paper; the core topic and contribution remain the same
How Much Online RL is Enough? Informative Rollouts for Offline Preference Optimization in RLVR
在线强化学习需要多少?用于RLVR中离线偏好优化的信息性回放
Richa Verma, Balaraman Ravindran
机构
*
TCS Research Department of CSE(TCS计算机科学系研究部)
;
IIT Madras(印度理工学院马德拉斯分校)
;
Department of Data Science & AI(数据科学与人工智能系)
;
Wadhwani School of Data Science & AI(Wadhwani数据科学与人工智能学院)
机构
*
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
Noise-corrected GRPO: From Noisy Rewards to Unbiased Gradients
噪声校正的GRPO:从噪声奖励到无偏梯度
Omar El Mansouri, Fathinah Asma Izzati, Mohamed El Amine Seddik, Salem Lahlou
机构
*
Department of Machine Learning, Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE
;
Technology Innovation Institute, Abu Dhabi, UAE
;
Department of Robotics, Khalifa University, Abu Dhabi, UAE
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
探究跨模态技能注入:场景、方法与超参数
Zhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun
机构
*
State Key Laboratory for Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机科学学院,北京大学)
;
WeChat AI, Tencent Inc., China(腾讯公司,中国)
;
The University of Hong Kong(香港大学)
机构
*
National Engineering Research Center for Multimedia Software, School of Computer Science, Wuhan University(国家多媒体软件工程研究中心,武汉大学计算机学院)
;
Meituan Longcat Team(美团Longcat团队)
;
The University of Sydney(悉尼大学)
;
University of Science and Technology of China(中国科学技术大学)
Minimal-Intervention KV Retention via Set-Conditioned Diversity
通过集合条件多样性实现最小干预的KV保留
Libo Sun, Po-wei Harn, Peixiong He, Xiao Qin
机构
*
Department of Computer Science and Software Engineering(计算机科学与软件工程系)
;
Department of Information Management(信息管理系)
;
Auburn University(阿伯丁大学)
;
National Central University(国立中央大学)
Strategic Over-Parameterization for Generalizable Low-Rank Adaptation
战略性过参数化以实现通用的低秩适应
Jing Gao, Zhong-Yi Lu, Pan Zhang, Ze-Feng Gao
机构
*
School of Fundamental Physics and Mathematical Sciences, Hangzhou Institute for Advanced Study, UCAS, Hangzhou 310024, China(1 基础物理与数学科学学院,杭州先进研究院,UCAS,杭州 310024,中国)
;
School of Physical Sciences, University of Chinese Academy of Sciences, No. 19A Yuquan Road, Beijing 100049, China(2 物理科学学院,中国科学院大学,玉泉路19A号,北京 100049,中国)
;
CAS Key Laboratory of Theoretical Physics, Institute of Theoretical Physics, Chinese Academy of Sciences, Beijing 100190, China(3 中国科学院理论物理重点实验室,理论物理研究所,中国科学院,北京 100190,中国)
;
School of Physics and Key Laboratory of Quantum State Construction and Manipulation (Ministry of Education), Renmin University of China, Beijing 100872, China(4 物理学院和量子态构造与操控(教育部)重点实验室,中国人民大学,北京 100872,中国)
机构
*
Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系)
;
Singapore-MIT Alliance for Research and Technology Centre(新加坡-麻省理工联盟研究技术中心)
;
The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳))
;
CSAIL, Massachusetts Institute of Technology(麻省理工学院计算机科学与人工智能实验室)
;
Institute of Data Science, National University of Singapore(新加坡国立大学数据科学研究院)
RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning
Sitao Cheng, Tianle Li, Xuhan Huang, Xunjian Yin, Difan Zou
机构
*
Department of XXX, University of YYY, Location, Country(XXX系,YYY大学,地点,国家)
;
School of ZZZ, Institute of WWW, Location, Country(ZZZ学院,WWW研究所,地点,国家)
Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
学习一个持续思考标记以增强测试时扩展
Liran Ringel, Elad Tolochinsky, Yaniv Romano
机构
*
Department of Computer Science, Technion – Israel Institute of Technology(计算机科学系,技术学院–以色列理工学院)
;
Department of Electrical and Computer Engineering, Technion – Israel Institute of Technology(电气与计算机工程系,技术学院–以色列理工学院)
ChatSR: Multimodal Large Language Models for Scientific Formula Discovery
ChatSR:用于科学公式发现的多模态大语言模型
Yanjie Li, Lina Yu, Weijun Li, Min Wu, Liping Zhang, Jingyi Liu, Yusong Deng, Mingzhu Wan, Xin Ning
机构
*
AnnLab, Institute of Semiconductors, Chinese Academy of Sciences, Beijing, China(安 lab,半导体研究所,中国科学院,北京,中国)
;
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing, China(电子、电气与通信工程学院,中国科学院大学,北京,中国)
;
Zhongguancun Academy, Beijing, China(中关村学院,北京,中国)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences, Beijing 101408, China(先进交叉科学学院,中国科学院大学,北京101408,中国)
;
College of Materials Science and Opto-Electronic Technology, University of Chinese Academy of Sciences, Beijing, 100049, China(材料科学与光电技术学院,中国科学院大学,北京100049,中国)
;
School of Integrated Circuits, University of Chinese Academy of Sciences, Beijing 100049, China(集成电路学院,中国科学院大学,北京100049,中国)