CommentsRevised version. This update includes substantial improvements to the methodology, ablation studies, temporal robustness analysis, and cross-city generalization experiments. Submitted to IEEE Transactions on Cognitive Communications and Networking. 15 pages, 14 figures
Improving Generalization Robustness of Multimodal RLVR
提升多模态RLVR的泛化鲁棒性
Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng, Chenrui Zhou, Lama Moukheiber, Yixing Ma, Bin Xu, Jiajun Song, Zhenglin Wan, Wangbo Zhao, Jiasheng Tang, Bohan Zhuang, Fan Wang, Yang You
机构
*
National University of Singapore(新加坡国立大学)
;
DAMO Academy Alibaba Group(阿里巴巴达摩院)
;
Hupan Lab(湖畔实验室)
;
Zhejiang University(浙江大学)
;
University of California Berkeley(加州大学伯克利分校)
;
Rochester Institute of Technology(罗切斯特理工学院)
;
Georgia Institute of Technology(佐治亚理工学院)
;
Renmin University of China(中国人民大学)
;
Hong Kong University of Science and Technology(香港科技大学)
专题命中
后训练与偏好优化
:large language model(abstract);language model(abstract);post-training(abstract);分类 cs.AI
Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
为扩散模型设计强化学习:统一的路径空间视角
Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He
机构
*
State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)
;
School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
;
ByteDance(字节跳动)