David Anugraha, Zilu Tang, Lester James V. Miranda, Hanyang Zhao, Mohammad Rifqi Farhansyah, Garry Kuwanto, Derry Wijaya, Genta Indra Winata
机构
*
Stanford University(斯坦福大学)
;
Boston University(波士顿大学)
;
Columbia University(哥伦比亚大学)
;
University of Toronto(多伦多大学)
;
Institut Teknologi Bandung(Bandung 工程技术大学)
;
Monash University Indonesia(墨尔本大学印尼分校)
;
Capital One
CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
Wenjie Li, Yujie Zhang, Haoran Sun, Yueqi Li, Fanrui Zhang, Mengzhe Xu, Victoria Borja Clausich, Sade Mellin, Renhao Yang, Chenrun Wang, Jethro Zih-Shuo Wang, Shiyi Yao, Gen Li, Yidong Xu, Hanyu Wang, Yilin Huang, Angela Lin Wang, Chen Shi, Yin Zhang, Jianan Guo, Luqi Yang, Renxuan Li, Yang Xu, Jiawei Liu, Yao Zhang, Lei Liu, Carlos Gutiérrez SanRomán, Lei Wang
机构
*
College of Health Science and Technology, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院健康科学与技术学院)
;
Shanghai Innovation Institute(上海创新研究院)
;
Clinical Center for Sports Medicine, Department of Orthopaedics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院骨科临床中心)
;
School of Basic Medical Sciences, Intelligent Medicine Institute, Fudan University(复旦大学基础医学学院)
;
Department of Hematology, The First Affiliated Hospital, College of Medicine, Zhejiang University(浙江大学医学院第一附属医院血液科)
;
MoE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室)
;
Department of Public Health and Primary Care, University of Cambridge(剑桥大学公共卫生与初级保健学院)
;
Department of Medicine, Faculty of Health Sciences, Universidad CEU Cardenal Herrera(CEU卡德纳尔-赫尔曼大学健康科学学院医学系)
;
Faculty of Medicine, University of Helsinki(赫尔辛基大学医学院)
;
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院X-LANCE实验室)
;
Department of Hepatobiliary Surgery, National Cancer Center / National Clinical Research Center for Cancer / Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国医学科学院肿瘤医院肝胆外科)
;
Department of Surgery, The Ohio State University Wexner Medical Center, The James Comprehensive Cancer Center(俄亥俄州立大学韦克斯纳医学中心外科部,詹姆斯综合癌症中心)
;
Ningbo Institute of Technology, Beihang University(北航宁波理工学院)
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院)
;
LLM Team, Shopee Pte. Ltd.(Shopee Pte. Ltd. 语言模型团队)
;
Beijing Key Laboratory of Research on Large Models and Intelligent Governance and Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(北京大型模型与智能治理研究重点实验室)
专题命中
偏好对齐
:DPO(abstract);分类 cs.CL、cs.AI、cs.LG
Comments24 pages, 11 figures. Accepted by ACL 2025 (main)
Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?
Kristian González Barman, Simon Lohse, Henk de Regt
专题命中
偏好对齐
:RLHF(abstract);分类 cs.CL、cs.AI、cs.CY
Journal refGonzález Barman, K., Lohse, S. & de Regt, H.W. Reinforcement Learning from Human Feedback in LLMs: Whose Culture, Whose Values, Whose Perspectives?. Philos. Technol. 38, 35 (2025)
M-RewardBench: Evaluating Reward Models in Multilingual Settings
Srishti Gureja, Lester James V. Miranda, Shayekh Bin Islam, Rishabh Maheshwary, Drishti Sharma, Gusti Winata, Nathan Lambert, Sebastian Ruder, Sara Hooker, Marzieh Fadaee
专题命中
偏好对齐
:safety(abstract);分类 cs.CL、cs.AI、cs.LG
Comments16 pages, 6 figures, 10 tables. Website: https://m-rewardbench.github.io/ , Updated results with latest models. Added more author information