机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)
;
Beijing Academy of Artificial Intelligence(北京人工智能研究院)
;
Tsinghua University(清华大学)
;
The Hong Kong Polytechnic University(香港理工大学)
机构
*
School of Computer Science, Peking University(北京大学计算机学院)
;
School of Software and Microelectronics, Peking University(北京大学软件与微电子学院)
;
School of Physics, Peking University(北京大学物理学院)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
Central Research Institute, ZTE Corporation(中兴公司中央研究院)
专题命中
后训练与偏好优化
:post-training(title,abstract);large language model(abstract);language model(abstract);分类 cs.AI
TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue
TD-DPO:用于减轻临床自闭症干预对话中谄媚行为的差异感知偏好优化
Shuzhong Lai, Junhong Lai, Chenxi Li, Qing Zhou, Haifeng Li, Gang Pan, Lin Yao, Yueming Wang
机构
*
Nanhu Brain-Computer Interface Institute(南湖脑机接口研究所)
;
MOE Frontiers Science Center for Brain and Brain-Machine Integration, Zhejiang University(浙江大学脑与脑机融合教育部前沿科学中心)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
;
Children’s Hospital Zhejiang University School of Medicine(浙江大学医学院附属儿童医院)
;
State Key Laboratory of Brain-Machine Intelligence(脑机智能技术国家重点实验室)
;
Department of Neurobiology, Affiliated Mental Health Center and Hangzhou Seventh People’s Hospital, Zhejiang University School of Medicine(浙江大学医学院附属精神卫生中心和杭州市第七人民医院神经生物学系)
专题命中
后训练与偏好优化
:preference optimization(title,abstract);large language model(abstract);language model(abstract);分类 cs.LG