arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3238 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3238 篇

2502.12148 2025-09-26 cs.CV 67%

HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and Generation

Ling Yang, Xinchen Zhang, Ye Tian, Chenming Shang, Minghao Xu, Wentao Zhang, Bin Cui

机构 * Peking University(北京大学) Tsinghua University(清华大学) Mila - Québec AI Institute(魁北克人工智能研究所)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

Comments NeurIPS 2025. Code: https://github.com/Gen-Verse/HermesFlow

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13388 2025-09-23 cs.CL cs.AI cs.LG 67%

R3: Robust Rubric-Agnostic Reward Models

David Anugraha, Zilu Tang, Lester James V. Miranda, Hanyang Zhao, Mohammad Rifqi Farhansyah, Garry Kuwanto, Derry Wijaya, Genta Indra Winata

机构 * Stanford University(斯坦福大学) Boston University(波士顿大学) Columbia University(哥伦比亚大学) University of Toronto(多伦多大学) Institut Teknologi Bandung(Bandung 工程技术大学) Monash University Indonesia(墨尔本大学印尼分校) Capital One

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15495 2025-09-18 cs.SE 67%

SynthCoder: A Synthetical Strategy to Tune LLMs for Code Completion

Dongjun Yu, Xiao Yan, Zhenrui Li, Jipeng Xiao, Haochuan He, Yongda Yu, Hao Zhang, Guoping Rong, Xiaobo Huang

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08826 2025-09-11 cs.CV 67%

RewardDance: Reward Scaling in Visual Generation

Jie Wu, Yu Gao, Zilyu Ye, Ming Li, Liang Li, Hanzhong Guo, Jie Liu, Zeyue Xue, Xiaoxia Hou, Wei Liu, Yan Zeng, Weilin Huang

机构 * ByteDance Seed(字节跳动种子)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract)

Comments Bytedance Seed Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00685 2025-09-03 eess.AS cs.SD 67%

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech

Kangxiang Xia, Xinfa Zhu, Jixun Yao, Lei Xie

机构 * School of Computer Science, Northwestern Polytechnical University(计算机科学学院)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

Comments Accepted by NCMMSC2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.03742 2025-09-03 cs.CL cs.AI cs.LG 67%

Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Ziyi Ye, Xiangsheng Li, Qiuchi Li, Qingyao Ai, Yujia Zhou, Wei Shen, Dong Yan, Yiqun Liu

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Baichuan AI(拜智科技) University of Copenhagen(哥本哈根大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref The Thirteenth International Conference on Learning Representations 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07818 2025-08-29 cs.CV 67%

DanceGRPO: Unleashing GRPO on Visual Generation

Zeyue Xue, Jie Wu, Yu Gao, Fangyuan Kong, Lingting Zhu, Mengzhao Chen, Zhiheng Liu, Wei Liu, Qiushan Guo, Weilin Huang, Ping Luo

机构 * The University of Hong Kong(香港大学)

专题命中 偏好对齐 :alignment(abstract);RLHF(abstract)

Comments Project Page: https://dancegrpo.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15805 2025-08-25 cs.CL cs.AI cs.LG 67%

ALAS: Autonomous Learning Agent for Self-Updating Language Models

Dhruv Atreja

机构 * Dhruv Atreja(独立研究者)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10652 2025-08-25 cs.CL cs.AI cs.CY 67%

Can Large Language Models Simulate Human Responses? A Case Study of Stated Preference Experiments in the Context of Heating-related Choices

Han Wang, Jacek Pawlak, Aruna Sivakumar

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00624 2025-08-12 cs.CV 67%

VideoSAVi: Self-Aligned Video Language Models without Human Supervision

Yogesh Kulkarni, Pooyan Fazli

机构 * Arizona State University(亚利桑那州立大学)

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03733 2025-08-07 cs.LG cs.AI cs.CL cs.CV 67%

CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning

Wenjie Li, Yujie Zhang, Haoran Sun, Yueqi Li, Fanrui Zhang, Mengzhe Xu, Victoria Borja Clausich, Sade Mellin, Renhao Yang, Chenrun Wang, Jethro Zih-Shuo Wang, Shiyi Yao, Gen Li, Yidong Xu, Hanyu Wang, Yilin Huang, Angela Lin Wang, Chen Shi, Yin Zhang, Jianan Guo, Luqi Yang, Renxuan Li, Yang Xu, Jiawei Liu, Yao Zhang, Lei Liu, Carlos Gutiérrez SanRomán, Lei Wang

机构 * College of Health Science and Technology, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院健康科学与技术学院) Shanghai Innovation Institute(上海创新研究院) Clinical Center for Sports Medicine, Department of Orthopaedics, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine(上海交通大学医学院骨科临床中心) School of Basic Medical Sciences, Intelligent Medicine Institute, Fudan University(复旦大学基础医学学院) Department of Hematology, The First Affiliated Hospital, College of Medicine, Zhejiang University(浙江大学医学院第一附属医院血液科) MoE Key Laboratory of Brain-Inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室) Department of Public Health and Primary Care, University of Cambridge(剑桥大学公共卫生与初级保健学院) Department of Medicine, Faculty of Health Sciences, Universidad CEU Cardenal Herrera(CEU卡德纳尔-赫尔曼大学健康科学学院医学系) Faculty of Medicine, University of Helsinki(赫尔辛基大学医学院) X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院X-LANCE实验室) Department of Hepatobiliary Surgery, National Cancer Center / National Clinical Research Center for Cancer / Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College(中国医学科学院肿瘤医院肝胆外科) Department of Surgery, The Ohio State University Wexner Medical Center, The James Comprehensive Cancer Center(俄亥俄州立大学韦克斯纳医学中心外科部,詹姆斯综合癌症中心) Ningbo Institute of Technology, Beihang University(北航宁波理工学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00222 2025-07-29 cs.CL cs.AI cs.LG 67%

Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training

Maximillian Chen, Ruoxi Sun, Tomas Pfister, Sercan Ö. Arık

机构 * Google(谷歌) Columbia University(哥伦比亚大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICLR 2025; Code: https://github.com/google-research/google-research/tree/master/learning_to_clarify

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15507 2025-07-22 cs.LG cs.AI cs.CL 67%

Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback

Johannes Ackermann, Takashi Ishida, Masashi Sugiyama

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accept at the Conference On Language Modeling (COLM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13158 2025-07-18 cs.LG cs.AI cs.CL 67%

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities

Hao Sun, Mihaela van der Schaar

机构 * Department of Applied Mathematics and Theoretical Physics(应用数学与理论物理系) University of Cambridge(剑桥大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13789 2025-07-17 cs.IR 67%

LEADRE: Multi-Faceted Knowledge Enhanced LLM Empowered Display Advertisement Recommender System

Fengxin Li, Yi Li, Yue Liu, Chao Zhou, Yuan Wang, Xiaoxiang Deng, Wei Xue, Dapeng Liu, Lei Xiao, Haijie Gu, Jie Jiang, Hongyan Liu, Biao Qin, Jun He

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

Comments Accepted by VLDB 2025 Industrial Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01299 2025-07-01 cs.CL cs.AI cs.LG 67%

The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation

Maja Pavlovic, Massimo Poesio

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments LREC-COLING NLPerspectives workshop

Journal ref https://aclanthology.org/2024.nlperspectives-1.11/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00845 2025-06-26 cs.CL cs.AI cs.LG 67%

Rewarding Graph Reasoning Process makes LLMs more Generalized Reasoners

Miao Peng, Nuo Chen, Zongrui Suo, Jia Li

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to KDD 2025 Research Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14574 2025-06-18 cs.LG cs.AI cs.CL 67%

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

Mingkang Zhu, Xi Chen, Zhongdao Wang, Bei Yu, Hengshuang Zhao, Jiaya Jia

机构 * The Chinese University of Hong Kong(香港中文大学) The University of Hong Kong(香港大学) The Hong Kong University of Science(香港科学大学) Huawei(华为)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.13238 2025-06-17 cs.LG cs.AI cs.CL 67%

Personalized Wireless Federated Learning for Large Language Models

Feibo Jiang, Li Dong, Siwei Tu, Yubo Peng, Kezhi Wang, Kun Yang, Cunhua Pan, Dusit Niyato

机构 * IEEE

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.18293 2025-06-10 cs.LG cs.AI cs.CL 67%

AMPO: Active Multi-Preference Optimization for Self-play Preference Selection

Taneesh Gupta, Rahul Madhavan, Xuchao Zhang, Chetan Bansal, Saravan Rajmohan

机构 * Microsoft(微软)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.06141 2025-06-05 cs.CV cs.AI cs.CL cs.LG 67%

MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization

Kangyu Zhu, Peng Xia, Yun Li, Hongtu Zhu, Sheng Wang, Huaxiu Yao

机构 * UNC Chapel-Hill(UNC夏洛特-希尔分校) Brown University(布朗大学) University of Washington(华盛顿大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01300 2025-06-03 cs.CV 67%

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding

Yiyang Zhou, Yangfan He, Yaofeng Su, Siwei Han, Joel Jang, Gedas Bertasius, Mohit Bansal, Huaxiu Yao

专题命中 偏好对齐 :alignment(abstract);DPO(abstract)

Comments 31 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01461 2025-06-03 cs.LG cs.AI cs.CL 67%

Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning Models

Huifeng Yin, Yu Zhao, Minghao Wu, Xuanfan Ni, Bo Zeng, Hao Wang, Tianqi Shi, Liangying Shao, Chenyang Lyu, Longyue Wang, Weihua Luo, Kaifu Zhang

机构 * Alibaba International Digital Commerce(阿里巴巴国际数字商业集团) Tsinghua University(清华大学) Monash University(墨尔本大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07067 2025-06-02 cs.CL cs.AI cs.LG 67%

DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs

Jongwoo Ko, Tianyi Chen, Sungnyun Kim, Tianyu Ding, Luming Liang, Ilya Zharkov, Se-Young Yun

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments ICML2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04350 2025-05-30 cs.CL cs.AI cs.LG cs.SC cs.SE 67%

CodeSteer: Symbolic-Augmented Language Models via Code/Text Guidance

Yongchao Chen, Yilun Hao, Yueying Liu, Yang Zhang, Chuchu Fan

机构 * Massachusetts Institute of Technology, Boston, MA, USA(麻省理工学院) Harvard University, Boston, MA, USA(哈佛大学) MIT-IBM Watson AI Lab, Boston, MA, USA(MIT-IBM Watson AI实验室) University of Illinois Urbana-Champaign, Urbana, IL, USA(伊利诺伊大学厄巴纳-香槟分校)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 28 pages, 12 figures

Journal ref International Conference on Machine Learning (ICML'2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18720 2025-05-27 cs.CL cs.AI cs.LG 67%

Optimal Transport-Based Token Weighting scheme for Enhanced Preference Optimization

Meng Li, Guangda Huzhang, Haibo Zhang, Xiting Wang, Anxiang Zeng

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) LLM Team, Shopee Pte. Ltd.(Shopee Pte. Ltd. 语言模型团队) Beijing Key Laboratory of Research on Large Models and Intelligent Governance and Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(北京大型模型与智能治理研究重点实验室)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 24 pages, 11 figures. Accepted by ACL 2025 (main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17482 2025-05-27 cs.CY cs.AI cs.CL cs.HC 67%

Reinforcement Learning from Human Feedback: Whose Culture, Whose Values, Whose Perspectives?

Kristian González Barman, Simon Lohse, Henk de Regt

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI、cs.CY

Journal ref González Barman, K., Lohse, S. & de Regt, H.W. Reinforcement Learning from Human Feedback in LLMs: Whose Culture, Whose Values, Whose Perspectives?. Philos. Technol. 38, 35 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15522 2025-05-21 cs.CL cs.AI cs.LG 67%

M-RewardBench: Evaluating Reward Models in Multilingual Settings

Srishti Gureja, Lester James V. Miranda, Shayekh Bin Islam, Rishabh Maheshwary, Drishti Sharma, Gusti Winata, Nathan Lambert, Sebastian Ruder, Sara Hooker, Marzieh Fadaee

专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 16 pages, 6 figures, 10 tables. Website: https://m-rewardbench.github.io/ , Updated results with latest models. Added more author information

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.12393 2025-05-16 cs.CL cs.AI cs.CY 67%

PersLLM: A Personified Training Approach for Large Language Models

Zheni Zeng, Jiayi Chen, Huimin Chen, Yukun Yan, Yuxuan Chen, Zhenghao Liu, Zhiyuan Liu, Maosong Sun

机构 * Tsinghua University(清华大学) Northeastern University(东北大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.CY

Comments 8 pages for main text, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.03618 2025-05-06 cs.CL cs.AI cs.LG 67%

A Logical Fallacy-Informed Framework for Argument Generation

Luca Mouchel, Debjit Paul, Shaobo Cui, Robert West, Antoine Bosselut, Boi Faltings

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏