arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3238 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3238 篇

2508.06026 2025-08-11 cs.CL cs.AI 62%

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future

Yidong Wang, Xin Wang, Cunxiang Wang, Junfeng Fang, Qiufeng Wang, Jianing Chu, Xuran Meng, Shuxun Yang, Libo Qin, Yue Zhang, Wei Ye, Shikun Zhang

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03990 2025-08-07 cs.CL cs.AI cs.HC 62%

Are Today's LLMs Ready to Explain Well-Being Concepts?

Bohan Jiang, Dawei Li, Zhen Tan, Chengshuai Zhao, Huan Liu

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 4 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.05273 2025-08-07 cs.RO cs.AI cs.LG 62%

Real-World Offline Reinforcement Learning from Vision Language Model Feedback

Sreyas Venkataraman, Yufei Wang, Ziyu Wang, Navin Sriram Ravie, Zackory Erickson, David Held

机构 * Indian Institute of Technology, Kharagpur(印度理工学院,克哈格浦尔分校) IIIS, Tsinghua University(清华大学人工智能研究所) Robotics Institute, Carnegie Mellon University(卡内基梅隆大学机器人研究所)

专题命中 偏好对齐 :safety(abstract);分类 cs.AI、cs.LG

Comments 7 pages. Accepted at the LangRob Workshop 2024 @ CoRL, 2024. Accepted at 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.13795 2025-08-05 cs.LG cs.AI 62%

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN

Pengxiang Li, Lu Yin, Shiwei Liu

机构 * Dalian University of Technology(大连理工大学) University of Surrey(Surrey大学) University of Oxford(牛津大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.04858 2025-07-31 cs.AI cs.LG 62%

Don't Lag, RAG: Training-Free Adversarial Detection Using RAG

Roie Kazoom, Raz Lapid, Moshe Sipper, Ofer Hadar

机构 * Electrical and Computer Engineering, Ben Gurion University, Beer Sheba 84105, Israel(电子与计算机工程系,本· Gurion 大学) Computer Science, Ben Gurion University, Beer Sheba 84105, Israel(计算机科学系,本· Gurion 大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

Comments Accepted at VecDB @ ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21931 2025-07-30 cs.CL cs.AI 62%

Post-Training Large Language Models via Reinforcement Learning from Self-Feedback

Carel van Niekerk, Renato Vukovic, Benjamin Matthias Ruppik, Hsien-chin Lin, Milica Gašić

机构 * Heinrich Heine Universität(海因里希·海因大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17107 2025-07-30 cs.LG cs.AI 62%

Reinforcement Learning Fine-Tunes a Sparse Subnetwork in Large Language Models

Andrii Balashov

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

Comments The manuscript has been withdrawn due to significant overlap in methodology and results with a prior work (arXiv:2505.11711) that we were not aware of at the time of submission. To maintain academic integrity and avoid redundancy in the literature, we have chosen to withdraw this version

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20956 2025-07-29 cs.CL cs.AI 62%

Mind the Gap: Conformative Decoding to Improve Output Diversity of Instruction-Tuned Large Language Models

Max Peeperkorn, Tom Kouwenhoven, Dan Brown, Anna Jordanous

机构 * School of Computing University of Kent(肯特大学计算机学院) Leiden Institute of Advanced Computer Science Universiteit Leiden(莱顿先进计算机科学研究所) Cheriton School of Computer Science University of Waterloo(滑铁卢大学查里顿计算机科学学院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06645 2025-07-29 cs.CL cs.AI 62%

FocalPO: Enhancing Preference Optimizing by Focusing on Correct Preference Rankings

Tong Liu, Xiao Yu, Wenxuan Zhou, Jindong Gu, Volker Tresp

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16252 2025-07-23 cs.CL cs.AI 62%

Efficient RL for optimizing conversation level outcomes with an LLM-based tutor

Hyunji Nam, Omer Gottesman, Amy Zhang, Dean Foster, Emma Brunskill, Lyle Ungar

机构 * Stanford University(斯坦福大学) Amazon(亚马逊公司) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Pennsylvania(宾夕法尼亚大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15907 2025-07-23 cs.LG cs.AI 62%

Dual Turing Test: A Framework for Detecting and Mitigating Undetectable AI

Alberto Messina

机构 * RAI - Radiotelevisione Italiana, Centre for Research, Technological Innovation and Experimentation (CRITS)(意大利广播电视台,研究中心、技术创新与实验中心(CRITS))

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15281 2025-07-22 cs.CL cs.AI 62%

A Novel Self-Evolution Framework for Large Language Models

Haoran Sun, Zekun Zhang, Shaoning Zeng

机构 * Haoran Sun, Zekun Zhang, Shaoning Zeng(作者)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12759 2025-07-18 cs.CL cs.AI 62%

Logit Arithmetic Elicits Long Reasoning Capabilities Without Training

Yunxiang Zhang, Muhammad Khalifa, Lechen Zhang, Xin Liu, Ayoung Lee, Xinliang Frederick Zhang, Farima Fatahi Bayat, Lu Wang

机构 * Computer Science and Engineering University of Michigan(计算机科学与工程系密歇根大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04099 2025-07-16 cs.CL cs.AI 62%

Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching

Thomas Savage

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07375 2025-07-11 cs.LG cs.CL 62%

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary

Zhiwei Zhang, Hui Liu, Xiaomin Li, Zhenwei Dai, Jingying Zeng, Fali Wang, Minhua Lin, Ramraj Chandradevan, Zhen Li, Chen Luo, Xianfeng Tang, Qi He, Suhang Wang

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) Amazon(亚马逊) Havard University(哈佛大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18071 2025-07-08 cs.CL cs.AI 62%

Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals

Jia-Nan Li, Jian Guan, Wei Wu, Rui Yan

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14363 2025-07-08 cs.LG cs.CL 62%

Improving RL Exploration for LLM Reasoning through Retrospective Replay

Shihan Dou, Muling Wu, Jingwen Xu, Rui Zheng, Tao Gui, Qi Zhang, Xuanjing Huang

机构 * Fudan University, Shanghai, China(复旦大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.LG

Comments 13 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.04517 2025-07-08 cs.LG cs.CL 62%

Towards Cost-Effective Reward Guided Text Generation

Ahmad Rashid, Ruotian Wu, Rongqi Fan, Hongliang Li, Agustinus Kristiadi, Pascal Poupart

机构 * University of Waterloo(滑铁卢大学) Vector Institute(向量研究所) Huawei Technologies(华为技术)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.LG

Comments 18 pages. Work accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.14774 2025-07-08 cs.LG cs.CL 62%

Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning

Simran Kaur, Simon Park, Anirudh Goyal, Sanjeev Arora

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.LG

Journal ref International Conference on Learning Representations (ICLR 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00316 2025-07-03 cs.LG cs.CL eess.IV 62%

$μ^2$Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation

Siyou Li, Pengyao Qin, Huanan Wu, Dong Nie, Arun J. Thirunavukarasu, Juntao Yu, Le Zhang

机构 * School of Electronic Engineering and Computer Science(电子工程与计算机科学学院) Queen Mary University of London(伦敦女王学院) School of Engineering, College of Engineering and Physical Sciences(工程学院,工程与物理科学学院) University of Birmingham(伯明翰大学) Guangdong University of Technology(广东工业大学) Meta Inc. US(Meta美国公司) Nuffield Department of Clinical Neurosciences(临床神经科学系) University of Oxford(牛津大学) William Harvey Research Institute, NIHR Barts Biomedical Research Centre, Queen Mary University London(威廉·哈里弗研究所在NIHR巴茨生物医学研究中心,伦敦女王学院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.LG

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22698 2025-07-02 cs.CL cs.AI 62%

Text Production and Comprehension by Human and Artificial Intelligence: Interdisciplinary Workshop Report

Emily Dux Speltz

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21463 2025-06-27 cs.CL cs.LG cs.SD eess.AS 62%

Aligning Spoken Dialogue Models from User Interactions

Anne Wu, Laurent Mazaré, Neil Zeghidour, Alexandre Défossez

机构 * Department of Computer Science, Cornell University. Work done at Kyutai.(计算机科学系,康奈尔大学。在Kyutai工作)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.LG

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00753 2025-06-27 cs.CL cs.LG 62%

LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey

Henry Peng Zou, Wei-Chieh Huang, Yaozu Wu, Yankai Chen, Chunyu Miao, Hoang Nguyen, Yue Zhou, Weizhi Zhang, Liancheng Fang, Langzhou He, Yangning Li, Dongyuan Li, Renhe Jiang, Xue Liu, Philip S. Yu

机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) University of Tokyo(东京大学) Tsinghua University(清华大学) MBZUAI McGill University(麦吉尔大学) University of California Irvine(加州大学 Irvine 分校)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.LG

Comments Paper lists and resources are available at https://github.com/HenryPengZou/Awesome-Human-Agent-Collaboration-Interaction-Systems

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17219 2025-06-26 cs.LG cs.AI 62%

No Free Lunch: Rethinking Internal Feedback for LLM Reasoning

Yanzhi Zhang, Zhaoxi Zhang, Haoxiang Guan, Yilin Cheng, Yitong Duan, Chen Wang, Yue Wang, Shuxin Zheng, Jiyan He

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18383 2025-06-24 cs.LG cs.AI 62%

LOGICPO: Efficient Translation of NL-based Logical Problems to FOL using LLMs and Preference Optimization

Koushik Viswanadha, Deepanway Ghosal, Somak Aditya

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16502 2025-06-23 cs.CL cs.AI 62%

Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples

Soumya Suvra Ghosal, Vaibhav Singh, Akash Ghosh, Soumyabrata Pal, Subhadip Baidya, Sriparna Saha, Dinesh Manocha

机构 * University of Maryland, College Park(马里兰大学 College Park 分校) IIT Bombay(印度理工学院班加罗尔分校) IIT Patna(印度理工学院帕坦分校) Adobe Research(Adobe 研究院) IIT Kanpur(印度理工学院坎普尔分校)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15706 2025-06-23 cs.LG cs.AI 62%

MDPO: Multi-Granularity Direct Preference Optimization for Mathematical Reasoning

Yunze Lin

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.21349 2025-06-23 cs.LG cs.AI cs.PF 62%

FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system

Zeyuan Li, Yangfan He, Lewei He, Jianhui Wang, Tianyu Shi, Bin Lei, Yuchen Li, Qiuwu Chen

机构 * School of Software, South China Normal University(南方科技大学软件学院) University of Minnesota - Twin Cities(明尼苏达大学双城分校) University of Electronic Science and Technology of China(电子科技大学) University of Toronto(多伦多大学) University of Connecticut(康涅狄格大学) AI center, AIGCode Inc.(AIGCode公司人工智能中心)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

Comments 20 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11098 2025-06-16 cs.LG cs.AI 62%

Debiasing Online Preference Learning via Preference Feature Preservation

Dongyoung Kim, Jinsung Yoon, Jinwoo Shin, Jaehyung Kim

机构 * KAIST(韩国科学技术院) Google Cloud AI Research(谷歌云人工智能研究) Yonsei University(延世大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

Comments 20 page, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07091 2025-06-13 cs.AI cs.LG 62%

AssistanceZero: Scalably Solving Assistance Games

Cassidy Laidlaw, Eli Bronstein, Timothy Guo, Dylan Feng, Lukas Berglund, Justin Svegliato, Stuart Russell, Anca Dragan

机构 * ucb(加州大学伯克利分校)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

Comments Presented at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏