arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3238 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3238 篇

2511.07896 2025-11-12 cs.AI cs.CL 62%

SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder

Dengcan Liu, Jiahao Li, Zheren Fu, Yi Tu, Jiajun Li, Zhendong Mao, Yongdong Zhang

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments 15pages,11figures,AAAI-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07483 2025-11-12 cs.AI cs.LG 62%

Beyond Correctness: Confidence-Aware Reward Modeling for Enhancing Large Language Model Reasoning

Qianxi He, Qingyu Ren, Shanzhe Lei, Xuhong Wang, Yingchun Wang

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Shanghai Key Laboratory of Data Science, School of Computer Science, Fudan University(上海数据科学关键实验室,计算机科学学院,复旦大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04286 2025-11-07 cs.LG cs.AI 62%

Efficient Reinforcement Learning from Human Feedback via Bayesian Preference Inference

Matteo Cercola, Valeria Capretti, Simone Formentin

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06915 2025-11-05 cs.CL cs.AI 62%

LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling

Zecheng Tang, Baibei Ji, Quantong Qiu, Haitian Wang, Xiaobo Liang, Juntao Li, Min Zhang

机构 * Soochow University(苏州大学) LCM Laboratory(LCM实验室)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01807 2025-11-04 cs.CL cs.AI 62%

Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining

Adewale Akinfaderin, Shreyas Subramanian, Akarsha Sehwag

机构 * Amazon Web Services(亚马逊网络服务)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments Presented at Workshop on Prompt Optimization, KDD 2025, Toronto, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13487 2025-11-03 cs.CL cs.AI 62%

Detecting Prefix Bias in LLM-based Reward Models

Ashwin Kumar, Yuzi He, Aram H. Markosyan, Bobbie Chern, Imanol Arrieta-Ibarra

机构 * Washington University in St Louis(华盛顿大学圣路易斯分校) Meta Platforms, Inc.(Meta平台公司)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01308 2025-10-30 cs.AI cs.CL cs.DB 62%

GradeSQL: Test-Time Inference with Outcome Reward Models for Text-to-SQL Generation from Large Language Models

Mattia Tritto, Giuseppe Farano, Dario Di Palma, Gaetano Rossiello, Fedelucio Narducci, Dharmashankar Subramanian, Tommaso Di Noia

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23751 2025-10-29 cs.LG cs.AI stat.ML 62%

Debiasing Reward Models by Representation Learning with Guarantees

Ignavier Ng, Patrick Blöbaum, Siddharth Bhandari, Kun Zhang, Shiva Kasiviswanathan

机构 * Carnegie Mellon University(卡内基梅隆大学) Amazon(亚马逊) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德智能大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.23217 2025-10-28 cs.CL cs.AI 62%

Process Reward Models for Sentence-Level Verification of LVLM Radiology Reports

Alois Thomas, Maya Varma, Jean-Benoit Delbrouck, Curtis P. Langlotz

机构 * AIMI Center Stanford University(AIMI中心 斯坦福大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.19050 2025-10-23 cs.AI cs.LG 62%

Rectifying Shortcut Behaviors in Preference-based Reward Learning

Wenqian Ye, Guangtao Zheng, Aidong Zhang

机构 * University of Virginia(弗吉尼亚大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15061 2025-10-23 cs.LG cs.CL 62%

Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models

Samuel Paech, Allen Roush, Judah Goldfeder, Ravid Shwartz-Ziv

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.LG

Comments 11 pages + appendices, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18849 2025-10-22 cs.CL cs.AI 62%

Towards Faithful and Controllable Personalization via Critique-Post-Edit Reinforcement Learning

Chenghao Zhu, Meiling Tao, Tiannan Wang, Dongyi Ding, Yuchen Eleanor Jiang, Wangchunshu Zhou

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Electronic Science and Technology of China(电子科技大学) South China Agricultural University(华南农业大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

Comments work in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01545 2025-10-17 cs.LG cs.AI cs.RO 62%

Predictive Preference Learning from Human Interventions

Haoyuan Cai, Zhenghao Peng, Bolei Zhou

机构 * Department of Computer Science, University of California, Los Angeles(计算机科学系,加州大学洛杉矶分校)

专题命中 偏好对齐 :safety(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2025 Spotlight. Project page: https://metadriverse.github.io/ppl

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10963 2025-10-14 cs.LG cs.AI 62%

APLOT: Robust Reward Modeling via Adaptive Preference Learning with Optimal Transport

Zhuo Li, Yuege Feng, Dandan Guo, Jinpeng Hu, Anningzhe Gao, Xiang Wan

机构 * Shenzhen International Center for Industrial and Applied Mathematics(深圳工业与应用数学国际中心) Shenzhen Research Institute of Big Data(深圳大数据研究 institute) The Chinese University of Hong Kong Shenzhen(香港中文大学(深圳)) Birmingham City University(伯明翰城市大学) Jilin University(吉林大学) KAUST(科威特大学) Hefei University of Technology(合肥工业大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

Comments EMNLP2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21776 2025-10-14 cs.CL cs.AI cs.IR 62%

WebThinker: Empowering Large Reasoning Models with Deep Research Capability

Xiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian, Yongkang Wu, Ji-Rong Wen, Yutao Zhu, Zhicheng Dou

机构 * Renmin University of China(中国人民大学) BAAI(北京人工智能研究院) Huawei Poisson Lab(华为Poisson实验室)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16567 2025-10-10 cs.LG cs.AI cs.CR 62%

Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning

Thibaud Gloaguen, Mark Vero, Robin Staab, Martin Vechev

机构 * ETH Zurich(苏黎世联邦理工学院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07497 2025-10-10 cs.CL cs.AI eess.AS 62%

Can Speech LLMs Think while Listening?

Yi-Jen Shih, Desh Raj, Chunyang Wu, Wei Zhou, SK Bong, Yashesh Gaur, Jay Mahadeokar, Ozlem Kalinli, Mike Seltzer

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) Meta Superintelligence Labs(Meta超智能实验室)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06391 2025-10-09 cs.CL cs.AI 62%

Reward Model Perspectives: Whose Opinions Do Reward Models Reward?

Elle

机构 * Elle University of Oxford(埃勒 奥克舍大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments Published at EMNLP 2025 under the full author name "Elle"

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21432 2025-10-08 cs.CL cs.AI 62%

Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour

Tareq Alsaleh, Bilal Farooq

机构 * Laboratory of Innovations in Transportation (LiTrans), Toronto Metropolitan University, Canada(交通创新实验室(LiTrans)、多伦多 Metropolitan 大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.16810 2025-10-07 cs.CL cs.AI 62%

Can GPT models Follow Human Summarization Guidelines? A Study for Targeted Communication Goals

Yongxin Zhou, Fabien Ringeval, François Portet

机构 * Univ. Grenoble Alpes, CNRS, Inria, Grenoble INP, LIG(格勒诺布尔阿尔卑斯大学、国家科学研究中心、法国国家信息与自动化研究所、格勒诺布尔INP、实验室)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

Comments INLG 2025, Hanoi, Vietnam, October 29 - November 2, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03231 2025-10-06 cs.CL cs.AI 62%

Reward Models are Metrics in a Trench Coat

Sebastian Gehrmann

机构 * Bloomberg(高盛)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02338 2025-10-06 cs.CL cs.AI 62%

Optimizing Long-Form Clinical Text Generation with Claim-Based Rewards

Samyak Jhaveri, Praphul Singh, Jangwon Kim, Tara Taghavi, Krishnaram Kenthapadi

机构 * Oracle Health AI(Oracle健康AI)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01394 2025-10-03 cs.LG cs.CL 62%

Optimal Stopping vs Best-of-$N$ for Inference Time Optimization

Yusuf Kalayci, Vinod Raman, Shaddin Dughmi

机构 * University of Southern California(南加州大学) University of Michigan(密歇根大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.LG

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01236 2025-10-03 cs.CL cs.LG 62%

GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings

Ismam Nur Swapnil, Aranya Saha, Tanvir Ahmed Khan, Mohammad Ariful Haque

机构 * Department of Electrical and Electronic Engineering, Bangladesh University of Engineering and Technology (BUET)(电气与电子工程系,孟加拉国工程与技术大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.LG

Comments Will be submitted at IEEE JBHI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03238 2025-10-03 cs.CL cs.AI 62%

FANS -- Formal Answer Selection for Natural Language Math Reasoning Using Lean4

Jiarui Yao, Ruida Wang, Tong Zhang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 偏好对齐 :trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23601 2025-09-30 cs.CL cs.AI 62%

Semantic-guided Diverse Decoding for Large Language Model

Weijie Shi, Yue Cui, Yaguang Wu, Jingzhi Fang, Shibo Zhang, Mengze Li, Sirui Han, Jia Zhu, Jiajie Xu, Xiaofang Zhou

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23715 2025-09-30 cs.CL cs.LG 62%

Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering

Eduard Barbu, Adrian Marius Dumitran

机构 * Faculty of Mathematics and Informatics, University of Bucharest(数学与信息学系,布加勒斯特大学)

专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.LG

Comments Accepted@ CONSILR 2025 Bucharest Romania 9-10 October

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23368 2025-09-30 cs.CL cs.AI 62%

MedCritical: Enhancing Medical Reasoning in Small Language Models via Self-Collaborative Correction

Xinchun Su, Chunxu Luo, Yixuan Li, Weidong Yang, Lipeng Ma

机构 * School of Computer Science, Fudan University, Shanghai, China(复旦大学计算机学院)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15845 2025-09-30 cs.CL cs.AI 62%

Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports

Chengbo Sun, Hui Yi Leong, Lei Li

机构 * University of Chicago(芝加哥大学) University of Washington(华盛顿大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12716 2025-09-29 cs.CL cs.AI 62%

Shadow-FT: Tuning Instruct Model via Training on Paired Base Model

Taiqiang Wu, Runming Yang, Jiayi Li, Pengfei Hu, Yik-Chung Wu, Ngai Wong, Yujiu Yang

机构 * The University of Hong Kong(香港大学) Tsinghua University(清华大学) Tencent(腾讯)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL、cs.AI

Comments 24 pages, 12 tables, 8 figures. Previous name: Shadow-FT: Tuning Instruct via Base

详情

展开后加载摘要…

URL PDF HTML 收藏