arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3243 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3243 篇

2509.05818 2025-10-28 cs.AI 57%

Chatbot To Help Patients Understand Their Health

Won Seok Jang, Hieu Tran, Manav Mistry, SaiKiran Gandluri, Yifan Zhang, Sharmin Sultana, Sunjae Kown, Yuan Zhang, Zonghai Yao, Hong Yu

机构 * Center for Healthcare Organization and Implementation Research, VA Bedford Health Care(VA贝德福德医疗中心健康组织与实施研究中心) Miner School of Computer and Information Sciences, University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校计算机与信息科学学院) Manning College of Information and Computer Sciences, University of Massachusetts Amherst(马萨诸塞大学阿姆赫斯特分校信息与计算机科学学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments Accepted in EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11336 2025-10-24 cs.CL 57%

XtraGPT: Context-Aware and Controllable Academic Paper Revision

Nuo Chen, Andre Lin HuiKai, Jiaying Wu, Junyi Hou, Zining Zhang, Qian Wang, Xidong Wang, Bingsheng He

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Preprint. The model report is available at https://arxiv.org/abs/2505.11336v1

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.20020 2025-10-24 cs.GT cs.AI 57%

Optimized Distortion in Linear Social Choice

Luise Ge, Gregory Kehne, Yevgeniy Vorobeychik

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18895 2025-10-23 cs.SE cs.AI cs.HC 57%

CosmoCore Affective Dream-Replay Reinforcement Learning for Code Generation

Santhosh Kumar Ravindran

机构 * Microsoft Corporation(微软公司)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18433 2025-10-22 cs.CV cs.AI cs.IR 57%

ImageGem: In-the-wild Generative Image Interaction Dataset for Generative Model Personalization

Yuanhe Guo, Linxi Xie, Zhuoran Chen, Kangrui Yu, Ryan Po, Guandao Yang, Gordon Wetztein, Hongyi Wen

机构 * NYU(纽约大学) Stanford(斯坦福大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12943 2025-10-21 cs.CL 57%

The Curious Case of Curiosity across Human Cultures and LLMs

Angana Borah, Zhijing Jin, Rada Mihalcea

机构 * University of Michigan - Ann Arbor(密歇根大学安阿伯分校) University of Toronto(多伦多大学) Vector Institute(向量研究所) MPI for Intelligent Systems, Tubingen, Germany(图宾根德国智能系统研究所)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Preprint (Paper under review)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.15993 2025-10-21 q-fin.PM cs.LG q-fin.ST 57%

Aligning Language Models with Investor and Market Behavior for Financial Recommendations

Fernando Spadea, Oshani Seneviratne

机构 * Rensselaer Polytechnic Institute(拉特兰理工学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11620 2025-10-20 cs.CL 57%

Enhancing Long Chain-of-Thought Reasoning through Multi-Path Plan Aggregation

Siheng Xiong, Ali Payani, Faramarz Fekri

机构 * Georgia Institute of Technology(佐治亚理工学院) Cisco Research(思科研究)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14526 2025-10-17 cs.CV cs.LG 57%

Noise Projection: Closing the Prompt-Agnostic Gap Behind Text-to-Image Misalignment in Diffusion Models

Yunze Tong, Didi Zhu, Zijing Hu, Jinluan Yang, Ziyu Zhao

机构 * Zhejiang University(浙江大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

Comments Appendix will be appended soon

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14200 2025-10-17 cs.CL 57%

RLSR: Reinforcement Learning with Supervised Reward Outperforms SFT in Instruction Following

Zhichao Wang, Andy Wong, Ruslan Belkin

机构 * Inflection AI

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13501 2025-10-16 cs.AI 57%

Confidence as a Reward: Transforming LLMs into Reward Models

He Du, Bowen Li, Chengxing Xie, Chang Gao, Kai Chen, Dacheng Tao

机构 * Fudan University(复旦大学) Shanghai AI Laboratory(上海人工智能实验室) Xidian University(西安电子科技大学) The Chinese University of Hong Kong(香港中文大学) Nanyang Technological University(南洋理工大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23144 2025-10-16 cs.AI cond-mat.stat-mech cs.MA nlin.AO physics.soc-ph 57%

Coordination Requires Simplification: Thermodynamic Bounds on Multi-Objective Compromise in Natural and Artificial Intelligence

Atma Anand

机构 * Department of Physics and Astronomy, University of Rochester(物理与天文学系,罗切斯特大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

Comments 15 pages, 1 figure, 9 pages supplementary material, submitted to Journal of Physics: Complexity

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10013 2025-10-16 cs.CV cs.CL 57%

Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect

Tom Kouwenhoven, Kiana Shahrasbi, Tessa Verhoef

机构 * Leiden Institute of Advanced Computer Science(莱顿先进计算机科学研究所) Leiden University(莱顿大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments Presented at the Thirty-Ninth Annual Conference on Neural Information Processing Systems (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.12014 2025-10-15 cs.IR cs.LG 57%

Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval

Eric He, Akash Gupta, Adian Liusie, Vatsal Raina, Piotr Molenda, Shirom Chabra, Vyas Raina

机构 * University of Cambridge(剑桥大学) Apta

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26041 2025-10-15 cs.CL 57%

Unspoken Hints: Accuracy Without Acknowledgement in LLM Reasoning

Arash Marioriyad, Shaygan Adim, Nima Alighardashi, Mahdieh Soleymani Banghshah, Mohammad Hossein Rohban

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments 5 Pages, 4 Figures, 4 Tables

Journal ref 39th Conference on Neural Information Processing Systems, 2025, Workshop: Reliable ML from Unreliable Data

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19223 2025-10-14 cs.LG 57%

LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models

Fengqi Zhu, Rongzhen Wang, Shen Nie, Xiaolu Zhang, Chunwei Wu, Jun Hu, Jun Zhou, Jianfei Chen, Yankai Lin, Ji-Rong Wen, Chongxuan Li

机构 * Gaoling School of AI, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心) Tsinghua University(清华大学) Ant Group(蚂蚁集团)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09103 2025-10-13 cs.LG 57%

AdaPM: a Partial Momentum Algorithm for LLM Training

Yimu Zhang, Yuanshi Liu, Cong Fang

机构 * State Key Lab of General AI, School of Intelligence Science and Technology, Peking University(通用人工智能国家重点实验室,智能科学与技术学院,北京大学) Institute for Artificial Intelligence, Peking University(人工智能研究院,北京大学)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13726 2025-10-13 cs.CL 57%

RPO: Retrieval Preference Optimization for Robust Retrieval-Augmented Generation

Shi-Qi Yan, Quan Liu, Zhen-Hua Ling

机构 * National Engineering Research Center of Speech and Language Information Processing, University of Science and Technology of China(语音与语言信息处理国家工程研究中心,中国科学技术大学) State Key Laboratory of Cognitive Intelligence, iFLYTEK Research(认知智能国家重点实验室,iFLYTEK研究院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.03438 2025-10-10 cs.AI 57%

BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving

Ran Xin, Chenguang Xi, Jie Yang, Feng Chen, Hang Wu, Xia Xiao, Yifan Sun, Shen Zheng, Kai Shen

机构 * ByteDance Seed(字节跳动种子) ByteDance Applied Machine Learning(字节跳动应用机器学习) Stanford University(斯坦福大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06870 2025-10-10 cs.CL 57%

$λ$-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences

Yining Wang, Jinman Zhao, Chuangxin Zhao, Shuhao Guan, Gerald Penn, Shinan Liu

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06627 2025-10-09 cs.LG 57%

POME: Post Optimization Model Edit via Muon-style Projection

Yong Liu, Di Fu, Yang Luo, Zirui Zhu, Minhao Cheng, Cho-Jui Hsieh, Yang You

机构 * National University of Singapore(新加坡国立大学) Penn State University(宾夕法尼亚州立大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25361 2025-10-07 cs.AI 57%

Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling

Xiaoyu Liu, Di Liang, Chang Dai, Hongyu Shan, Peiyang Liu, Yonghao Liu, Muling Wu, Yuntao Li, Xianjie Wu, LI Miao, Jiangrong Shen, Minlong Peng

机构 * Northeastern University, Boston(东北大学,波士顿) Independent Developer(独立开发者) Peiking University(北京大学) Jilin University(吉林大学) Beihang University(北航) Xi’an Jiaotong University(西安交通大学) Baidu Inc(百度公司)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22558 2025-10-03 cs.AI 57%

StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models

Chenyu Zhou, Tianyi Xu, Jianghao Lin, Dongdong Ge

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.00263 2025-10-02 cs.CL 57%

Judging with Confidence: Calibrating Autoraters to Preference Distributions

Zhuohang Li, Xiaowei Li, Chengyu Huang, Guowang Li, Katayoon Goshvadi, Bo Dai, Dale Schuurmans, Paul Zhou, Hamid Palangi, Yiwen Song, Palash Goyal, Murat Kantarcioglu, Bradley A. Malin, Yuan Xue

机构 * Google(谷歌) Vanderbilt University(范德比大学) Cornell University(康奈尔大学) DeepMind(深Mind) University of Alberta(阿尔伯塔大学) Virginia Tech(弗吉尼亚理工学院) Scale AI

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24361 2025-10-01 cs.CV cs.AI cs.HC 57%

UI-UG: A Unified MLLM for UI Understanding and Generation

Hao Yang, Weijie Qiu, Ru Zhang, Zhou Fang, Ruichao Mao, Xiaoyu Lin, Maji Huang, Zhaosong Huang, Teng Guo, Shuoyang Liu, Hai Rao

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23285 2025-10-01 cs.AI 57%

Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning

Yifei Chen, Guanting Dong, Zhicheng Dou

机构 * Renmin University of China(中国人民大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24781 2025-09-30 cs.CL 57%

SeaPO: Strategic Error Amplification for Robust Preference Optimization of Large Language Models

Jun Rao, Yunjie Liao, Xuebo Liu, Zepeng Lin, Lian Lian, Dong Jin, Shengjun Cheng, Jun Yu, Min Zhang

机构 * Institute of Computing and Intelligence, Harbin Institute of Technology, Shenzhen(计算与智能学院,哈尔滨工业大学,深圳) Huawei Cloud Computing Technologies Co., Ltd.(华为云计算技术有限公司) School of Intelligence Science and Engineering, Harbin Institute of Technology, Shenzhen(智能科学与工程学院,哈尔滨工业大学,深圳)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24342 2025-09-30 cs.AI 57%

Fin-Ally: Pioneering the Development of an Advanced, Commonsense-Embedded Conversational AI for Money Matters

Sarmistha Das, Priya Mathur, Ishani Sharma, Sriparna Saha, Kitsuchart Pasupa, Alka Maurya

机构 * Department of Computer Science and Engineering, Indian Institute of Technology Patna, India(计算机科学与工程系,印度理工学院帕纳布分校) School of Information Technology, King Mongkut's Institute of Technology Ladkrabang, Thailand(信息科技学院,拉差班国王技术学院) CRISIL Limited, India(CRISIL有限公司)

专题命中 偏好对齐 :DPO(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23140 2025-09-30 cs.CL 57%

Tagging the Thought: Unlocking Personalization Reasoning via Reinforcement Learning

Song Jin, Juntian Zhang, Yong Liu, Xun Zhang, Yufei Zhang, Fei Jiang, Guojun Yin, Wei Lin, Rui Yan

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) Meituan(美团) Wuhan University(武汉大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22713 2025-09-30 cs.CL 57%

RAR$^2$: Retrieval-Augmented Medical Reasoning via Thought-Driven Retrieval

Kaishuai Xu, Wenjun Hou, Yi Cheng, Wenjie Li

机构 * Department of Computing, The Hong Kong Polytechnic University(计算机系,香港理工大学) Research Institute of Trustworthy Autonomous Systems and Department of Computer Science and Engineering, Southern University of Science and Technology(可信自主系统研究院和计算机科学与工程系,南方科技大学)

专题命中 偏好对齐 :DPO(abstract);分类 cs.CL

Comments Accepted by EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏