arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 30608 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 偏好对齐 3243 篇

2305.10425 2023-05-18 cs.CL cs.AI 62%

SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Yao Zhao, Rishabh Joshi, Tianqi Liu, Misha Khalman, Mohammad Saleh, Peter J. Liu

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.07036 2023-05-15 cs.LG cs.AI 62%

GFlowNets with Human Feedback

Yinchuan Li, Shuang Luo, Yunfeng Shao, Jianye Hao

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2304.11490 2023-04-27 cs.AI cs.CL 62%

Boosting Theory-of-Mind Performance in Large Language Models via Prompting

Shima Rahimi Moghaddam, Christopher J. Honey

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL、cs.AI

Comments 27 pages, 4 main figures, 2 supplementary figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.15906 2023-03-01 cs.AI cs.HC cs.LG 62%

Relative Behavioral Attributes: Filling the Gap between Symbolic Goal Specification and Reward Learning from Human Preferences

Lin Guan, Karthik Valmeekam, Subbarao Kambhampati

专题命中 偏好对齐 :RLHF(abstract);分类 cs.AI、cs.LG

Comments ICLR 2023 Camera Ready

详情

展开后加载摘要…

URL PDF HTML 收藏
2302.10291 2023-02-22 cs.CL cs.LG 62%

Can Large Language Models Change User Preference Adversarially?

Varshini Subhash

专题命中 偏好对齐 :red teaming(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2211.15006 2022-11-29 cs.LG cs.CL 62%

Fine-tuning language models to find agreement among humans with diverse preferences

Michiel A. Bakker, Martin J. Chadwick, Hannah R. Sheahan, Michael Henry Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matthew M. Botvinick, Christopher Summerfield

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2202.09662 2022-07-28 cs.CL cs.AI 62%

Reward Modeling for Mitigating Toxicity in Transformer-based Language Models

Farshid Faal, Ketra Schmitt, Jia Yuan Yu

专题命中 偏好对齐 :safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.01969 2021-07-06 cs.LG cs.AI 62%

The MineRL BASALT Competition on Learning from Human Feedback

Rohin Shah, Cody Wild, Steven H. Wang, Neel Alex, Brandon Houghton, William Guss, Sharada Mohanty, Anssi Kanervisto, Stephanie Milani, Nicholay Topin, Pieter Abbeel, Stuart Russell, Anca Dragan

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI、cs.LG

Comments NeurIPS 2021 Competition Track

详情

展开后加载摘要…

URL PDF HTML 收藏
2102.09388 2021-05-04 cs.IR cs.AI cs.LG 62%

ELIXIR: Learning from User Feedback on Explanations to Improve Recommender Models

Azin Ghazimatin, Soumajit Pramanik, Rishiraj Saha Roy, Gerhard Weikum

专题命中 偏好对齐 :trustworthy(abstract);分类 cs.AI、cs.LG

Comments WWW 2021, 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1912.13037 2020-01-01 cs.LG cs.AI stat.ML 62%

A New Framework for Query Efficient Active Imitation Learning

Daniel Hsu

专题命中 偏好对齐 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.17220 2025-10-30 cs.CL 61%

RLAIF-V: Open-Source AI Feedback Leads to Super GPT-4V Trustworthiness

Tianyu Yu, Haoye Zhang, Qiming Li, Qixin Xu, Yuan Yao, Da Chen, Xiaoman Lu, Ganqu Cui, Yunkai Dang, Taiwen He, Xiaocheng Feng, Jun Song, Bo Zheng, Zhiyuan Liu, Tat-Seng Chua, Maosong Sun

机构 * Tsinghua University(清华大学) Shanghai Qi Zhi Institute(上海启智研究院) Harbin Institute of Technology(哈尔滨工业大学) Taobao & Tmall Group of Alibaba(阿里巴巴淘宝与天猫集团) Peng Cheng Laboratory(鹏城实验室) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL;RLHF(comments)

Comments Project Website: https://github.com/RLHF-V/RLAIF-V

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13621 2025-03-19 cs.AI 61%

Superalignment with Dynamic Human Values

Florian Mai, David Kaczér, Nicholas Kluge Corrêa, Lucie Flek

专题命中 偏好对齐 :alignment(abstract,comments);分类 cs.AI

Comments Published at the ICLR 2025 Workshop on Bidirectional Human-AI Alignment (BiAlign)

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18248 2024-07-26 cs.CL 61%

Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning

Tianduo Wang, Shichen Li, Wei Lu

专题命中 偏好对齐 :DPO(abstract,comments);分类 cs.CL

Comments ACL 2024. Code and data are available at https://github.com/TianduoWang/DPO-ST

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.16839 2024-02-07 cs.CV cs.CL 61%

Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Zhiyuan Zhao, Bin Wang, Linke Ouyang, Xiaoyi Dong, Jiaqi Wang, Conghui He

专题命中 偏好对齐 :DPO(abstract,comments);分类 cs.CL

Comments Project Website: https://opendatalab.github.io/HA-DPO, Code: https://github.com/opendatalab/HA-DPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05199 2023-11-30 cs.CL 61%

Loose lips sink ships: Mitigating Length Bias in Reinforcement Learning from Human Feedback

Wei Shen, Rui Zheng, Wenyu Zhan, Jun Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL;RLHF(comments)

Comments EMNLP 2023 findings, Length Bias in RLHF, Mitigate bias in reward modeling

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.18770 2026-08-20 cs.LG cs.RO 新提交 57%

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

欲行远,需同行:多样化偏好引出奖励优化的课程设置

Taehyung Kim, Jongeun Choi

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

AI总结 该研究针对AI对齐中服务不足用户的问题,提出CurriPO方法,通过构建树状课程设置适配多样化用户目标,在个性化连续控制任务上使群体满意度达最强基线的1.2-2.1倍且缩短训练时间。

Comments 14 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14473 2026-08-20 cs.CL 版本更新 57%

AI Can Learn Scientific Taste

AI 可以学习科学品味

Jingqi Tong, Mingzhe Li, Hangcheng Li, Yongzhuo Yang, Yurong Mou, Weijie Ma, Hongji Chen, Xiaoran Liu, Qinyuan Cheng, Ming Zhang, Qiguang Chen, Weifeng Ge, Qipeng Guo, Tianlei Ying, Tianxiang Sun, Yining Zheng, Zhiheng Xi, Xinchi Chen, Jun Zhao, Ning Ding, Xuanjing Huang, Yu-Gang Jiang, Xipeng Qiu

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

AI总结 本文提出RLCF框架,通过社区反馈学习科学品味,使AI能提出高潜力研究想法,实验显示其优于现有模型并具备泛化能力。

Comments 47 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17795 2026-08-19 cs.CL 新提交 57%

TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification

TraceSQL:无参考文本到SQL验证的可追踪答案性估计

Neelesh Kumar Shukla, Debasmita Panda, Srutanik Bhaduri, Aditya Banerjee, Viji Krishnamurthy

机构 * Oracle Corporation(甲骨文公司)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

AI总结 针对无参考文本到SQL验证的可追踪性不足问题,提出基于67种诊断特征的轻量模型TraceSQL,在BIRD数据集上优于基线模型,可提供可追溯的预测证据。

Comments 9 pages main paper with 6 pages supplementary material

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02637 2026-08-18 cs.CL 版本更新 57%

Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training

通过角色扮演训练LLM:探索AI素养对说服力的影响

Qihui Fan, Min Ge, Chenyan Jia, Weiyan Shi

机构 * Institute for Clarity in Documentation(文档清晰研究所) Inria Paris-Rocquencourt(法国国家信息与自动化研究所巴黎-罗康库尔中心) Rajiv Gandhi University(拉吉夫·甘地大学) Tsinghua University(清华大学) Palmer Research Laboratories(帕默研究实验室)

专题命中 偏好对齐 :RLHF(abstract);分类 cs.CL

AI总结 本文通过角色扮演训练LLMimic,提升用户AI素养,减少AI说服力效果,并增强诚实与社会责任感。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15035 2026-08-18 cs.AI cs.HC 版本更新 57%

Calibrated Generative AI as Meta-Reviewer: A Systemic Functional Linguistics Discourse Analysis of Reviews of Peer Reviews

作为元评审员的校准生成式AI:对同行评审的评审的系统功能语言学话语分析

Gabriela C. Zapata, Bill Cope, Mary Kalantzis, Duane Searsmith

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

AI总结 本研究以美国公立大学研究生在线课程为对象,基于系统功能语言学分析120份元评审,发现校准生成式AI可近似有效人类反馈特征,能提升学习者对同行评审的参与度。

Comments 39 pages, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14522 2026-08-17 cs.AI 新提交 57%

Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers

参与式道德人工智能并非中立:开发者的隐形之手

Taenyun Kim, Edyta Bogucka, Daniele Quercia

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

AI总结 该研究发现,道德AI偏好征集流程中的特征范围、投票者构成、问题措辞三项关键选择会显著影响AI决策,仅靠投票汇总无法实现公平透明的AI,需对流程各阶段审计披露。

Comments 35 pages, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11604 2026-08-13 cs.AI 新提交 57%

Learning from Online User Feedback for Shopping Agents

从在线用户反馈中学习购物智能体

Haobo Zhang, Kelong Mao, Sulong Xu, Simiu Gu, Zhicheng Dou

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

AI总结 提出LOFA框架,结合强化学习与反馈感知的策略内蒸馏,利用真实在线用户反馈改进购物智能体,在电商日志实验中提升了推荐质量、回复有用性及用户满意度对齐效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11354 2026-08-13 cs.AI 新提交 57%

Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

用于内容推荐的反向心智理论建模:从网页浏览到动态智能界面

Mengyu Chen, Feiyu Lu, Chun-Fu Chen, Lucas Vinh Tran, Jay Katukuri

机构 * JPMorganChase(摩根大通)

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

AI总结 本研究提出反向心智理论(IToM)流程,从用户交互反向推理推断信念与偏好,在OPeRA数据集上验证其画像推断效果,并通过VisionOS应用展示跨模态迁移能力,提升内容推荐的用户理解精度。

Comments Preprint for conference full paper at 20th ACM Conference on Recommender Systems (RecSys '26), Minneapolis, MN, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08326 2026-08-12 cs.AI 版本更新 57%

StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning

StructReward:用于自修正多模态推理的高效结构化过程奖励

Yifan Li, Ruxin Sun, Tongzhou Zhao

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

AI总结 本研究提出StructReward框架,通过结构化步骤级奖励对齐结合GRPO目标等方式,在无额外验证器的情况下降低多模态强化学习开销,实现自修正多模态推理。

Comments 9 pages, 3 figures. Yifan Li and Ruxin Sun contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08058 2026-08-11 cs.CY 新提交 57%

Representational Equality in Cross-country Value Simulation: A Systematic Analysis of Large Language Models

跨国价值观模拟中的表征平等:对大型语言模型的系统分析

Xiaowen Jian, Xinyi Mou, Daisong Gong, Chen Qian, Huimin Chen, Maosong Sun

专题命中 偏好对齐 :alignment(abstract);分类 cs.CY

AI总结 本研究分析59国的LLM跨国价值观模拟,发现富裕技术先进国家人群模拟更准确,且两种干预路径的准确率提升未必带来表征平等,为构建更具包容性的LLM模拟提供指导。

Comments Accepted for publication in Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07557 2026-08-11 cs.RO cs.AI cs.CV 新提交 57%

AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization

AeroDPO:释放具备高保真感知与自动偏好优化的轻量级无人机导航能力

Peng Xu, Chengcheng Wang, Shaohua Wan

专题命中 偏好对齐 :DPO(abstract_cn);分类 cs.AI

AI总结 本文提出AeroDPO,通过高保真感知与自动偏好优化,使20亿参数轻量级无人机导航模型达到70亿参数模型的成功率,同时降低分布外场景碰撞率,成为自主空中智能体新SOTA。

Comments 7 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07509 2026-08-11 cs.HC cs.AI 新提交 57%

PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering

PIVOT:用于教学导师引导的基于偏好的干预向量

Fares Fawzi, Jiaxu Zhao, Tanya Nazaretsky, Tanja Käser

专题命中 偏好对齐 :alignment(abstract);分类 cs.AI

AI总结 研究针对LLM对话辅导缺乏教学策略推理时控制的问题,提出PIVOT框架,通过偏好干预向量实现导师动作控制,在用户研究中获多数教师认可。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00754 2026-08-11 cs.SE cs.LG 版本更新 57%

Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring

Themis: 训练鲁棒的多语言代码奖励模型以实现灵活的多标准评分

Indraneil Paul, Goran Glavaš, Iryna Gurevych

机构 * UKP Lab, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE(UKP实验室,德累斯顿理工大学及应用网络安全国家研究中心ATHENE)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

AI总结 本文提出Themis-CodeRewardBench基准,评估多语言多标准代码奖励模型,训练了Themis-RM系列模型,展示了其在多语言迁移和多标准训练中的优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00997 2026-08-11 cs.CL 版本更新 57%

Uncertainty-Aware Variational Reward Factorization via Probabilistic Preference Bases for LLM Personalization

基于概率偏好基的不确定性感知变分奖励分解用于大语言模型个性化

Gyuseok Lee, Wonbin Kweon, Zhenrui Yue, SeongKu Kang, Jiawei Han, Dong Wang

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Korea University(高丽大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.CL

AI总结 本文提出变分奖励分解框架,通过概率偏好基和变分分布实现用户偏好建模,提升LLM个性化效果。

Comments COLM'26

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07419 2026-08-10 cs.LG 新提交 57%

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration

超越事后温度缩放:用于大语言模型校准的双层优化方法

Ruochen Jin, Zhanliang Wang, Zongyu Dai, Jiancong Xiao, Bojian Hou

机构 * Dartmouth College(达特茅斯学院) University of Pennsylvania(宾夕法尼亚大学) National University of Singapore(新加坡国立大学)

专题命中 偏好对齐 :alignment(abstract);分类 cs.LG

AI总结 针对LLM偏好对齐导致的过度自信与校准问题,提出基于双层优化的校准方法,通过最大化预测分布熵实现,在多项选择与开放式问答任务中提升了校准效果与域外泛化能力。

Comments Third Conference on Language Modeling (COLM 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏