arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3281 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3281 篇

2406.08697 2025-07-08 stat.ML cs.LG math.OC stat.ME 57%

Structured Difference-of-Q via Orthogonal Learning

Defu Cao, Angela Zhou

机构 * Department of Computer Science, University of Southern California(计算机科学系,南加州大学) Department of Data Sciences and Operations and Computer Science(数据科学与运营与计算机科学系,南加州大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.08271 2025-07-02 cs.LG 57%

LangTime: A Language-Guided Unified Model for Time Series Forecasting with Proximal Policy Optimization

Wenzhe Niu, Zongxia Xie, Yanru Sun, Wei He, Man Xu, Chao Hao

机构 * Tianjin University(天津大学)

专题命中 安全训练 :alignment(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23844 2025-07-01 cs.AI 57%

A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents

Hang Su, Jun Luo, Chang Liu, Xiao Yang, Yichi Zhang, Yinpeng Dong, Jun Zhu

机构 * Dept. of Comp. Sci. & Tech.(计算机科学与技术系) College of AI(人工智能学院) Institute for AI(人工智能研究院) BNRist Center(BNRist中心) THBI Lab(THBI实验室) Tsinghua-Bosch Joint Center for ML(清华大学-博世联合机器学习中心) Tsinghua University(清华大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.07460 2025-07-01 cs.CR cs.AI cs.SY eess.SY 57%

KnowSafe: Combined Knowledge and Data Driven Hazard Mitigation in Artificial Pancreas Systems

Xugui Zhou, Maxfield Kouzel, Chloe Smith, Homa Alemzadeh

机构 * Louisiana State University(路易斯安那州立大学) University of Virginia(弗吉尼亚大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 17 pages, 11 figures, 11 tables, to appear in the IEEE Transactions on Dependable and Secure Computing (TDSC'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14968 2025-06-30 cs.RO cs.AI 57%

FEAST: A Flexible Mealtime-Assistance System Towards In-the-Wild Personalization

Rajat Kumar Jenamani, Tom Silver, Ben Dodson, Shiqin Tong, Anthony Song, Yuting Yang, Ziang Liu, Benjamin Howe, Aimee Whitneck, Tapomayukh Bhattacharjee

机构 * Cornell University(康奈尔大学) University of Michigan(密歇根大学) Independent researcher(独立研究者)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments RSS 2025 - Best Paper Award

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16952 2025-06-23 cs.CY 57%

Modeling and Visualization Reasoning for Stakeholders in Education and Industry Integration Systems: Research on Structured Synthetic Dialogue Data Generation Based on NIST Standards

Wei Meng

专题命中 安全训练 :alignment(abstract);分类 cs.CY

Comments This paper presents an innovative and rigorous framework for stakeholder modelling in education-industry integration, combining NIST-compliant synthetic data generation with interpretable visual reasoning

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15522 2025-06-19 cs.CL 57%

Lessons from Training Grounded LLMs with Verifiable Rewards

Shang Hong Sim, Tej Deep Pala, Vernon Toh, Hai Leong Chieu, Amir Zadeh, Chuan Li, Navonil Majumder, Soujanya Poria

专题命中 安全训练 :trustworthy(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15028 2025-06-19 cs.CR cs.ET cs.LG 57%

Systems-Theoretic and Data-Driven Security Analysis in ML-enabled Medical Devices

Gargi Mitra, Mohammadreza Hallajiyan, Inji Kim, Athish Pranav Dharmalingam, Mohammed Elnawawy, Shahrear Iqbal, Karthik Pattabiraman, Homa Alemzadeh

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments 32 pages, 6 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.13978 2025-06-18 cs.CL 57%

AI shares emotion with humans across languages and cultures

Xiuwen Wu, Hao Wang, Zhiang Yan, Xiaohan Tang, Pengfei Xu, Wai-Ting Siok, Ping Li, Jia-Hong Gao, Bingjiang Lyu, Lang Qin

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12600 2025-06-17 cs.MA cs.AI cs.ET cs.GT cs.RO 57%

Trust-MARL: Trust-Based Multi-Agent Reinforcement Learning Framework for Cooperative On-Ramp Merging Control in Heterogeneous Traffic Flow

Jie Pan, Tianyi Wang, Christian Claudel, Jing Shi

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 34 pages, 7 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.16600 2025-06-17 cs.GT cs.AI cs.MA 57%

Convex Markov Games: A New Frontier for Multi-Agent Reinforcement Learning

Ian Gemp, Andreas Haupt, Luke Marris, Siqi Liu, Georgios Piliouras

机构 * College of Computing, MIT, Cambridge, MA, USA(麻省理工学院计算机学院) Google DeepMind, London, UK(谷歌深Mind)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Published at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12290 2025-06-17 cs.AI 57%

Ontology Enabled Hybrid Modeling and Simulation

John Beverley, Andreas Tolk

专题命中 安全训练 :alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05981 2025-06-11 cs.AI 57%

CrimeMind: Simulating Urban Crime with Multi-Modal LLM Agents

Qingbin Zeng, Ruotong Zhao, Jinzhu Mao, Haoyang Li, Fengli Xu, Yong Li

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) Department of Sociology, Hong Kong Baptist University(香港 Baptist 大学社会学系)

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments Typos corrected

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05599 2025-06-10 cs.CV cs.CL 57%

Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Yi Peng, Peiyu Wang, Xiaokun Wang, Yichen Wei, Jiangbo Pei, Weijie Qiu, Ai Jian, Yunzhuo Hao, Jiachun Pan, Tianyidan Xie, Li Ge, Rongxian Zhuang, Xuchen Song, Yang Liu, Yahui Zhou

机构 * Skywork AI(Skywork人工智能) Kunlun Inc.(昆仑公司)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05560 2025-06-09 cs.CL 57%

Improving LLMs with a knowledge from databases

Petr Máša

机构 * Department of Information and Knowledge Engineering(信息与知识工程系) Prague University of Economics and Business(布拉格经济与商业大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04828 2025-06-06 cs.AI 57%

Safe Planning and Policy Optimization via World Model Learning

Artem Latyshev, Gregory Gorbov, Aleksandr I. Panov

机构 * AIRI, Moscow, Russia(俄罗斯莫斯科AIRI) MIPT, Moscow, Russia(俄罗斯莫斯科MIPT) FRC CSC RAS, Moscow, Russia(俄罗斯莫斯科FRC CSC RAS)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03568 2025-06-06 cs.RO cs.AI 57%

Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving

Li Zeqiao, Wang Yijing, Wang Haoyu, Li Zheng, Li Peng, Zuo zhiqiang, Hu Chuan

机构 * Tianjin Key Laboratory of Intelligent Unmanned Swarm Technology and System, School of Electrical and Information Engineering, Tianjin University(天津智能无人群技术与系统重点实验室,电气与信息工程学院,天津大学) Key Laboratory of System Control and Information Processing, Ministry of Education of China(系统控制与信息处理重点实验室,中华人民共和国教育部) School of Mechanical Engineering, Shanghai Jiao Tong University(机械工程学院,上海交通大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.06053 2025-06-04 cs.LG math.OC stat.ML 57%

Safe-EF: Error Feedback for Nonsmooth Constrained Optimization

Rustem Islamov, Yarden As, Ilyas Fatkhullin

机构 * University of Basel(巴塞尔大学) ETH Zürich(苏黎世联邦理工学院) ETH AI Center(苏黎世联邦理工学院人工智能中心)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Journal ref International Conference on Machine Learning 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01333 2025-06-03 cs.CR cs.AI cs.ET 57%

ETDI: Mitigating Tool Squatting and Rug Pull Attacks in Model Context Protocol (MCP) by using OAuth-Enhanced Tool Definitions and Policy-Based Access Control

Manish Bhatt, Vineeth Sai Narajala, Idan Habler

机构 * Amazon Independent Researcher(亚马逊独立研究者) Project Kuiper Security (KPES)(项目Kuiper安全) Proactive Security OWASP, Amazon Web Services(主动安全OWASP,亚马逊网络服务) Intuit Adversarial AI Security reSearch (A2RS)(Intuit对抗AI安全reSearch)

专题命中 安全训练 :trustworthy(abstract);分类 cs.AI

Comments 11 Pages, 10 figures, Github links in introduction

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23556 2025-05-30 cs.CL 57%

Understanding Refusal in Language Models with Sparse Autoencoders

Wei Jie Yeo, Nirmalendu Prakash, Clement Neo, Roy Ka-Wei Lee, Erik Cambria, Ranjan Satapathy

机构 * Nanyang Technological University(南洋理工大学) Singapore University of Technology and Design(新加坡科技设计大学) Digital Trust Centre(数字信任中心) Institute of High Performance Computing (IHPC)(高性能计算研究所)

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09187 2025-05-30 cs.LG 57%

GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning

Zhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong, Qinbin Li, Han Xie, Jiawei Zhang, Zidi Xiong, Chulin Xie, Carl Yang, Dawn Song, Bo Li

机构 * University of Georgia(佐治亚大学) University of Chicago(芝加哥大学) UIUC(伊利诺伊大学香槟分校) University of Texas at Austin(德克萨斯大学奥斯汀分校) University of California, Berkeley(加州大学伯克利分校) Emory University(埃默里大学) Virtue AI

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Accepted to ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12486 2025-05-29 cs.CL 57%

EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning

Xiaoqian Liu, Ke Wang, Yongbin Li, Yuchuan Wu, Wentao Ma, Aobo Kong, Fei Huang, Jianbin Jiao, Junge Zhang

机构 * University of Chinese Academy of Sciences(中国科学院大学) Tongyi Lab(通义实验室) The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(复杂系统认知与决策智能重点实验室,中国科学院自动化研究所)

专题命中 安全训练 :alignment(abstract);分类 cs.CL

Comments ACL2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11057 2025-05-29 cs.RO cs.AI cs.SY eess.SY 57%

A Physics-Informed Machine Learning Framework for Safe and Optimal Control of Autonomous Systems

Manan Tayal, Aditya Singh, Shishir Kolathaya, Somil Bansal

专题命中 安全训练 :safety(abstract);分类 cs.AI

Comments 22 Pages, 12 Figures. First two authors have contributed equally. Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13831 2025-05-28 cs.AI 57%

TelePlanNet: An AI-Driven Framework for Efficient Telecom Network Planning

Zongyuan Deng, Yujie Cai, Qing Liu, Shiyao Mu, Bin Lyu, Zhen Yang

机构 * Nanjing University of Posts and Telecommunications(南京邮电大学) China Telecom Corporation Limited Jiangsu Branch(中国电信江苏分公司)

专题命中 安全训练 :alignment(abstract);分类 cs.AI

Comments 6 pages, 5 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19000 2025-05-27 cs.CL cs.CV 57%

VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization

Yunxin Li, Xinyu Chen, Zitao Li, Zhenyu Liu, Longyue Wang, Wenhan Luo, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology, Shenzhen, China(哈尔滨工业大学(深圳)) Alibaba International Group(阿里巴巴国际集团) Division of AMC and Department of ECE, HKUST(HKUST电子工程系与AMC division)

专题命中 安全训练 :DPO(abstract);分类 cs.CL

Comments 19 pages, 9 figures, Project Link: https://github.com/HITsz-TMG/VerIPO

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15793 2025-05-23 cs.RO cs.LG 57%

HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving

Zhiwen Chen, Bo Leng, Zhuoren Li, Hanming Deng, Guizhe Jin, Ran Yu, Huanxi Wen

机构 * Tongji University(同济大学) SenseTime Research(商汤科技研究院)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14112 2025-05-21 cs.CL cs.CR 57%

Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking

Tianle Gu, Zongqi Wang, Kexin Huang, Yuanqi Yao, Xiangliang Zhang, Yujiu Yang, Xiuying Chen

机构 * Mohamed bin Zayed University of Artificial Intelligence(莫扎德·本·扎耶德人工智能大学) Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院) Tsinghua University(清华大学) Fudan University(复旦大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) University of Notre Dame(Notre Dame 大学)

专题命中 安全训练 :safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11049 2025-05-19 cs.AI cs.CR 57%

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning

Yue Liu, Shengfang Zhai, Mingzhe Du, Yulin Chen, Tri Cao, Hongcheng Gao, Cheng Wang, Xinfeng Li, Kun Wang, Junfeng Fang, Jiaheng Zhang, Bryan Hooi

机构 * National University of Singapore(新加坡国立大学) Nanyang Technological University(南洋理工大学)

专题命中 安全训练 :safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09427 2025-05-16 cs.LG cs.RO 57%

SafePath: Conformal Prediction for Safe LLM-Based Autonomous Navigation

Achref Doula, Max Mühlhäuser, Alejandro Sanchez Guinea

机构 * Technical University of Darmstadt(德累斯顿技术大学) NTT Data(NTT数据)

专题命中 安全训练 :safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09734 2025-05-16 eess.SY cs.LG cs.RO cs.SY math.OC 57%

Risk-Aware Safe Reinforcement Learning for Control of Stochastic Linear Systems

Babak Esmaeili, Nariman Niknejad, Hamidreza Modares

机构 * Department of Mechanical Engineering, Michigan State University(机械工程系,密歇根州立大学)

专题命中 安全训练 :safety(abstract);分类 cs.LG

Comments Submitted to Asian Journal of Control

详情

展开后加载摘要…

URL PDF HTML 收藏