arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9248 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9248 篇

2411.00069 2024-11-04 cs.CR cs.AI 83%

Meta-Sealing: A Revolutionizing Integrity Assurance Protocol for Transparent, Tamper-Proof, and Trustworthy AI System

Mahesh Vaijainthymala Krishnamoorthy

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 24 pages, 3 figures and 10 Code blocks, to be presented in the conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.20087 2024-11-01 cs.LG cs.AI cs.CL cs.CY cs.HC 83%

ProgressGym: Alignment with a Millennium of Moral Progress

Tianyi Qiu, Yang Zhang, Xuchuan Huang, Jasmine Xinze Li, Jiaming Ji, Yaodong Yang

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments NeurIPS 2024 Track on Datasets and Benchmarks (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.04418 2024-10-28 cs.LG cs.CR cs.NI eess.SP 83%

Trustworthy Federated Learning via Blockchain

Zhanpeng Yang, Yuanming Shi, Yong Zhou, Zixin Wang, Kai Yang

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.LG

Comments This work has been submitted to the IEEE Internet of Things Journal for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.19521 2024-10-01 cs.CR cs.LG 83%

GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks

Rongchang Li, Minjie Chen, Chang Hu, Han Chen, Wenpeng Xing, Meng Han

专题命中 安全评测 :prompt injection(title,abstract);safety(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.04313 2024-07-15 cs.LG cs.AI cs.CL cs.CV cs.CY 83%

Improving Alignment and Robustness with Circuit Breakers

Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, Dan Hendrycks

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Code and models are available at https://github.com/GraySwanAI/circuit-breakers

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.03575 2024-06-21 cs.AI cs.HC 83%

Toward Human-AI Alignment in Large-Scale Multi-Player Games

Sugandha Sharma, Guy Davidson, Khimya Khetarpal, Anssi Kanervisto, Udit Arora, Katja Hofmann, Ida Momennejad

专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19524 2024-05-31 cs.CR cs.AI 83%

AI Risk Management Should Incorporate Both Safety and Security

Xiangyu Qi, Yangsibo Huang, Yi Zeng, Edoardo Debenedetti, Jonas Geiping, Luxi He, Kaixuan Huang, Udari Madhushani, Vikash Sehwag, Weijia Shi, Boyi Wei, Tinghao Xie, Danqi Chen, Pin-Yu Chen, Jeffrey Ding, Ruoxi Jia, Jiaqi Ma, Arvind Narayanan, Weijie J Su, Mengdi Wang, Chaowei Xiao, Bo Li, Dawn Song, Peter Henderson, Prateek Mittal

专题命中 安全评测 :safety(title,abstract);AI safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14580 2024-02-23 cs.AI cs.SY eess.SY 83%

Savvy: Trustworthy Autonomous Vehicles Architecture

Ali Shoker, Rehana Yasmin, Paulo Esteves-Verissimo

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.05818 2023-10-10 cs.CL 83%

SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese

Liang Xu, Kangkang Zhao, Lei Zhu, Hang Xue

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract);分类 cs.CL

Comments 20 pages, 8 tables, 16 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.06421 2023-10-06 cs.AI cs.HC 83%

AI Alignment Dialogues: An Interactive Approach to AI Alignment in Support Agents

Pei-Yu Chen, Myrthe L. Tielman, Dirk K. J. Heylen, Catholijn M. Jonker, M. Birna van Riemsdijk

专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract);分类 cs.AI

Comments Withdraw because the content of the paper has been largely revised. The newest version is very different than the submitted one

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16144 2023-08-09 cs.RO cs.AI 83%

Towards trustworthy multi-modal motion prediction: Holistic evaluation and interpretability of outputs

Sandra Carrasco Limeros, Sylwia Majchrowska, Joakim Johnander, Christoffer Petersson, Miguel Ángel Sotelo, David Fernández Llorca

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 16 pages, 7 figures, 6 tables

Journal ref CAAI Transactions on Intelligence Technology 24686557 (ISSN) 24682322 (eISSN) 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.09705 2023-07-20 cs.CL 83%

CValues: Measuring the Values of Chinese Large Language Models from Safety to Responsibility

Guohai Xu, Jiayi Liu, Ming Yan, Haotian Xu, Jinghui Si, Zhuoran Zhou, Peng Yi, Xing Gao, Jitao Sang, Rong Zhang, Ji Zhang, Chao Peng, Fei Huang, Jingren Zhou

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL

Comments Working in Process

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.10007 2023-03-20 cs.CY 83%

Information Governance as a Socio-Technical Process in the Development of Trustworthy Healthcare AI

Nigel Rees, Kelly Holding, Mark Sujan

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.CY

Journal ref Frontiers in Computer Science. 2023;5

详情

展开后加载摘要…

URL PDF HTML 收藏
2208.14451 2022-09-13 cs.CY 83%

Foreseeing the Impact of the Proposed AI Act on the Sustainability and Safety of Critical Infrastructures

Francesco Sovrano, Giulio Masetti

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CY

Comments 10 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2201.08580 2022-01-24 cs.IR cs.AI cs.DB 83%

Trustworthy Knowledge Graph Completion Based on Multi-sourced Noisy Data

Jiacheng Huang, Yao Zhao, Wei Hu, Zhen Ning, Qijin Chen, Xiaoxia Qiu, Chengfu Huo, Weijun Ren

专题命中 安全评测 :trustworthy(title,abstract);alignment(abstract);分类 cs.AI

Comments Accepted in the ACM Web Conference (WWW 2022)

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.14019 2021-10-28 cs.LG 83%

Reliable and Trustworthy Machine Learning for Health Using Dataset Shift Detection

Chunjong Park, Anas Awadalla, Tadayoshi Kohno, Shwetak Patel

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.LG

Comments Neu

详情

展开后加载摘要…

URL PDF HTML 收藏
2107.06641 2021-08-20 cs.AI 83%

Trustworthy AI: A Computational Perspective

Haochen Liu, Yiqi Wang, Wenqi Fan, Xiaorui Liu, Yaxin Li, Shaili Jain, Yunhao Liu, Anil K. Jain, Jiliang Tang

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 55 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2105.04408 2021-05-11 cs.RO cs.AI 83%

The Challenges and Opportunities of Human-Centered AI for Trustworthy Robots and Autonomous Systems

Hongmei He, John Gray, Angelo Cangelosi, Qinggang Meng, T. Martin McGinnity, Jörn Mehnen

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.05501 2020-09-14 stat.ML cs.LG 83%

Towards a More Reliable Interpretation of Machine Learning Outputs for Safety-Critical Systems using Feature Importance Fusion

Divish Rengasamy, Benjamin Rothwell, Grazziela Figueredo

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1904.01540 2019-04-03 cs.AI 83%

Augmented Utilitarianism for AGI Safety

Nadisha-Marie Aliman, Leon Kester

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.07744 2024-02-15 cs.AI cs.CL cs.LG 83%

Towards Unified Alignment Between Agents, Humans, and Environment

Zonghan Yang, An Liu, Zijun Liu, Kaiming Liu, Fangzhou Xiong, Yile Wang, Zeyuan Yang, Qingyuan Hu, Xinrui Chen, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li, Yang Liu

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Project webpage: https://agent-force.github.io/unified-alignment-for-agents.html

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12382 2026-04-28 cs.LG cs.AI cs.CR 82%

Exploring the Secondary Risks of Large Language Models

探索大型语言模型的二次风险

Jiawei Chen, Zhengwei Fang, Yu Tian, Jiawei Du, Chao Yu, Zhaoxia Yin, Hang Su

机构 * Shenzhen International Graduate School, Tsinghua University(深圳国际研究生院,清华大学) Beijing Zhongguancun Academy(北京中关村学院) Department of Computer Science and Technology, THBI Lab, Tsinghua University(计算机科学与技术系,清华大学THBI实验室) CFAR, A*STAR, Singapore(新加坡A*STAR CFAR)

专题命中 安全评测 :jailbreak(abstract,abstract_cn);alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出二次风险作为新型失败模式,通过SecLens框架系统评估,揭示了大型语言模型在良性交互中存在广泛且转移性强的有害行为,强调需加强安全机制。

Comments 18 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10796 2026-08-12 cs.CV 新提交 82%

E$^3$mo-Bench: A Scalable Benchmark for Multimodal Evoked and Expressed Emotion Understanding via Bayesian Pairwise Alignment

E³mo-Bench:基于贝叶斯成对对齐的可扩展多模态诱发与表达情感理解基准

Lancheng Gao, Ziheng Jia, Shengyan Li, Zixuan Xing, Jiarui Wang, Huiyu Duan, Xiongkuo Min

专题命中 安全评测 :alignment(title,abstract)

AI总结 该研究针对现有多模态情感理解基准的不足,推出E³mo-Bench基准,提出贝叶斯成对对齐方法与E³mo-Score智能体,验证了框架有效性并指出MLLMs在情感任务中的缺陷,为多模态情感智能发展指明方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28969 2026-08-12 cs.CV 版本更新 82%

SafeNexus: Discovering and Steering Modality-Universal Safety Neurons in MLLMs

SafeNexus:在多模态大语言模型(MLLMs)中发现与调控模态通用安全神经元

Jian Yu, Fei Shen, Cong Wang, Jian Wang, Lu Jin, Xiaoyu Du, Jinhui Tang, Tat-Seng Chua

专题命中 安全评测 :safety(title,abstract);alignment(abstract)

AI总结 SafeNexus是一种跨模态安全对齐框架,通过定位并调控模态通用安全神经元,提升多模态大语言模型在跨模态威胁下的安全性,且能保留模型效用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.08485 2026-08-11 cs.AI cs.CL cs.LG 新提交 82%

HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails

HoloAegis:冻结表示、拓扑推理:用于零样本大语言模型(LLM)护栏的最小参数安全流形

Tak Ho Alex Li, Kaijie Liu, Lik-Hang Lee, Kin Chung Ho, Ping Shum, Michael K. Ng

机构 * Hong Kong Baptist University(香港浸会大学) Guangdong Polytechnic Normal University(广东技术师范大学) Guangdong Institute of Digital Industry(广东数字产业研究院) The Hong Kong Polytechnic University(香港理工大学) The Education University of Hong Kong(香港教育大学) Southern University of Science and Technology(南方科技大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 HoloAegis是一种最小参数拓扑推理框架,通过冻结语义表示的纯几何推理实现零样本LLM安全护栏,在8个基准测试中达到最先进准确率,兼具低延迟、零冷启动数据和跨语言迁移能力。

Comments Preprint, August 2026. 10 tables, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.08044 2026-08-05 cs.LG cs.AI cs.CL 版本更新 82%

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

当行为安全评估失败时:表征层面的视角

Enyi Jiang, Anders Gjølbye, Yibo Jacky Zhang, Sanmi Koyejo

机构 * Stanford University(斯坦福大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Technical University of Denmark(丹麦技术大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出行为安全与干预鲁棒性之间的“审计差距”,通过构建解离模型和引入潜在脆弱性评分(LVS),证明行为安全指标不足以衡量表征层面的鲁棒性。

Comments Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29199 2026-08-03 cs.CR 新提交 82%

Alignment Is Local: A Paired Diagnostic for GUI Agents under User-Side Persuasion

对齐是局部的:用户侧说服下GUI智能体的配对诊断

Haoxin An, Yunpeng Song, Zihao Bai, Zhongmin Cai, Guojun Xiong, Chenhao Lin, Wentao Chen, Feng Wei, Chao Shen

专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract)

AI总结 本文通过配对诊断发现GUI智能体的提示级对齐是局部现象,一行防护栏可大幅降低单次ASR,但四轮升级链会使其防护效果下降,静态单轮ASR高估了实际部署的鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.26518 2026-07-31 cs.CV 版本更新 82%

EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

EgoSafe:用于视觉安全理解的第一人称移动采集基准

Yuyun Chen, Tianao Li, TianQuan Feng, Cen Chen, Huiping Zhuang, Hao Peng, Ziqian Zeng

机构 * South China University of Technology(华南理工大学) Beihang University(北京航空航天大学)

专题命中 安全评测 :safety(title,abstract);alignment(abstract)

AI总结 该研究推出第一人称移动采集的视觉安全理解基准 EgoSafe-Bench,含12000个样本,用分层推理评估(HRE)协议测试,发现现有大视觉语言模型存在感知-推理解耦问题,为逻辑鲁棒视频理解系统提供评估框架

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21134 2026-07-28 cs.CL cs.CY cs.LG 版本更新 82%

TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law

TRIDENT:评估金融、医学和法律领域大语言模型的安全性

Zheng Hui, Yijiang River Dong, Ehsan Shareghi, Nigel Collier

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.CY、cs.LG

AI总结 研究针对大语言模型在金融、医学和法律领域的安全评估问题,基于相关伦理准则定义安全原则并引入Trident-Bench基准进行评估,揭示了不同模型的安全差距,为LLM安全研究提供了系统资源和研究基础。

Comments COLM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20485 2026-07-24 cs.AI cs.CL cs.LG 新提交 82%

Expectation Alignment of Language Models for Real-World User Expectations

语言模型与现实世界用户期望的期望对齐

Miaomiao Li, Yang Wang, Bin Liang, Shudong Liu, Zhiwei Zhang, Kam-Fai Wong

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究LLMs是否满足用户期望,提出提取期望程序并引入ExpectBench基准,分析发现LLMs存在问题,进而提出LENS框架,可让模型内化期望以生成更契合的响应,凸显明确建模用户期望对实现现实人机对齐的重要性。

Comments Accepted by ICML 2026

详情

展开后加载摘要…

URL PDF HTML 收藏