arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2302.06613 2023-02-15 eess.IV 78%

Explainable artificial intelligence toward usable and trustworthy computer-aided early diagnosis of multiple sclerosis from Optical Coherence Tomography

Monica Hernandez, Ubaldo Ramon-Julvez, Elisa Vilades, Beatriz Cordon, Elvira Mayordomo, Elena Garcia-Martin

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2301.11136 2023-01-27 stat.ML 78%

Conformal Prediction for Trustworthy Detection of Railway Signals

Léo Andéol, Thomas Fel, Florence De Grancey, Luca Mossina

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.01599 2022-11-11 cs.HC 78%

A Gaze Data-based Comparative Study to Build a Trustworthy Human-AI Collaboration in Crash Anticipation

Yu Li, Muhammad Monjurul Karim, Ruwen Qin

专题命中 安全评测 :trustworthy(title);safety(abstract)

Comments Revised and submitted to International Conference on Transportation and Development (ICTD 2023) on Nov 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.17291 2022-11-01 cs.IT math.IT 78%

SIX-Trust for 6G: Towards a Secure and Trustworthy 6G Network

Yiying Wang, Xin Kang, Tieyan Li, Haiguang Wang, Cheng-Kang Chu, Zhongding Lei

专题命中 安全评测 :trustworthy(title,abstract)

Comments 7 pages, 3 figures, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2206.09682 2022-11-01 cs.RO 78%

SafeBench: A Benchmarking Platform for Safety Evaluation of Autonomous Vehicles

Chejian Xu, Wenhao Ding, Weijie Lyu, Zuxin Liu, Shuai Wang, Yihan He, Hanjiang Hu, Ding Zhao, Bo Li

专题命中 安全评测 :safety(title,abstract)

Comments Published as a conference paper at NeurIPS 2022 (Track on Datasets and Benchmarks)

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.07198 2022-10-14 cs.CL cs.AI cs.LG 78%

Towards Trustworthy Automatic Diagnosis Systems by Emulating Doctors' Reasoning with Deep Reinforcement Learning

Arsene Fansi Tchango, Rishab Goel, Julien Martel, Zhi Wen, Gaetan Marceau Caron, Joumana Ghosn

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI、cs.LG

Comments Camera ready. NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.02403 2022-09-07 cs.HC 78%

Guidelines to Develop Trustworthy Conversational Agents for Children

Marina Escobar-Planas, Emilia Gómez, Carlos-D Martínez-Hinarejos

专题命中 安全评测 :trustworthy(title,abstract)

Comments 19 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.03083 2022-07-26 cs.CV cs.RO 78%

Construction Site Safety Monitoring and Excavator Activity Analysis System

Sibo Zhang, Liangjun Zhang

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2205.15757 2022-06-01 cs.DC cs.CR 78%

Dropbear: Machine Learning Marketplaces made Trustworthy with Byzantine Model Agreement

Alex Shamis, Peter Pietzuch, Antoine Delignat-Lavaud, Andrew Paverd, Manuel Costa

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.00688 2022-04-05 cs.HC 78%

Designing AI for Online-to-Offline Safety Risks with Young Women: The Context of Social Matching

Douglas Zytko, Hanan Aljasim

专题命中 安全评测 :safety(title,abstract)

Comments Accepted to the ACM CSCW 2021 workshop "MOSafely: Building an Open-Source HCAI Community to Make the Internet a Safer Place For Youth"

详情

展开后加载摘要…

URL PDF HTML 收藏
2110.04983 2021-10-14 eess.SY cs.SY 78%

Understanding the Safety Requirements for Learning-based Power Systems Operations

Yize Chen, Daniel Arnold, Yuanyuan Shi, Sean Peisert

专题命中 安全评测 :safety(title,abstract)

Comments In submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.00964 2021-03-02 cs.SE 78%

Practices for Engineering Trustworthy Machine Learning Applications

Alex Serban, Koen van der Blom, Holger Hoos, Joost Visser

专题命中 安全评测 :trustworthy(title,abstract)

Comments Published at WAIN'21 - 1st Workshop on AI Engineering - Software Engineering for AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2009.00802 2020-09-03 cs.LG cs.AI cs.CV cs.CY cs.SE stat.ML 78%

Estimating the Brittleness of AI: Safety Integrity Levels and the Need for Testing Out-Of-Distribution Performance

Andrew J. Lohn

专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.02893 2020-08-31 cs.SE 78%

Making Fair ML Software using Trustworthy Explanation

Joymallya Chakraborty, Kewen Peng, Tim Menzies

专题命中 安全评测 :trustworthy(title,abstract)

Comments New Ideas and Emerging Results (NIER) track; The 35th IEEE/ACM International Conference on Automated Software Engineering; Melbourne, Australia

Journal ref ASE 2020: The 35th IEEE/ACM International Conference on Automated Software Engineering, Melbourne, Australia, Mon 21 - Fri 25 September 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
2007.10243 2020-07-21 cs.CV 78%

Inter-Homines: Distance-Based Risk Estimation for Human Safety

Matteo Fabbri, Fabio Lanzi, Riccardo Gasparini, Simone Calderara, Lorenzo Baraldi, Rita Cucchiara

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2005.09999 2020-05-21 cs.RO 78%

Model Predictive Instantaneous Safety Metric for Evaluation of Automated Driving Systems

Bowen Weng, Sughosh J. Rao, Eeshan Deosthale, Scott Schnelle, Frank Barickman

专题命中 安全评测 :safety(title,abstract)

Comments Accepted at IEEE Intelligent Vehicles Symposium (IV), 2020

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.04374 2019-11-13 q-bio.GN q-bio.QM 78%

Combining human cell line transcriptome analysis and Bayesian inference to build trustworthy machine learning models for prediction of animal toxicity in drug development

Laura-Jayne Gardiner, Anna Paola Carrieri, Jenny Wilshaw, Stephen Checkley, Edward O Pyzer-Knapp, Ritesh Krishna

专题命中 安全评测 :trustworthy(title,abstract)

Comments Machine Learning for Health (ML4H) at NeurIPS 2019 - Extended Abstract

详情

展开后加载摘要…

URL PDF HTML 收藏
1911.00608 2019-11-05 cs.FL cs.LO cs.SC cs.SY eess.SY 78%

Multi-Agent Safety Verification using Symmetry Transformations

Hussein Sibai, Navid Mokhlesi, Chuchu Fan, Sayan Mitra

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
1805.06427 2018-09-11 cs.HC 78%

MOABB: Trustworthy algorithm benchmarking for BCIs

Vinay Jayaram, Alexandre Barachant

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00969 2025-06-03 cs.LG cs.AI 77%

Data Heterogeneity Modeling for Trustworthy Machine Learning

Jiashuo Liu, Peng Cui

机构 * Tsinghua University(清华大学)

专题命中 安全评测 :trustworthy(title,comments);分类 cs.AI、cs.LG

Comments Survey paper for tutorial "Data Heterogeneity Modeling for Trustworthy Machine Learning" in KDD'25

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.03166 2026-08-05 cs.AI 新提交 77%

Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation

基于多智能体评估的角色扮演语言智能体的对抗性压力测试

Saqib Shouqi, Abdullah Nazly, Januki Wanniarachchi, Ravisha De Alwis

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本研究提出模块化多智能体平台,通过多轮对话对角色扮演语言智能体开展对抗性压力测试,揭示单策略测试无法发现的故障模式,相关成果作为开源平台发布以支持AI安全。

Comments 8 pages, 1 figure, 7 tables; accepted and presented at ADScAI Conference 2026, University of Moratuwa, Sri Lanka

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.07881 2026-07-10 cs.SE cs.CR cs.LG 新提交 77%

Functional and Secure Code Generation with Task Vectors

使用任务向量进行功能和安全代码生成

Felix Wang, Anudeep Das, Mei Nagappan, N. Asokan

机构 * KTH Royal Institute of Technology(皇家理工学院)

专题命中 安全评测 :alignment(abstract);harmlessness(abstract);trustworthy(abstract);分类 cs.LG

AI总结 研究旨在解决大语言模型生成安全功能代码的问题,提出SecVecCoder方法,利用任务向量生成可信代码,在CodeGuard+基准测试中提升了可信代码完成率,且解码延迟低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.12716 2026-06-26 cs.CL 新提交 77%

Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review

AI审稿人是否看到全貌?攻击与防御多模态同行评审

Xinyu Zhao, Rana Muhammad Shahroz Khan, Zhen Xu, Zhen Tan, Tianlong Chen

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)

专题命中 安全评测 :safety(abstract);prompt injection(abstract);trustworthy(abstract);分类 cs.CL

AI总结 针对AI同行评审易受多模态对抗攻击的问题,提出PaperGuard基准,包含多领域数据集、统一攻击套件和基于分块嵌入搜索的实用防御方法。

Comments Accepted to ICML 2026, Project Page: https://paper-guard.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.25057 2026-06-25 cs.CL 新提交 77%

LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

基于LLM的科学同行评审:方法、基准与可靠性挑战

Thi Huyen Nguyen, Zahra Ahmadi

机构 * L3S Research Center, Leibniz University Hannover(莱布尼茨汉诺威大学L3S研究中心) Peter L. Reichertz Institute for Medical Informatics of TU Braunschweig and Hannover Medical School(布伦瑞克工业大学与汉诺威医学院彼得·L·赖歇茨医学信息学研究所) Lower Saxony Center for AI and Causal Methods in Medicine (CAIMed)(下萨克森州医学人工智能与因果方法中心(CAIMed))

专题命中 安全评测 :alignment(abstract);prompt injection(abstract);trustworthy(abstract);分类 cs.CL

AI总结 本文综述了基于大语言模型的科学同行评审,聚焦批评生成与分数预测两大功能,分析了建模方法、基准评估、数据局限及鲁棒性风险,并提出了未来研究方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.17114 2026-06-17 cs.CR cs.AI 新提交 77%

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios

现实场景中工具使用LLM代理的数据泄露风险评估

Hankyul Baek, Jaewon Noh, Sang Seo, Yongsu Kim, Gabriel Waikin Loh Matienzo, Young Il Kim, Ee Wei Seah, Akriti Vij

机构 * Korea AI Safety Institute(韩国人工智能安全研究所) Singapore AI Safety Institute(新加坡人工智能安全研究所)

专题命中 安全评测 :safety(abstract);prompt injection(abstract);AI safety(abstract);分类 cs.AI

AI总结 评估了12个非对抗性任务中AI代理的数据泄露风险,发现所有代理均存在数据安全意识不足、信息过度访问等问题,表明操作数据泄露是独立于对抗性窃取的一阶安全风险。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.14987 2026-05-22 cs.CL cs.DB 77%

Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI

超越基准岛屿:面向代理AI的代表性可信度评估

Jinhu Qi, Yifan Li, Minghao Zhao, Wentao Zhang, Zijian Zhang, Yaoman Li, Irwin King

机构 * The Chinese University of Hong Kong(香港中文大学) Macao Polytechnic University(澳门理工学院) Jilin University(吉林大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);trustworthy(abstract);分类 cs.CL

AI总结 本文提出了一种基于五属性的代理可信度定义,并引入了Holographic Agent Assessment Framework(HAAF)框架,通过场景 manifold 的静态策略分析、沙盒模拟、社会伦理对齐评估和分布感知采样,实现对代理系统在社会技术场景中的可信度评估,展示了其在13个模型家族上的跨家族迁移实验结果。

Comments 9 pages, 3 figures, 8 tables. Submitted to the Agent4IR Workshop at KDD 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08812 2026-05-21 cs.LG 77%

TRAM: Test-Time Risk Adaptation with Mixture of Agents

TRAM: 测试时风险适应与代理混合

Mohamad Fares El Hajj Chehade, Amrit Singh Bedi, Amy Zhang, Hao Zhu

机构 * UT Austin(得克萨斯大学) University of Central Florida(中央佛罗里达大学) MIT(麻省理工学院) UMD(大学公园分校)

专题命中 安全评测 :safety(abstract,abstract_cn);alignment(abstract);分类 cs.LG

AI总结 本文研究了在部署时无需更新的零更新适应问题,提出TRAM方法通过混合代理评估源策略的风险调整分数,以降低部署风险并保持奖励。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03992 2026-04-30 cs.CV cs.AI 77%

Value-Guided Iterative Refinement and the DIQ-H Benchmark for Evaluating VLM Robustness

基于价值引导的迭代精炼与DIQ-H基准:评估VLM鲁棒性的新方法

Hanwen Wan, Zexin Lin, Yixuan Deng, Xiaoqiang Ji

机构 * The School of Science and Engineering, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)科学与工程学院) The School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China(香港中文大学(深圳)人工智能学院) The Shenzhen Institute of Artificial Intelligence and Robotics for Society, Shenzhen, China(深圳人工智能与机器人社会研究院)

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文提出DIQ-H基准和VIR框架,通过模拟真实世界压力源评估VLM在连续序列中的鲁棒性,提升误差恢复和伦理一致性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07341 2026-04-24 econ.GN cs.AI q-fin.EC 77%

The Economics of p(doom): Scenarios of Existential Risk and Economic Growth in the Age of Transformative AI

transformative AI 的经济影响:转型人工智能时代存在风险与经济增长的场景

Jakub Growiec, Klaus Prettner

机构 * SGH Warsaw School of Economics(SGH华沙经济学院)

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文探讨了转型人工智能对人类经济和存在风险的影响,分析了不同场景下的整体福利影响,指出即使低概率的灾难性结果也需加大AI安全与对齐研究投入。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.17008 2026-04-21 cs.CL 77%

BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories

BIASEDTALES-ML:一个多语言数据集用于分析LLM生成故事中的叙述属性分布

Yuxuan Ouyang, yingfeng luo, JingBo Zhu, Tong Xiao

机构 * School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院,中国沈阳) NiuTrans Research, Shenyang, China(NiuTrans研究院,中国沈阳)

专题命中 安全评测 :alignment(abstract);safety(abstract);AI safety(abstract);分类 cs.CL

AI总结 本文提出BiasedTales-ML数据集,通过多语言生成分析叙述属性分布,揭示跨语言生成模式的差异,强调英语中心评估的局限性。

Comments Accepted to ACL 2026 Findings. Data are available at https://huggingface.co/spaces/Linyuana/BIASEDTALES-ML

详情

展开后加载摘要…

URL PDF HTML 收藏