arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9248 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9248 篇

2404.03027 2024-11-26 cs.CR cs.AI cs.CL 84%

JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks

Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, Chaowei Xiao

专题命中 安全评测 :jailbreak(title,abstract);alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.10301 2024-11-06 stat.ML cs.AI cs.LG 84%

Conformal Alignment: Knowing When to Trust Foundation Models with Guarantees

Yu Gui, Ying Jin, Zhimei Ren

专题命中 安全评测 :alignment(title,abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.07000 2024-10-29 cs.CL cs.AI 84%

Alignment for Honesty

Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, Pengfei Liu

专题命中 安全评测 :alignment(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI

Comments NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.12855 2024-10-21 cs.CL cs.AI 84%

JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework

Fan Liu, Yue Feng, Zhao Xu, Lixin Su, Xinyu Ma, Dawei Yin, Hao Liu

专题命中 安全评测 :jailbreak(title,abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.10019 2024-10-08 cs.CL cs.AI 84%

R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, Gongshen Liu

专题命中 安全评测 :safety(title,abstract);harmlessness(abstract);分类 cs.CL、cs.AI

Comments EMNLP Findings 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.02127 2024-09-05 cs.LG cs.AI 84%

Enabling Trustworthy Federated Learning in Industrial IoT: Bridging the Gap Between Interpretability and Robustness

Senthil Kumar Jagatheesaperumal, Mohamed Rahouti, Ali Alfatemi, Nasir Ghani, Vu Khanh Quy, Abdellah Chehri

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.LG

Comments 7 pages, 2 figures

Journal ref IEEE Internet of Things Magazine, Year: 2024, Volume: 7, Issue: 5

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.00029 2024-08-20 cs.CR cs.AI cs.CL 84%

Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework

Matthew Pisano, Peter Ly, Abraham Sanders, Bingsheng Yao, Dakuo Wang, Tomek Strzalkowski, Mei Si

专题命中 安全评测 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.09658 2024-07-23 cs.CV cs.AI cs.LG 84%

FUTURE-AI: Guiding Principles and Consensus Recommendations for Trustworthy Artificial Intelligence in Medical Imaging

Karim Lekadir, Richard Osuala, Catherine Gallin, Noussair Lazrak, Kaisar Kushibar, Gianna Tsakou, Susanna Aussó, Leonor Cerdá Alberich, Kostas Marias, Manolis Tsiknakis, Sara Colantonio, Nickolas Papanikolaou, Zohaib Salahuddin, Henry C Woodruff, Philippe Lambin, Luis Martí-Bonmatí

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.LG

Comments Please refer to arXiv:2309.12325 for the latest FUTURE-AI framework for healthcare

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.07666 2024-07-11 cs.CL cs.AI 84%

A Proposed S.C.O.R.E. Evaluation Framework for Large Language Models : Safety, Consensus, Objectivity, Reproducibility and Explainability

Ting Fang Tan, Kabilan Elangovan, Jasmine Ong, Nigam Shah, Joseph Sung, Tien Yin Wong, Lan Xue, Nan Liu, Haibo Wang, Chang Fu Kuo, Simon Chesterman, Zee Kin Yeong, Daniel SW Ting

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.10415 2024-06-18 cs.CY cs.AI cs.SE 84%

PRISM: A Design Framework for Open-Source Foundation Model Safety

Terrence Neumann, Bryan Jones

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.07594 2024-06-18 cs.CL cs.AI cs.CR 84%

MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models

Tianle Gu, Zeyang Zhou, Kexin Huang, Dandan Liang, Yixu Wang, Haiquan Zhao, Yuanqi Yao, Xingge Qiao, Keqing Wang, Yujiu Yang, Yan Teng, Yu Qiao, Yingchun Wang

专题命中 安全评测 :safety(title,abstract);red teaming(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20956 2024-06-05 cs.AI cs.CL 84%

A Robot Walks into a Bar: Can Language Models Serve as Creativity Support Tools for Comedy? An Evaluation of LLMs' Humour Alignment with Comedians

Piotr Wojciech Mirowski, Juliette Love, Kory W. Mathewson, Shakir Mohamed

专题命中 安全评测 :alignment(title,abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 15 pages, 1 figure, published at ACM FAccT 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.09447 2024-04-03 cs.CL cs.AI 84%

How Trustworthy are Open-Source LLMs? An Assessment under Malicious Demonstrations Shows their Vulnerabilities

Lingbo Mo, Boshi Wang, Muhao Chen, Huan Sun

专题命中 安全评测 :trustworthy(title);alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments NAACL 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.13289 2024-01-30 cs.AI cs.CR cs.CY 84%

Secure and Trustworthy Artificial Intelligence-Extended Reality (AI-XR) for Metaverses

Adnan Qayyum, Muhammad Atif Butt, Hassan Ali, Muhammad Usman, Osama Halabi, Ala Al-Fuqaha, Qammer H. Abbasi, Muhammad Ali Imran, Junaid Qadir

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.CY

Comments 24 pages, 11 figures

Journal ref ACM Computing Surveys (2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.00633 2023-10-03 cs.LG cs.AI 84%

A Survey of Robustness and Safety of 2D and 3D Deep Learning Models Against Adversarial Attacks

Yanjie Li, Bin Xie, Songtao Guo, Yuanyuan Yang, Bin Xiao

专题命中 安全评测 :safety(title,abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments Submitted to CSUR

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.11391 2023-08-29 cs.AI cs.LG 84%

A Survey of Safety and Trustworthiness of Large Language Models through the Lens of Verification and Validation

Xiaowei Huang, Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, Kaiwen Cai, Yanghao Zhang, Sihao Wu, Peipei Xu, Dengyu Wu, Andre Freitas, Mustafa A. Mustafa

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03193 2023-07-10 cs.CY cs.AI 84%

Finding differences in perspectives between designers and engineers to develop trustworthy AI for autonomous cars

Gustav Jonelid, K. R. Larsson

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2306.00380 2023-06-13 cs.AI cs.CY 84%

Survey of Trustworthy AI: A Meta Decision of AI

Caesar Wu, Yuan-Fang Lib, Pascal Bouvry

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.10117 2022-09-22 cs.IR cs.AI cs.CR cs.LG 84%

A Comprehensive Survey on Trustworthy Recommender Systems

Wenqi Fan, Xiangyu Zhao, Xiao Chen, Jingran Su, Jingtong Gao, Lin Wang, Qidong Liu, Yiqi Wang, Han Xu, Lei Chen, Qing Li

专题命中 安全评测 :trustworthy(title,abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.10044 2026-06-24 cs.AI cs.CL cs.CY cs.LG 84%

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety

脚手架下的安全性:评估条件如何影响测量的安全性

David Gringras

机构 * Harvard University(哈佛大学) MIT(麻省理工学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本研究通过62,808次盲法预注册评估,测试了六种前沿模型在四种部署配置下的安全性,发现脚手架架构对安全性影响较小,而格式转换(如选择题与开放式问题)可导致5-20个百分点的测量差异,且模型-脚手架间存在显著异质性,质疑了单一综合安全性分数的实用性。

Comments 74 pages including appendices. 6 frontier models, 62,808 primary observations (~89k total). Pre-registered: OSF DOI 10.17605/OSF.IO/CJW92. Code and data: https://github.com/davidgringras/safety-under-scaffolding

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.04992 2026-03-09 cs.CL 84%

ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contexts

ThaiSafetyBench: 评估泰语文化背景下语言模型的安全性

Trapoom Ukarapol, Nut Chukamphaeng, Kunat Pipatanakul, Pakhapoom Sarapat

机构 * SCB DataX(SCB数据X) Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) SCBX R&D(SCBX研发部) SCB 10X

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL;trustworthy(comments)

AI总结 本研究提出ThaiSafetyBench,通过泰语文化背景下的恶意提示评估语言模型安全性,发现闭源模型安全性更强,同时揭示了文化特定攻击的高风险。

Comments ICLR 2026 Workshop on Principled Design for Trustworthy AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.15796 2026-03-05 cs.AI 84%

From Privacy to Trust in the Agentic Era: A Taxonomy of Challenges in Trustworthy Federated Learning Through the Lens of Trust Report 2.0

从隐私到信任在代理时代:通过信任报告2.0的视角,对可信联邦学习中挑战的分类

Nuria Rodríguez-Barroso, Mario García-Márquez, M. Victoria Luzón, Francisco Herrera

机构 * Department of Computer Science and Artificial Intelligence, Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI) University of Granada(计算机科学与人工智能系,数据科学与计算智能安达卢西亚研究 institute,格拉纳达大学) Department of Software Engineering, Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI) University of Granada(软件工程系,数据科学与计算智能安达卢西亚研究 institute,格拉纳达大学)

专题命中 安全评测 :trustworthy(title,abstract);alignment(abstract);分类 cs.AI

AI总结 本文通过信任报告2.0提出可信联邦学习的挑战分类,强调信任作为持续维持的操作条件,并引入协调蓝图以处理跨要求的权衡和治理对齐。

Comments Already published in Information Fusion

Journal ref Rodríguez-Barroso, et. al. (2026). From Privacy to Trust in the Agentic Era: A Taxonomy of Challenges in Trustworthy Federated Learning Through the Lens of Trust Report 2.0. Information Fusion, 104236

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11903 2026-01-21 cs.AI 84%

AEMA: Verifiable Evaluation Framework for Trustworthy and Controlled Agentic LLM Systems

AEMA:可验证评估框架用于可信和受控的代理LLM系统

YenTing Lee, Keerthi Koneru, Zahra Moslemi, Sheethal Kumar, Ramesh Radhakrishnan

机构 * University of California, San Diego(加州大学圣地亚哥分校) Center for Advanced AI, Accenture(Accenture高级人工智能中心) University of California, Irvine(加州大学伊拉斯姆斯分校)

专题命中 安全评测 :trustworthy(title,abstract);alignment(abstract,comments);分类 cs.AI

AI总结 AEMA提出了一种可验证的评估框架,用于评估基于LLM的多代理系统,通过人类监督实现稳定、可追溯的自动化评估。

Comments Workshop on W51: How Can We Trust and Control Agentic AI? Toward Alignment, Robustness, and Verifiability in Autonomous LLM Agents at AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16903 2025-07-10 cs.CL cs.CR 84%

GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods

Ruixuan Huang, Xunguang Wang, Zongjie Li, Daoyuan Wu, Shuai Wang

专题命中 安全评测 :jailbreak(title,abstract);safety(abstract,comments);分类 cs.CL

Comments Homepage: https://sproutnan.github.io/AI-Safety_Benchmark/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11528 2026-08-13 cs.CL 新提交 83%

Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

群体对齐诱导的谄媚性:可引导的多元对齐的双向评估

Haokai Zhao, Yunze Xiao, Weihao Xuan, Flora Salim, Benjamin Tag, Aditya Joshi

机构 * University of New South Wales(新南威尔士大学) Carnegie Mellon University(卡内基梅隆大学) University of Tokyo(东京大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

AI总结 该研究提出GAS指标,评估3种方法、4个模型在13个人口统计群体上的对齐效果,发现群体对齐的观点收益和谄媚性变化存在群体异质性,建议采用双向多维度报告方式。

Comments 9 pages main text, 23 pages in total, under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10954 2026-08-12 cs.CV cs.AI 新提交 83%

Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes

复杂城市场景中基于证据的可信多模态推理与评估基准

Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao

机构 * University of Chinese Academy of Sciences (UCAS)(中国科学院大学) Tencent CDG(腾讯云与智慧产业事业群) Institute of Information Engineering, CAS(中国科学院信息工程研究所)

专题命中 安全评测 :trustworthy(title,abstract);alignment(abstract);分类 cs.AI

AI总结 针对复杂城市场景中多模态大语言模型的推理可靠性问题,提出AD2-Bench基准与EGVOR模型,提升了不利条件下的多模态推理稳定性。

Comments Accepted by IJCV

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10176 2026-08-12 cs.AI 新提交 83%

TRACE: Trustworthy Retrieval-Augmented Conversational Engine

TRACE:可信赖的检索增强对话引擎

Touseef Hasan, Laila Cure, Souvika Sarkar

机构 * Wichita State University(威奇托州立大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

AI总结 研究人员提出TRACE框架,通过强化检索提升公共服务对话系统的约束满足度、减少幻觉,降低对模型规模的依赖,证实检索质量是系统稳健性的关键。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.29243 2026-08-12 cs.LG 版本更新 83%

KrishokChat: A Provenance-Traceable Multi-Task Bengali Agricultural Benchmark with Safety-Critical Chemical Advisory

KrishokChat: 一个基于引用的孟加拉语农业咨询数据集与基准

Khan Raiyan Ibne Reza, Sumaiya Tabassum Nimi, Omar Ibne Shahid

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.LG

AI总结 提出KrishokChat,首个基于引用的孟加拉语农业指令微调数据集,通过分层知识节点扩展生成14.55万QA对,并引入农夫基准评估,发现微调改善格式但化学剂量泛化仍困难,强调其作为RAG知识库的价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07862 2026-08-11 cs.CL 新提交 83%

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

SurakshaEval:面向多语言大语言模型的印度语安全基准

Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah, Xudong Han, Yuxia Wang, Parameswari Krishnamurthy, Utkarsh Agarwal, Atharva Kulkarni, Swaran Lata, Ayush Munot, Dhruv Sahnan, Aaryamonvikram Singh, Preslav Nakov, Monojit Choudhury

专题命中 安全评测 :safety(title,abstract);alignment(abstract);分类 cs.CL

AI总结 针对现有LLM安全评估数据集忽视印度语言文化安全风险的问题,本文推出SurakshaEval基准,测试发现多语言LLM在印度语场景下安全表现不足,凸显了适配区域数据的安全评估框架的必要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06202 2026-08-07 cs.HC cs.AI 新提交 83%

What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)

当前AI基准测试未测量的内容:模态、搜索、引用及其对安全评估的启示

Ro Encarnación, Tina Behzad, Emma Lurie, Danaé Metaxa

专题命中 安全评测 :safety(title,abstract);AI safety(abstract);分类 cs.AI

AI总结 该研究发现AI基准测试未涵盖模态、搜索等关键因素,以ChatGPT为例对比两种模态及搜索条件,发现这些因素会影响模型准确率、一致性等,主张AI安全评估需纳入这些维度。

Comments 18 pages

详情

展开后加载摘要…

URL PDF HTML 收藏