arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2503.18172 2025-09-23 cs.CL cs.AI 73%

Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering

Zixin Chen, Sicheng Song, Kashun Shum, Yanna Lin, Rui Sheng, Weiqi Wang, Huamin Qu

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments 34 pages in total, EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08022 2025-09-17 cs.CL cs.AI 73%

MVPBench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

Yao Liang, Dongcheng Zhao, Feifei Zhao, Guobin Shen, Yuwei Wang, Dongqi Liang, Yi Zeng

专题命中 安全评测 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments Some parts of the paper need to be revised. We would therefore like to withdraw the paper and resubmit it after making the necessary changes

详情

展开后加载摘要…

URL PDF HTML 收藏
2101.12100 2025-09-10 cs.LG cs.AI 73%

Increasing the Confidence of Deep Neural Networks by Coverage Analysis

Giulio Rossolini, Alessandro Biondi, Giorgio Buttazzo

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Journal ref IEEE Transactions on Software Engineering ( Volume: 49, Issue: 2, 01 February 2023)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07992 2025-09-09 cs.CL cs.LG 73%

Concept Bottleneck Large Language Models

Chung-En Sun, Tuomas Oikarinen, Berk Ustun, Tsui-Wei Weng

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.LG

Comments Accepted to ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04549 2025-09-08 cs.CL cs.AI 73%

Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions

Faruk Alpay, Taylan Alpay

机构 * Lightcap Department of Future(未来系) Turkish Aeronautical Association(土耳其航空航天协会)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 13 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01185 2025-09-05 cs.CL cs.AI 73%

Modular Techniques for Synthetic Long-Context Data Generation in Language Model Training and Evaluation

Seganrasan Subramanian, Abhigya Verma

机构 * ServiceNow

专题命中 安全评测 :alignment(abstract);DPO(abstract);分类 cs.CL、cs.AI

Comments 26 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15370 2025-08-22 cs.CL cs.AI 73%

Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation

Yichi Zhang, Yao Huang, Yifan Wang, Yitong Sun, Chang Liu, Zhe Zhao, Zhengwei Fang, Huanran Chen, Xiao Yang, Xingxing Wei, Hang Su, Yinpeng Dong, Jun Zhu

机构 * Department of Computer Science and Technology, College of AI, Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University(计算机科学与技术系、人工智能学院、人工智能研究所、清华-博世联合机器学习中心、THBI实验室、BNRist中心、清华大学) Institute of Artificial Intelligence, Beihang University(人工智能研究院、北航) RealAI

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments For Appendix, please refer to arXiv:2406.07057

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11957 2025-08-19 cs.MA cs.AI cs.LG 73%

A Comprehensive Review of AI Agents: Transforming Possibilities in Technology and Beyond

Xiaodong Qu, Andrews Damoah, Joshua Sherwood, Peiyan Liu, Christian Shun Jin, Lulu Chen, Minjie Shen, Nawwaf Aleisa, Zeyuan Hou, Chenyu Zhang, Lifu Gao, Yanshu Li, Qikai Yang, Qun Wang, Cristabelle De Souza

机构 * University of Maryland, College Park(马里兰大学学院公园分校) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) Brown University(布朗大学) San Francisco State University(旧金山州立大学) Stanford University(斯坦福大学)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06296 2025-08-14 cs.AI cs.LG 73%

LLM Robustness Leaderboard v1 --Technical report

Pierre Peigné - Lefebvre, Quentin Feuillade-Montixi, Tom David, Nicolas Miailhe

机构 * PRISM Eval(PRISM评估)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00381 2025-08-11 cs.CV cs.AI cs.CE cs.LG 73%

Advancing Welding Defect Detection in Maritime Operations via Adapt-WeldNet and Defect Detection Interpretability Analysis

Kamal Basha S, Athira Nambiar

机构 * Department of Computational Intelligence, Faculty of Engineering and Technology, SRM Institute of Science and Technology(计算智能系,工程与技术学院,SRM科学与技术学院)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.03936 2025-08-07 cs.CR cs.CL cs.LG cs.SE 73%

ASTRA: Autonomous Spatial-Temporal Red-teaming for AI Software Assistants

Xiangzhe Xu, Guangyu Shen, Zian Su, Siyuan Cheng, Hanxi Guo, Lu Yan, Xuan Chen, Jiasheng Jiang, Xiaolong Jin, Chengpeng Wang, Zhuo Zhang, Xiangyu Zhang

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.LG

Comments The first two authors (Xiangzhe Xu and Guangyu Shen) contributed equally to this work

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16206 2025-07-23 cs.LG cs.AI 73%

METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark

Xu Yang, Qi Zhang, Shuming Jiang, Yaowen Xu, Zhaofan Zou, Hao Sun, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(电信人工智能研究院) Institute of Artificial Intelligence and Robotics(IAIR), Xi’an Jiaotong University(人工智能与机器人研究院) Advanced Technique of Artificial Intelligence(ATAI), Chongqing University of Technology(人工智能先进技术研究院)

专题命中 安全评测 :DPO(abstract);safety(abstract);分类 cs.AI、cs.LG

Comments 9 pages,3 figures ICCV format

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14615 2025-07-22 cs.CL cs.AI 73%

Retrieval-Augmented Clinical Benchmarking for Contextual Model Testing in Kenyan Primary Care: A Methodology Paper

Fred Mutisya, Shikoh Gitau, Christine Syovata, Diana Oigara, Ibrahim Matende, Muna Aden, Munira Ali, Ryan Nyotu, Diana Marion, Job Nyangena, Nasubo Ongoma, Keith Mbae, Elizabeth Wamicha, Eric Mibuari, Jean Philbert Nsengemana, Talkmore Chidede

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 29 pages, 6 figs, 6 tables. Companion methods paper forthcoming

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.21042 2025-07-17 cs.CR cs.AI cs.LG 73%

What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift

Jiamin Chang, Haoyang Li, Hammond Pearce, Ruoxi Sun, Bo Li, Minhui Xue

机构 * University of New South Wales \& CSIRO's Data61 Sydney Australia Macquarie University Sydney Australia University of New South Wales Sydney Australia University of Illinois at Urbana–Champaign Champaign University of New South Wales \& CSIRO's Data61 Macquarie University University of New South Wales University of Illinois at Urbana–Champaign

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

Comments Accepted to The ACM Conference on Computer and Communications Security (CCS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01786 2025-07-10 cs.CL cs.AI 73%

Probing and Steering Evaluation Awareness of Language Models

Jord Nguyen, Khiem Hoang, Carlo Leonardo Attubato, Felix Hofstätter

机构 * Pivotal Research Waseda University(早稻田大学) Apollo Research

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments Actionable Interpretability Workshop (Poster) and Workshop on Technical AI Governance (Poster) at ICML 2025, Vancouver, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04142 2025-07-08 cs.CL cs.AI 73%

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies

Mael Jullien, Marco Valentino, Leonardo Ranaldi, Andre Freitas

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02788 2025-07-04 cs.AI cs.CY 73%

Moral Responsibility or Obedience: What Do We Want from AI?

Joseph Boland

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22808 2025-07-01 cs.CL cs.AI 73%

MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs

Jianhui Wei, Zijie Meng, Zikai Xiao, Tianxiang Hu, Yang Feng, Zhijie Zhou, Jian Wu, Zuozhu Liu

机构 * Zhejiang University(浙江大学) Angelalign Technology Inc.(Angelalign技术公司)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05489 2025-07-01 cs.CL cs.AI 73%

Mechanistic Interpretability of Emotion Inference in Large Language Models

Ala N. Tak, Amin Banayeeanzade, Anahita Bolourani, Mina Kian, Robin Jia, Jonathan Gratch

机构 * Institute for Creative Technologies, University of Southern California (USC)(创意技术研究所,南加州大学) Information Sciences Institute, University of Southern California (USC)(信息科学研究所,南加州大学) Department of Computer Science, University of Southern California (USC)(计算机科学系,南加州大学) Department of Statistics and Data Science, University of California, Los Angeles (UCLA)(统计与数据科学系,加州大学洛杉矶分校)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 camera-ready version. First two authors contributed equally

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09251 2025-06-23 cs.RO cs.AI cs.LG 73%

V2X-VLM: End-to-End V2X Cooperative Autonomous Driving Through Large Vision-Language Models

Junwei You, Haotian Shi, Zhuoyu Jiang, Zilin Huang, Rui Gan, Keshu Wu, Xi Cheng, Xiaopeng Li, Bin Ran

机构 * organization= Department of Civil Environmental Engineering, University of Wisconsin–Madison , city= Madison , state= WI , postcode= 53706 , country= USA organization= College of Computing Data Science, Nanyang Technological University , city= Singapore , postcode= 639798 , country= Singapore organization= College of Transportation, Tongji University , city= Shanghai , postcode= 201804 , country= China organization= Zachry Department of Civil Environmental Engineering, Texas A\&M University , city= College Station , state= TX , postcode= 77840 , country= USA organization= School of Civil Environmental Engineering, Cornell University , city= Ithaca , state= NY , postcode= 14853 , country= USA

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16645 2025-06-19 cs.CL cs.AI cs.SE 73%

CODESYNC: Synchronizing Large Language Models with Dynamic Code Evolution at Scale

Chenlong Wang, Zhaoyang Chu, Zhengxiang Cheng, Xuyi Yang, Kaiyue Qiu, Yao Wan, Zhou Zhao, Xuanhua Shi, Dongping Chen

机构 * National Engineering Research Center for Big Data Technology and Systems(大数据技术与系统国家工程研究中心) Services Computing Technology and System Lab(服务计算技术与系统实验室) Cluster and Grid Computing Lab(集群与网格计算实验室) School of Computer Science and Technology(计算机科学与技术学院) Huazhong University of Science and Technology(华中科技大学) Wuhuan University(吴奂大学) Zhejiang University(浙江大学)

专题命中 安全评测 :DPO(abstract);safety(abstract);分类 cs.CL、cs.AI

Journal ref International Conference of Machine Learning, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09408 2025-06-12 cs.CL cs.AI 73%

Token Constraint Decoding Improves Robustness on Question Answering for Large Language Models

Jui-Ming Yao, Hao-Yuan Chen, Zi-Xian Tang, Bing-Jia Tan, Sheng-Wei Peng, Bing-Cheng Xie, Shun-Feng Su

机构 * Department of Computer Science and Information Engineering(计算机科学与信息工程系) National Taiwan University of Science and Technology(台湾科技大学) Bachelor’s Program in Computer Science(计算机科学学士班) University of London(伦敦大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03381 2025-06-05 eess.SY cs.AI cs.LG cs.SY 73%

Automated Traffic Incident Response Plans using Generative Artificial Intelligence: Part 1 -- Building the Incident Response Benchmark

Artur Grigorev, Khaled Saleh, Jiwon Kim, Adriana-Simona Mihaita

机构 * University of Technology Sydney(技术学院悉尼大学) University of Newcastle(新castle大学) The University of Queensland(昆士兰大学)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11642 2025-05-28 cs.MA cs.AI cs.LG 73%

PeerGuard: Defending Multi-Agent Systems Against Backdoor Attacks Through Mutual Reasoning

Falong Fan, Xi Li

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of Alabama at Birmingham(阿拉巴马大学伯明翰分校)

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.AI、cs.LG

Comments This paper has been accepted to IEEE IRI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15114 2025-05-28 cs.LG cs.AI 73%

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Josh Clymer, Jai Dhyani, Elena Ericheva, Katharyn Garcia, Brian Goodrich, Nikola Jurkovic, Holden Karnofsky, Megan Kinniment, Aron Lajko, Seraphina Nix, Lucas Sato, William Saunders, Maksym Taran, Ben West, Elizabeth Barnes

机构 * METR AI R&D Evaluation(METR AI R&D评估)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19443 2025-05-27 cs.SE cs.AI cs.CL 73%

Vibe Coding vs. Agentic Coding: Fundamentals and Practical Implications of Agentic AI

Ranjan Sapkota, Konstantinos I. Roumeliotis, Manoj Karkee

专题命中 安全评测 :safety(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

Comments 35 Pages, 8 Figures, 6 Tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16100 2025-05-23 cs.AI cs.CL 73%

BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research

Zifeng Wang, Benjamin Danek, Jimeng Sun

机构 * University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(abstract);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14943 2025-05-22 cs.LG cs.AI 73%

Soft Prompts for Evaluation: Measuring Conditional Distance of Capabilities

Ross Nordby

机构 * Ross Nordby(独立研究者)

专题命中 安全评测 :alignment(abstract);red teaming(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12692 2025-05-20 cs.AI cs.CL 73%

Bullying the Machine: How Personas Increase LLM Vulnerability

Ziwei Xu, Udit Sanghi, Mohan Kankanhalli

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10670 2025-05-19 cs.AI cs.CY cs.GT 73%

Interpretable Risk Mitigation in LLM Agent Systems

Jan Chojnacki

机构 * Department of Physics, University of Warsaw(华沙大学物理系) Samsung R&D Institute Poland(三星波兰研发中心)

专题命中 安全评测 :alignment(abstract);safety(abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏