arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9311 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9311 篇

2507.16033 2025-07-23 cs.HC cs.AI 79%

"Just a strange pic": Evaluating 'safety' in GenAI Image safety annotation tasks from diverse annotators' perspectives

Ding Wang, Mark Díaz, Charvi Rastogi, Aida Davani, Vinodkumar Prabhakaran, Pushkar Mishra, Roma Patel, Alicia Parrish, Zoe Ashwood, Michela Paganini, Tian Huey Teh, Verena Rieser, Lora Aroyo

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments Accepted to AAAI/ACM Conference on Artificial Intelligence, Ethics, and Society 2025 (AIES 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14693 2025-07-22 cs.CL cs.AI cs.CY cs.LG 79%

Rethinking Suicidal Ideation Detection: A Trustworthy Annotation Framework and Cross-Lingual Model Evaluation

Amina Dzafic, Merve Kavut, Ulya Bayram

专题命中 安全评测 :trustworthy(title);分类 cs.CL、cs.AI、cs.CY

Comments This manuscript has been submitted to the IEEE Journal of Biomedical and Health Informatics

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.19819 2025-07-17 cs.AR cs.AI 79%

ChipAlign: Instruction Alignment in Large Language Models for Chip Design via Geodesic Interpolation

Chenhui Deng, Yunsheng Bai, Haoxing Ren

机构 * NVIDIA(英伟达)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments Accepted to DAC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10905 2025-07-15 cs.HC cs.AI cs.ET 79%

The impact of labeling automotive AI as "trustworthy" or "reliable" on user evaluation and technology acceptance

John Dorsch, Ophelia Deroy

机构 * Faculty of Philosophy, Philosophy of Science and the Study of Religion, Ludwig-Maximilians-Universität München(哲学学院、科学哲学与宗教研究学院,慕尼黑路德维希-马克西米利安大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 36 pages, 12 figures

Journal ref Nature Scientific Reports, 15, 1481. 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08109 2025-07-14 cs.CL 79%

Audit, Alignment, and Optimization of LM-Powered Subroutines with Application to Public Comment Processing

Reilly Raab, Mike Parker, Dan Nally, Sadie Montgomery, Anastasia Bernat, Sai Munikoti, Sameera Horawalavithana

机构 * Pacific Northwest National Laboratory(太平洋西北国家实验室)

专题命中 安全评测 :alignment(title);safety(abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07787 2025-07-14 cs.AI 79%

Measuring AI Alignment with Human Flourishing

Elizabeth Hilliard, Akshaya Jagadeesh, Alex Cook, Steele Billings, Nicholas Skytland, Alicia Llewellyn, Jackson Paull, Nathan Paull, Nolan Kurylo, Keatra Nesbitt, Robert Gruenewald, Anthony Jantzi, Omar Chavez

机构 * Gloo Valkyrie Intelligence Biblica

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06378 2025-07-10 cs.CL 79%

Evaluating Morphological Alignment of Tokenizers in 70 Languages

Catherine Arnett, Marisa Hudspeth, Brendan O'Connor

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments 6 pages, 3 figures. Accepted to the Tokenization Workshop at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06949 2025-07-09 cs.SE cs.CL 79%

Towards Exception Safety Code Generation with Intermediate Representation Agents Framework

Xuanming Zhang, Yuxuan Chen, Yuan Yuan, Minlie Huang

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Stanford University(斯坦福大学) Tsinghua University(清华大学) Beihang University(北航大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04976 2025-07-08 cs.CV cs.CL 79%

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

Eunseop Yoon, Hee Suk Yoon, Mark A. Hasegawa-Johnson, Chang D. Yoo

机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) University of Illinois at Urbana-Champaign (UIUC)(伊利诺伊大学厄巴纳-香槟分校)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

Comments ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.01274 2025-07-03 cs.HC cs.AI 79%

AI Meets Maritime Training: Precision Analytics for Enhanced Safety and Performance

Vishakha Lall, Yisi Liu

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments Accepted and Presented at 11th International Maritime Science Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00096 2025-07-02 cs.CR cs.AI 79%

AI-Governed Agent Architecture for Web-Trustworthy Tokenization of Alternative Assets

Ailiya Borjigin, Wei Zhou, Cong He

机构 * Probe Group Pte. Ltd.(Probe集团)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 8 Pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20832 2025-07-01 cs.CV cs.AI 79%

Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models

Cansu Korkmaz, Ahmet Murat Tekalp, Zafer Dogan

机构 * Department of Electrical and Electronics Engineering and KUIS AI Center, Koç University(电气与电子工程系和KUIS人工智能中心,科克大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments 14 pages, 9 figures, 5 tables, accepted to IEEE Transactions on Circuits and Systems for Video Technology

Journal ref IEEE Transactions on Circuits and Systems for Video Technology 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17404 2025-07-01 cs.AI 79%

Super Co-alignment of Human and AI for Sustainable Symbiotic Society

Yi Zeng, Feifei Zhao, Yuwei Wang, Enmeng Lu, Yaodong Yang, Lei Wang, Chao Liu, Yitao Liang, Dongcheng Zhao, Bing Han, Haibo Tong, Yao Liang, Dongqi Liang, Kang Sun, Boyuan Chen, Jinyu Fan

机构 * Beijing Key Laboratory of Safe AI and Superalignment(北京安全人工智能与超对齐关键实验室) Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院) Brain-inspired Cognitive AI Lab(脑启发认知人工智能实验室) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Wenge Technology Co., Ltd.(Wenger 技术有限公司) Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) Long-term AI(长期人工智能) State Key Laboratory of Cognitive Neuroscience and Learning, Beijing Normal University(北京师范大学认知神经科学与学习国家重点实验室)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13232 2025-06-30 cs.AI cs.CV 79%

StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment

Younghyun Kim, Jongheon Jeong, Sangkyung Kwak, Kyungmin Lee, Juho Lee, Jinwoo Shin

机构 * Samsung(三星) Korea University(韩国大学) General Robotics(通用机器人) KAIST(韩国科学技术院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments IJCAI 2025; Code is available at https://github.com/alinlab/StarFT

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21017 2025-06-27 cs.CV cs.AI 79%

Multimodal Prompt Alignment for Facial Expression Recognition

Fuyan Ma, Yiran He, Bin Sun, Shutao Li

机构 * Chinese Academy of Military Science(中国军事科学院) Changchun University of Science and Technology(长春理工大学) Hunan University(湖南大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

Comments To appear in ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19702 2025-06-25 cs.AI 79%

LLM-Driven Medical Document Analysis: Enhancing Trustworthy Pathology and Differential Diagnosis

Lei Kang, Xuanshuo Fu, Oriol Ramos Terrades, Javier Vazquez-Corral, Ernest Valveny, Dimosthenis Karatzas

机构 * Computer Vision Center, Universitat Autònoma de Barcelona(计算机视觉中心,巴塞罗那自治大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

Comments Accepted at ICDAR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15035 2025-06-24 cs.CL 79%

LLMs Lost in Translation: M-ALERT uncovers Cross-Linguistic Safety Inconsistencies

Felix Friedrich, Simone Tedeschi, Patrick Schramowski, Manuel Brack, Roberto Navigli, Huu Nguyen, Bo Li, Kristian Kersting

机构 * TU Darmstadt(图宾根大学) Sapienza University of Rome(罗马萨皮恩扎大学) DFKI(德意志联邦防务研究院) CERTAIN(CERTAIN公司) University of Chicago(芝加哥大学) UIUC(伊利诺伊大学香槟分校)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16322 2025-06-23 cs.CL 79%

PL-Guard: Benchmarking Language Model Safety for Polish

Aleksandra Krasnodębska, Karolina Seweryn, Szymon Łukasik, Wojciech Kusa

机构 * NASK – National Research Institute(国家研究 institute)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments Accepted to the 10th Workshop on Slavic Natural Language Processing

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15699 2025-06-23 cs.AI 79%

Advancing Embodied Agent Security: From Safety Benchmarks to Input Moderation

Ning Wang, Zihan Yan, Weiyang Li, Chuan Ma, He Chen, Tao Xiang

机构 * College of computer science, Chongqing University(重庆大学计算机学院) Department of Information Engineering, The Chinese University of Hong Kong(香港中文大学信息工程系)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.06809 2025-06-19 cs.CL cs.CR 79%

Root Defence Strategies: Ensuring Safety of LLM at the Decoding Level

Xinyi Zeng, Yuying Shang, Jiawei Chen, Jingyuan Zhang, Yu Tian

机构 * Aerospace Information Research Institute, Chinese Academy of Sciences(航天信息研究所,中国科学院) Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学) Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University(上海多信息处理重点实验室,华东师范大学) Kuaishou-Inc(快手公司)

专题命中 安全评测 :safety(title);jailbreak(abstract);分类 cs.CL

Comments 19 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02958 2025-06-18 cs.CL 79%

Position: Editing Large Language Models Poses Serious Safety Risks

Paul Youssef, Zhixue Zhao, Daniel Braun, Jörg Schlötterer, Christin Seifert

机构 * Marburg University(马尔堡大学) University of Sheffield(谢菲尔德大学) University of Mannheim(曼海姆大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments Accepted at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00840 2025-06-11 cs.CR cs.AI 79%

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense

Jiawen Zhang, Kejia Chen, Lipeng He, Jian Lou, Dan Li, Zunlei Feng, Mingli Song, Jian Liu, Kui Ren, Xiaohu Yang

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.10903 2025-06-10 cs.CL cs.HC 79%

BEYOND DIALOGUE: A Profile-Dialogue Alignment Framework Towards General Role-Playing Language Model

Yeyong Yu, Runsheng Yu, Haojie Wei, Zhanqiu Zhang, Quan Qian

机构 * School of Computer Engineering & Science, Shanghai University(上海大学计算机工程与科学学院) LIGHTSPEED The Hong Kong University of Science and Technology(香港科技大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.08469 2025-06-10 cs.AI cs.ET 79%

Building Trustworthy AI: Transparent AI Systems via Large Language Models, Ontologies, and Logical Reasoning (TranspNet)

Fadi Al Machot, Martin Thomas Horsch, Habib Ullah

机构 * Department of Data Science, Faculty of Science and Technology, Norwegian University of Life Sciences(数据科学系,科学与技术学院,挪威生命科学大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04142 2025-06-05 cs.CL 79%

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis

Kejian Zhu, Shangqing Tu, Zhuoran Jin, Lei Hou, Juanzi Li, Jun Zhao

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences, Beijing, China(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院,北京,中国) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) Tsinghua University(清华大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL

Comments Accepted to ACL 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09295 2025-06-05 cs.CL cs.CV 79%

AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models

Yuhang Wu, Wenmeng Yu, Yean Cheng, Yan Wang, Xiaohan Zhang, Jiazheng Xu, Ming Ding, Yuxiao Dong

机构 * Tsinghua University(清华大学) Zhipu AI(智谱AI) Peking University(北京大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02573 2025-06-04 cs.CL 79%

IndoSafety: Culturally Grounded Safety for LLMs in Indonesian Languages

Muhammad Falensi Azmi, Muhammad Dehan Al Kautsar, Alfan Farizki Wicaksono, Fajri Koto

机构 * Faculty of Computer Science, Universitas Indonesia(印度尼西亚大学计算机科学学院) Department of Natural Language Processing, MBZUAI(自然语言处理部门,MBZUAI)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

Comments 25 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02372 2025-06-04 cs.CL 79%

AnswerCarefully: A Dataset for Improving the Safety of Japanese LLM Output

Hisami Suzuki, Satoru Katsumata, Takashi Kodama, Tetsuro Takahashi, Kouta Nakayama, Satoshi Sekine

机构 * NII-LLMC(日本国立信息与通信技术研究所语言模型中心) Retrieva, Inc.(Retrieva公司) Kagoshima University(鹿儿岛大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10930 2025-06-03 cs.LG 79%

Physics-informed Temporal Alignment for Auto-regressive PDE Foundation Models

Congcong Zhu, Xiaoyan Xu, Jiayue Han, Jingrun Chen

专题命中 安全评测 :alignment(title,abstract);分类 cs.LG

Comments Accepted as a conference paper in ICML2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23793 2025-06-02 cs.CR cs.AI 79%

USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models

Baolin Zheng, Guanlin Chen, Hongqiong Zhong, Qingyang Teng, Yingshui Tan, Zhendong Liu, Weixun Wang, Jiaheng Liu, Jian Yang, Huiyun Jing, Jincheng Wei, Wenbo Su, Xiaoyong Zhu, Bo Zheng, Kaifu Zhang

专题命中 安全评测 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏