arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9248 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9248 篇

2510.09710 2025-10-16 cs.CL cs.AI 81%

SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG

Xiaonan Si, Meilin Zhu, Simeng Qin, Lijia Yu, Lijun Zhang, Shuaitong Liu, Xinfeng Li, Ranjie Duan, Yang Liu, Xiaojun Jia

机构 * Institute of Software Chinese Academy of Sciences Beijing China(中国科学院软件研究所) Key Laboratory of System Software (Chinese Academy of Sciences) and State Key Laboratory of Computer Science, Institute of Software, Chinese Academy of Sciences, Beijing, China(中国科学院系统软件重点实验室和计算机科学国家重点实验室) University of Chinese Academy of Sciences, Beijing, China(中国科学院大学) Northeast University China(东北大学) Institute of Ai For industries Nanjing China(人工智能产业研究院) Southwest University China(西南大学) Nanyang Technological University Singapore(新加坡南洋理工大学) Alibaba China(阿里巴巴(中国))

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17244 2025-10-16 cs.CL cs.AI 81%

ReasoningShield: Safety Detection over Reasoning Traces of Large Reasoning Models

Changyi Li, Jiayi Wang, Xudong Pan, Geng Hong, Min Yang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08983 2025-10-13 cs.SE cs.AI cs.LG 81%

Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations

David N. Palacio, Daniel Rodriguez-Cardenas, Alejandro Velasco, Dipin Khati, Kevin Moran, Denys Poshyvanyk

机构 * University of Central Florida(佛罗里达中央大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments Under Review to appear in ACM Transactions on Software Engineering and Methodology (TOSEM)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10663 2025-10-08 q-bio.NC cs.AI cs.CV cs.LG 81%

Optimal Transport for Brain-Image Alignment: Unveiling Redundancy and Synergy in Neural Information Processing

Yang Xiao, Wang Lu, Jie Ji, Ruimeng Ye, Gen Li, Xiaolong Ma, Bo Hui

机构 * University of Tulsa(图拉大学) Tsinghua University(清华大学) Clemson University(克莱姆森大学) The University of Arizona(亚利桑那大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 14pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01606 2025-10-03 cs.IR cs.AI cs.CL 81%

Bridging Collaborative Filtering and Large Language Models with Dynamic Alignment, Multimodal Fusion and Evidence-grounded Explanations

Bo Ma, LuYao Liu, Simon Lau, Chandler Yuan, and XueY Cui, Rosie Zhang

机构 * Department of Software \& Microelectronics, Peking University, Beijing, China Economic Law School, China University of Political Science Financial Media, Peking University, ChangSha, China

专题命中 安全评测 :alignment(title);trustworthy(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01231 2025-10-03 cs.CL cs.AI stat.ML 81%

Trustworthy Summarization via Uncertainty Quantification and Risk Awareness in Large Language Models

Shuaidong Pan, Di Wu

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15463 2025-10-01 cs.HC cs.AI cs.CL 81%

Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?

Hua Shen, Nicholas Clark, Tanushree Mitra

机构 * University of Washington(华盛顿大学) NYU Shanghai(纽约大学上海校区) New York University(纽约大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025 Main Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13742 2025-09-23 cs.LG cs.AI math.OC 81%

Search-Optimized Quantization in Biomedical Ontology Alignment

Oussama Bouaggad, Natalia Grabar

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in Frontiers in Artificial Intelligence - Medicine and Public Health (Original Research)

Journal ref Front. Artif. Intell. 8:1662984 (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16394 2025-09-23 cs.CL cs.AI cs.HC 81%

Evaluating Behavioral Alignment in Conflict Dialogue: A Multi-Dimensional Comparison of LLM Agents and Humans

Deuksin Kwon, Kaleen Shrestha, Bin Han, Elena Hayoung Lee, Gale Lucas

机构 * University of Southern California(南加州大学) USC for Institute of Creative Technologies(南加州大学创意技术研究所)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted to EMNLP 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15932 2025-09-22 cs.LG cs.AI cs.IT math.IT stat.ML 81%

The Alignment Bottleneck

Wenjun Cao

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08380 2025-09-18 cs.AI cs.LG 81%

Co-Investigator AI: The Rise of Agentic AI for Smarter, Trustworthy AML Compliance Narratives

Prathamesh Vasudeo Naik, Naresh Kumar Dintakurthi, Zhanghao Hu, Yue Wang, Robby Qiu

专题命中 安全评测 :trustworthy(title);alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.01166 2025-09-03 cs.CL cs.AI 81%

Enhancing Large Language Model for Knowledge Graph Completion via Structure-Aware Alignment-Tuning

Yu Liu, Yanan Cao, Xixun Lin, Yanmin Shang, Shi Wang, Shirui Pan

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络空间安全学院) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Griffith University(格里菲斯大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025, Main, Long Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02531 2025-08-28 cs.CY cs.CL 81%

Towards New Benchmark for AI Alignment & Sentiment Analysis in Socially Important Issues: A Comparative Study of Human and LLMs in the Context of AGI

Ljubisa Bojic, Dylan Seychell, Milan Cabarkapa

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments 34 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.15192 2025-08-22 cs.AI cs.CL 81%

LLM4Sweat: A Trustworthy Large Language Model for Hyperhidrosis Support

Wenjie Lin, Jin Wei-Kocsis

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22146 2025-08-22 cs.CV cs.AI cs.CL q-bio.NC 81%

Flexible Tool Selection through Low-dimensional Attribute Alignment of Vision and Language

Guangfu Hao, Haojie Wen, Liangxuan Guo, Yang Chen, Yanchao Bi, Shan Yu

机构 * Laboratory of Brain Atlas and Brain-inspired Intelligence, Institute of Automation Chinese Academy of Sciences (CASIA)(中国科学院自动化研究所脑图谱与类脑智能实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院) School of Systems Science, Beijing Normal University(北京师范大学系统科学学院) School of Psychological and Cognitive Sciences & Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京大学心理与认知科学学院) IDG/McGovern Institute for Brain Research, Peking University(北京大学IDG/ McGovern脑科学研究院) Institute for Artificial Intelligence & Key Laboratory of Machine Perception (Ministry of Education), Peking University(北京大学人工智能研究所) School of Future Technology, University of Chinese Academy of Sciences (UCAS)(中国科学院大学未来技术学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14735 2025-08-21 cs.CL cs.AI 81%

Evaluating Multilingual and Code-Switched Alignment in LLMs via Synthetic Natural Language Inference

Samir Abdaljalil, Erchin Serpedin, Khalid Qaraqe, Hasan Kurban

机构 * Texas A\&M University, College Station, TX., USA(德克萨斯大学) Hamad Bin Khalifa University, Doha, Qatar(哈马德·本·卡伊夫大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11247 2025-08-19 cs.AI cs.LG cs.RO 81%

LD-Scene: LLM-Guided Diffusion for Controllable Generation of Adversarial Safety-Critical Driving Scenarios

Mingxing Peng, Yuting Xie, Xusen Guo, Ruoyu Yao, Hai Yang, Jun Ma

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 18 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.12624 2025-08-19 cs.CL cs.AI 81%

Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, Dieuwke Hupkes

机构 * University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校) Meta

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments https://aclanthology.org/2025.gem-1.33/

Journal ref Proceedings of the Fourth Workshop on Generation Evaluation and Metrics GEM2 2025 pages 404 to 430; July 31 August 1 2025; 2025 Association for Computational Linguistics

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13983 2025-08-13 cs.CL cs.AI 81%

AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models

Yang Fan

专题命中 安全评测 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments There are serious academic problems in this paper, such as data falsification and plagiarism in the method of the paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.02078 2025-08-13 cs.CV cs.AI cs.LG 81%

From Lab to Field: Real-World Evaluation of an AI-Driven Smart Video Solution to Enhance Community Safety

Shanle Yao, Babak Rahimi Ardabili, Armin Danesh Pazho, Ghazal Alinezhad Noghre, Christopher Neff, Lauren Bourque, Hamed Tabkhi

机构 * Department of Electrical and Computer Engineering, University of North Carolina at Charlotte(电气与计算机工程系,北卡罗来纳大学夏洛特分校) Department of Public Policy, University of North Carolina at Charlotte(公共政策系,北卡罗来纳大学夏洛特分校)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08273 2025-08-13 cs.CL cs.LG 81%

TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning

Kristian Miok, Blaz Škrlj, Daniela Zaharie, Marko Robnik Šikonja

机构 * Faculty of Computer and Information Science, University of Ljubljana, Slovenia(卢布尔雅那大学计算机与信息科学学院) ICAM - Advanced Environmental Research Institute, West University of Timisoara, Romania(蒂米șoara西大学先进环境研究所) Department of Computer Science, West University of Timisoara, Romania(蒂米șoa拉西大学计算机科学系)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.16591 2025-07-31 cs.LG cs.AI cs.CR 81%

Bridging Privacy and Robustness for Trustworthy Machine Learning

Xiaojin Zhang, Wei Chen

机构 * Huazhong University of Science and Technology(华中科技大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17616 2025-07-24 cs.CV cs.AI cs.LG 81%

Vision Transformer attention alignment with human visual perception in aesthetic object evaluation

Miguel Carrasco, César González-Martín, José Aranda, Luis Oliveros

机构 * Escuela de Informática y Telecomunicaciones, Universidad Diego Portáles(戴维·波特莱斯大学信息与电信学院) Department of Specific Didactics, University of Cordoba(科尔多瓦大学特定教学系) Facultad de Ingeniería y Ciencias, Universidad Adolfo Ibáñez(阿道弗·伊巴涅斯大学工程与科学学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 25 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12404 2025-07-24 cs.LG cs.AI 81%

EXGnet: a single-lead explainable-AI guided multiresolution network with train-only quantitative features for trustworthy ECG arrhythmia classification

Tushar Talukder Showrav, Soyabul Islam Lincoln, Md. Kamrul Hasan

机构 * Dept. of Electrical and Electronic Engineering(电子与电气工程系) Bangladesh University of Engineering and Technology(孟加拉工程与技术大学) Dept. of Electronics & Communication Engineering(电子与通信工程系) Khulna University of Engineering and Technology(库尔纳工程与技术大学)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

Comments 17 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02145 2025-07-21 cs.AI cs.CL cs.RO 81%

From Words to Collisions: LLM-Guided Evaluation and Adversarial Generation of Safety-Critical Driving Scenarios

Yuan Gao, Mattia Piccinini, Korbinian Moller, Amr Alanwar, Johannes Betz

机构 * Professorship of Autonomous Vehicle Systems, TUM School of Engineering and Design, Technical University of Munich(自主车辆系统教授职位,技术大学慕尼黑工程与设计学院) Munich Institute of Robotics and Machine Intelligence (MIRMI)(慕尼黑机器人与机器智能研究所) TUM School of Computation, Information and Technology, Department of Computer Engineering, Technical University of Munich(技术大学慕尼黑计算、信息与技术学院,计算机工程系)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments Final Version and Paper Accepted at IEEE ITSC 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17352 2025-07-03 cs.CL cs.AI 81%

Towards Safety Evaluations of Theory of Mind in Large Language Models

Tatsuhiro Aoshima, Mitsuaki Akiyama

机构 * NTT(日本电报电话株式会社)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19564 2025-07-01 cs.LG cs.AI 81%

FedMM-X: A Trustworthy and Interpretable Framework for Federated Multi-Modal Learning in Dynamic Environments

Sree Bhargavi Balija

机构 * University of California, San Diego(加州大学圣地亚哥分校)

专题命中 安全评测 :trustworthy(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22446 2025-07-01 cs.LG cs.AI 81%

EAGLE: Efficient Alignment of Generalized Latent Embeddings for Multimodal Survival Prediction with Interpretable Attribution Analysis

Aakash Tripathi, Asim Waqas, Matthew B. Schabath, Yasin Yilmaz, Ghulam Rasool

机构 * Dept. of Machine Learning Moffitt Cancer Center(机器学习系莫菲特癌症中心) Dept. of Cancer Epidemiology Moffitt Cancer Center(癌症流行病学系莫菲特癌症中心) Dept. of Electrical Engineering University of South Florida(电气工程系佛罗里达州立大学)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21252 2025-06-27 cs.CL cs.AI 81%

Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents

Tianyi Men, Zhuoran Jin, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao

机构 * The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences(认知与决策智能复杂系统重点实验室,自动化研究所,中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学)

专题命中 安全评测 :safety(title,abstract);分类 cs.CL、cs.AI

Comments ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20251 2025-06-26 cs.LG cs.AI 81%

Q-resafe: Assessing Safety Risks and Quantization-aware Safety Patching for Quantized Large Language Models

Kejia Chen, Jiawen Zhang, Jiacong Hu, Yu Wang, Jian Lou, Zunlei Feng, Mingli Song

机构 * The State Key Laboratory of Blockchain and Data Security, Zhejiang University(区块链与数据安全国家重点实验室,浙江大学) Sun Yat-sen University(中山大学) Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security(杭州高新技术区(滨江)区块链与数据安全研究院)

专题命中 安全评测 :safety(title,abstract);分类 cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏