arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2511.17730 2025-11-25 eess.SY cs.SY 78%

Safety and Risk Pathways in Cooperative Generative Multi-Agent Systems: A Telecom Perspective

合作生成多智能体系统的安全性和风险路径:电信视角

Zeinab Nezami, Shehr Bano, Abdelaziz Salama, Maryam Hafeez, Syed Ali Raza Zaidi

专题命中 安全评测 :safety(title,abstract)

AI总结 本文从电信视角探讨生成多智能体系统中的安全性和风险路径,提出模块化安全评估框架,揭示智能体多样性对系统稳定性的影响,并展示通过模拟验证的改进与持续存在的漏洞。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24143 2025-11-21 cs.DC 78%

Enhancing Traffic Safety with AI and 6G: Latency Requirements and Real-Time Threat Detection

用AI和6G提升交通安全性:延迟需求与实时威胁检测

Kurt Horvath, Dragi Kimovski, Stojan Kitanov, Radu Prodan

专题命中 安全评测 :safety(title,abstract)

AI总结 本文提出基于6G和AI的交通安全框架,通过实时威胁检测和低延迟通信提升交通安全性。

Comments Sumbitted/Accepted ICINT 2025 (PrePrint)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.15206 2025-11-20 cs.CR cs.IT math.IT 78%

Trustworthy GenAI over 6G: Integrated Applications and Security Frameworks

Bui Duc Son, Trinh Van Chien, Dong In Kim

专题命中 安全评测 :trustworthy(title,abstract)

Comments 8 pages, 5 figures. Submitted for publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.03565 2025-11-12 cs.CV 78%

Bridged Semantic Alignment for Zero-shot 3D Medical Image Diagnosis

Haoran Lai, Zihang Jiang, Qingsong Yao, Rongsheng Wang, Zhiyang He, Xiaodong Tao, Weifu Lv, Wei Wei, S. Kevin Zhou

机构 * University of Science and Technology of China(中国科学技术大学) Suzhou Institute for Advanced Research(苏州先进研究所) Stanford University(斯坦福大学) iFlytek Co. Ltd.(iFlytek公司) The First Affiliated Hospital of USTC, Division of Life Sciences and Medicine, USTC(中国科学技术大学第一附属医院)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00775 2025-10-28 cs.HC 78%

Efficiency with Rigor! A Trustworthy LLM-powered Workflow for Qualitative Data Analysis

Jie Gao, Zhiyao Shu, Shun Yi Yeo, Alok Prakash, Chien-Ming Huang, Mark Dredze, Ziang Xiao

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21606 2025-10-27 cs.CV 78%

Modest-Align: Data-Efficient Alignment for Vision-Language Models

Jiaxiang Liu, Yuan Wang, Jiawei Du, Joey Tianyi Zhou, Mingkun Xu, Zuozhu Liu

机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) ZJU-Angelalign R&D Center for Intelligence Healthcare(浙大天使align智能医疗研发中心) Centre for Frontier AI Research (CFAR)(前沿人工智能研究中心) Agency for Science, Technology and Research (A*STAR)(科技研究局) Institute of High Performance Computing (IHPC)(高性能计算研究所)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21120 2025-10-27 cs.CV 78%

SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation

Alec Helbling, Shruti Palaskar, Kundan Krishna, Polo Chau, Leon Gatys, Joseph Yitan Cheng

机构 * Georgia Tech(佐治亚理工学院) Apple(苹果公司)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18550 2025-10-22 cs.NI 78%

JAUNT: Joint Alignment of User Intent and Network State for QoE-centric LLM Tool Routing

Enhan Li, Hongyang Du

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.10086 2025-10-14 cs.RO 78%

Beyond ADE and FDE: A Comprehensive Evaluation Framework for Safety-Critical Prediction in Multi-Agent Autonomous Driving Scenarios

Feifei Liu, Haozhe Wang, Zejun Wei, Qirong Lu, Yiyang Wen, Xiaoyu Tang, Jingyan Jiang, Zhijian He

机构 * South China Normal University(华南师范大学) Shenzhen Technology University(深圳技术大学)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.17040 2025-10-14 cs.CV 78%

Multimodal Alignment and Fusion: A Survey

Songtao Li, Hao Tang

机构 * Peking University(北京大学) Northeastern University(东北大学) Sydney Smart Technology College(悉尼智能技术学院) School of Computer Science, Peking University(北京大学计算机学院) The State Key Laboratory of Multimedia Information Processing(多媒体信息处理国家重点实验室)

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted to IJCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.07652 2025-10-10 cs.CV 78%

Dual-Stream Alignment for Action Segmentation

Harshala Gammulle, Clinton Fookes, Sridha Sridharan, Simon Denman

机构 * Signal Processing, Artificial Intelligence and Vision Technologies (SAIVT) Lab(信号处理、人工智能与视觉技术实验室) Queensland University of Technology(昆士兰理工大学)

专题命中 安全评测 :alignment(title,abstract)

Comments Journal Submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02987 2025-10-06 cs.CV 78%

TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency

Juntong Wang, Huiyu Duan, Jiarui Wang, Ziheng Jia, Guangtao Zhai, Xiongkuo Min

机构 * Institute of Image Communication and Network Engineering(图像通信与网络工程研究所) MoE Key Lab of Artificial Intelligence, AI Institute(人工智能关键实验室) Shanghai Jiao Tong University(上海交通大学)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26122 2025-10-01 math.NA cs.NA 78%

Trustworthy AI in numerics: On verification algorithms for neural network-based PDE solvers

Emil Haugen, Alexei Stepanenko, Anders C. Hansen

专题命中 安全评测 :trustworthy(title,abstract)

Comments 25 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04974 2025-10-01 eess.AS cs.SD 78%

From Voice to Safety: Language AI Powered Pilot-ATC Communication Understanding for Airport Surface Movement Collision Risk Assessment

Yutian Pang, Andrew Paul Kendall, Alex Porcayo, Mariah Barsotti, Anahita Jain, John-Paul Clarke

机构 * Department of Aerospace Engineering and Engineering Mechanics(航空航天工程与工程力学系)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23760 2025-09-30 cs.CV 78%

UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception

Xinyang Song, Libin Wang, Weining Wang, Shaozhen Liu, Dandan Zheng, Jingdong Chen, Qi Li, Zhenan Sun

机构 * Ant Group(蚂蚁集团)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20393 2025-09-26 cs.CY cs.AI cs.LG 78%

The Secret Agenda: LLMs Strategically Lie and Our Current Safety Tools Are Blind

Caleb DeLeeuw, Gaurav Chawla, Aniket Sharma, Vanessa Dietze

机构 * Independent Researcher(独立研究者)

专题命中 安全评测 :safety(title);分类 cs.AI、cs.CY、cs.LG

Comments 9 pages plus citations and appendix, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.14446 2025-09-19 q-bio.NC 78%

Mouse vs. AI: A Neuroethological Benchmark for Visual Robustness and Neural Alignment

Marius Schneider, Joe Canzano, Jing Peng, Yuchen Hou, Spencer LaVere Smith, Michael Beyeler

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08997 2025-09-12 cs.HC 78%

YouthSafe: A Youth-Centric Safety Benchmark and Safeguard Model for Large Language Models

Yaman Yu, Yiren Liu, Jacky Zhang, Yun Huang, Yang Wang

专题命中 安全评测 :safety(title,abstract)

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00700 2025-09-10 cs.CV 78%

Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision

Raehyuk Jung, Seungjun Yu, Hyunjung Shim

专题命中 安全评测 :alignment(title,abstract)

Comments Link to publicly available codes is added

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.10610 2025-09-03 cs.RO cs.SY eess.SY 78%

Safety-Critical Human-Machine Shared Driving for Vehicle Collision Avoidance based on Hamilton-Jacobi reachability

Shiyue Zhao, Junzhi Zhang, Rui Zhou, Neda Masoud, Jianxiong Li, Helai Huang, Shijie Zhao

专题命中 安全评测 :safety(title,abstract)

Comments 36 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14527 2025-08-26 cs.CV 78%

Adversarial Generation and Collaborative Evolution of Safety-Critical Scenarios for Autonomous Vehicles

Jiangfan Liu, Yongkang Guo, Fangzhi Zhong, Tianyuan Zhang, Zonglei Jing, Siyuan Liang, Jiakai Wang, Mingchuan Zhang, Aishan Liu, Xianglong Liu

机构 * Beihang University(北京航空航天大学) Nanyang Technological University(南洋理工大学) Zhongguancun Laboratory(中关村实验室) Henan University of Science and Technology(河南科技大学)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16213 2025-08-25 cs.CV 78%

MedOmni-45°: A Safety-Performance Benchmark for Reasoning-Oriented LLMs in Medicine

Kaiyuan Ji, Yijin Guo, Zicheng Zhang, Xiangyang Zhu, Yuan Tian, Ning Liu, Guangtao Zhai

专题命中 安全评测 :safety(title,abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01408 2025-08-19 cs.RO cs.CV 78%

From Shadows to Safety: Occlusion Tracking and Risk Mitigation for Urban Autonomous Driving

Korbinian Moller, Luis Schwarzmeier, Johannes Betz

专题命中 安全评测 :safety(title,abstract)

Comments 8 Pages. Submitted to the IEEE Intelligent Vehicles Symposium (IV 2025), Romania

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.16867 2025-08-19 cs.CV 78%

ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

Kaisi Guan, Zhengfeng Lai, Yuchong Sun, Peng Zhang, Wei Liu, Kieran Liu, Meng Cao, Ruihua Song

机构 * Renmin University of China(中国人民大学) Apple(苹果公司)

专题命中 安全评测 :alignment(title,abstract)

Comments International Conference on Computer Vision 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.00399 2025-08-15 cs.CV 78%

iSafetyBench: A video-language benchmark for safety in industrial environment

Raiyaan Abdullah, Yogesh Singh Rawat, Shruti Vyas

机构 * University of Central Florida(中央佛罗里达大学)

专题命中 安全评测 :safety(title,abstract)

Comments Accepted to VISION'25 - ICCV 2025 workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07560 2025-08-12 cs.RO cs.CV 78%

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

Yan Gong, Naibang Wang, Jianli Lu, Xinyu Zhang, Yongsheng Gao, Jie Zhao, Zifan Huang, Haozhi Bai, Nanxin Zeng, Nayu Su, Lei Yang, Ziying Song, Xiaoxi Hu, Xinmin Jiang, Xiaojuan Zhang, Susanto Rahardja

机构 * State Key Laboratory of Robotics and System(机器人系统国家重点实验室) Harbin Institute of Technology(哈尔滨工业大学) State Key Laboratory of Intelligent Green Vehicle and Mobility(智能绿色车辆与移动性国家重点实验室) Tsinghua University(清华大学) the School of Mechanical and Aerospace Engineering(机械与航空航天工程学院) Nanyang Technological University(南洋理工大学) Beijing Key Laboratory of Traffic Data Mining and Embodied Intelligence(北京交通数据挖掘与具身智能重点实验室) Beijing Jiaotong University(北京交通大学) the Institute for Infocomm Research(信息通信研究所) A*STAR the Engineering Cluster(工程集群) the Singapore Institute of Technology(新加坡理工学院)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03834 2025-08-11 cs.RO cs.CV 78%

CARE: Enhancing Safety of Visual Navigation through Collision Avoidance via Repulsive Estimation

Joonkyung Kim, Joonyeol Sim, Woojun Kim, Katia Sycara, Changjoo Nam

机构 * Department of Electronic Engineering, Sogang University(电子工程系,首尔大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学)

专题命中 安全评测 :safety(title,abstract)

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04240 2025-08-07 eess.SP 78%

ChineseEEG-2: An EEG Dataset for Multimodal Semantic Alignment and Neural Decoding during Reading and Listening

Sitong Chen, Beiqianyi Li, Cuilin He, Dongyang Li, Mingyang Wu, Xinke Shen, Song Wang, Xuetao Wei, Xindi Wang, Haiyan Wu, Quanying Liu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10817 2025-08-01 stat.AP 78%

Is Your Model Risk ALARP? Evaluating Prospective Safety-Critical Applications of Complex Models

Domenic Di Francesco, Alan Forrest, Fiona McGarry, Nicholas Hall, Adam Sobey

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.22389 2025-07-31 cs.RO cs.SY eess.SY 78%

Safety Evaluation of Motion Plans Using Trajectory Predictors as Forward Reachable Set Estimators

Kaustav Chakraborty, Zeyuan Feng, Sushant Veer, Apoorva Sharma, Wenhao Ding, Sever Topan, Boris Ivanovic, Marco Pavone, Somil Bansal

机构 * Department of Electrical Engineering, University of Southern California(电气工程系,美国南加州大学) Department of Aeronautics and Astronautics, Stanford University(航空与宇航系,斯坦福大学) NVIDIA Research(NVIDIA研究)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏