arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 9324 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 9324 篇

2507.21741 2025-07-30 cs.CV cs.MM 78%

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

Shaojun E, Yuchen Yang, Jiaheng Wu, Yan Zhang, Tiejun Zhao, Ziyan Chen

机构 * Global Tone Communication Technology Co., Ltd.(全球 tone 通信技术有限公司) Faculty of computing, Harbin Institute of Technology(哈尔滨工业大学计算机学院) School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机科学与技术学院)

专题命中 安全评测 :alignment(title,abstract)

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21145 2025-07-30 cs.CR 78%

Leveraging Trustworthy AI for Automotive Security in Multi-Domain Operations: Towards a Responsive Human-AI Multi-Domain Task Force for Cyber Social Security

Vita Santa Barletta, Danilo Caivano, Gabriel Cellammare, Samuele del Vescovo, Annita Larissa Sciacovelli

专题命中 安全评测 :trustworthy(title,abstract)

Comments 13 pages, 6 figures, 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09816 2025-07-21 cs.CV 78%

Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment

Angelos Zavras, Dimitrios Michail, Begüm Demir, Ioannis Papoutsis

机构 * organization= Orion Lab, National Observatory of Athens \& National Technical University of Athens , country= Greece organization= Department of Informatics \& Telematics, Harokopio University of Athens , country= Greece organization= Faculty of Electrical Engineering organization= BIFOLD - Berlin Institute for the Foundations of Learning

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted at the ISPRS Journal of Photogrammetry and Remote Sensing. Our code implementation and weights for all experiments are publicly available at https://github.com/Orion-AI-Lab/MindTheModalityGap

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12377 2025-07-15 cs.CV 78%

Alignment and Adversarial Robustness: Are More Human-Like Models More Secure?

Blaine Hoak, Kunyang Li, Patrick McDaniel

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted to International Workshop on Security and Privacy-Preserving AI/ML (SPAIML) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02137 2025-07-11 cs.SE 78%

Towards Trustworthy Sentiment Analysis in Software Engineering: Dataset Characteristics and Tool Selection

Martin Obaidi, Marc Herrmann, Jil Klünder, Kurt Schneider

专题命中 安全评测 :trustworthy(title,abstract)

Comments This paper has been accepted at the RETRAI workshop of the 33rd IEEE International Requirements Engineering Workshop (REW 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19052 2025-06-25 cs.CR 78%

Trustworthy Artificial Intelligence for Cyber Threat Analysis

Shuangbao Paul Wang, Paul Mullin

专题命中 安全评测 :trustworthy(title,abstract)

Journal ref Springer Lecture Note in Networks and Systems. 978-3-031-16071-4,Vol I, LNNS 542. pp 493-504. 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18938 2025-06-25 cs.CV cs.SY eess.SY 78%

Bird's-eye view safety monitoring for the construction top under the tower crane

Yanke Wang, Yu Hin Ng, Haobo Liang, Ching-Wei Chang, Hao Chen

机构 * Hong Kong Center for Construction Robotics(香港建设机器人中心) The Hong Kong University of Science and Technology(香港理工大学)

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17812 2025-06-24 cs.SE 78%

Is Your Automated Software Engineer Trustworthy?

Noble Saji Mathews, Meiyappan Nagappan

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13929 2025-06-11 cs.HC 78%

Awake at the Wheel: Enhancing Automotive Safety through EEG-Based Fatigue Detection

Gourav Siddhad, Sayantan Dey, Partha Pratim Roy, Masakazu Iwamura

专题命中 安全评测 :safety(title,abstract)

Comments 7 Pages, 2 Figure, 1 Table

Journal ref International Conference on Pattern Recognition (ICPR) 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07633 2025-06-10 cs.RO 78%

Blending Participatory Design and Artificial Awareness for Trustworthy Autonomous Vehicles

Ana Tanevska, Ananthapathmanabhan Ratheesh Kumar, Arabinda Ghosh, Ernesto Casablanca, Ginevra Castellano, Sadegh Soudjani

机构 * Uppsala University(乌普萨拉大学) Max Planck Institute for Software Systems(马克斯·普朗克软件系统研究所) Newcastle University(新castle大学)

专题命中 安全评测 :trustworthy(title,abstract)

Comments Submitted to IEEE RO-MAN 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.01667 2025-06-10 cs.CV 78%

VProChart: Answering Chart Question through Visual Perception Alignment Agent and Programmatic Solution Reasoning

Muye Huang, Lingling Zhang, Lai Han, Wenjun Wu, Xinyu Zhang, Jun Liu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11077 2025-05-30 eess.SY cs.SY 78%

LLM-Enhanced Symbolic Control for Safety-Critical Applications

Amir Bayat, Alessandro Abate, Necmiye Ozay, Raphael M. Jungers

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20136 2025-05-27 cs.SE cs.CR 78%

Engineering Trustworthy Machine-Learning Operations with Zero-Knowledge Proofs

Filippo Scaramuzza, Giovanni Quattrocchi, Damian A. Tamburri

专题命中 安全评测 :trustworthy(title,abstract)

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18730 2025-05-27 cs.CV 78%

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

Wenchao Zhang, Jiahe Tian, Runze He, Jizhong Han, Jiao Dai, Miaomiao Feng, Wei Mi, Xiaodan Zhang

机构 * Institute of Information Engineering, Chinese Academy of Sciences(信息工程研究所,中国科学院)

专题命中 安全评测 :alignment(title,abstract)

Comments Code: https://github.com/smile365317/ABP

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17727 2025-05-26 cs.CV 78%

SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain

Jiawei Zhou, Linye Lyu, Zhuotao Tian, Cheng Zhuo, Yu Li

专题命中 安全评测 :safety(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10917 2025-05-20 cs.CV 78%

VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization

Mingxiao Li, Na Su, Fang Qu, Zhizhou Zhong, Ziyang Chen, Yuan Li, Zhaopeng Tu, Xiaolong Li

机构 * Tencent(腾讯)

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11845 2025-05-20 cs.CV 78%

ElderFallGuard: Real-Time IoT and Computer Vision-Based Fall Detection System for Elderly Safety

Tasrifur Riahi, Md. Azizul Hakim Bappy, Md. Mehedi Islam

机构 * Institute of Information and Communicaton Technology, Bangladesh University of Engineering Technology(信息与通信技术学院,孟加拉国工程科技大学) Dept. of Electronics and Communication Engineering, Hajee Mohammad Danesh Science and Technology University(电子与通信工程系,海杰穆罕默德丹尼什科学与技术大学)

专题命中 安全评测 :safety(title,abstract)

Comments 9 page, 1 table, 5 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.13707 2025-05-13 eess.SY cs.RO cs.SY 78%

Safety-Critical Formation Control of Non-Holonomic Multi-Robot Systems in Communication-Limited Environments

Vishrut Bohara, Siavash Farzan

机构 * Robotics Engineering Department, Worcester Polytechnic Institute(沃斯彻斯特理工学院机器人工程系) Electrical Engineering Department, California Polytechnic State University(加州州立大学弗雷斯诺分校电子工程系)

专题命中 安全评测 :safety(title,abstract)

Comments Under review. Video demonstration: https://vimeo.com/1075016147/41612f2f8c

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02598 2025-05-06 cs.RO cs.SY eess.SY 78%

LiDAR-Inertial SLAM-Based Navigation and Safety-Oriented AI-Driven Control System for Skid-Steer Robots

Mehdi Heydari Shahna, Eemil Haaparanta, Pauli Mustalahti, Jouni Mattila

机构 * Faculty of Engineering and Natural Sciences, Tampere University(工程与自然科学学院,塔尔库大学)

专题命中 安全评测 :safety(title,abstract)

Comments This paper has been submitted in the IEEE CDC 2025 for potential presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17968 2025-04-28 cs.RO 78%

Virtual Roads, Smarter Safety: A Digital Twin Framework for Mixed Autonomous Traffic Safety Analysis

Hao Zhang, Ximin Yue, Kexin Tian, Sixu Li, Keshu Wu, Zihao Li, Dominique Lord, Yang Zhou

机构 * Zachry Department of Civil and Environmental Engineering, Texas A&M University(土木与环境工程系,德克萨斯A&M大学) Department of Landscape Architecture and Urban Planning(景观建筑与城市规划系)

专题命中 安全评测 :safety(title,abstract)

Comments 14 pages, 18 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12018 2025-04-17 cs.CV 78%

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching

Xinli Yue, JianHui Sun, Junda Lu, Liangchao Yao, Fan Xia, Tianyi Wang, Fengyun Rao, Jing Lyu, Yuetang Deng

专题命中 安全评测 :alignment(title,abstract)

Comments Accepted to CVPR 2025 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09757 2025-04-15 cs.CR 78%

Alleviating the Fear of Losing Alignment in LLM Fine-tuning

Kang Yang, Guanhong Tao, Xun Chen, Jun Xu

专题命中 安全评测 :alignment(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07556 2025-04-11 cs.CV 78%

TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs

Zijian Zhang, Xuhui Zheng, Xuecheng Wu, Chong Peng, Xuezhi Cao

专题命中 安全评测 :alignment(title,abstract)

Comments 10 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02406 2025-04-04 cs.NI 78%

Lifecycle Management of Trustworthy AI Models in 6G Networks: The REASON Approach

Juan Parra-Ullauri, Xueqing Zhou, Shadi Moazzeni, Rasheed Hussain, Xenofon Vasilakos, Yulei Wu, Renjith Baby, M M Hassan Mahmud, Gabriele Incorvaia, Darryl Hond, Hamid Asgari, Andrea Tassi, Daniel Warren, Dimitra Simeonidou

专题命中 安全评测 :trustworthy(title,abstract)

Journal ref IEEE Wireless Communications, vol. 32, no. 2, pp. 42-51, April 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02017 2025-04-04 cs.SE 78%

Enhancing LLMs in Long Code Translation through Instrumentation and Program State Alignment

Li Xin-Ye, Du Ya-Li, Li Ming

专题命中 安全评测 :alignment(title,abstract)

Comments 20 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.18609 2025-03-28 cs.CV 78%

Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models

Jinhui Yi, Syed Talal Wasim, Yanan Luo, Muzammal Naseer, Juergen Gall

专题命中 安全评测 :alignment(title,abstract)

Comments CVPR 2025 camera-ready version

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.03534 2025-03-26 cs.SE cs.SY eess.SY 78%

Simulation-Based Application of Safety of The Intended Functionality to Mitigate Foreseeable Misuse in Automated Driving Systems

Milin Patel, Rolf Jung

专题命中 安全评测 :safety(title,abstract)

Comments SAE MobilityRxiv Preprint, 2023

Journal ref SAE MobilityRxiv Preprint, 2023

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15522 2025-03-21 cs.HC 78%

"I don't like things where I do not have control": Participants' Experience of Trustworthy Interaction with Autonomous Vehicles

Ana Tanevska, Katie Winkle, Ginevra Castellano

专题命中 安全评测 :trustworthy(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09594 2025-03-13 cs.CV cs.RO 78%

SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

Katrin Renz, Long Chen, Elahe Arani, Oleg Sinavski

专题命中 安全评测 :alignment(title,abstract)

Comments CVPR 2025. 1st Place @ CARLA Challenge 2024. Challenge tech report (preliminary version of SimLingo): arXiv:2406.10165

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11087 2025-03-05 cs.CV 78%

Locality Alignment Improves Vision-Language Models

Ian Covert, Tony Sun, James Zou, Tatsunori Hashimoto

专题命中 安全评测 :alignment(title,abstract)

Comments ICLR 2025 Camera-Ready

详情

展开后加载摘要…

URL PDF HTML 收藏