arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 3281 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全训练 3281 篇

2509.08732 2025-09-11 econ.TH 50%

Incentives for Digital Twins: Task-Based Productivity Enhancements with Generative AI

Catherine Wu, Arun Sundararajan

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07304 2025-09-10 eess.SY cs.SY 50%

Distributed Leader-Follower Consensus for Uncertain Multiagent Systems with Time-Triggered Switching of the Communication Network

Armel Koulong, Ali Pakniyat

专题命中 安全训练 :safety(abstract)

Comments Joint submission paper MECC-JDSMC. Accepted for the 2025 Modeling, Estimation and Control Conference (MECC). Currently under review by the ASME Journal of Dynamic Systems, Measurement, and Control (JDSMC)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.04714 2025-09-08 cs.SI 50%

ThumbnailTruth: A Multi-Modal LLM Approach for Detecting Misleading YouTube Thumbnails Across Diverse Cultural Settings

Wajiha Naveed, Zartash Afzal Uzmi, Zafar Ayyub Qazi

专题命中 安全训练 :trustworthy(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.00918 2025-09-03 cs.CR 50%

PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement

Xubin Yue, Zhenhua Xu, Wenpeng Xing, Jiahui Yu, Mohan Li, Meng Han

专题命中 安全训练 :harmlessness(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14235 2025-08-21 cs.RO 50%

SLAM-based Safe Indoor Exploration Strategy

Omar Mostafa, Nikolaos Evangeliou, Anthony Tzes

机构 * Center for Artificial Intelligence \& Robotics (CAIR) New York University Abu Dhabi (NYUAD) United Arab Emirates Robotics \& Intelligent Systems Control Lab NYUAD United Arab Emirates

专题命中 安全训练 :safety(abstract)

Comments 5 pages, 8 figures. Published in the 2025 11th International Conference on Automation, Robotics, and Applications (ICARA)

Journal ref 2025 11th International Conference on Automation, Robotics, and Applications (ICARA), pp. 375-379

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11799 2025-08-19 math.OC cs.RO 50%

Scaling Robust Optimization for Swarms: A Distributed Perspective

Arshiya Taj Abdul, Augustinos D. Saravanos, Evangelos A. Theodorou

机构 * Georgia Institute of Technology(佐治亚理工学院)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10378 2025-08-15 cs.RO 50%

A Semantic-Aware Framework for Safe and Intent-Integrative Assistance in Upper-Limb Exoskeletons

Yu Chen, Shu Miao, Chunyu Wu, Jingsong Mu, Bo OuYang, Xiang Li

机构 * Department of Automation, Tsinghua University(自动化系,清华大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01272 2025-08-15 cs.CV 50%

PromptSafe: Gated Prompt Tuning for Safe Text-to-Image Generation

Zonglei Jing, Xiao Yang, Xiaoqian Li, Siyuan Liang, Aishan Liu, Mingchuan Zhang, Xianglong Liu

机构 * Beihang University(北航) Beijing University of Posts and Telecommunications(北京邮电大学) Taishan University(泰山大学) Nanyang Technological University(南洋理工大学) Henan University of Science and Technology(河南科技大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2209.00334 2025-08-15 cs.RO 50%

Traversability analysis with vision and terrain probing for safe legged robot navigation

Garen Haddeler, Meng Yee Michael Chuah, Yangwei You, Jianle Chan, Albertus H. Adiwahono, Wei Yun Yau, Chee-Meng Chew

专题命中 安全训练 :safety(abstract)

Journal ref Frontiers in Robotics and AI, Volume 9 - 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.16608 2025-08-14 cs.RO cs.SY eess.SY 50%

Barriers on the EDGE: A scalable CBF architecture over EDGE for safe aerial-ground multi-agent coordination

Viswa Narayanan Sankaranarayanan, Achilleas Santi Seisa, Akshit Saradagi, Sumeet Satpute, George Nikolakopoulos

机构 * Robotics and Artificial Intelligence Group of the Department of Computer Science, Electrical and Space Engineering at Luleå University of Technology(鲁内斯大学机器人与人工智能小组)

专题命中 安全训练 :safety(abstract)

Comments 6 pages, 2 figures, first draft of a paper currently under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22792 2025-08-12 cs.CV 50%

Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

Yuxi Zhang, Yueting Li, Xinyu Du, Sibo Wang

机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) University of California, Berkeley(加州大学伯克利分校)

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.24152 2025-08-12 cs.RO 50%

Language-Driven Policy Distillation for Cooperative Driving in Multi-Agent Reinforcement Learning

Jiaqi Liu, Chengkai Xu, Peng Hang, Jian Sun, Wei Zhan, Masayoshi Tomizuka, Mingyu Ding

机构 * Department of Computer Science at University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校计算机科学系) College of Transportation and Key Laboratory of Road and Traffic Engineering, Ministry of Education, Tongji University(同济大学交通学院及交通工程教育部重点实验室) Department of Mechanical Engineering at the University of California, Berkeley(加州大学伯克利分校机械工程系)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21619 2025-07-30 cs.CV 50%

EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO

Wei Guan, Jun Lan, Jian Cao, Hao Tan, Huijia Zhu, Weiqiang Wang

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.21547 2025-07-30 math.OC cs.RO cs.SY eess.SY 50%

Decentralized Modeling of Vehicular Maneuvers and Interactions at Urban Junctions

Saeed Rahmani, Simeon C. Calvert, Bart van Arem

机构 * Delft University of Technology(代尔夫特理工大学)

专题命中 安全训练 :safety(abstract)

Comments Manuscript under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.17868 2025-07-25 eess.SY cs.SY 50%

Safe Reinforcement Learning-based Automatic Generation Control

Amr S. Mohamed, Emily Nguyen, Deepa Kundur

专题命中 安全训练 :safety(abstract)

Comments 5 pages, conference: IEEE Power and Energy Systems General Meeting 2025

Journal ref Mohamed, Amr, Emily Nguyen, and Deepa Kundur. "Safe Reinforcement Learning-based Automatic Generation Control." 2025 IEEE Power & Energy Society General Meeting (PESGM). IEEE, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2207.00477 2025-07-23 cs.CV 50%

Vision-based Conflict Detection within Crowds based on High-Resolution Human Pose Estimation for Smart and Safe Airport

Karan Kheta, Claire Delgove, Ruolin Liu, Adeola Aderogba, Marc-Olivier Pokam, Muhammed Mehmet Unal, Yang Xing, Weisi Guo

专题命中 安全训练 :safety(abstract)

Comments One of the authors has expressed privacy concerns and made a related request

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.06574 2025-07-23 cs.RO 50%

AI Space Cortex: An Experimental System for Future Era Space Exploration

Thomas Touma, Ersin Daş, Erica Tevere, Martin Feather, Ksenia Kolcio, Maurice Prather, Alberto Candela, Ashish Goel, Erik Kramer, Hari Nayar, Lorraine Fesq, Joel W. Burdick

机构 * California Institute of Technology(加州理工学院) Jet Propulsion Laboratory(喷气推进实验室)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14438 2025-07-22 physics.ins-det hep-ex 50%

A GEANT4-Based Simulation of Directional Neutron Detectors Using Liquid Scintillators and Boron Carbide Moderators

J. -H. Chen, M. Mirzakhani, R. Mahapatra, S. Sahoo

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.17646 2025-07-22 cs.PL 50%

Portability of Optimizations from SC to TSO

Akshay Gopalakrishnan, Clark Verbrugge

专题命中 安全训练 :safety(abstract)

Comments Submitted Manuscript. This pre-print has not undergone any post-review modifications/improvements

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12977 2025-07-18 cs.RO 50%

Non-differentiable Reward Optimization for Diffusion-based Autonomous Motion Planning

Giwon Lee, Daehee Park, Jaewoo Jeong, Kuk-Jin Yoon

机构 * Department of Mechanical Engineering, KAIST(韩国科学技术院机械工程系) Department of Electrical Engineering and Computer Science, DGIST(韩国科学技术院电子工程与计算机科学系)

专题命中 安全训练 :safety(abstract)

Comments Accepted at IROS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.12083 2025-07-17 cs.CV cs.RO 50%

Foresight in Motion: Reinforcing Trajectory Prediction with Reward Heuristics

Muleilan Pei, Shaoshuai Shi, Xuesong Chen, Xu Liu, Shaojie Shen

机构 * HKUST(香港科技大学) Voyager Research, Didi Chuxing(维嘉尔研究,滴滴出行) Zhuoyu Technology(筑宇科技)

专题命中 安全训练 :safety(abstract)

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09537 2025-07-15 cs.RO 50%

Self-supervised Pretraining for Integrated Prediction and Planning of Automated Vehicles

Yangang Ren, Guojian Zhan, Chen Lv, Jun Li, Fenghua Liang, Keqiang Li

机构 * Changan Automobile(长安汽车) Tsinghua University(清华大学) Nanyang Technological University(南洋理工大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07019 2025-07-10 econ.GN q-fin.EC 50%

The Post Science Paradigm of Scientific Discovery in the Era of Artificial Intelligence: Modelling the Collapse of Ideation Costs, Epistemic Inversion, and the End of Knowledge Scarcity

Christian William Callaghan

专题命中 安全训练 :alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.23739 2025-07-01 cs.RO cs.CE cs.HC 50%

Validation of AI-Based 3D Human Pose Estimation in a Cyber-Physical Environment

Lisa Marie Otto, Michael Kaiser, Daniel Seebacher, Steffen Müller

专题命中 安全训练 :alignment(abstract)

Comments 6 pages, 5 figures, Preprint for 2025 IEEE IAVVC (International Automated Vehicle Validation Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21693 2025-06-30 cs.SE cs.RO 50%

The DevSafeOps Dilemma: A Systematic Literature Review on Rapidity in Safe Autonomous Driving Development and Operation

Ali Nouri, Beatriz Cabrero-Daniel, Fredrik Törner, Christian Berger

机构 * Chalmers University of Technology, Department of Computer Science(查尔姆斯理工大学计算机科学系) University of Gothenburg, Department of Computer Science(哥德堡大学计算机科学系)

专题命中 安全训练 :safety(abstract)

Comments Accepted for publication in the Journal of Systems and Software (JSS)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16748 2025-06-23 cs.RO cs.MA 50%

A Scalable Post-Processing Pipeline for Large-Scale Free-Space Multi-Agent Path Planning with PiBT

Arjo Chakravarty, Michael X. Grey, M. A. Viraj J. Muthugala, Mohan Rajesh Elara

机构 * Intrinsic Innovation LLC ROAR Lab(ROAR 实验室) Singapore University of Technology and Design(新加坡科技设计大学)

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09042 2025-06-19 cs.CV 50%

Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

Xuanchi Ren, Yifan Lu, Tianshi Cao, Ruiyuan Gao, Shengyu Huang, Amirmojtaba Sabour, Tianchang Shen, Tobias Pfaff, Jay Zhangjie Wu, Runjian Chen, Seung Wook Kim, Jun Gao, Laura Leal-Taixe, Mike Chen, Sanja Fidler, Huan Ling

专题命中 安全训练 :safety(abstract)

Comments Only the core contributors are listed. The full list of contributors can be found in Appendix A of this paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14749 2025-06-18 eess.SY cs.SY 50%

Swarm-STL: A Framework for Motion Planning in Large-Scale, Multi-Swarm Systems

Shiyu Cheng, Luyao Niu, Bhaskar Ramasubramanian, Andrew Clark, Radha Poovendran

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14018 2025-06-18 cs.HC 50%

"I Cannot Write This Because It Violates Our Content Policy": Understanding Content Moderation Policies and User Experiences in Generative AI Products

Lan Gao, Oscar Chen, Rachel Lee, Nick Feamster, Chenhao Tan, Marshini Chetty

专题命中 安全训练 :safety(abstract)

Comments Preprint for USENIX Security 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23549 2025-06-16 cs.SE 50%

LLM-based Property-based Test Generation for Guardrailing Cyber-Physical Systems

Khashayar Etemadi, Marjan Sirjani, Mahshid Helali Moghadam, Per Strandberg, Paul Pettersson

专题命中 安全训练 :safety(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏