arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7945 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7945 篇

2505.17024 2025-05-26 cs.AI q-bio.NC 79%

An Affective-Taxis Hypothesis for Alignment and Interpretability

Eli Sennesh, Maxwell Ramstead

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16792 2025-05-23 cs.CV cs.AI 79%

REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training

Ziqiao Wang, Wangbo Zhao, Yuhao Zhou, Zekai Li, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Pengfei Zhou, Kaipeng Zhang, Zhangyang Wang, Kai Wang, Yang You

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15682 2025-05-22 cs.CL 79%

The Representational Alignment between Humans and Language Models is implicitly driven by a Concreteness Effect

Cosimo Iaia, Bhavin Choksi, Emily Wiebers, Gemma Roig, Christian J. Fiebach

机构 * Goethe University Frankfurt(弗赖堡歌德大学) Center for Brains, Minds and Machines, MIT Hessian.AI(大脑、心智与机器中心,MIT 荷尔斯泰因人工智能) Brain Imaging Center(脑成像中心)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 13 pages, 4 Figures, 1 Table

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15158 2025-05-22 cs.CV cs.CL 79%

ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving

Yunsheng Ma, Burhaneddin Yaman, Xin Ye, Mahmut Yurt, Jingru Luo, Abhirup Mallik, Ziran Wang, Liu Ren

机构 * Bosch Research North America & Bosch Center for Artificial Intelligence (BCAI)(博世北美研究部及博世人工智能中心) Purdue University(普渡大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14341 2025-05-21 cs.CV cs.AI 79%

Replace in Translation: Boost Concept Alignment in Counterfactual Text-to-Image

Sifan Li, Ming Tao, Hao Zhao, Ling Shao, Hao Tang

机构 * Liaoning University(辽宁大学) Nanjing University of Posts and Telecommunications(南京邮电大学) Tsinghua University(清华大学) University of Chinese Academy of Sciences(中国科学院大学) Peking University(北京大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13844 2025-05-21 cs.CL 79%

Improve Language Model and Brain Alignment via Associative Memory

Congchi Yin, Yongpeng Zhang, Xuyun Wen, Piji Li

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学) The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(脑机智能技术重点实验室,教育部)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted by Findings of ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13175 2025-05-20 cs.AI 79%

Enhancing LLMs for Time Series Forecasting via Structure-Guided Cross-Modal Alignment

Siming Sun, Kai Zhang, Xuejun Jiang, Wenchao Meng, Qinmin Yang

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12322 2025-05-20 cs.LG cs.CV 79%

Model alignment using inter-modal bridges

Ali Gholamzadeh, Noor Sajid

机构 * MPI for Biological Cybernetics & University of Tübingen(生物感知研究所及图宾根大学) Kempner Institute, Harvard University(凯普纳研究所及哈佛大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01711 2025-05-20 cs.IR cs.CL 79%

MAPS: Motivation-Aware Personalized Search via LLM-Driven Consultation Alignment

Weicong Qin, Yi Xu, Weijie Yu, Chenglei Shen, Ming He, Jianping Fan, Xiao Zhang, Jun Xu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) University of International Business and Economics(国际商务经济大学) AI Lab at Lenovo Research, Lenovo Group Limited(联想集团研究院人工智能实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments accepted to ACL 2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04379 2025-05-08 cs.MA cs.AI cs.SY eess.SY 79%

Consensus-Aware AV Behavior: Trade-offs Between Safety, Interaction, and Performance in Mixed Urban Traffic

Mohammad Elayan, Wissam Kontar

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

Comments 7 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00322 2025-05-02 cs.RO cs.AI 79%

AI2-Active Safety: AI-enabled Interaction-aware Active Safety Analysis with Vehicle Dynamics

Keshu Wu, Zihao Li, Sixu Li, Xinyue Ye, Dominique Lord, Yang Zhou

机构 * AI2

专题命中 其他安全 :safety(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15785 2025-04-23 cs.AI 79%

WALL-E 2.0: World Alignment by NeuroSymbolic Learning improves World Model-based LLM Agents

Siyu Zhou, Tianyi Zhou, Yijun Yang, Guodong Long, Deheng Ye, Jing Jiang, Chengqi Zhang

机构 * Australian AI Institute, Faculty of Engineering and IT, University of Technology Sydney(澳大利亚人工智能研究所,工程与信息学院,悉尼大学) Department of Computer Science, University of Maryland, College Park(计算机科学系,马里兰大学,College Park) Tencent, China(腾讯,中国)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments Code is available at https://github.com/elated-sawyer/WALL-E

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.00246 2025-04-15 cs.CV cs.LG 79%

ResiDual Transformer Alignment with Spectral Decomposition

Lorenzo Basile, Valentino Maiorca, Luca Bortolussi, Emanuele Rodolà, Francesco Locatello

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments Published in Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23534 2025-04-01 cs.CV cs.AI 79%

BiPVL-Seg: Bidirectional Progressive Vision-Language Fusion with Global-Local Alignment for Medical Image Segmentation

Rafi Ibn Sultan, Hui Zhu, Chengyin Li, Dongxiao Zhu

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.19508 2025-03-26 cs.CV cs.LG 79%

Improved Alignment of Modalities in Large Vision Language Models

Kartik Jangra, Aman Kumar Singh, Yashwani Mann, Geetanjali Rathee

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.14358 2025-03-19 cs.CV cs.LG 79%

RFMI: Estimating Mutual Information on Rectified Flow for Text-to-Image Alignment

Chao Wang, Giulio Franzese, Alessandro Finamore, Pietro Michiardi

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments to appear at ICLR 2025 Workshop on Deep Generative Model in Machine Learning: Theory, Principle and Efficacy

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10211 2025-03-14 cs.CL cs.SD eess.AS 79%

Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation

Henglyu Liu, Andong Chen, Kehai Chen, Xuefeng Bai, Meizhi Zhong, Yuan Qiu, Min Zhang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 12 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06948 2025-03-11 cs.CV cs.AI 79%

Large Language Model Guided Progressive Feature Alignment for Multimodal UAV Object Detection

Wentao Wu, Chenglong Li, Xiao Wang, Bin Luo, Qi Liu

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06708 2025-03-11 cs.CL 79%

Alignment for Efficient Tool Calling of Large Language Models

Hongshen Xu, Zihan Wang, Zichen Zhu, Lei Pan, Xingyu Chen, Lu Chen, Kai Yu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04675 2025-03-07 cs.CL 79%

LLM-guided Plan and Retrieval: A Strategic Alignment for Interpretable User Satisfaction Estimation in Dialogue

Sangyeop Kim, Sohhyung Park, Jaewon Jung, Jinseok Kim, Sungzoon Cho

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments Accepted by NAACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.15259 2025-03-04 cs.CV cs.AI 79%

StarVid: Enhancing Semantic Alignment in Video Diffusion Models via Spatial and SynTactic Guided Attention Refocusing

Yuanhang Li, Qi Mao, Lan Chen, Zhen Fang, Lei Tian, Xinyan Xiao, Libiao Jin, Hua Wu

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.11718 2025-03-04 cs.CL 79%

Converging to a Lingua Franca: Evolution of Linguistic Regions and Semantics Alignment in Multilingual Large Language Models

Hongchuan Zeng, Senyu Han, Lu Chen, Kai Yu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

Comments 16 pages, 11 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.20572 2025-03-03 cs.CV cs.CL 79%

HazardNet: A Small-Scale Vision Language Model for Real-Time Traffic Safety Detection at Edge Devices

Mohammad Abu Tami, Mohammed Elhenawy, Huthaifa I. Ashqar

专题命中 其他安全 :safety(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2307.03373 2025-03-03 cs.CV cs.AI 79%

All in One: Exploring Unified Vision-Language Tracking with Multi-Modal Alignment

Chunhui Zhang, Xin Sun, Yiqian Yang, Li Liu, Qiong Liu, Xi Zhou, Yanfeng Wang

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI

Comments In this version, we corrected some typos

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15140 2025-02-24 cs.CL cs.HC 79%

Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns

Naiming Liu, Shashank Sonkar, Richard G. Baraniuk

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.20727 2025-02-20 cs.LG stat.ML 79%

Faster WIND: Accelerating Iterative Best-of-$N$ Distillation for LLM Alignment

Tong Yang, Jincheng Mei, Hanjun Dai, Zixin Wen, Shicong Cen, Dale Schuurmans, Yuejie Chi, Bo Dai

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.12732 2025-02-19 cs.LG 79%

Circuit Representation Learning with Masked Gate Modeling and Verilog-AIG Alignment

Haoyuan Wu, Haisheng Zheng, Yuan Pu, Bei Yu

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.05454 2025-02-14 cs.RO cs.LG 79%

Temporal Representation Alignment: Successor Features Enable Emergent Compositionality in Robot Instruction Following

Vivek Myers, Bill Chunyuan Zheng, Anca Dragan, Kuan Fang, Sergey Levine

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.01084 2025-02-14 cs.LG cs.SD eess.AS 79%

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis

Weiwei Lin, Chenghan He

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments ICLR 2025

Journal ref ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.20759 2025-02-12 cs.LG cs.CV 79%

Information Theoretic Text-to-Image Alignment

Chao Wang, Giulio Franzese, Alessandro Finamore, Massimo Gallo, Pietro Michiardi

专题命中 其他安全 :alignment(title,abstract);分类 cs.LG

Comments to appear at ICLR25

详情

展开后加载摘要…

URL PDF HTML 收藏