arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7945 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7945 篇

2505.12884 2025-07-01 cs.LG cs.AI cs.CV 81%

TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks

Yuanze Hu, Zhaoxin Fan, Xinyu Wang, Gen Li, Ye Qiu, Zhichao Yang, Wenjun Wu, Kejian Wu, Yifan Sun, Xiaotie Deng, Jin Dong

机构 * Beijing Advanced Innovation Center for Future Blockchain and Privacy Computing(北京未来区块链与隐私计算先进创新中心) Beihang University(北京航空航天大学) Hangzhou International Innovation Institute(杭州国际创新研究院) Xreal Renmin University(中国人民大学) Peking University(北京大学) Beijing Academy of Blockchain and Edge Computing (BABEC)(北京区块链与边缘计算研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.19202 2025-07-01 cs.LG cs.AI 81%

Vulnerable Road User Detection and Safety Enhancement: A Comprehensive Survey

Renato M. Silva, Gregorio F. Azevedo, Matheus V. V. Berto, Jean R. Rocha, Eduardo C. Fidelis, Matheus V. Nogueira, Pedro H. Lisboa, Tiago A. Almeida

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 60 pages, 18 tables, 8 figures, citing 370 (up-to-date) papers. Expert Systems With Applications (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.18185 2025-06-24 cs.CL cs.AI 81%

CareLab at #SMM4H-HeaRD 2025: Insomnia Detection and Food Safety Event Extraction with Domain-Aware Transformers

Zihan Liang, Ziwen Pan, Sumon Kanti Dey, Azra Ismail

机构 * CareLab

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

Comments In the Proceedings of the 10th Social Media Mining for Health and Health Real-World Data Workshop and Shared Tasks, co-located with AAAI ICWSM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16473 2025-06-23 cs.HC cs.AI cs.CL 81%

Do We Talk to Robots Like Therapists, and Do They Respond Accordingly? Language Alignment in AI Emotional Support

Sophie Chiang, Guy Laban, Hatice Gunes

机构 * Department of Computer Science Technology, University of Cambridge Cambridge UK Technology, University of Cambridge

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18799 2025-06-19 cs.CL cs.AI 81%

ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models

Hao Chen, Haoze Li, Zhiqing Xiao, Lirong Gao, Qi Zhang, Xiaomeng Hu, Ningtao Wang, Xing Fu, Junbo Zhao

机构 * Zhejiang University(浙江大学) Ant Group(蚂蚁集团)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted@ACL25-findings, 17 pages, 8 figures, 14 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.05929 2025-06-19 cs.LG cs.AI 81%

M3-JEPA: Multimodal Alignment via Multi-gate MoE based on the Joint-Embedding Predictive Architecture

Hongyang Lei, Xiaolong Cheng, Qi Qin, Dan Wang, Kun Fan, Huazhen Huang, Qingqing Gu, Yetao Wu, Zhonglin Jiang, Yong Chen, Luo Ji

机构 * Geely AI Lab(Geely人工智能实验室) Shenzhen Institutes of Advanced Technology(深圳先进技术研究院) Peking University(北京大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 16 pages, 5 figures. ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14786 2025-06-19 cs.LG cs.AI cs.CV 81%

PIPE: Physics-Informed Position Encoding for Alignment of Satellite Images and Time Series

Haobo Li, Eunseo Jung, Zixin Chen, Zhaowei Wang, Yueya Wang, Huamin Qu, Alexis Kai Hon Lau

机构 * Department of Computer Science & Engineering(香港理工大学计算机科学与工程系) Division of Environment & Sustainability(香港理工大学环境与可持续发展学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11769 2025-06-16 cs.CL cs.LG 81%

Long-Short Alignment for Effective Long-Context Modeling in LLMs

Tianqi Du, Haotian Huang, Yifei Wang, Yisen Wang

机构 * State Key Lab of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University, China(人工智能国家重点实验室,智能科学与技术学院,北京大学) MIT CSAIL, USA(麻省理工学院计算机科学与人工智能实验室) Institute for Artificial Intelligence, Peking University, China(人工智能研究院,北京大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.08990 2025-06-11 cs.CV cs.AI cs.LG 81%

Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models

Chenyu Lian, Hong-Yu Zhou, Dongyun Liang, Jing Qin, Liansheng Wang

机构 * School of Informatics, Xiamen University(厦门大学信息学院) Center for Smart Health, School of Nursing, The Hong Kong Polytechnic University(香港理工大学护理学院智能健康中心) Department of Biomedical Informatics, Harvard Medical School(哈佛医学院生物医学信息学系) Department of Radiology, Zhongshan Hospital (Xiamen), Fudan University(复旦大学中山医院放射科) National Institute for Data Science in Health and Medicine(医学与数据科学国家研究院) Department of Computer Science, School of Informatics, Xiamen University(厦门大学信息学院计算机科学系)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments TMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07276 2025-06-10 cs.LG cs.AI 81%

Tokenized Bandit for LLM Decoding and Alignment

Suho Shin, Chenghao Yang, Haifeng Xu, Mohammad T. Hajiaghayi

机构 * University of Maryland(马里兰大学) University of Chicago(芝加哥大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments To appear at ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05451 2025-06-09 cs.SE cs.AI cs.CL 81%

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety

Seongmin Lee, Aeree Cho, Grace C. Kim, ShengYun Peng, Mansi Phute, Duen Horng Chau

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

Comments 31 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14830 2025-06-03 cs.CL cs.AI 81%

Middle-Layer Representation Alignment for Cross-Lingual Transfer in Fine-Tuned LLMs

Danni Liu, Jan Niehues

机构 * Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.00242 2025-06-03 cs.AI cs.CL 81%

Whispers of Many Shores: Cultural Alignment through Collaborative Cultural Expertise

Shuai Feng, Wei-Chuang Chan, Srishti Chouhan, Junior Francisco Garcia Ayala, Srujananjali Medicherla, Kyle Clark, Mingwei Shi

机构 * Arizona State University(亚利桑那州立大学) National Taiwan University(台湾国立大学) Carnegie Mellon University(卡内基梅隆大学) New York University(纽约大学) Indian Institute of Technology Hyderabad(印度海得拉巴理工学院) Minitab(Minitab公司) Trinity College Dublin(都柏林圣三一学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 14 main pages;8 page appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22907 2025-05-30 cs.CY cs.CL 81%

Conversational Alignment with Artificial Intelligence in Context

Rachel Katharine Sterken, James Ravi Kirkpatrick

机构 * University of Hong Kong(香港大学) University of Oxford(牛津大学) Magdalen College, Oxford(牛津大学玛格丽特学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.CY

Comments 20 pages, to be published in Philosophical Perspectives

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22104 2025-05-29 cs.AI cs.LG cs.LO cs.RO cs.SY eess.SY 81%

Efficient Dynamic Shielding for Parametric Safety Specifications

Davide Corsi, Kaushik Mallik, Andoni Rodriguez, Cesar Sanchez

机构 * University of California, Irvine(加州大学尔湾分校) IMDEA Software Institute(IMDEA软件研究所)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09949 2025-05-16 cs.LG cs.CL stat.AP 81%

Advanced Crash Causation Analysis for Freeway Safety: A Large Language Model Approach to Identifying Key Contributing Factors

Ahmed S. Abdelrahman, Mohamed Abdel-Aty, Samgyu Yang, Abdulrahman Faden

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17692 2025-05-13 cs.CL cs.LG 81%

From Distributional to Overton Pluralism: Investigating Large Language Model Alignment

Thom Lake, Eunsol Choi, Greg Durrett

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

Comments NAACL 2025 (Main Conference)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.04939 2025-05-09 cs.LG cs.AI 81%

Structural Alignment in Link Prediction

Jeffrey Seathrún Sardina

机构 * Trinity College Dublin, the University of Dublin(三一学院都柏林,都柏林大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Ph.D. thesis submitted to Trinity College Dublin

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01162 2025-05-05 cs.CL cs.AI 81%

On the Limitations of Steering in Language Model Alignment

Chebrolu Niranjan, Kokil Jaidka, Gerard Christopher Yeo

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.00350 2025-05-02 cs.LG cs.AI 81%

Optimizing Deep Neural Networks using Safety-Guided Self Compression

Mohammad Zbeeb, Mariam Salman, Mohammad Bazzi, Ammar Mohanna

机构 * Department of Electrical and Computer Engineering, American University of Beirut (AUB)(电气与计算机工程系,贝鲁特美国大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments A Preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12717 2025-04-18 cs.CV cs.AI cs.LG 81%

Post-pre-training for Modality Alignment in Vision-Language Foundation Models

Shin'ya Yamaguchi, Dewei Feng, Sekitoshi Kanai, Kazuki Adachi, Daiki Chijiwa

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted to CVPR 2025; Code: https://github.com/yshinya6/clip-refine

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09063 2025-04-15 cs.LG cs.AI 81%

A Practical Approach to using Supervised Machine Learning Models to Classify Aviation Safety Occurrences

Bryan Y. Siow

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 9 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.17875 2025-04-09 cs.CL cs.AI 81%

Understanding Layer Significance in LLM Alignment

Guangyuan Shi, Zexin Lu, Xiaoyu Dong, Wenlong Zhang, Xuanyu Zhang, Yujie Feng, Xiao-Ming Wu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.01551 2025-04-04 cs.MA cs.AI cs.LG cs.NI cs.SY eess.SY 81%

Safety-Aware Multi-Agent Learning for Dynamic Network Bridging

Raffaele Galliera, Konstantinos Mitsopoulos, Niranjan Suri, Raffaele Romagnoli

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 8 pages, 18 equations, 4 figures, 1 algorithm, and 1 table

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01916 2025-04-03 cs.CV cs.AI cs.CL 81%

FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs

Mothilal Asokan, Kebin Wu, Fatima Albreiki

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.18119 2025-03-28 cs.CV cs.AI cs.LG 81%

Multi-View and Multi-Scale Alignment for Contrastive Language-Image Pre-training in Mammography

Yuexi Du, John Onofrey, Nicha C. Dvornek

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments This paper is accepted by IPMI 2025 for Oral Presentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.09960 2025-03-14 cs.LG cs.AI 81%

Optimizing Fire Safety: Reducing False Alarms Using Advanced Machine Learning Techniques

Muhammad Hassan Jamal, Abdulwahab Alazeb, Shahid Allah Bakhsh, Wadii Boulila, Syed Aziz Shah, Aizaz Ahmad Khattak, Muhammad Shahbaz Khan

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07751 2025-03-11 cs.AI cs.CY 81%

Rethinking AI Cultural Alignment

Michal Bravansky, Filip Trhlik, Fazl Barez

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.04792 2025-03-10 cs.CL cs.AI 81%

Cross-linguistic disagreement as a conflict of semantic alignment norms in multilingual AI~Linguistic Diversity as a Problem for Philosophy, Cognitive Science, and AI~

Masaharu Mizumoto, Dat Tien Nguyen, Justin Sytsma, Mark Alfano, Yu Izumi, Koji Fujita, Nguyen Le Minh

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01881 2025-03-05 cs.LG cs.AI 81%

Mapping representations in Reinforcement Learning via Semantic Alignment for Zero-Shot Stitching

Antonio Pio Ricciardi, Valentino Maiorca, Luca Moschella, Riccardo Marin, Emanuele Rodolà

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments 11 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏