arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7945 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7945 篇

2510.04919 2025-10-07 cs.CL cs.AI cs.DB 81%

Do LLMs Align with My Task? Evaluating Text-to-SQL via Dataset Alignment

Davood Rafiei, Morgan Lindsay Heisler, Weiwei Zhang, Mohammadreza Pourreza, Yong Zhang

机构 * University of Alberta(阿尔伯塔大学) Huawei Tech. Canada(华为技术加拿大)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.25253 2025-10-01 cs.LG cs.AI 81%

Knowledge distillation through geometry-aware representational alignment

Prajjwal Bhattarai, Mohammad Amjad, Dmytro Zhylko, Tuka Alhanai

机构 * New York University Abu Dhabi(纽约大学阿布扎赫德分校) New York University(纽约大学) Tandon School of Engineering(Tandon工程学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11939 2025-09-30 eess.SP cs.AI cs.LG 81%

Fine-grained Contrastive Learning for ECG-Report Alignment with Waveform Enhancement

Haitao Li, Che Liu, Zhengyao Ding, Ziyi Liu, Wenqi Shao, Zhengxing Huang

机构 * Zhejiang University(浙江大学) Imperial College London(伦敦帝国学院) Transtek Medical Electronics Co., Ltd(Transtek医疗电子有限公司) Shanghai AI Lab(上海AI实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22697 2025-09-30 cs.CV cs.AI cs.LG 81%

Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment

Abhiroop Chatterjee, Susmita Ghosh

机构 * Jadavpur University(贾瓦帕尔大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted at the IEEE/CVF International Conference on Computer Vision (ICCV 2025), Workshop on Curated Data for Efficient Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20751 2025-09-26 cs.CV cs.AI cs.CL 81%

Seeing Through Words, Speaking Through Pixels: Deep Representational Alignment Between Vision and Language Models

Zoe Wanying He, Sean Trott, Meenakshi Khosla

机构 * Department of Cognitive Science University of California, San Diego(认知科学系,加州大学圣地亚哥分校)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted at EMNLP 2025 (camera-ready)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.19657 2025-09-25 cs.CL cs.AI cs.SI 81%

Large Language Models for Pedestrian Safety: An Application to Predicting Driver Yielding Behavior at Unsignalized Intersections

Yicheng Yang, Zixian Li, Jean Paul Bizimana, Niaz Zafri, Yongfeng Dong, Tianyi Li

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06144 2025-09-24 cs.CL cs.AI 81%

Language Models Resist Alignment: Evidence From Data Compression

Jiaming Ji, Kaile Wang, Tianyi Qiu, Boyuan Chen, Jiayi Zhou, Changye Li, Hantao Lou, Juntao Dai, Yunhuai Liu, Yaodong Yang

机构 * Institute for Artificial Intelligence, Peking University(北京大学人工智能研究院) Beijing Academy of Artificial Intelligence(北京人工智能研究院) School of Computer Science, Peking University(北京大学计算机科学学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12613 2025-09-22 cs.HC cs.AI cs.CY cs.MA 81%

Negotiative Alignment: Embracing Disagreement to Achieve Fairer Outcomes -- Insights from Urban Studies

Rashid Mushkani, Hugo Berard, Shin Koseki

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

Comments 16 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13706 2025-09-18 cs.CL cs.AI 81%

Automated Triaging and Transfer Learning of Incident Learning Safety Reports Using Large Language Representational Models

Peter Beidler, Mark Nguyen, Kevin Lybarger, Ola Holmberg, Eric Ford, John Kang

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.15633 2025-09-16 cs.CL cs.AI 81%

GATEAU: Selecting Influential Samples for Long Context Alignment

Shuzheng Si, Haozhe Zhao, Gang Chen, Yunshui Li, Kangyang Luo, Chuancheng Lv, Kaikai An, Fanchao Qi, Baobao Chang, Maosong Sun

机构 * Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) Peking University(北京大学) DeepLang AI Institute for AI, Tsinghua University(清华大学人工智能研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09931 2025-09-16 cs.LG cs.AI 81%

Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications

Yoon Pyo Lee

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments Accepted for publication in Nuclear Technology. 24 pages, 2 tables, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10344 2025-09-15 cs.CV cs.AI cs.LG 81%

GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography

Yuexi Du, Lihui Chen, Nicha C. Dvornek

机构 * Department of Biomedical Engineering(生物医学工程系) Department of Radiology & Biomedical Imaging(放射科与生物医学成像系) Yale University(耶鲁大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.07588 2025-09-10 cs.CL cs.AI 81%

BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment

Andrey Sakhovskiy, Elena Tutubalina

机构 * AIRI Sber AI ISP RAS Research Center for Trusted AI(俄罗斯科学院信息与系统研究所可信人工智能研究中心)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 9 pages, 1 figure, published in "The 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2025)"

Journal ref Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval (2025). Association for Computing Machinery, 1152-1164

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.03934 2025-09-05 cs.CL cs.AI 81%

SelfAug: Mitigating Catastrophic Forgetting in Retrieval-Augmented Generation via Distribution Self-Alignment

Yuqing Huang, Rongyang Zhang, Qimeng Wang, Chengqiang Lu, Yan Gao, Yi Wu, Yao Hu, Xuyang Zhi, Guiquan Liu, Xin Li, Hao Wang, Enhong Chen

机构 * University of Science and Technology of China(中国科学技术大学) Xiaohongshu Inc.(小红书公司)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21188 2025-09-03 cs.LG cs.CL 81%

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions

Haoze Wu, Cheng Wang, Wenshuo Zhao, Junxian He

机构 * Zhejiang University(浙江大学) National University of Singapore(国立新加坡大学) HKUST(香港科技大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14232 2025-09-03 cs.AI cs.CL 81%

Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment

Antoun Yaacoub, Jérôme Da-Rugna, Zainab Assaghir

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments This paper was presented in the 17th Int. Conf. on Computer Science and Information Technology (ICCSIT 2024), Dubai, United Arab Emirates, 2024, Oct. 23-25. IT was published in the International Journal of Computer Theory and Engineering, vol. 17, no. 3, pp. 114-125, 2025

Journal ref International Journal of Computer Theory and Engineering, vol. 17, no. 3, pp. 114-125, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.20130 2025-08-29 q-bio.QM cs.AI cs.LG 81%

Artificial Intelligence for CRISPR Guide RNA Design: Explainable Models and Off-Target Safety

Alireza Abbaszadeh, Armita Shahlai

机构 * Department of Computer Engineering, Ma.C., Islamic Azad University(计算机工程系,伊斯兰阿兹德大学) Department of Biological Sciences and Technologies, Faculty of Basic Sciences, Islamic Azad University(基础科学学院生物科学与技术系,伊斯兰阿兹德大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 29 pages, 5 figures, 2 tables, 42 cited references

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.16770 2025-08-15 cs.CL cs.AI 81%

LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint

Qianli Ma, Dongrui Liu, Qian Chen, Linfeng Zhang, Jing Shao

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) East China Normal University(华东师范大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL2025 main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08131 2025-08-12 cs.CL cs.AI 81%

Optimal Transport Regularization for Speech Text Alignment in Spoken Language Models

Wenze Xu, Chun Wang, Jiazhen Yu, Sheng Chen, Liang Gao, Weihong Deng

机构 * Mashang Consumer Finance Co., Ltd.(Mashang消费金融有限公司) The University of Sydney(悉尼大学) Macau University of Science and Technology(澳门科学技术大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments To be presented at ACPR 2025 Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07195 2025-08-12 cs.CL cs.AI 81%

Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Semantic Alignment

Yanru Sun, Emadeldeen Eldele, Zongxia Xie, Yucheng Wang, Wenzhe Niu, Qinghua Hu, Chee Keong Kwoh, Min Wu

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06592 2025-08-12 cs.CY cs.AI 81%

Towards Integrated Alignment

Ben Y. Reis, William La Cava

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02762 2025-08-07 cs.LG cs.AI 81%

Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment

Dahun Kim, Anelia Angelova

机构 * Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(title,abstract);分类 cs.AI、cs.LG

Comments COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16104 2025-07-23 cs.CL cs.CV cs.LG 81%

Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models

Yue Li, Xin Yi, Dongsheng Shi, Gerard de Melo, Xiaoling Wang, Linlin Wang

机构 * East China Normal University(华东师范大学) Hasso Plattner Institute/University of Potsdam(哈索普劳特纳研究所/波茨坦大学)

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.LG

Comments ACL 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.16575 2025-07-21 cs.AI cs.LG cs.RO 81%

Self-Improving Safety Performance of Reinforcement Learning Based Driving with Black-Box Verification Algorithms

Resul Dagdanov, Halil Durmus, Nazim Kemal Ure

机构 * ITU Artificial Intelligence and Data Science Research Center(伊斯坦布尔技术大学人工智能与数据科学研究中心) Department of Aeronautical Engineering(航空工程系) Eatron Technologies(Eatron技术公司) Department of Electronics and Communication Engineering(电子与通信工程系) ITU Artificial Intelligence and Data Science Application and Research Center(伊斯坦布尔技术大学人工智能与数据科学应用与研究中心) Department of Computer Engineering(计算机工程系)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

Comments 7 pages, 7 figures, 2 tables, published in IEEE International Conference on Robotics and Automation (ICRA), June 2, 2023, London, UK

Journal ref IEEE International Conference on Robotics and Automation (ICRA), 2023, pp. 5631-5637

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.09037 2025-07-15 cs.CL cs.AI 81%

ALIGN: Prompt-based Attribute Alignment for Reliable, Responsible, and Personalized LLM-based Decision-Making

Bharadwaj Ravichandran, David Joy, Paul Elliott, Brian Hu, Jadie Adams, Christopher Funk, Emily Veenhuis, Anthony Hoogs, Arslan Basharat

机构 * kitware(Kitware公司)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments 10 pages total (including appendix), ICML 2025 Workshop on Reliable and Responsible Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01403 2025-07-14 cs.IR cs.AI cs.CL 81%

Generative Retrieval and Alignment Model: A New Paradigm for E-commerce Retrieval

Ming Pang, Chunyuan Yuan, Xiaoyu He, Zheng Fang, Donghao Xie, Fanyi Qu, Xue Jiang, Changping Peng, Zhangang Lin, Ching Law, Jingping Shao

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by WWW2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07599 2025-07-11 cs.AI cs.CL 81%

Enhancing Vaccine Safety Surveillance: Extracting Vaccine Mentions from Emergency Department Triage Notes Using Fine-Tuned Large Language Models

Sedigh Khademi, Jim Black, Christopher Palmer, Muhammad Javed, Hazel Clothier, Jim Buttery, Gerardo Luis Dimaguila

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI

Comments 5 pages

Journal ref Medinfo 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07572 2025-07-11 cs.CL cs.AI cs.CV 81%

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

Yupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

Comments Accepted by ACL 2025 Main

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00068 2025-07-02 cs.CV cs.AI cs.CL 81%

MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding

Ziqi Zhong, Daniel Tang

机构 * London School of Economics(伦敦经济学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12576 2025-07-01 cs.CL cs.AI 81%

Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders

Ananya Joshi, Celia Cintas, Skyler Speakman

机构 * Carnegie Mellon University(卡内基梅隆大学) IBM Research(IBM研究院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏