arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7978 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7978 篇

2505.18956 2025-06-11 cs.CV cs.AI cs.LG cs.MM 62%

How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation

Yining Pan, Qiongjie Cui, Xulei Yang, Na Zhao

机构 * Singapore University of Technology and Design (SUTD)(新加坡科技设计大学) Institute for Infocomm Research (I2R), A*STAR, Singapore(信息与通信研究院(I2R))

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted at the 2025 International Conference on Machine Learning (ICML)

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14898 2025-06-11 cs.CL cs.AI cs.IR 62%

Retrieval-augmented systems can be dangerous medical communicators

Lionel Wong, Ayman Ali, Raymond Xiong, Shannon Zeijang Shen, Yoon Kim, Monica Agrawal

机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室) Duke University(杜克大学) Stanford University(斯坦福大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments Position paper in Proceedings of the 42 nd International Conference on Machine Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10997 2025-06-10 cs.CL cs.AI cs.CV 62%

RONA: Pragmatically Diverse Image Captioning with Coherence Relations

Aashish Anantha Ramakrishnan, Aadarsh Anantha Ramakrishnan, Dongwon Lee

机构 * The Pennsylvania State University(宾夕法尼亚州立大学) National Institute of Technology, Tiruchirappalli(特里奇里帕利理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted in the NAACL Fourth Workshop on Intelligent and Interactive Writing Assistants (In2Writing), Albuquerque, New Mexico, May 2025, https://in2writing.glitch.me

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.06355 2025-06-10 cs.CY cs.CE cs.CL cs.CV 62%

LLMs as World Models: Data-Driven and Human-Centered Pre-Event Simulation for Disaster Impact Assessment

Lingyao Li, Dawei Li, Zhenhui Ou, Xiaoran Xu, Jingxiao Liu, Zihui Ma, Runlong Yu, Min Deng

机构 * University of South Florida(佛罗里达州立大学) Arizona State University(亚利桑那州立大学) Massachusetts Institute of Technology(麻省理工学院) New York University(纽约大学) University of Alabama(阿拉巴马大学) Texas Tech University(德克萨斯科技大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11586 2025-06-10 cs.AI cs.CL 62%

Broaden your SCOPE! Efficient Multi-turn Conversation Planning for LLMs with Semantic Space

Zhiliang Chen, Xinyuan Niu, Chuan-Sheng Foo, Bryan Kian Hsiang Low

机构 * Department of Computer Science, National University of Singapore(新加坡国立大学计算机科学系) Institute for Infocomm Research (I2R), A*STAR, Singapore(新加坡A*STAR信息与通信研究所) Centre for Frontier AI Research (CFAR), A*STAR, Singapore(新加坡A*STAR前沿人工智能研究中心)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI

Comments ICLR 2025 Spotlight

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05057 2025-06-06 cs.CL cs.AI 62%

TALL -- A Trainable Architecture for Enhancing LLM Performance in Low-Resource Languages

Moshe Ofer, Orel Zamler, Amos Azaria

机构 * Ariel University(阿里尔大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03576 2025-06-05 cs.CL cs.AI 62%

KG-BiLM: Knowledge Graph Embedding via Bidirectional Language Models

Zirui Chen, Xin Wang, Zhao Li, Wenbin Guo, Dongxiao He

机构 * College of Intelligence and Computing(智能与计算学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03174 2025-06-05 cs.CV cs.AI cs.LG 62%

Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition Tasks

Koki Matsuishi, Kosuke Ukita, Tsuyoshi Okita

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 25 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16933 2025-06-05 cs.LG cs.CL cs.CV 62%

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Zebin You, Shen Nie, Xiaolu Zhang, Jun Hu, Jun Zhou, Zhiwu Lu, Ji-Rong Wen, Chongxuan Li

机构 * Gaoling School of AI, Renmin University of China(中国人民大学人工智能学院) Beijing Key Laboratory of Research on Large Models and Intelligent Governance(北京大型模型与智能治理研究重点实验室) Engineering Research Center of Next-Generation Intelligent Search and Recommendation, MOE(下一代智能搜索与推荐工程技术研究中心,教育部) Ant Group(蚂蚁集团)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

Comments Project page and codes: \url{https://ml-gsai.github.io/LLaDA-V-demo/}

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10360 2025-06-05 cs.LG cs.AI stat.AP 62%

FactsR: A Safer Method for Producing High Quality Healthcare Documentation

Victor Petrén Bach Hansen, Lasse Krogsbøll, Jonas Lyngsø, Mathias Baltzersen, Andreas Motzfeldt, Kevin Pelgrims, Lars Maaløe

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.02165 2025-06-05 cs.AI cs.CL 62%

A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization

Yucheng Chu, Hang Li, Kaiqi Yang, Harry Shomer, Hui Liu, Yasemin Copur-Gencturk, Jiliang Tang

机构 * Michigan State University(密歇根州立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments EDM 2025 Long Paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03143 2025-06-04 cs.CL cs.AI cs.CV 62%

GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

Qianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang, Jianwei Yang, Huiqiang Jiang, Jian Mu, Baolin Peng, Bo Qiao, Reuben Tan, Si Qin, Lars Liden, Qingwei Lin, Huan Zhang, Tong Zhang, Jianbing Zhang, Dongmei Zhang, Jianfeng Gao

机构 * Microsoft(微软公司) Nanjing University(南京大学) University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.02669 2025-06-03 cs.CV cs.CL cs.LG 62%

Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?

Simon Park, Abhishek Panigrahi, Yun Cheng, Dingli Yu, Anirudh Goyal, Sanjeev Arora

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01450 2025-06-03 cs.LG cs.AI 62%

ShaTS: A Shapley-based Explainability Method for Time Series Artificial Intelligence Models applied to Anomaly Detection in Industrial Internet of Things

Manuel Franco de la Peña, Ángel Luis Perales Gómez, Lorenzo Fernández Maimó

机构 * Departamento de Ingeniería y Tecnología de Computadores, University of Murcia(计算机工程与技术部门,穆尔西亚大学)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

Comments 22 pages;16 figures;Submitted to Elsevier (Information Fusion)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.16344 2025-06-03 eess.AS cs.AI cs.CL cs.SD 62%

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning

Rajath Rao, Adithya Ganesan, Oscar Kjell, Jonah Luby, Akshay Raghavan, Scott Feltman, Whitney Ringwald, Ryan L. Boyd, Benjamin Luft, Camilo Ruggero, Neville Ryant, Roman Kotov, H. Andrew Schwartz

机构 * Stony Brook University(石溪大学) University of Minnesota(明尼苏达大学) University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 16 pages, 8 figures, ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08838 2025-06-03 cs.CL cs.AI 62%

SD$^2$: Self-Distilled Sparse Drafters

Mike Lasby, Nish Sinnadurai, Valavan Manohararajah, Sean Lie, Yani Ioannou, Vithursan Thangarasa

机构 * Cerebras Systems Inc.(Cerebras系统公司) Schulich School of Engineering, University of Calgary(卡莱尔大学施乐工程学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 24 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.17645 2025-06-03 cs.SD cs.AI cs.CL eess.AS 62%

SongComposer: A Large Language Model for Lyric and Melody Generation in Song Composition

Shuangrui Ding, Zihan Liu, Xiaoyi Dong, Pan Zhang, Rui Qian, Junhao Huang, Conghui He, Dahua Lin, Jiaqi Wang

机构 * The Chinese University of Hong Kong(中国香港大学) Beihang University(北航) Shanghai AI Laboratory(上海人工智能实验室) CPII under InnoHK(创新香港下的CPII)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments ACL 2025 main. project page: https://pjlab-songcomposer.github.io/ code: https://github.com/pjlab-songcomposer/songcomposer

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24255 2025-06-02 cs.CL cs.AI cs.HC 62%

Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games

Neemesh Yadav, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim

机构 * Singapore Management University(新加坡管理大学) Australian National University(澳大利亚国立大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments 17 pages, 1 figure, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.05211 2025-06-02 cs.CV cs.AI cs.CL 62%

VITA: Towards Open-Source Interactive Omni Multimodal LLM

Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Yuhang Dai, Meng Zhao, Yi-Fan Zhang, Shaoqi Dong, Yangze Li, Xiong Wang, Haoyu Cao, Di Yin, Long Ma, Xiawu Zheng, Rongrong Ji, Yunsheng Wu, Ran He, Caifeng Shan, Xing Sun

机构 * NJU(南京大学) Tencent Youtu Lab(腾讯计算机系统有限公司视觉实验室) XMU(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Project Page: https://vita-home.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23034 2025-05-30 cs.AI cs.LG 62%

Case-Based Reasoning Enhances the Predictive Power of LLMs in Drug-Drug Interaction

Guangyi Liu, Yongqi Zhang, Xunyuan Liu, Quanming Yao

机构 * Department of Electronic Engineering, Tsinghua University(清华大学电子工程系) The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.15272 2025-05-30 cs.CL cs.AI cs.IR 62%

SimGRAG: Leveraging Similar Subgraphs for Knowledge Graphs Driven Retrieval-Augmented Generation

Yuzheng Cai, Zhenyue Guo, Yiwen Pei, Wanrui Bian, Weiguo Zheng

机构 * Fudan University(复旦大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments accepted by ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.16014 2025-05-30 cs.LG cs.AI 62%

OmniArch: Building Foundation Model For Scientific Computing

Tianyu Chen, Haoyi Zhou, Ying Li, Hao Wang, Chonghan Gao, Rongye Shi, Shanghang Zhang, Jianxin Li

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21591 2025-05-29 cs.LG cs.AI 62%

Pioneering 4-Bit FP Quantization for Diffusion Models: Mixup-Sign Quantization and Timestep-Aware Fine-Tuning

Maosen Zhao, Pengtao Chen, Chong Yu, Yan Wen, Xudong Tan, Tao Chen

机构 * School of Information Science and Technology, Fudan University(信息科学与技术学院,复旦大学) Academy for Engineering and Technology, Fudan University(工程与技术学院,复旦大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21513 2025-05-29 cs.CV cs.AI cs.LG 62%

Enhancing Vision Transformer Explainability Using Artificial Astrocytes

Nicolas Echevarrieta-Catalan, Ana Ribas-Rodriguez, Francisco Cedron, Odelia Schwartz, Vanessa Aguiar-Pulido

机构 * University of Miami(迈阿密大学) University of A Coruña(阿罗乌纳大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments LXCV Workshop at IEEE / CVF Computer Vision and Pattern Recognition Conference (CVPR) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.11107 2025-05-29 cs.LG cs.AI 62%

Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL

Wei Yao, Wenkai Yang, Ziqiao Wang, Yankai Lin, Yong Liu

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学人工智能学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments Accepted by ACL 2025 (Findings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21362 2025-05-28 cs.CL cs.AI cs.HC 62%

Evaluating LLM Adaptation to Sociodemographic Factors: User Profile vs. Dialogue History

Qishuai Zhong, Zongmin Li, Siqi Fan, Aixin Sun

机构 * Nanyang Technological University, Singapore(南洋理工大学,新加坡) University of Electronic Science and Technology of China, Chengdu, China(电子科技大学,中国)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.21301 2025-05-28 cs.CL cs.AI 62%

How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian

Andrea Pedrotti, Giulia Rambelli, Caterina Villani, Marianna Bolognesi

机构 * Università di Bologna(博洛尼亚大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

Comments Accepted at ACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20405 2025-05-28 cs.CV cs.AI cs.CL cs.MM 62%

What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models

Lorenzo Baraldi, Davide Bucciarelli, Federico Betti, Marcella Cornia, Lorenzo Baraldi, Nicu Sebe, Rita Cucchiara

机构 * University of Modena and Reggio Emilia(摩德纳和雷吉奥艾米利亚大学) University of Pisa(比萨大学) University of Trento(特伦托大学) IIT-CNR

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20053 2025-05-27 cs.CV cs.AI cs.CL cs.MM 62%

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion

Zheqi Lv, Junhao Chen, Qi Tian, Keting Yin, Shengyu Zhang, Fei Wu

机构 * Zhejiang University(浙江大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20029 2025-05-27 q-bio.NC cs.AI cs.LG 62%

Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)

Subba Reddy Oota, Akshett Jindal, Ishani Mondal, Khushbu Pahwa, Satya Sai Srinath Namburi, Manish Shrivastava, Maneesh Singh, Bapi S. Raju, Manish Gupta

机构 * Technische Universität Berlin(柏林技术大学) IIIT Hyderabad(海得拉巴国家理工学院) Univ of Maryland(马里兰大学) Rice Univ(Rice 大学) Univ of Wisconsin - Madison(威斯康星大学麦迪逊分校) Spector Inc(Spector 公司) Microsoft(微软公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

Comments 30 pages, 22 figures, The Thirteenth International Conference on Learning Representations, ICLR-2025, Singapore. https://openreview.net/pdf?id=xkgfLXZ4e0

Journal ref ICLR-2025, Singapore

详情

展开后加载摘要…

URL PDF HTML 收藏