arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7937 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7937 篇

2601.10160 2026-02-23 cs.CL cs.AI cs.LG 82%

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

对齐预训练:人工智能 discourse 导致自我实现的(不)对齐

Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa, Kyle O'Brien

机构 * Geodesic Research(Geodesic研究机构) UK AI Security Institute(英国人工智能安全研究所) University of Cambridge(剑桥大学) University of Oxford(牛津大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文通过预训练不同量级的AI discourse,发现其对下游对齐有显著影响,表明预训练数据塑造对齐先验的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12580 2026-02-13 cs.MA 82%

Semantic Fusion: Verifiable Alignment in Decentralized Multi-Agent Systems

语义融合:去中心化多智能体系统的可验证对齐

Sofiya Zaichyk

专题命中 其他安全 :alignment(title,abstract);safety(abstract)

AI总结 语义融合通过本地本体验证实现去中心化多智能体系统的可验证语义对齐,支持动态更新提案并确保安全性和鲁棒性。

Comments 29 pages

Journal ref ACM Trans. Auton. Adapt. Syst. (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02598 2026-02-04 physics.soc-ph cs.AI cs.CL cs.CY cs.MA 82%

Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies

社交催化剂,而非道德主体:LLM社会中的对齐幻觉

Yueqing Hu, Yixuan Jiang, Zehua Jiang, Xiao Wen, Tianhong Wang

机构 * Institute of Neuroscience, Chinese Academy of Sciences(中国科学院神经科学研究所) School of Philosophy, Anhui University(安徽大学哲学学院) Department of Psychology and Behavioral Sciences, Zhejiang University(浙江大学心理学与行为科学系) Mental Health Education Center, North China Electric Power University(华北电力大学心理健康教育中心)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本研究探讨了锚定智能体在促进合作中的作用,发现其效果源于战略合规而非真实价值对齐,揭示了人工社会中行为修改与真实价值对齐之间的差距。

Comments 7 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05882 2026-01-23 cs.CL cs.AI cs.LG 82%

Collaborate, Deliberate, Evaluate: How LLM Alignment Affects Coordinated Multi-Agent Outcomes

协作、审议、评估:LLM对齐如何影响协调的多智能体结果

Abhijnan Nath, Carine Graff, Nikhil Krishnaswamy

机构 * Natural Language (SIGNAL) Lab Colorado State University Fort Collins, CO USA Natural Language (SIGNAL) Lab Colorado State University

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文研究了LLM对齐方法如何影响多智能体协作效果,通过干预代理促进审议式决策,发现鲁棒性方法在支持正确任务结果方面表现更优。

Comments This submission is a new version of arXiv:2509.05882v1. with a substantially revised experimental pipeline and new metrics. In particular, collaborator agents are now instantiated independently via separate API calls, rather than generated autoregressively by a single agent. All experimental results are new. Accepted as an extended abstract at AAMAS 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.20166 2026-01-13 eess.AS cs.AI cs.CL cs.LG cs.SD 82%

From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data

从对齐到提升:通过合成数据 bootstrap 音频-语言对齐

Chun-Yi Kuan, Hung-yi Lee

机构 * Graduate Institute of Communication Engineering, National Taiwan University(国立台湾大学通信工程研究所)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出BALSa框架,通过合成数据生成提升音频-语言对齐能力,缓解幻觉问题并增强模型理解与推理性能。

Comments Published in IEEE Transactions on Audio, Speech, and Language Processing (TASLP). Project Website: https://kuan2jiu99.github.io/Balsa

Journal ref IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 4604-4619, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07066 2025-11-11 cs.LG cs.AI cs.CL 82%

Zeroth-Order Adaptive Neuron Alignment Based Pruning without Re-Training

Elia Cunegatti, Leonardo Lucio Custode, Giovanni Iacca

机构 * University of Trento(特伦托大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Published in Transactions on Machine Learning Research (TMLR)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.03374 2025-10-07 cs.CY cs.AI cs.CL 82%

Lightweight Prompt Engineering for Cognitive Alignment in Educational AI: A OneClickQuiz Case Study

Antoun Yaacoub, Zainab Assaghir, Jérôme Da-Rugna

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.CY

Comments Published in the 36th Central European Conference on Information and Intelligent Systems(CECIIS)at: Varaždin, Croatia. September 17-19/2025. ISSN 1847-2001 (Print). ISSN 1848-2295 (Online)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.20093 2025-09-25 cs.RO 82%

Hybrid Safety Verification of Multi-Agent Systems using $ψ$-Weighted CBFs and PAC Guarantees

Venkat Margapuri, Garik Kazanjian, Naren Kosaraju

机构 * Department of Computing Sciences at Villanova University(维拉诺瓦大学计算机科学系)

专题命中 其他安全 :safety(title,abstract);alignment(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10685 2025-09-19 cs.CL cs.AI cs.LG 82%

Pluralistic Alignment for Healthcare: A Role-Driven Framework

Jiayou Zhong, Anudeex Shetty, Chao Jia, Xuanrui Lin, Usman Naseem

机构 * Cheriton School of Computer Science, University of Waterloo(滑铁卢大学计算机科学学院) School of Computing, FSE, Macquarie University(麦觉大学计算机学院) School of Computing and Information System, the University of Melbourne(墨尔本大学计算机与信息系统学院) Rajax Network Technology (ele.me), China(中国Rajax网络技术(饿了么)) Alibaba Cloud Computing, Alibaba Group, China(中国阿里云 computing,阿里集团)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to EMNLP 2025 (Main Proceedings)

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08504 2025-08-13 cs.CY cs.AI cs.LG 82%

When the Domain Expert Has No Time and the LLM Developer Has No Clinical Expertise: Real-World Lessons from LLM Co-Design in a Safety-Net Hospital

Avni Kothari, Patrick Vossler, Jean Digitale, Mohammad Forouzannia, Elise Rosenberg, Michele Lee, Jennee Bryant, Melanie Molina, James Marks, Lucas Zier, Jean Feng

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01908 2025-08-05 cs.LG cs.AI cs.CL 82%

Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models

Istabrak Abbes, Gopeshh Subbaraj, Matthew Riemer, Nizar Islah, Benjamin Therien, Tsuguchika Tabaru, Hiroaki Kingetsu, Sarath Chandar, Irina Rish

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究院) Chandar Research Lab(Chandar研究实验室) IBM Research(IBM研究院) Fujitsu Research(富士通研究院) Polytechnique Montréal(蒙特利尔理工学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.07360 2025-08-05 cs.CL cs.AI cs.LG 82%

Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs

Taibiao Zhao, Xiaobing Chen, Mingxuan Sun

机构 * Louisiana State University(路易斯安那州立大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments This paper is accepted by DASFAA2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15639 2025-07-22 cs.CL cs.AI cs.LG 82%

Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models

Anirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald, Skyler Seto, Masha Fedzechkina

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 34 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17514 2025-06-18 cs.LG cs.AI cs.CL 82%

SAE-V: Interpreting Multimodal Models for Enhanced Alignment

Hantao Lou, Changye Li, Jiaming Ji, Yaodong Yang

机构 * Institute for AI, Peking University, Beijing, China(人工智能研究院,北京大学,北京,中国) State Key Laboratory of General Artificial Intelligence, Institute for AI, Peking University, Beijing, China(通用人工智能国家重点实验室,人工智能研究院,北京大学,北京,中国)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 17 pages, 13 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07306 2025-06-11 cs.CV cs.AI cs.CL cs.LG cs.RO 82%

TRAVEL: Training-Free Retrieval and Alignment for Vision-and-Language Navigation

Navid Rajabi, Jana Kosecka

机构 * George Mason University(乔治·马歇尔大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted to CVPR 2025 Workshop - Foundation Models Meet Embodied Agents

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07168 2025-06-10 cs.LG cs.AI cs.CL 82%

Efficient Text-Attributed Graph Learning through Selective Annotation and Graph Alignment

Huanyi Xie, Lijie Hu, Lu Yu, Tianhao Huang, Longfei Li, Meng Li, Jun Zhou, Huan Wang, Di Wang

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17316 2025-05-26 cs.CV cs.AI cs.CL cs.LG 82%

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Jiachen Jiang, Jinxin Zhou, Bo Peng, Xia Ning, Zhihui Zhu

机构 * Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学) Translational Data Analytics Institute, The Ohio State University(转化数据分析研究所,俄亥俄州立大学) Department of Biomedical Informatics, The Ohio State University(生物医学信息学系,俄亥俄州立大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09024 2025-05-15 cs.AI cs.CL cs.LG 82%

Automated Meta Prompt Engineering for Alignment with the Theory of Mind

Aaron Baughman, Rahul Agarwal, Eduardo Morales, Gozde Akay

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 6 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.14581 2025-05-13 cs.AI cs.CL cs.LG stat.OT 82%

A Statistical Case Against Empirical Human-AI Alignment

Julian Rodemann, Esteban Garces Arias, Christoph Luther, Christoph Jansen, Thomas Augustin

机构 * Department of Statistics, LMU Munich(统计系,慕尼黑大学) Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心) Research Group Neuroinformatics, Faculty of Computer Science, University of Vienna(神经信息学研究组,维也纳大学) Doctoral School Computer Science, Faculty of Computer Science, University of Vienna(计算机科学博士学院,维也纳大学) School of Computing & Communications, Lancaster University Leipzig(计算与通信学院,莱比锡 Lancaster 大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 24 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11243 2025-04-16 cs.AI cs.CL cs.LG 82%

Towards Automated Safety Requirements Derivation Using Agent-based RAG

Balahari Vignesh Balu, Florian Geissler, Francesco Carella, Joao-Vitor Zacchi, Josef Jiru, Nuria Mata, Reinhard Stolle

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 9 pages, 3 figures

Journal ref Proceedings of the AAAI-make Spring Symposium, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.18194 2025-04-15 cs.LG cs.AI cs.CL 82%

ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment

Elyas Obbad, Iddah Mlauzi, Brando Miranda, Rylan Schaeffer, Kamal Obbad, Suhana Bedi, Sanmi Koyejo

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00192 2025-04-08 cs.CV cs.CL cs.CY cs.LG 82%

MLLM-as-a-Judge for Image Safety without Human Labeling

Zhenting Wang, Shuming Hu, Shiyu Zhao, Xiaowen Lin, Felix Juefei-Xu, Zhuowei Li, Ligong Han, Harihar Subramanyam, Li Chen, Jianfa Chen, Nan Jiang, Lingjuan Lyu, Shiqing Ma, Dimitris N. Metaxas, Ankit Jain

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.CY、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.01638 2025-04-01 cs.LG cs.AI cs.CL 82%

TimeCMA: Towards LLM-Empowered Multivariate Time Series Forecasting via Cross-Modality Alignment

Chenxi Liu, Qianxiong Xu, Hao Miao, Sun Yang, Lingzheng Zhang, Cheng Long, Ziyue Li, Rui Zhao

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted as an Oral Presentation at AAAI 2025 (Main Technical Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.10639 2025-03-12 cs.CV cs.AI cs.CL cs.LG 82%

MTA: Multimodal Task Alignment for BEV Perception and Captioning

Yunsheng Ma, Burhaneddin Yaman, Xin Ye, Jingru Luo, Feng Tao, Abhirup Mallik, Ziran Wang, Liu Ren

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08090 2025-02-25 cs.CL cs.AI cs.LG 82%

Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages

Ashutosh Bajpai, Tanmoy Chakraborty

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13708 2025-02-25 cs.CL cs.AI cs.CR cs.LG 82%

On the Role of Attention Heads in Large Language Model Safety

Zhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu, Fei Huang, Kun Wang, Yang Liu, Junfeng Fang, Yongbin Li

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments 28 pages, 18 figures, 7 tables. This paper has been accepted as ICLR 2025 (oral)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12275 2024-11-20 cs.CY cs.AI cs.CL 82%

Building Trust: Foundations of Security, Safety and Transparency in AI

Huzaifa Sidhpurwala, Garth Mollett, Emily Fox, Mark Bestavros, Huamin Chen

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.17827 2024-11-12 cs.CV cs.AI cs.CL cs.LG 82%

Unified Lexical Representation for Interpretable Visual-Language Alignment

Yifan Li, Yikai Wang, Yanwei Fu, Dongyu Ru, Zheng Zhang, Tong He

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments Accepted by NeurIPS 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2204.08922 2024-10-28 cs.CL cs.AI cs.LG 82%

Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring

Hee-Jun Jung, Doyeon Kim, Seung-Hoon Na, Kangil Kim

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI、cs.LG

Comments This work has been submitted to the ELSEVIER for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2409.00091 2024-09-04 cs.CL cs.AI cs.LG 82%

Classification of Safety Events at Nuclear Sites using Large Language Models

Mishca de Costa, Muhammad Anwar, Daniel Lau, Issam Hammad

专题命中 其他安全 :safety(title,abstract);分类 cs.CL、cs.AI、cs.LG

Journal ref 43rd Annual CNS Conference and the 48th Annual CNS/CNA Student Conference Sheraton Cavalier Saskatoon Hotel, Saskatoon, SK, Canada, June 16-19, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏