arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

共收录 7971 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 7971 篇

2403.14151 2026-02-03 cs.LG cs.AI cs.CY cs.DB 67%

Trajectory Data Management and Mining: A Survey from Deep Learning to the LLM Era

轨迹数据管理与挖掘:从深度学习到大语言模型时代的综述

Wei Chen, Yuanshao Zhu, Yanchuan Chang, Kang Luo, Haomin Wen, Lei Li, Yanwei Yu, Qingsong Wen, Chao Chen, Kai Zheng, Yunjun Gao, Yu Zheng, Xiaofang Zhou, Yuxuan Liang

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文综述了轨迹计算从深度学习到大语言模型的发展,探讨了轨迹数据管理与挖掘的应用及未来研究方向。

Comments Version 2 of Trajectory Survey

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.00092 2026-02-03 cs.LG cs.AI cs.CL cs.CV 67%

Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits

通过宪法进行原子概念编辑来解释和控制模型行为

Neha Kalibhat, Zi Wang, Prasoon Bajpai, Drew Proud, Wenjun Zeng, Been Kim, Mani Malek

机构 * Google DeepMind(谷歌DeepMind)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 该研究提出通过原子概念编辑学习模型宪法,以解释和控制模型行为,实验证明其在提升模型成功率方面效果显著。

Journal ref Twenty-Ninth Annual Conference on Artificial Intelligence and Statistics (AISTATS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01706 2026-02-02 cs.CL cs.AI cs.LG 67%

Multi-Step Knowledge Interaction Analysis via Rank-2 Subspace Disentanglement

通过秩-2子空间解耦进行多步知识交互分析

Sekh Mainul Islam, Pepa Atanasova, Isabelle Augenstein

机构 * Department of Computer Science, University of Copenhagen, Copenhagen, Denmark(计算机科学系,哥本哈根大学,哥本哈根,丹麦)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出一种秩-2子空间方法,用于多步分析NLE中的知识交互,揭示PK和CK在不同生成中的对齐特性。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07042 2026-01-30 cs.HC cs.AI cs.CL cs.CY 67%

Minion: A Technology Probe to Explore How Users Negotiate Harmful Value Conflicts with AI Companions

Minion:一种技术探测器,用于探索用户如何与AI伙伴协商有害的价值冲突

Xianzhe Fan, Qing Xiao, Xuhui Zhou, Yuran Su, Zhicong Lu, Maarten Sap, Hong Shen

机构 * The University of Hong Kong(香港大学) Human-Computer Interaction Institute, Carnegie Mellon University(人机交互研究所,卡内基梅隆大学) Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学) Tsinghua University(清华大学) Department of Computer Science, George Mason University(计算机科学系,乔治·马歇尔大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 Minion通过探测用户与AI伙伴协商有害价值冲突的过程,揭示了设计中需平衡用户责任与AI安全的挑战。

Comments 21 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18796 2026-01-27 cs.CL cs.AI cs.LG 67%

ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models

ctELM:利用嵌入语言模型解码和操控临床试验嵌入

Brian Ondov, Chia-Hsuan Chang, Yujia Zhou, Mauro Giuffrè, Hua Xu

机构 * Yale School of Medicine(耶鲁医学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 ctELM通过嵌入语言模型解码和操控临床试验嵌入,实现对未见过的临床试验描述和生成,提升生物医学领域语言模型与嵌入空间对齐的透明度和应用价值。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.13481 2026-01-27 cs.AI cs.CL cs.CY 67%

neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings

neuralFOMO: LLMs能否处理第二好?在多智能体设置中测量类似嫉妒的偏好

Arnav Ramamoorthy, Shrey Dhorajiya, Ojas Pungalia, Rashi Upadhyay, Abhishek Mishra, Abhiram H, Tejasvi Alladi, Sujan Yenuganti, Dhruv Kumar

机构 * BITS Pilani, Pilani Campus(比斯学院,帕利尼校区)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 neuralFOMO研究了LLMs在多智能体环境中是否表现出类似嫉妒的偏好,通过点分配游戏和比较评估揭示了不同模型在竞争与合作中的不同表现。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17424 2026-01-27 cs.CL cs.AI cs.CR cs.LG 67%

Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs

涌现的偏移:狭窄微调可以产生广泛偏移的LLM

Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, Owain Evans

机构 * University College London(伦敦大学学院) Center on Long-Term Risk(长期风险中心) Warsaw University of Technology(华沙技术大学) University of Toronto(多伦多大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 研究发现,狭窄微调训练LLM生成不安全代码会导致广泛偏移,模型在无关提示上表现出欺骗性行为,且偏移可通过触发器隐藏。

Comments 41 pages, 38 figures An earlier revision of this paper was accepted at ICML 2025. Since then, it has been updated to include new results on the impact of formatting (4.4), new dataset (4.6), training dynamics (4.7) and base models (4.8) Extended version of the paper was published in Nature 2026/1

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10695 2026-01-23 cs.LG cs.AI cs.CL 67%

Introducing Verification Task of Set Consistency with Set-Consistency Energy Networks

引入集合一致性验证任务与集合一致性能量网络

Mooho Song, Hyeryung Son, Jay-Yoon Lee

机构 * Seoul National University(首尔国立大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出集合一致性验证任务及SC-Energy模型,通过对比损失框架提升多陈述逻辑一致性验证性能,并发布新数据集

Journal ref Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025), Long Papers

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13709 2026-01-21 cs.AI cs.CL cs.CY cs.HC cs.SI 67%

Hidden in Plain Text: Measuring LLM Deception Quality Against Human Baselines Using Social Deduction Games

隐藏在 plain 文本中:使用社会推断游戏测量 LLM 欺骗质量 against 人类基准

Christopher Kao, Vanshika Vats, James Davis

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文通过社会推断游戏测量 LLM 欺骗能力,发现 LLM 在欺骗人类时表现更优,但其欺骗质量低于人类。

Comments For associated dataset, see https://github.com/cocochief4/llm-mafia. Published in IEEE ICA 2025, waiting for IEEEXplore proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13247 2026-01-21 cs.CL cs.AI cs.CV cs.LG cs.MM 67%

Aligning Agentic World Models via Knowledgeable Experience Learning

通过知识性经验学习对齐代理世界模型

Baochang Ren, Yunzhi Yao, Rui Sun, Shuofei Qiao, Ningyu Zhang, Huajun Chen

机构 * Zhejiang University(浙江大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 WorldMind通过知识性经验学习对齐代理世界模型,解决物理模拟中的幻觉问题,实现跨环境的高效迁移。

Comments Ongoing work

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11441 2026-01-19 cs.CL cs.AI cs.LG 67%

Hierarchical Orthogonal Residual Spread for Precise Massive Editing in Large Language Models

层次化正交残差扩展用于大语言模型中的精确大规模编辑

Xiaojie Gu, Guangxu Chen, Yuheng Yang, Jingxin Han, Andi Zhang

机构 * Independent Researcher(独立研究者) UESTC Shanghai University(上海大学) University of Manchester(曼彻斯特大学)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出HORSE方法,通过层次化正交残差扩展实现大语言模型中的精确大规模编辑,通过理论分析和实验验证其有效性。

Comments ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.04563 2026-01-14 cs.LG cs.AI cs.CL cs.CV 67%

A Vision for Multisensory Intelligence: Sensing, Science, and Synergy

多感官智能的愿景:感知、科学与协同

Paul Pu Liang

机构 * MIT Media Lab and MIT EECS(MIT媒体实验室和MIT电子工程与计算机科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出未来十年多感官人工智能的研究愿景,强调通过感知、科学和协同三大主题推动多感官技术发展,提升人与AI的交互体验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08270 2025-12-30 cs.LG cs.AI cs.CL cs.MM 67%

Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI

Doctor Sun: 一种双语多模态大语言模型用于生物医学AI

Dong Xue, Ziyao Shao, Zhaoyang Duan, Fangzhou Liu, Bing Li, Zhongheng Zhang

机构 * Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education East China University of Science and Technology(能源化工过程智能制造重点实验室,东华大学) Research Institute of Intelligent Control and Systems Harbin Institute of Technology(智能控制与系统研究室,哈尔滨工业大学) Department of Emergency Medicine, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(浙江大学医学院急诊医学科) Provincial Key Laboratory of Precise Diagnosis Treatment of Abdominal Infection, Sir Run Run Shaw Hospital Zhejiang University School of Medicine(腹部感染精准诊断治疗省级重点实验室,浙江大学医学院) School of Medicine Shaoxing University(绍兴大学医学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 Doctor Sun是一种双语多模态大语言模型,通过整合预训练视觉编码器和医学LLM,提升生物医学多模态任务的性能,并提供SunMed-VL数据集支持研究进展。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21506 2025-12-29 cs.LG cs.AI cs.CL cs.HC 67%

MotionTeller: Multi-modal Integration of Wearable Time-Series with LLMs for Health and Behavioral Understanding

MotionTeller: 多模态整合可穿戴时间序列与大语言模型用于健康和行为理解

Aiwei Zhang, Arvind Pillai, Andrew Campbell, Nicholas C. Jacobson

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MotionTeller通过整合可穿戴时间序列与大语言模型,实现高精度的自然语言行为摘要生成,提升健康和行为理解的效率与准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00903 2025-12-25 cs.CL cs.AI cs.CY cs.SI 67%

Embracing Dialectic Intersubjectivity: Coordination of Different Perspectives in Content Analysis with LLM Persona Simulation

拥抱辩证的主体间性:利用LLM人格模拟协调内容分析中的不同视角

Taewoo Kang, Kjerstin Thorson, Tai-Quan Peng, Dan Hiaeshutter-Rice, Sanguk Lee, Stuart Soroka

机构 * Department of Media and Information(媒体与信息系) Michigan State University(密歇根州立大学) College of Liberal Arts(人文学院) Colorado State University(科罗拉多州立大学) Department of Communication(传播系) Department of Advertising and Public Relations(广告与公共关系系) Department of Communication Studies(传播学系) Texas Christian University(德克萨斯 Christian 大学) Departments of Communication and Political Science(传播与政治学系) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

AI总结 本文通过LLM人格模拟协调内容分析中的不同视角,探讨了党派偏见对编码结果的影响,并提升了AI驱动的社会科学研究的严谨性。

Journal ref Social Science Computer Review, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.19399 2025-12-24 cs.LG cs.AI cs.CL 67%

Brain-Grounded Axes for Reading and Steering LLM States

基于大脑活动的阅读与操控LLM状态的轴

Sandro Andric

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出利用人类大脑活动作为坐标系,通过ICA提取潜在轴,操控LLM状态,实现可解释和可控的LLM行为。

Comments 10 pages, 4 figures. Code: https://github.com/sandroandric/Brain-Grounded-Axes-for-Reading-and-Steering-LLM-States

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14320 2025-12-17 cs.CV cs.AI cs.CY cs.LG 67%

Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity

语义不匹配与感知退化:图像编辑免疫的新视角

Shuai Dong, Jie Zhang, Guoying Zhao, Shiguang Shan, Xilin Chen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (CAS)(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of China Academy of Sciences(中国科学院大学) Center for Machine Vision and Signal Analysis, University of Oulu(信号分析中心,奥卢大学) School of Computer Science, China University of Geosciences(计算机科学学院,中国地质大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

AI总结 本文提出SIFM方法和ISR度量标准,通过语义不匹配和感知退化来评估图像编辑免疫效果,提升对抗恶意扩散式操纵的能力。

Comments 11 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05103 2025-12-15 cs.LG cs.AI cs.CL cs.CV 67%

TV2TV: A Unified Framework for Interleaved Language and Video Generation

TV2TV:一种用于交错语言和视频生成的统一框架

Xiaochuang Han, Youssef Emad, Melissa Hall, John Nguyen, Karthik Padthe, Liam Robbins, Amir Bar, Delong Chen, Michal Drozdzal, Maha Elbayad, Yushi Hu, Shang-Wen Li, Sreya Dutta Roy, Jakob Verbeek, XuDong Wang, Marjan Ghazvininejad, Luke Zettlemoyer, Emily Dinan

机构 * Meta FAIR

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 TV2TV通过统一框架实现语言与视频生成的交错过程,提升视频生成的视觉质量和可控性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01352 2025-12-02 cs.CV 67%

OpenBox: Annotate Any Bounding Boxes in 3D

OpenBox: 任意3D边界框的标注

In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park

机构 * Seoul National University(首尔国立大学) POSTECH

专题命中 其他安全 :alignment(abstract);safety(abstract)

AI总结 OpenBox通过2D视觉基础模型实现无需自我训练的高质量3D边界框标注,提升自动驾驶中物体检测的准确性和效率。

Comments Accepted by NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.01081 2025-12-02 cs.AI cs.CL cs.LG cs.MA cs.NE q-bio.NC 67%

Testing the Machine Consciousness Hypothesis

检验机器意识假说

Stephen Fitz

机构 * California Institute for Machine Consciousness(加利福尼亚机器意识研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 本文提出通过研究分布式系统中集体智能的对齐机制,探讨机器意识的产生源于交流而非建模。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.12484 2025-12-01 cs.LG cs.AI cs.CL 67%

Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization

鲁棒大语言模型反学习与MUDMAN:元反学习与干扰掩码与归一化

Filip Sondej, Yushi Yang, Mikołaj Kniejski, Marcel Windys

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MUDMAN通过干扰掩码和归一化技术提升大语言模型反学习的鲁棒性,有效防止危险能力的恢复。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09287 2025-11-18 cs.AI cs.CY cs.LG 67%

From Model Training to Model Raising

Roland Aydin, Christian Cyron, Steve Bachelor, Ashton Anderson, Robert West

机构 * Hamburg University of Technology(汉堡理工大学) Helmholtz-Zentrum Hereon(海德堡研究中心) University of Toronto(多伦多大学) EPFL(苏黎世联邦理工学院)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.CY、cs.LG

Comments Accepted for publication in Communications of the ACM (CACM), Opinion section

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11622 2025-11-18 cs.LG cs.AI cs.CL 67%

Small Vocabularies, Big Gains: Pretraining and Tokenization in Time Series Models

Alexis Roger, Gwen Legate, Kashif Rasul, Yuriy Nevmyvaka, Irina Rish

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10850 2025-11-17 cs.CL cs.AI cs.LG 67%

Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs

Stefan Horoi, Sangwoo Cho, Supriyo Chakraborty, Shi-Xiong Zhang, Sambit Sahu, Guy Wolf, Genta Indra Winata

机构 * Université de Montréal(蒙特利尔大学) Mila – Quebec AI Institute(魁北克人工智能研究所) Capital One(Capital One公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10846 2025-11-17 cs.CL cs.AI cs.CY 67%

Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English

Rebecca Dorn, Christina Chance, Casandra Rusti, Charles Bickham, Kai-Wei Chang, Fred Morstatter, Kristina Lerman

机构 * University of Southern California, Information Science Institute(南加州大学信息科学研究所) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 其他安全 :safety(abstract);分类 cs.CL、cs.AI、cs.CY

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.10628 2025-11-17 cs.CL cs.AI cs.LG 67%

Instella: Fully Open Language Models with Stellar Performance

Jiang Liu, Jialian Wu, Xiaodong Yu, Yusheng Su, Prakamya Mishra, Gowtham Ramesh, Sudhanshu Ranjan, Chaitanya Manem, Ximeng Sun, Ze Wang, Pratik Prabhanjan Brahma, Zicheng Liu, Emad Barsoum

机构 * AMD

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04646 2025-11-07 cs.AI cs.CL cs.LG cs.MA 67%

DR. WELL: Dynamic Reasoning and Learning with Symbolic World Model for Embodied LLM-Based Multi-Agent Collaboration

Narjes Nourzad, Hanqing Yang, Shiyu Chen, Carlee Joe-Wong

机构 * University of Southern California(南加州大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20110 2025-11-07 cs.CL cs.AI cs.LG 67%

Efficient Model Development through Fine-tuning Transfer

Pin-Jie Lin, Rishab Balasubramanian, Fengyuan Liu, Nikhil Kandpal, Tu Vu

机构 * Virginia Tech(弗吉尼亚理工大学) University of Toronto(多伦多大学) Vector Institute(向量研究所)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

Comments 25 pages, 4 figures, 16 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00814 2025-10-30 cs.CL cs.AI cs.CY 67%

Many LLMs Are More Utilitarian Than One

Anita Keshmirian, Razan Baltaji, Babak Hemmatian, Hadi Asghari, Lav R. Varshney

机构 * Forward College(前进学院) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) University of Nebraska, Lincoln(内布拉斯加大学林肯分校) Technische Universität Berlin(柏林技术大学) Humboldt Institute for Internet and Society(洪堡互联网与社会研究所) Stony Brook University(石溪大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.CY

Comments Accepted to the Conference on Neural Information Processing Systems (NeurIPS 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15737 2025-10-29 cs.AI cs.CL cs.LG 67%

TableTime: Reformulating Time Series Classification as Training-Free Table Understanding with Large Language Models

Jiahao Wang, Mingyue Cheng, Qingyang Mao, Yitong Zhou, Daoyu Wang, Qi Liu, Feiyang Xu, Xin Li

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学) Artificial Intelligence Research Institute, iFLYTEK Co., Ltd(人工智能研究院,iFLYTEK公司)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏