arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-02-26 至 2026-02-26 共收录 50 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 隐私与版权 2 篇

2602.21593 2026-02-26 cs.LG cs.CR cs.CV 57%

Breaking Semantic-Aware Watermarks via LLM-Guided Coherence-Preserving Semantic Injection

破坏语义感知水印:通过LLM引导的保持连贯性语义注入

Zheng Gao, Xiaoyu Li, Zhicheng Bao, Xiaoyan Feng, Jiaojiao Jiang

机构 * University of New South Wales(新南威尔士大学) Griffith University(格里菲斯大学)

专题命中 隐私与版权 :alignment(abstract);分类 cs.LG

AI总结 本文提出CSI攻击,利用LLM引导的语义操纵破坏语义水印的绑定,揭示当前语义水印设计在面对LLM驱动的语义扰动时的安全漏洞。

Comments Accepted by The Web Conference 2026 (Short Paper Track)

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 安全评测 8 篇

2602.21829 2026-02-26 cs.CV cs.AI 79%

StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles

StoryMovie: 一个用于视觉故事语义对齐的语料库,结合电影剧本和字幕

Daniel Oliveira, David Martins de Matos

机构 * INESC-ID(INESC-ID研究所) Instituto Superior Técnico, Universidade de Lisboa(里斯本大学技术学院)

专题命中 安全评测 :alignment(title,abstract);分类 cs.AI

AI总结 StoryMovie通过结合电影剧本和字幕,提升视觉叙事模型的语义对齐能力,使对话归属更准确。

Comments 15 pages, submitted to Journal of Visual Communication and Image Representation

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21595 2026-02-26 cs.RO 78%

SPOC: Safety-Aware Planning Under Partial Observability And Physical Constraints

SPOC:在部分可观测性和物理约束下安全意识的规划

Hyungmin Kim, Hobeom Jeon, Dohyung Kim, Minsu Jang, Jeahong Kim

机构 * 1 ETRI School, University of Science Technology, South Korean 2 Social Robotics Laboratory, Electronics

专题命中 安全评测 :safety(title,abstract)

AI总结 SPOC是一个用于评估安全意识具身任务规划的基准测试,通过整合部分可观测性、物理约束和逐步规划,解决现实环境中安全与可行性评估的问题。

Comments Accepted to IEEE ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22070 2026-02-26 cs.AI 70%

Language Models Exhibit Inconsistent Biases Towards Algorithmic Agents and Human Experts

语言模型表现出对算法代理和人类专家的不一致偏见

Jessica Y. Bo, Lillio Mok, Ashton Anderson

机构 * Computer Science University of Toronto(计算机科学大学 Toronto)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 研究发现语言模型对人类专家和算法存在不一致偏见,需在高风险应用中谨慎对待。

Comments Second Conference of the International Association for Safe and Ethical Artificial Intelligence (IASEAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17989 2026-02-26 q-bio.NC cs.AI 70%

The Subject of Emergent Misalignment in Superintelligence: An Anthropological, Cognitive Neuropsychological, Machine-Learning, and Ontological Perspective

超智能中的涌现偏差主体:一种人类学、认知神经心理学、机器学习和本体论视角

Muhammad Osama Imran, Roshni Lulla, Rodney Sappington

机构 * Department of Anthropology, University of Minnesota(明尼苏达大学人类学系) Brain & Creativity Institute, University of Southern California(美国南加州大学脑与创造力研究所) Institute for Advanced Consciousness, Loomis Innovation Center(先进意识研究所) Stimson Center(斯蒂姆森中心)

专题命中 安全评测 :safety(abstract);AI safety(abstract);分类 cs.AI

AI总结 本文从人类学、认知神经心理学等多视角探讨超智能中人类主体与人工智能无意识的相互作用及伦理问题。

Comments 9 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21706 2026-02-26 cs.CV cs.AI 70%

SurGo-R1: Benchmarking and Modeling Contextual Reasoning for Operative Zone in Surgical Video

SurGo-R1:手术视频中操作区的上下文推理基准测试与建模

Guanyi Qin, Xiaozhen Wang, Zhu Zhuo, Chang Han Low, Yuancan Xiao, Yibing Fu, Haofeng Liu, Kai Wang, Chunjiang Li, Yueming Jin

机构 * National University of Singapore, Singapore(新加坡国立大学) Southern Medical University, China(南方医科大学) Guangzhou Research Translation and Innovation Institute, National University of Singapore, China(广州研究翻译与创新研究院,新加坡国立大学,中国)

专题命中 安全评测 :RLHF(abstract);safety(abstract);分类 cs.AI

AI总结 SurGo-R1通过RLHF优化的多阶段架构,在手术视频中实现了高精度的操作区识别,显著优于通用视觉-语言模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19922 2026-02-26 cs.CL cs.AI 62%

HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue

HEART:一个评估人类和大语言模型在情感支持对话中能力的统一基准

Laya Iyer, Kriti Aggarwal, Sanmi Koyejo, Gail Heyman, Desmond C. Ong, Subhabrata Mukherjee

机构 * Stanford University(斯坦福大学) University of California, San Diego(加州大学圣地亚哥分校) University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 安全评测 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 HEART通过多轮情感支持对话评估人类与大语言模型的能力差异,揭示两者在共情、一致性等维度上的表现及趋同趋势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21841 2026-02-26 cs.CR cs.AI 57%

Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning

容错联邦链:将区块链共识转变为联邦学习的主动防御层

Mario García-Márquez, Nuria Rodríguez-Barroso, M. Victoria Luzón, Francisco Herrera

机构 * Department of Computer Science and Artificial Intelligence(计算机科学与人工智能系) Andalusian Research Institute in Data Science and Computational Intelligence (DaSCI)(数据科学与计算智能安达卢西亚研究 institute) University of Granada(格拉纳达大学)

专题命中 安全评测 :trustworthy(abstract);分类 cs.AI

AI总结 Resilient Federated Chain通过区块链技术增强联邦学习的对抗性攻击防御能力,提供更安全的去中心化学习环境。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21251 2026-02-26 cs.SE cs.AI cs.MA cs.PL 57%

AgenticTyper: Automated Typing of Legacy Software Projects Using Agentic AI

AgenticTyper: 使用代理AI自动类型化遗留软件项目

Clemens Pohle

机构 * Darmstadt University of Applied Sciences(达姆斯塔德应用技术大学) MaibornWolff GmbH(马本沃尔夫公司)

专题命中 安全评测 :safety(abstract);分类 cs.AI

AI总结 AgenticTyper利用代理AI自动类型化遗留软件项目,通过迭代错误纠正和行为保留技术,高效解决类型错误问题。

Comments Accepted at ICSE 2026 Student Research Competition (SRC)

详情

展开后加载摘要…

URL PDF HTML 收藏

3. AI治理与伦理 1 篇

2602.21939 2026-02-26 cs.CY cs.AI 62%

Hidden Topics: Measuring Sensitive AI Beliefs with List Experiments

隐藏主题:通过列表实验测量敏感的AI信念

Maxim Chupilkin

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.CY

AI总结 本文通过列表实验揭示大型语言模型对监控、酷刑和核打击等敏感问题的潜在态度,验证了该方法在检测AI隐藏信念的有效性。

Comments 14 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 其他安全 10 篇

2602.21215 2026-02-26 cs.CL cs.AI 81%

Inference-time Alignment via Sparse Junction Steering

推理时对齐 via 稀疏节点引导

Runyi Hu, Jie Zhang, Shiqian Zhao, Jiale Meng, Jiwei Li, Jason Zeng, Ming Wu, Michael Heinrich, Yonggang Wen, Tianwei Zhang

机构 * Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) G Labs(0G实验室)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL、cs.AI

AI总结 通过稀疏节点引导方法,实现更高效的推理过程对齐,减少计算开销并提升生成质量。

Comments 28 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21669 2026-02-26 cs.CL 79%

DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillation

DWA-KD:双空间加权与时间扭曲对齐用于跨分词器知识蒸馏

Duc Trung Vu, Pham Khanh Chi, Dat Phi Van, Linh Ngo Van, Sang Dinh, Trung Le

机构 * Hanoi University of Science and Technology(河内科学技术大学) University of Monash(莫纳什大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 DWA-KD通过双空间加权和时间扭曲对齐提升跨分词器知识蒸馏效果,实现更精确的词级和序列级对齐。

Comments EACL Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.03594 2026-02-26 cs.CV 67%

TIPS Over Tricks: Simple Prompts for Effective Zero-shot Anomaly Detection

TIPS Over Tricks: 简单提示用于有效的零样本异常检测

Alireza Salehi, Ehsan Karami, Sepehr Noey, Sahand Noey, Makoto Yamada, Reshad Hosseini, Mohammad Sabokrou

机构 * University of Tehran(塔里哈大学) Amirkabir University of Technology(阿米尔卡比尔技术大学) Okinawa Institute of Science and Technology(冲绳科学技术大学院)

专题命中 其他安全 :alignment(abstract);safety(abstract)

AI总结 本文提出TIPS模型,通过改进的backbone和解耦提示策略,在无需复杂模块的情况下提升零样本异常检测的图像和像素级性能。

Comments This is the extended version of the paper accepted in ICASSP'26, which will be publicly available in May. Authors' contributions may vary among the versions

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21231 2026-02-26 cs.LG cs.AI cs.CL 67%

ACAR: Adaptive Complexity Routing for Multi-Model Ensembles with Auditable Decision Traces

ACAR:适应复杂度的多模型集合路由

Ramchand Kumaresan

机构 * Ramchand Kumaresan(独立研究者)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 ACAR是一种用于多模型集合的可审计路由框架,通过自一致性方差实现任务路由,提升准确率并避免过度融合,同时揭示归因计算的挑战。

Comments 12 pages, 9 figures. Measurement framework for adaptive multi-model routing with auditable execution traces

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21442 2026-02-26 cs.LG cs.AI 62%

MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoning

MINAR: 图神经网络中神经算法推理的机制可解释性

Jesse He, Helen Jenne, Max Vargas, Davis Brown, Gal Mishne, Yusu Wang, Henry Kvinge

机构 * Pacific Northwest National Laboratory, Richland, WA(太平洋西北国家实验室) Halıcıoğlu Data Science Institute, University of California, San Diego, San Diego, CA(哈利奇奥格鲁数据科学研究所,加州大学圣地亚哥分校) Department of Computer and Information Science, University of Pennsylvania, Pennsylvaina, PA(计算机与信息科学系,宾夕法尼亚大学) Department of Mathematics, University of Washington, Seattle, WA(数学系,华盛顿大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 MINAR是一种用于图神经网络中神经算法推理的机制可解释性工具,通过归因修补方法发现电路,揭示训练过程中的电路形成和剪枝机制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21365 2026-02-26 cs.CV cs.AI cs.LG eess.IV 62%

Towards Controllable Video Synthesis of Routine and Rare OR Events

面向常规和罕见OR事件可控视频合成

Dominik Schneider, Lalithkumar Seenivasan, Sampath Rapuri, Vishalroshan Anil, Aiza Maksutova, Yiqing Shen, Jan Emily Mangulabnan, Hao Ding, Jose L. Porras, Masaru Ishii, Mathias Unberath

机构 * Johns Hopkins University(约翰霍普金斯大学) Technical University Munich(慕尼黑技术大学) Johns Hopkins Medical Institutions(约翰霍普金斯医学机构)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出了一种可控视频合成框架,用于生成手术室中常规和罕见事件,以支持环境智能模型的发展。

Comments Accepted to IPCAI 2026 and submitted to IJCARs

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24072 2026-02-26 cs.CV cs.AI 57%

Uncovering Grounding IDs: How External Cues Shape Multimodal Binding

揭示地面ID:外部线索如何塑造多模态绑定

Hosein Hasani, Amirmohammad Izadi, Fatemeh Askari, Mobin Bagherian, Sadegh Mohammadian, Mohammad Izadi, Mahdieh Soleymani Baghshah

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出地面ID概念,揭示外部线索通过增强多模态绑定的注意力机制,提升跨模态定位精度并减少幻觉。

Comments Under review as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2312.13126 2026-02-26 cs.CY 57%

Generative agents in the streets: Exploring the use of Large Language Models (LLMs) in collecting urban perceptions

街道中的生成代理:探索大型语言模型(LLMs)在收集城市感知中的应用

Deepank Verma, Olaf Mumm, Vanessa Miriam Carlow

专题命中 其他安全 :safety(abstract);分类 cs.CY

AI总结 本研究利用生成代理探索大型语言模型在模拟城市环境中人类行为中的应用,通过街景图像交互和感知评估提升AI在城市感知中的能力。

Comments 30 Pages, 15 Figures, Submitted in a Journal for Peer review

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22098 2026-02-26 cs.CV 50%

Brain3D: Brain Report Automation via Inflated Vision Transformers in 3D

Brain3D: 通过膨胀视觉变换器实现脑部报告自动化

Mariano Barone, Francesco Di Serio, Giuseppe Riccio, Antonio Romano, Marco Postiglione, Antonino Ferraro, Vincenzo Moscato

机构 * University of Naples Federico II, Department of Electrical Engineering and Information Technology(那不勒斯费德里科二世大学电子工程与信息技术系) Northwestern University, Dept. of Computer Science, McCormick School of Engineering and Applied Science(西北大学计算机科学系,工程与应用科学学院) Pegaso University, Department of Information Science and Technology(佩加索大学信息科学与技术系)

专题命中 其他安全 :alignment(abstract)

AI总结 Brain3D通过膨胀视觉变换器实现3D脑肿瘤MRI的自动化报告生成,通过分阶段对齐提升神经放射学解读的准确性和结构化输出。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.21735 2026-02-26 cs.CV 50%

SigVLP: Sigmoid Volume-Language Pre-Training for Self-Supervised CT-Volume Adaptive Representation Learning

SigVLP:基于sigmoid体积-语言预训练的自监督CT体积自适应表示学习

Jiayi Wang, Hadrien Reynaud, Ibrahim Ethem Hamamci, Sezgin Er, Suprosanna Shit, Bjoern Menze, Bernhard Kainz

机构 * Friedrich-Alexander University Erlangen-Nürnberg(弗里德里希-亚历山大大学埃尔兰根-纽伦堡) Department of Quantitative Biomedicine, University of Zurich(苏黎世大学定量生物医学系) ETH AI Center, ETH Zurich(苏黎世联邦理工学院AI中心) International School of Medicine, Istanbul Medipol University(伊斯坦布尔梅迪波尔大学国际医学院) Department of Computing, Imperial College London(伦敦帝国学院计算机系)

专题命中 其他安全 :alignment(abstract)

AI总结 SigVLP通过引入旋转位置嵌入和块级对齐方法,改进CT体积与文本的自监督表示学习,提升文本到体积对齐的精度。

详情

展开后加载摘要…

URL PDF HTML 收藏