arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-15 至 2026-04-15 共收录 19 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 19 篇

2604.12012 2026-04-15 cs.CV 78%

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

TIPSv2:通过增强的补丁-文本对齐推进视觉-语言预训练

Bingyi Cao, Koert Chen, Kevis-Kokitsi Maninis, Kaifeng Chen, Arjun Karpur, Ye Xia, Sahil Dua, Tanmaya Dabral, Guangxing Han, Bohyung Han, Joshua Ainslie, Alex Bewley, Mithun Jacob, René Wagner, Washington Ramos, Krzysztof Choromanski, Mojtaba Seyedhosseini, Howard Zhou, André Araujo

机构 * Google(谷歌) DeepMind(深度Mind)

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出TIPSv2,通过改进的预训练方法提升视觉-语言模型的补丁-文本对齐能力,实验表明其在多种任务和数据集上表现优异。

Comments CVPR2026 camera-ready + appendix

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11964 2026-04-15 cs.HC cs.MM 78%

When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs

当绘画不够时:探索通过草图进行意图对齐的自发性言语在多模态大语言模型中的应用

Weiyan Shi, Dorien Herremans, Kenny Tsu Wei Choo

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文探讨了在多模态大语言模型中,通过草图与自发性言语的结合来提升设计初期意图对齐的效果,通过实验表明加入自发性言语能显著提高生成图像的意图匹配度。

Comments Accepted at DIS 2026 PWiP

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13789 2026-04-15 cs.CR cs.AI 70%

Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks

揭示并对齐异常注意力头以防御NLP后门攻击

Haotian Jin, Yang Li, Haihui Fan, Lin Shen, Xiangfang Li, Bo Li

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) State Key Laboratory of Cyberspace Security Defense(网络空间安全防御国家重点实验室) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)

专题命中 其他安全 :alignment(abstract);safety(abstract);分类 cs.AI

AI总结 本文提出基于注意力相似性的后门检测方法,通过检测异常注意力头相似性来防御后门攻击,结合头部微调修复污染的注意力头,有效降低攻击成功率并保持下游任务性能。

Journal ref "Proceedings of the 40th AAAI Conference on Artificial Intelligence (AAAI 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23061 2026-04-15 cs.IR cs.AI cs.CL cs.DB cs.LG 67%

MoDora: Tree-Based Semi-Structured Document Analysis System

MoDora:基于树的半结构化文档分析系统

Bangrui Xu, Qihang Yao, Zirui Tang, Xuanhe Zhou, Yeye He, Shihan Yu, Qianqian Xu, Bin Wang, Guoliang Li, Conghui He, Fan Wu

机构 * Shanghai Jiao Tong University(上海交通大学) Microsoft Research(微软研究院) Beihang University(北航) Shanghai AI Lab(上海人工智能实验室) Tsinghua University(清华大学) Institute for Clarity in Documentation(文档清晰性研究所) Inria Paris-Rocquencourt(Inria巴黎-罗quentcourt研究中心) East China Normal University(华东师范大学) Scientific Writing Academy(科学写作学院)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI、cs.LG

AI总结 MoDora通过局部对齐聚合策略和组件相关树结构,提升半结构化文档的分析能力,实验表明其在准确性上优于基线方法。

Comments Extension of our SIGMOD 2026 paper. Please refer to source code available at https://github.com/weAIDB/MoDora

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12227 2026-04-15 cs.AI cs.CL 62%

Designing Reliable LLM-Assisted Rubric Scoring for Constructed Responses: Evidence from Physics Exams

设计可靠的LLM辅助评分系统用于构造响应:来自物理考试的证据

Xiuxiu Tang, G. Alex Ambrose, Ying Cheng

机构 * Department of Psychology, University of Notre Dame(诺特大学心理学系) Notre Dame Learning, University of Notre Dame(诺特大学学习中心)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 本文研究了GPT-4o在物理考试构造响应评分中的可靠性,发现细粒度检查清单式评分优于整体评分,提示清晰的评分标准对AI辅助评分至关重要。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11835 2026-04-15 cs.LG cs.AI 62%

Schema-Adaptive Tabular Representation Learning with LLMs for Generalizable Multimodal Clinical Reasoning

基于LLM的Schema自适应表格表示学习用于通用多模态临床推理

Hongxi Mao, Wei Zhou, Mengting Jia, Tao Fang, Huan Gao, Bin Zhang, Shangyang Li

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) Boston University(波士顿大学) University of Southern California(南加州大学) GDIIST Renyixun Health Technology Co., Ltd.(仁心迅健康科技有限公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文提出基于LLM的Schema自适应表格表示学习方法,通过将结构化变量转换为语义自然语言并编码,实现零样本跨schema对齐,结合表格和MRI数据在多模态 dementia 诊断中取得优于临床基准的性能。

Comments 11 pages, 4 figures

Journal ref ACL 2026, Main conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22440 2026-04-15 cs.HC cs.AI cs.CL 62%

AI and My Values: User Perceptions of LLMs' Ability to Extract, Embody, and Explain Human Values from Casual Conversations

AI与我的价值观:用户对LLMs从闲聊中提取、体现和解释人类价值观的能力的看法

Bhada Yun, Renn Su, April Yi Wang

机构 * Stanford University(斯坦福大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨LLMs从闲聊中提取、体现和解释人类价值观的能力,通过参与者与聊天机器人互动并完成评估访谈,发现13名参与者认为AI能理解人类价值观,警示'武器化共情'风险,并提出VAPT工具用于评估AI价值观对齐。

Comments To appear in CHI '26

Journal ref Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26), April 13--17, 2026, Barcelona, Spain

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12776 2026-04-15 cs.CL 57%

EvoSpark: Endogenous Interactive Agent Societies for Unified Long-Horizon Narrative Evolution

EvoSpark:内生交互代理社会的统一长周期叙述进化

Shiyu He, Minchi Kuang, Mengxian Wang, Bin Hu, Tingxiang Gu

机构 * School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院) Department of Precision Instrument, Tsinghua University(清华大学精密仪器系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 EvoSpark通过内生交互代理社会框架,解决LLM多代理系统中长周期叙述演化的矛盾,通过分层叙述记忆和生成场景机制实现逻辑一致的持续叙事。

Comments Accepted to the Main Conference of ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12721 2026-04-15 cs.CL 57%

InsightFlow: LLM-Driven Synthesis of Patient Narratives for Mental Health into Causal Models

InsightFlow: 基于LLM的患者心理健康叙事合成至因果模型

Shreya Gupta, Prottay Kumar Adhikary, Bhavyaa Dave, Salam Michael Singh, Aniket Deroy, Tanmoy Chakraborty

机构 * Department of Mathematics, IIT Delhi(印度德里理工学院数学系) Department of Electrical Engineering, IIT Delhi(印度德里理工学院电气工程系) Department of Computer Science and Engineering, IIIT Manipur(曼尼普尔理工学院计算机科学与工程系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 InsightFlow利用LLM自动从患者-治疗师对话生成与5P框架对齐的因果图,通过结构、语义和专家评估验证,证明其在临床案例构建中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12663 2026-04-15 cs.AI 57%

Human-Centric Topic Modeling with Goal-Prompted Contrastive Learning and Optimal Transport

以人为中心的主题建模:基于目标提示对比学习与最优传输

Rui Wang, Yi Zheng, Dongxin Wang, Haiping Huang, Yuanzhi Yao, Yuxiang Zhou, Jialin Yu, Philip Torr

机构 * School of Computer Science, Nanjing University of Posts and Telecommunications(南京邮电大学计算机科学学院) School of Computer Science and Information Engineering, Hefei University of Technology(合肥工业大学计算机科学与信息工程学院) School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦玛丽女王大学电子工程与计算机科学学院) Department of Engineering Science, University of Oxford(牛津大学工程科学系)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出Human-TM任务框架,结合人类提供的目标进行主题建模,通过GCTM-OT方法提升主题的可解释性、多样性和目标导向性,实验表明其在主题连贯性和多样性上优于现有方法。

Comments 11 Pages, 6 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12537 2026-04-15 cs.CV cs.AI 57%

MODIX: A Training-Free Multimodal Information-Driven Positional Index Scaling for Vision-Language Models

MODIX:一种无需训练的多模态信息驱动的位置索引缩放

Ruoxiang Huang, Zhen Yuan

机构 * Peking University(北京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 MODIX通过动态调整位置步长,基于模态特定贡献优化多模态模型的位置编码,提升多模态推理能力。

Comments Accepted by CVPR 2026 (Highlight). 10 pages, 2 figures, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12506 2026-04-15 cs.CL cs.SD 57%

Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs

超越转录:面向感知意识的统一音频架构

Linhao Zhang, Yuhan Song, Aiwei Liu, Chuhan Wu, Sijun Zhang, Wei Jia, Yuan Liu, Houfeng Wang, Xiao Zhou

机构 * Basic Model Technology Center, WeChat AI, Tencent Inc.(腾讯基本模型技术中心、微信AI、腾讯公司) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室、计算机学院、北京大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出统一音频架构(UAS),通过将音频信息划分为转录、语调和非语言事件三个组件,提升音频细粒度感知性能,同时保持推理能力。

Comments Accepted to ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12113 2026-04-15 cs.CV cs.AI 57%

PR-MaGIC: Prompt Refinement Via Mask Decoder Gradient Flow For In-Context Segmentation

PR-MaGIC:通过掩码解码器梯度流进行上下文分割的提示精炼

Minjae Lee, Sungwoo Hur, Soojin Hwang, Won Hwa Kim

机构 * Pohang University of Science and Technology(浦项科学技术大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 PR-MaGIC通过掩码解码器梯度流精炼提示,提升上下文分割性能,无需额外训练,有效缓解提示不足问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.11948 2026-04-15 cs.LG cs.AR 57%

Active Imitation Learning for Thermal- and Kernel-Aware LFM Inference on 3D S-NUCA Many-Cores

主动模仿学习用于热感知和核意识的LFM推理在3D S-NUCA多核系统

Yixian Shen, Chaoyao Shen, Jan Deen, George Floros, Andy Pimentel, Anuj Pathania

机构 * University of Amsterdam, Netherlands(阿姆斯特丹大学) Southeast University, China(东南大学) University of Thessaly, Greece(塞萨洛尼基大学)

专题命中 其他安全 :safety(abstract);分类 cs.LG

AI总结 本文提出AILFM框架,通过主动模仿学习实现热感知调度,考虑核心性能异质性和LFM核行为,提升性能并保障热安全。

Comments Accepted for publication at the 63rd ACM/IEEE Design Automation Conference (DAC 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14004 2026-04-15 cs.CL 57%

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

定位、引导与改进:大型语言模型中可操作机制可解释性的一项实用调查

Hengyuan Zhang, Zhihao Zhang, Mingyang Wang, Zunhai Su, Yiwei Wang, Qianli Wang, Shuzhou Yuan, Ercong Nie, Xufeng Duan, Feijiang Han, Qibo Xue, Zeping Yu, Chenming Shang, Xiao Liang, Jing Xiong, Hui Shen, Chaofan Tao, Zhengwu Liu, Senjie Jin, Zhiheng Xi, Dongdong Zhang, Sophia Ananiadou, Tao Gui, Ruobing Xie, Hayden Kwok-Hay So, Hinrich Schütze, Xuanjing Huang, Qi Zhang, Ngai Wong

机构 * The University of Hong Kong(香港大学) Fudan University(复旦大学) LMU Munich(慕尼黑大学) Tsinghua University(清华大学) Technische Universität Darmstadt(达姆施塔特技术大学) Technische Universität Berlin(柏林技术大学) Technische Universität Dresden(德累斯顿技术大学) The Chinese University of Hong Kong(香港中文大学) University of Pennsylvania(宾夕法尼亚大学) Nanjing University(南京大学) University of Manchester(曼彻斯特大学) Dartmouth College(达特茅斯学院) University of California Los Angeles(加州大学洛杉矶分校) University of Michigan(密歇根大学) Microsoft(微软) Tencent(腾讯)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 本文提出一个实用调查,围绕'定位、引导与改进'流程,系统分类定位和引导方法,展示如何通过该框架提升模型对齐、能力和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.09087 2026-04-15 cs.AI 57%

The Stackelberg Speaker: Optimizing Persuasive Communication in Social Deduction Games

Stackelberg发言者:优化社会推断游戏中的说服性沟通

Zhang Zheng, Deheng Ye, Peilin Zhao, Hao Wang

机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) Tencent(腾讯) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出基于Stackelberg竞争的强化学习框架,用于优化社会推断游戏中说服性沟通,通过实验展示其在三种不同游戏中的优越性。

Comments Accepted by ACL 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12747 2026-04-15 cs.SE 50%

Short Version of VERIFAI2026 Paper -- Learning Infused Formal Reasoning: Contract Synthesis, Artefact Reuse and Semantic Foundations

VERIFAI2026论文简报 -- 学习融合形式推理:合同合成、验证成果重用与语义基础

Arshad Beg, Diarmuid O'Donoghue, Rosemary Monahan

专题命中 其他安全 :safety(abstract)

AI总结 本文提出学习融合形式推理框架,通过自然语言要求自动合成合同、利用图匹配和学习嵌入重用验证成果,并基于UTP和机构理论建立数学语义基础,推动验证从孤立证明转向知识驱动过程。

Comments 2 pages. Accepted at ADAPT Annual Scientific Conference (AASC) 2026. To be held on 14th of May, 2026 at Dublin City University, Ireland

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23178 2026-04-15 cs.CV 50%

Intelligent bear deterrence system based on computer vision: Reducing human-bear conflicts in remote areas

基于计算机视觉的智能熊驱赶系统:减少偏远地区人熊冲突

Pengyu Chen, Teng Fei, John A. Kupfer, Yunyan Du, Jiawei Yi, Yi Li

机构 * School of Resources and Environmental Sciences, Wuhan University(武汉大学资源与环境科学学院) Department of Geography, University of South Carolina(南卡罗来纳大学地理系) State Key Laboratory of Resources and Environmental Information System, Beijing(北京资源与环境信息系统国家重点实验室) Institute of Zoology, Chinese Academy of Sciences(中国科学院动物研究所)

专题命中 其他安全 :safety(abstract)

AI总结 本文提出一种低功耗、网络无关的驱赶系统,结合计算机视觉与物联网硬件,通过YOLOv5-MobileNet模型和太阳能驱赶装置,有效减少偏远地区人熊冲突,提升安全与保育。

Journal ref Ursus 2026(37e6), 1-11 (2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.12105 2026-04-15 cs.SE 50%

Automated BPMN Model Generation from Textual Process Descriptions: A Multi-Stage LLM-Driven Approach

从文本过程描述自动生成BPMN模型:一种多阶段大语言模型驱动的方法

Ion Matei, Maksym Zhenirovskyy, Praveen Kumar Menaka Sekar, Hon Yung Wong

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出一种多阶段大语言模型驱动的方法,通过多语言BPMN XML文件翻译、验证和修复生成可靠数据,从而自动生成可执行的BPMN 2.0模型,实现高相似度的重建。

详情

展开后加载摘要…

URL PDF HTML 收藏