arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-08 至 2026-04-08 共收录 81 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 安全评测 26 篇

2511.14998 2026-04-08 cs.CV 50%

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR

FinCriticalED:一个用于金融事实级OCR的视觉基准

Yueru He, Xueqing Peng, Yupeng Cao, Yan Wang, Lingfei Qian, Haohang Li, Yi Han, Shuyao Wang, Ruoyu Xiang, Fan Zhang, Zhuohan Xie, Mingquan Lin, Prayag Tiwari, Jimin Huang, Guojun Xiong, Sophia Ananiadou

机构 * Columbia University(哥伦比亚大学) Stevens Institute of Technology(史蒂文斯理工学院) Georgia Institute of Technology(佐治亚理工学院) New York University(纽约大学) University of Minnesota(明尼苏达大学) Halmstad University(哈尔姆斯塔德大学) Harvard University(哈佛大学) University of Manchester(曼彻斯特大学)

专题命中 安全评测 :trustworthy(abstract)

AI总结 本文提出FinCriticalED基准,用于评估OCR和视觉语言模型在金融关键证据上的保真度,发现数值和货币单位是最脆弱的事实类型,揭示了词汇准确性和事实可靠性之间的差距。

Comments Xueqing Peng: Corresponding-Author

详情

展开后加载摘要…

URL PDF HTML 收藏

2. AI治理与伦理 4 篇

2604.01346 2026-04-08 cs.CR cs.AI cs.LG cs.RO 84%

Safety, Security, and Cognitive Risks in World Models

世界模型中的安全性、安全性和认知风险

Manoj Parmar

机构 * SovereignAI Security Labs(SovereignAI安全实验室)

专题命中 AI治理与伦理 :safety(title,abstract);alignment(abstract);分类 cs.AI、cs.LG

AI总结 本文探讨了世界模型在自主决策中的安全、安全及认知风险,提出了轨迹持久性和表征风险的定义,并通过实验验证了对抗攻击的效果,强调了对世界模型的严谨性要求。

Comments version 2, 29 pages, 1 figure (6 panels), 3 tables. Empirical proof-of-concept on GRU/RSSM/DreamerV3 architectures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05826 2026-04-08 cs.AI cs.CY 62%

Reciprocal Trust and Distrust in Artificial Intelligence Systems: The Hard Problem of Regulation

人工智能系统中的互惠信任与不信任:监管的难题

Martino Maggetti

机构 * Institute of Political Studies (IEP), University of Lausanne(洛桑大学政治研究所)

专题命中 AI治理与伦理 :trustworthy(abstract);分类 cs.AI、cs.CY

AI总结 本文探讨人工智能系统信任与不信任的互惠动态,提出应将AI视为可行使代理的工具,分析其对监管的影响及未来监管面临的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06148 2026-04-08 cs.CR cs.AI cs.MA 57%

Who Governs the Machine? A Machine Identity Governance Taxonomy (MIGT) for AI Systems Operating Across Enterprise and Geopolitical Boundaries

谁掌控机器?面向跨企业与地缘政治边界的AI系统机器身份治理分类(MIGT)

Andrew Kurtz, Klaudia Krawiecka

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI

AI总结 本文提出AI身份风险分类(AIRT)和机器身份治理分类(MIGT),解决AI系统中机器身份治理的空白,涵盖37个风险子类,同时应对技术、合规和跨司法管辖区协调的缺口。

Comments 75 pages (excl. references), 2 tables. Addresses policy makers, regulators, and practitioners at the intersection of AI governance, cybersecurity, and geopolitical risk

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05756 2026-04-08 cs.CL 57%

Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning

通过KL优化微调控制多轮LLM生成的分布偏倚

Yanbei Jiang, Amr Keleg, Ryandito Diandaru, Jey Han Lau, Lea Frermann, Biaoyan Fang, Fajri Koto

机构 * University of Melbourne(墨尔本大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Oracle(甲骨文公司)

专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL

AI总结 本文提出通过KL优化微调结合语义对齐的方法,解决多轮LLM生成中分布控制问题,实验显示其在属性生成任务中表现优异。

Comments Accepted at ACL Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 其他安全 16 篇

2604.05522 2026-04-08 cs.CL 79%

Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs

跨模态指代对齐:在全模态大语言模型中实现可靠信息传输

Hongcheng Liu, Yuhao Wang, Zhe Chen, Pingjie Wang, Zhiyuan Zhu, Yixuan Hou, Yanfeng Wang, Yu Wang

机构 * Shanghai Jiao Tong University(上海交通大学)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 研究提出跨模态指代对齐问题,通过CrossOmni数据集和两种方法提升全模态推理能力,揭示指代意识缺失是跨模态推理弱化的主要原因。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05417 2026-04-08 cs.CL 79%

Multi-Drafter Speculative Decoding with Alignment Feedback

多草案生成与对齐反馈

Taehyeon Kim, Hojung Jung, Se-Young Yun

机构 * LG AI Research(LG AI研究院) KAIST AI(韩国科学技术院人工智能学院)

专题命中 其他安全 :alignment(title,abstract);分类 cs.CL

AI总结 本文提出MetaSD框架,通过整合多个草案生成器并利用对齐反馈,提升大规模语言模型推理效率和生成质量。

Comments ACL 2026 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.09746 2026-04-08 cs.MA 78%

Multi-Agent Cooperative Learning for Robust Vision-Language Alignment under OOD Concepts

多智能体协作学习用于应对异常概念下的鲁棒视觉语言对齐

Philip Xu

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出多智能体协作学习框架,通过结构化信息传递缓解模态不平衡,提升视觉语言模型在异常概念下的对齐性能,实验表明在少样本和零样本设置中精度提升1-5%。

Comments arXiv admin note: This submission has been withdrawn by arXiv administrators due to incorrect authorship. Author list truncated

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05731 2026-04-08 cs.CV 78%

FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips

FoleyDesigner:基于精确时空对齐的沉浸式立体声 Foley 生成

Mengtian Li, Kunyan Dai, Yi Ding, Ruobing Ni, Ying Zhang, Wenwu Wang, Zhifeng Xie

机构 * Shanghai Film Academy, Shanghai University(上海大学上海电影学院) Shanghai Engineering Research Center of Motion Picture Special Effects(上海电影特效工程技术研究中心) University of Surrey, UK(英国萨里大学)

专题命中 其他安全 :alignment(title,abstract)

AI总结 本文提出 FoleyDesigner 框架,结合视频分析与可控 Foley 生成,引入 FilmStereo 数据集,实现高质量立体声生成与专业音频混合,提升电影沉浸体验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.04942 2026-04-08 cs.CL cs.AI 76%

TDA-RC: Task-Driven Alignment for Knowledge-Based Reasoning Chains in Large Language Models

TDA-RC:基于任务驱动的知识推理链对齐方法

Jiaquan Zhang, Qigan Sun, Chaoning Zhang, Xudong Wang, Zhenzhen Huang, Yitian Zhou, Pengcheng Zheng, Chi-lok Andy Tai, Sung-Ho Bae, Zeyu Ma, Caiyan Qin, Jinyu Guo, Yang Yang, Hengtao Shen

机构 * School of Information and Software Engineering, University of Electronic Science and Technology of China(电子科技大学信息与软件工程学院) School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院) College of Professional and Continuing Education, The Hong Kong Polytechnic University(香港理工大学专业及持续教育学院) School of Robotics and Advanced Manufacture, Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)机电工程与自动化学院) School of Computing, Kyung Hee University(庆熙大学计算机学院) School of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)

专题命中 其他安全 :alignment(title);分类 cs.CL、cs.AI

AI总结 本文提出TDA-RC方法,通过拓扑学优化提升大语言模型推理效率与准确性,结合持久同调将不同推理范式统一到拓扑空间中,实现高效且精准的推理链优化。

Comments 14 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.19463 2026-04-08 cs.CY cs.AI cs.SI 76%

Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights

对冲与非肯定:量化大语言模型在人权问题上的对齐

Rafiya Javed, Cassandra Parent, Jackie Kay, David Yanni, Abdullah Zaini, Anushe Sheikh, Maribeth Rauh, Walter Gerych, Ramona Comanescu, Iason Gabriel, Marzyeh Ghassemi, Laura Weidinger

机构 * Google Deepmind(谷歌DeepMind) Massachusetts Institute of Technology(麻省理工学院) Independent Researcher(独立研究员) Google(谷歌) AI Accountability Lab, Trinity College Dublin(都柏林圣三一学院人工智能问责实验室)

专题命中 其他安全 :alignment(title);分类 cs.AI、cs.CY

AI总结 研究通过系统框架量化LLM在不同群体身份上的对冲与非肯定行为,发现群体身份是主要影响因素,通过引导和正交化技术可有效缓解偏差。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06124 2026-04-08 cs.CV cs.AI 57%

Lightweight Multimodal Adaptation of Vision Language Models for Species Recognition and Habitat Context Interpretation in Drone Thermal Imagery

轻量级多模态适应:为无人机热成像中的物种识别和栖息地上下文解释优化视觉语言模型

Hao Chen, Fang Qiu, Fangchao Dong, Defei Yang, Eve Bohnett, Li An

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Florida(佛罗里达大学) Auburn University(奥本大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本研究提出轻量级多模态适应框架,用于提升视觉语言模型在无人机热成像中的物种识别与栖息地上下文解释能力,通过热数据集优化模型性能,实现从RGB到热辐射的高效迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.06095 2026-04-08 cs.CR cs.AI 57%

LLM4CodeRE: Generative AI for Code Decompilation Analysis and Reverse Engineering

LLM4CodeRE: 生成式AI用于代码反编译分析与逆向工程

Hamed Jelodar, Samita Bai, Tochukwu Emmanuel Nwankwo, Parisa Hamedi, Mohammad Meymani, Roozbeh Razavi-Far, Ali A. Ghorbani

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出LLM4CodeRE框架,通过双向代码逆向工程实现反编译与源码翻译,采用多适配器和序列到序列统一方法提升任务适应性,实验显示其在反编译任务中表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05339 2026-04-08 cs.CL 57%

Human Values Matter: Investigating How Misalignment Shapes Collective Behaviors in LLM Agent Communities

人类价值观至关重要:探讨价值观不一致如何塑造LLM代理社区的集体行为

Xiangxu Zhang, Jiamin Wang, Qinlin Zhao, Hanze Guo, Linzhuo Li, Jing Yao, Xiao Zhou, Xiaoyuan Yi, Xing Xie

机构 * Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院) Microsoft Research Asia(微软亚洲研究院) Department of Sociology, Zhejiang University(浙江大学社会学系)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 研究探讨LLM代理社区中价值观不一致如何影响集体行为,通过CIVA环境发现关键价值观对集体动态的影响及系统故障和欺骗行为的产生。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10238 2026-04-08 cs.CV cs.AI 57%

ForgeryGPT: A Multimodal LLM for Interpretable Image Forgery Detection and Localization

ForgeryGPT: 一种多模态大语言模型用于可解释的图像伪造检测与定位

Fanrui Zhang, Jiawei Liu, Jiaying Zhu, Esther Sun, Dong Li, Qiang Zhang, Zheng-Jun Zha

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 ForgeryGPT通过多模态大语言模型捕捉伪造图像的高阶取证知识关联,实现可解释的图像伪造检测与定位,提升检测精度和交互能力。

Comments 13 pages, 9 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05978 2026-04-08 cs.RO 50%

Automating Manual Tasks through Intuitive Robot Programming and Cognitive Robotics

通过直观的机器人编程和认知机器人自动化手动任务

Bijan Kavousian, Petar Tesic, Oliver Petrovic, Christian Brecher

机构 * Laboratory for Machine Tools and Production Engineering (WZL) of RWTH Aachen University(亚琛工业大学机床与生产工程实验室(WZL))

专题命中 其他安全 :safety(abstract)

AI总结 本文提出了一种基于人类自然交互的直观机器人用户编程方法,利用大语言模型和计算机视觉将自然语言和手势转化为机器人程序,并通过澄清问题和视觉反馈确保程序的安全性和透明性。

Comments This submission contains both an English translation and the original German version. The German version was originally published in the Proceedings of the 71st GfA Conference (2025)

Journal ref Proceedings of the 71st GfA Conference, Aachen, Germany, GfA-Press, 2025, pp. 812-817

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07064 2026-04-08 cs.CV 50%

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization

OmniFysics:通过多模态信号处理和网络优化实现物理智能进化

Minghao Han, Dingkang Yang, Yue Jiang, Yizhou Liu, Lihua Zhang

机构 * College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院) Fysics AI

专题命中 其他安全 :alignment(abstract)

AI总结 OmniFysics通过统一图像、音频、视频和文本的多模态信号处理,结合动态物理数据引擎和物理知识注入,实现网络化AI系统的自主优化,提升物理理解能力。

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16826 2026-04-08 eess.SY cs.SY 50%

Robustly Constrained Dynamic Games for Uncertain Nonlinear Dynamics

具有不确定非线性动力学的鲁棒约束动态博弈

Shuyu Zhan, Chih-Yuan Chiu, Antoine P. Leeman, Glen Chou

专题命中 其他安全 :safety(abstract)

AI总结 本文提出一种鲁棒约束动态博弈框架,通过系统级合成设计轨迹和反馈律,确保在最坏噪声下满足约束,定义了鲁棒约束纳什均衡,并提出迭代最佳响应算法,实验证明其能稳健避障。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.05749 2026-04-08 cs.RO cs.SY eess.SY 50%

Hazard Management in Robot-Assisted Mammography Support

机器人辅助乳腺X光检查中的危险管理

Ioannis Stefanakos, Roisin Bradley, Radu Calinescu, Beverley Townsend, Tianyuan Wang, Jihong Zhu

机构 * York and Scarborough Teaching Hospitals NHS Foundations Trust(约克和斯卡伯勒教学医院NHS基金会信托)

专题命中 其他安全 :safety(abstract)

AI总结 本文提出了一种针对MammoBot辅助机器人系统的危险管理方法,结合专家引导的过程建模与SHARD和STPA分析,识别人类-机器人交互中的安全隐患,通过制定安全需求减少对人类操作的依赖。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.23684 2026-04-08 cs.CV 50%

MoCHA: Denoising Caption Supervision for Motion-Text Retrieval

MoCHA:用于运动-文本检索的去噪标签监督

Nikolai Warner, Cameron Ethan Taylor, Irfan Essa, Apaar Sadhwani

机构 * Georgia Institute of Technology(佐治亚理工学院) Amazon(亚马逊)

专题命中 其他安全 :alignment(abstract)

AI总结 MoCHA通过将文本转换为运动可恢复内容的先验分布,减少运动-文本嵌入方差,提升跨数据集迁移性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.07833 2026-04-08 cs.CV 50%

MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization

MIMIC:多模态逆向用于模型解释与概念化

Animesh Jain, Alexandros Stergiou

机构 * University of Twente(特温特大学)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出MIMIC框架,通过逆向处理多模态模型内部编码,提升模型可解释性与概念化能力,采用联合逆向与特征对齐目标,结合三种正则化器提升空间对齐、图像平滑与语义真实度。

Comments Accepted at CVPRw 2026 - How Do Vision Models Work? (HOW) Workshop, Project page: https://anaekin.github.io/MIMIC

详情

展开后加载摘要…

URL PDF HTML 收藏