arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

大模型对齐与安全

大模型对齐、安全、越狱、红队、提示注入和可信评测。

2026-04-03 至 2026-04-03 共收录 14 信号源:cs.CL, cs.AI, cs.CY, cs.LG

1. 其他安全 14 篇

2604.02198 2026-04-03 cs.AI cs.LG 81%

From High-Dimensional Spaces to Verifiable ODD Coverage for Safety-Critical AI-based Systems

从高维空间到可验证的ODD覆盖以保障安全关键的AI系统

Thomas Stefani, Johann Maximilian Christensen, Elena Hoemann, Frank Köster, Sven Hallerbach

机构 * Institute for AI Safety and Security, German Aerospace Center (DLR)(德国航空航天中心(DLR)人工智能安全与安保研究所)

专题命中 其他安全 :safety(title,abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种整合参数离散化、约束过滤和关键性驱动降维的方法,用于验证高维空间中AI系统的ODD覆盖,满足EASA对完整性的要求。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00277 2026-04-03 eess.SY cs.AI cs.LG cs.SY math.DS 62%

Hybrid Energy-Based Models for Physical AI: Provably Stable Identification of Port-Hamiltonian Dynamics

基于能量的混合模型用于物理AI:可证明稳定的端哈密顿动力学识别

Simone Betteti, Luca Laurenti

机构 * RIAS Lab The Italian Institute of Artificial Intelligence for Industry(意大利人工智能工业研究所 RIAS 实验室) Delft Center for Systems and Control, TU Delft(代尔夫特理工大学代尔夫特系统与控制中心)

专题命中 其他安全 :safety(abstract);分类 cs.AI、cs.LG

AI总结 本文提出一种混合能量模型框架,用于系统识别,通过可证明稳定的耗散吸收不变动态,扩展了能量模型理论,验证了端哈密顿能量模型的稳定性与安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02280 2026-04-03 cs.AI cs.CV 57%

Novel Memory Forgetting Techniques for Autonomous AI Agents: Balancing Relevance and Efficiency

新颖的记忆遗忘技术用于自主AI代理:在相关性和效率之间平衡

Payal Fofadiya, Sunil Tiwari

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出一种自适应预算遗忘框架,通过相关性引导评分和有界优化,提升长周期对话任务的F1分数和记忆一致性,减少虚假记忆。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02211 2026-04-03 cs.IR cs.AI cs.MA 57%

Multi-Agent Video Recommenders: Evolution, Patterns, and Open Challenges

多智能体视频推荐系统:演进、模式与开放挑战

Srivaths Ranganathan, Abhishek Dharmaratnakar, Anushree Sinha, Debanshu Das

机构 * Google LLC(谷歌公司)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文探讨多智能体视频推荐系统的演进、协作模式及挑战,分析了从早期MARL到LLM驱动架构的发展,并提出可扩展性、多模态理解等开放问题。

Comments Accepted for publication in The Nineteenth ACM International Conference on Web Search and Data Mining (WSDM Companion 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.00261 2026-04-03 cs.CL 57%

Can Large Language Models Self-Correct in Medical Question Answering? An Exploratory Study

大语言模型能否在医疗问答中自我纠正?一项探索性研究

Zaifu Zhan, Mengyuan Cui, Rui Zhang

机构 * Department of Electrical and Computer Engineering, University of Minnesota(明尼苏达大学电气与计算机工程系) Division of Computational Health Sciences, Department of Surgery, University of Minnesota(明尼苏达大学外科学系计算健康科学部) Department of Software Engineering, University of St. Thomas(圣托马斯大学软件工程系)

专题命中 其他安全 :safety(abstract);分类 cs.CL

AI总结 本文探讨大语言模型在医疗问答中的自我反思能力,通过比较标准Co-T提示与迭代自我反思循环,发现自我反思并不总能提升准确性,且效果依赖于数据集和模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06763 2026-04-03 cs.CV cs.AI 57%

SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding

SafePLUG: 通过像素级洞察和时间锚定赋能多模态大语言模型以理解交通事故

Zihao Sheng, Zilin Huang, Yansong Qu, Jiancong Chen, Yuhao Luo, Yen-Jung Chen, Yue Leng, Sikai Chen

机构 * University of Wisconsin-Madison(威斯康星大学麦迪逊分校) Purdue University(普渡大学) Google(谷歌)

专题命中 其他安全 :safety(abstract);分类 cs.AI

AI总结 本文提出SafePLUG框架,通过像素级理解和时间锚定提升多模态大语言模型在交通事故分析中的能力,引入新数据集并实现区域问答、像素分割、时间事件定位等任务的高性能表现。

Comments The code, dataset, and model checkpoints will be made publicly available at: https://zihaosheng.github.io/SafePLUG

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01630 2026-04-03 cs.CL 57%

Grounding AI-in-Education Development in Teachers' Voices: Findings from a National Survey in Indonesia

将教育中的AI发展扎根于教师的声音:来自印度尼西亚全国调查的结果

Nurul Aisyah, Muhammad Dehan Al Kautsar, Arif Hidayat, Fajri Koto

机构 * Quantic School of Business and Technology(昆蒂克商业与技术学院) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Indonesia University of Education(印度尼西亚教育大学)

专题命中 其他安全 :alignment(abstract);分类 cs.CL

AI总结 研究通过全国性调查揭示印尼教师对AI在教育中的应用现状,发现AI在教学、内容开发和教学媒体中日益普及,但应用不均,教师更倾向于利用AI减轻教学准备负担。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01234 2026-04-03 cs.CV cs.AI eess.IV 57%

CLPIPS: A Personalized Metric for AI-Generated Image Similarity

CLPIPS:一种用于AI生成图像相似度的个性化度量标准

Khoi Trinh, Jay Rothenberger, Scott Seidenberger, Dimitrios Diochnos, Anindya Maiti

机构 * University of Oklahoma(俄克拉荷马大学)

专题命中 其他安全 :alignment(abstract);分类 cs.AI

AI总结 本文提出CLPIPS,一种基于LPIPS的个性化度量标准,通过人类反馈优化提升感知对齐,展示有限的人工调整可显著增强文本到图像工作流中的感知对齐。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22294 2026-04-03 cs.CV cs.LG 57%

Structure is Supervision: Multiview Masked Autoencoders for Radiology

结构是监督:用于放射学的多视图掩码自编码器

Sonia Laguna, Andrea Agostini, Alain Ryser, Samuel Ruiperez-Campillo, Irene Cannistraci, Moritz Vandenhirtz, Stephan Mandt, Nicolas Deperrois, Farhad Nooralahzadeh, Michael Krauthammer, Thomas M. Sutter, Julia E. Vogt

机构 * Department of Computer Science, ETH Zurich(苏黎世联邦理工学院计算机科学系) Department of Computer Science, UC Irvine(加州大学尔湾分校计算机科学系) Department of Quantitative Biomedicine, University of Zurich(苏黎世大学定量生物医学系)

专题命中 其他安全 :alignment(abstract);分类 cs.LG

AI总结 本文提出多视图掩码自编码器(MVMAE),通过利用放射学研究的多视图结构学习视图不变和疾病相关表征,并在三个公开数据集上验证了其在疾病分类任务中的优越性能。

Journal ref Transactions on Machine Learning Research (TMLR) 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.02039 2026-04-03 cs.SE 50%

APITestGenie: Generating Web API Tests from Requirements and API Specifications with LLMs

APITestGenie: 从需求和API规范生成Web API测试

André Pereira, Bruno Lima, João Pascoal Faria

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出APITestGenie工具,利用LLM、RAG和提示工程从业务需求和OpenAPI规范自动生成API测试,测试覆盖率达89%,发现API缺陷,提升测试与需求的一致性。

Journal ref 7th ACM/IEEE International Conference on Automation of Software Test (AST 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01659 2026-04-03 cs.RO 50%

AURA: Multimodal Shared Autonomy for Real-World Urban Navigation

AURA:多模态共享自主控制用于真实世界城市导航

Yukai Ma, Honglin He, Selina Song, Wayne Wu, Bolei Zhou

机构 * University of California, Los Angeles(加州大学洛杉矶分校) Zhejiang University(浙江大学)

专题命中 其他安全 :safety(abstract)

AI总结 AURA通过将城市导航分解为高层人类指令和低层AI控制,减少人工操作负担,提升导航稳定性,并在真实世界中实现44%的接管减少。

Comments 17 pages, 18 figures, 4 tables, conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01605 2026-04-03 cs.CV cs.RO 50%

F3DGS: Federated 3D Gaussian Splatting for Decentralized Multi-Agent World Modeling

F3DGS:联邦3D高斯点云用于去中心化多智能体世界建模

Morui Zhu, Mohammad Dehghani Tezerjani, Mátyás Szántó, Márton Vaitkus, Song Fu, Qing Yang

机构 * University of North Texas(北德克萨斯大学) Budapest University of Technology and Economics(布达佩斯技术与经济大学)

专题命中 其他安全 :alignment(abstract)

AI总结 F3DGS通过联邦学习实现去中心化多智能体3D重建,利用共享几何框架和可见性感知聚合解决部分观测问题,实现分布式优化。

Comments Accepted to the CVPR 2026 SPAR-3D Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.01274 2026-04-03 cs.GR cs.CV 50%

Non-Rigid 3D Shape Correspondences: From Foundations to Open Challenges and Opportunities

非刚性3D形状对应:从基础到开放挑战与机遇

Aleksei Zhuravlev, Lennart Bastian, Dongliang Cao, Nafie El Amrani, Paul Roetzer, Viktoria Ehm, Riccardo Marin, Hiroki Nishizawa, Shigeo Morishima, Christian Theobalt, Nassir Navab, Daniel Cremers, Florian Bernard, Zorah Lähner, Vladislav Golyanik

机构 * MPI for Informatics(马克斯·普朗克信息学研究所) University of Bonn(波恩大学) Lamarr Institute(拉马尔研究所) Waseda University(早稻田大学) Waseda Research Institute for Science and Engineering(早稻田大学理工学研究所) Imperial College London(帝国理工学院) Technical University of Munich(慕尼黑工业大学) Munich Center for Machine Learning(慕尼黑机器学习中心)

专题命中 其他安全 :alignment(abstract)

AI总结 本文探讨了非刚性3D形状对应问题,分析了三种主要方法:谱方法、组合方法和基于变形的方法,并讨论了各自优缺点及最新发展,最后总结了该领域的挑战与机遇。

Comments 35 pages and 15 figures; Eurographics 2026 STAR; Project page: https://nonrigid-shape-correspondences.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23205 2026-04-03 cs.CV 50%

EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents

EmbodMocap: 野外4D人体-场景重建用于具身智能体

Wenjia Wang, Liang Pan, Huaijin Pi, Yuke Lou, Xuqian Ren, Yifan Wu, Zhouyingcheng Liao, Lei Yang, Rishabh Dabral, Christian Theobalt, Taku Komura

机构 * The University of Hong Kong(香港大学) Tampere University(坦佩雷大学) The Chinese University of Hong Kong(香港中文大学) Max-Planck Institute for Informatics(马克斯·普朗克信息学研究所)

专题命中 其他安全 :alignment(abstract)

AI总结 本文提出EmbodMocap,利用两部移动iPhone实现低成本、便携式的人体-场景重建,通过联合校准双RGB-D序列,在无静态相机和标记的情况下实现日常环境中的度量尺度和场景一致的重建,提升具身智能研究。

详情

展开后加载摘要…

URL PDF HTML 收藏