arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-04-29 至 2026-04-29 共收录 6 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2604.07802 2026-04-29 cs.CV cs.AI 62%

Latent Anomaly Knowledge Excavation: Unveiling Sparse Sensitive Neurons in Vision-Language Models

潜在异常知识挖掘:揭示视觉-语言模型中的稀疏敏感神经元

Shaotian Li, Shangze Li, Chuancheng Shi, Wenhua Wu, Yanqiu Wu, Xiaohan Yu, Fei Shen, Tat-Seng Chua

机构 * Macquarie University(麦考瑞大学) Nanjing University of Science and Technology(南京理工大学) The University of Sydney(悉尼大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出LAKE框架,通过挖掘视觉-语言模型中稀疏敏感神经元,实现异常检测的内在可解释性,实验表明其在工业基准上表现优异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25884 2026-04-29 quant-ph cs.CV 57%

QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding

QCalEval:用于量子校准图理解的视觉-语言模型基准测试

Shuxiang Cao, Zijian Zhang, Abhishek Agarwal, Grace Bratrud, Niyaz R. Beysengulov, Daniel C. Cole, Alejandro Gómez Frieiro, Elena O. Glen, Hao Hsu, Gang Huang, Raymond Jow, Greshma Shaji, Tom Lubowe, Ligeng Zhu, Luis Mantilla Calderón, Nicola Pancotti, Joel Pendleton, Brandon Severin, Charles Etienne Staub, Sara Sussman, Antti Vepsäläinen, Neel Rajeshbhai Vora, Yilun Xu, Varinia Bernales, Daniel Bowring, Elica Kyoseva, Ivan Rungger, Giulia Semeghini, Sam Stanwyck, Timothy Costa, Alán Aspuru-Guzik, Krysta Svore

机构 * NVIDIA University of Toronto(多伦多大学) IQM Quantum Computers(IQM量子计算机) Lawrence Berkeley National Laboratory(伯克利国家实验室) Conductor Quantum(Conductor量子) National Physical Laboratory(国家物理实验室) Infleqtion Harvard University(哈佛大学) Fermi National Accelerator Laboratory(费米国家加速器实验室) Northwestern University(西北大学) EeroQ Corporation(EeroQ公司) Royal Holloway University of London(伦敦皇家霍洛威大学) Vector Institute for Artificial Intelligence(人工智能向量研究所)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出QCalEval,首个用于评估视觉-语言模型理解量子校准图能力的基准测试,包含243个样本和87种场景类型,测试零样本和上下文学习下的六种问题类型,展示了不同模型的性能差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.25562 2026-04-29 cs.CR cs.AI 57%

SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents

SnapGuard: 基于截图的轻量级提示注入检测方法

Mengyao Du, Han Fang, Haokai Ma, Jiahao Chen, Kai Xu, Quanjun Yin, Ee-Chien Chang

机构 * National University of Defense Technology(国防科技大学) University of Science and Technology of China(中国科学技术大学) National University of Singapore(新加坡国立大学) Zhejiang University(浙江大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

AI总结 针对基于截图的网络代理面临的提示注入攻击问题,提出SnapGuard方法,通过视觉稳定指标和文本信号分析实现高效检测,达到F1得分0.75,速度提升8倍。

Comments 10 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.23214 2026-04-29 cs.CL 57%

DARC-CLIP: Dynamic Adaptive Refinement with Cross-Attention for Meme Understanding

DARC-CLIP:动态自适应细化与跨注意力机制用于表情包理解

Qiyuan Jin

机构 * The Hong Kong University of Science(香港科技大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

AI总结 本文提出DARC-CLIP框架,通过层次化细化栈实现自适应多模态融合,提升表情包中多模态线索的建模精度,尤其在仇恨检测任务中取得显著提升。

Comments Accepted to IEEE ICASSP 2026. 5 pages, 3 figures, 4 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.29080 2026-04-29 cs.CV cs.LG 57%

Is the Modality Gap a Bug or a Feature? A Robustness Perspective

模态间隙是bug还是feature?从鲁棒性视角

Rhea Chowers, Oshri Naparstek, Udi Barzelay, Yair Weiss

机构 * Hebrew University(希伯来大学) IBM Research(IBM研究院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文从鲁棒性角度探讨模态间隙的存在原因,发现减少间隙可提升模型鲁棒性而不影响清洁准确率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13302 2026-04-29 cs.CL 57%

Images Amplify Misinformation Sharing in Vision-Language Models

图像在视觉-语言模型中放大虚假信息的传播

Alice Plebe, Timothy Douglas, Diana Riazi, R. Maria del Rio-Chanona

机构 * Department of Industrial Engineering, University of Trento(特伦托大学工业工程系) Computer Science Department, University College London(伦敦大学学院计算机科学系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL

AI总结 研究探讨了图像如何影响视觉-语言模型分享新闻内容的倾向,发现图像能提高虚假新闻的分享率,且不同模型对图像的反应存在差异。

Comments Accepted for oral presentation at ICWSM 2026

详情

展开后加载摘要…

URL PDF HTML 收藏