arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-05-04 至 2026-05-04 共收录 3 信号源:cs.CV, cs.AI, cs.LG

1. 视觉问答 3 篇

2506.22982 2026-05-04 cs.CV 87%

Revisiting CroPA: A Reproducibility Study and Enhancements for Cross-Prompt Adversarial Transferability in Vision-Language Models

重新审视CroPA:面向视觉-语言模型跨提示对抗转移性的可重复性研究与改进

Atharv Mittal, Agam Pandey, Amritanshu Tiwari, Sukrit Jindal, Swadesh Swain

机构 * Mehta Family School of Data Science and Artificial Intelligence(梅hta家族数据科学与人工智能学院) Indian Institute of Technology, Roorkee(印度理工学院罗奥克学院) Department of Civil Engineering(土木工程系) Department of Electronics and Communication(电子与通信系)

专题命中 视觉问答 :vision-language model(title,abstract);LLaVA(abstract,abstract_cn);visual question answering(abstract);分类 cs.CV

AI总结 本文重新审视CroPA,验证其跨提示转移性,并提出改进方法,包括新的初始化策略、跨图像迁移性研究及针对视觉编码器的损失函数,提升对抗有效性。

Comments Accepted to MLRC 2025

Journal ref Transactions on Machine Learning Research (TMLR), 2025. Available at OpenReview: https://openreview.net/forum?id=5L90cl0xtf

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.21199 2026-05-04 cs.LG cs.CV 82%

ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response

ARFBench: 用于软件事件响应的时间序列问题回答能力基准测试

Stephan Xie, Ben Cohen, Mononito Goswami, Junhong Shen, Emaad Khwaja, Chenghao Liu, David Asker, Othmane Abou-Amal, Ameet Talwalkar

机构 * Machine Learning Department, Carnegie Mellon University(卡内基梅隆大学机器学习系) Datadog AI Research(Datadog AI研究) Amazon Web Services(亚马逊网络服务)

专题命中 视觉问答 :VLM(summary_cn,abstract);分类 cs.CV、cs.LG

AI总结 本文提出ARFBench基准测试,评估多模态基础模型对软件事件数据中时间序列异常的理解能力,发现前沿VLM表现优异,提出混合模型并建立模型-专家 oracle,达到超人类水平。

Comments Updated author affiliation

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.22285 2026-05-04 cs.CV 57%

VideoDetective: Clue Hunting via both Extrinsic Query and Intrinsic Relevance for Long Video Understanding

VideoDetective:通过外在查询和内在相关性进行长视频理解的线索搜索

Ruoliu Yang, Chu Wu, Caifeng Shan, Ran He, Chaoyou Fu

机构 * Nanjing University(南京大学) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 视觉问答 :multimodal large language model(abstract);分类 cs.CV

AI总结 本文提出VideoDetective框架,通过结合查询与片段的相关性及片段间亲和力,有效定位长视频问答中的关键片段,提升多模态大语言模型在长视频理解中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏