arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1565 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1565 篇

2209.07511 2022-09-16 cs.CV 79%

Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models

Manli Shu, Weili Nie, De-An Huang, Zhiding Yu, Tom Goldstein, Anima Anandkumar, Chaowei Xiao

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments NeurIPS 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2111.13333 2022-03-25 cs.CV 79%

Predict, Prevent, and Evaluate: Disentangled Text-Driven Image Manipulation Empowered by Pre-Trained Vision-Language Model

Zipeng Xu, Tianwei Lin, Hao Tang, Fu Li, Dongliang He, Nicu Sebe, Radu Timofte, Luc Van Gool, Errui Ding

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments To appear in CVPR 2022

详情

展开后加载摘要…

URL PDF HTML 收藏
2108.09661 2021-08-24 cs.CV 79%

From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network

Yuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang, Shenggao Zhu, Yongdong Zhang

专题命中 其他VLM :visual language model(title,abstract);分类 cs.CV

Comments Accept by ICCV2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2103.08849 2021-04-16 cs.CV cs.CL 79%

Multilingual Multimodal Pre-training for Zero-Shot Cross-Lingual Transfer of Vision-Language Models

Po-Yao Huang, Mandela Patrick, Junjie Hu, Graham Neubig, Florian Metze, Alexander Hauptmann

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

Comments accepted by NAACL 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
2003.01163 2020-03-04 cs.CV cs.RO 79%

Understanding Contexts Inside Robot and Human Manipulation Tasks through a Vision-Language Model and Ontology System in a Video Stream

Chen Jiang, Masood Dehghan, Martin Jagersand

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11108 2025-06-30 cs.CL 79%

Benchmarking Vision Language Models on German Factual Data

René Peinl, Vincent Tischler

专题命中 其他VLM :vision language model(title,abstract)

Comments Peinl, René; Tischler, Vincent (2025): Benchmarking Vision Language Models on German Factual Data. 21st International Conference on Artificial Intelligence Applications and Innovations, 26-29 June, 2025, Limassol, Cyprus (accepted)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07543 2026-08-11 cs.CV cs.AI 新提交 79%

Performance of large language models in the optical diagnosis of colorectal polyps

大型语言模型在结直肠息肉光学诊断中的性能

Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan, Robert Bechara, Asher C. Wiggins, Celine N. Rousan, Kaitlyn V. G. L. Morgado, Angie Ibrahim, Kevin H. M. Kuo, Daniel von Renteln, Alexander Hann, Dennis L. Shung, Michael A. Scaffidi, Charles Ménard, Joshua Landy, Samir C. Grover

专题命中 其他VLM :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI

AI总结 本研究评估Claude Opus 4等5款大型语言模型对结直肠息肉的光学诊断性能,发现其区分息肉亚型的准确率接近专家共识,但灵敏度与特异度未达ESGE标准,需进一步研究方可临床应用。

Comments 22 pages, 1 figure, 5 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.18695 2026-08-11 cs.CV cs.AI cs.LG eess.IV 版本更新 78%

Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models

属性应来自图像,而非类名:视觉语言模型的分布条件属性选择

Gautam Rajendrakumar Gare, Jia Shi, Zhiqiu Lin, Deepak Pathak, John Galeotti, Deva Ramanan

机构 * Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他VLM :vision-language model(title);分类 cs.CV、cs.AI、cs.LG

AI总结 研究视觉语言模型可解释零样本分类,指出基于类名的描述符视觉证据不足。提出从目标图像集选属性的方法,该方法提升了准确率,性能优于CoOp,用时短,所选属性还能描述数据分布,是一种有效的分布条件属性选择策略。

Comments Accepted at the PFATCV Workshop, ECCV 2026. Project page: https://ggare-cmu.github.io/AttributeSelect/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.16318 2026-07-21 cs.HC cs.RO 新提交 78%

The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models

社交机器人眼中的世界——用视觉语言模型增强人机对话

Thomas Sievers

机构 * Institute of Information Systems, University of Lübeck(信息系统研究所,吕贝克大学)

专题命中 其他VLM :vision language model(title,abstract)

AI总结 研究探讨如何用视觉语言模型增强人机对话,通过将米斯特拉尔人工智能语言模型与胡椒机器人结合用于人机交互对话,并研究视觉信息对响应时间的影响,发现纳入视觉信息可增添对话背景,且使用欧洲托管的语言模型便于实际应用。

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.21036 2026-06-16 cs.CL 版本更新 78%

GePBench: Evaluating Fundamental Geometric Perception for Multimodal Large Language Models

GePBench:评估多模态大语言模型的基础几何感知能力

Shangyu Xing, Changhao Xiang, Yuteng Han, Yifan Yue, Zhen Wu, Xinyu Liu, Zhangtai Wu, Fei Zhao, Xinyu Dai

机构 * University of Science and Technology of China(中国科学技术大学) Tsinghua University(清华大学)

专题命中 其他VLM :multimodal large language model(title,abstract)

AI总结 提出GePBench基准,系统评估多模态大语言模型的几何形状识别与空间关系感知能力,发现现有模型存在显著缺陷,而基于该基准训练可提升下游任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.25195 2026-03-27 cs.HC 78%

On-Demand Instructional Material Providing Agent Based on MLLM for Tutoring Support

基于MLLM的按需教学材料提供代理用于辅导支持

Takumi Kato, Masato Kikuchi, Tadachika Ozono

专题命中 其他VLM :MLLM(title);multimodal large language model(abstract)

AI总结 本文提出基于多模态大语言模型的代理,用于在一对一辅导中按需提供教学材料,通过分析对话自动检索相关图片,实验显示检索时间减少44.4秒,85.7%的试次提供可接受质量的图片。

Comments The 20th International Conference on E-Service and Knowledge Management (ESKM 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.08668 2026-03-10 cs.RO 78%

Exp-Force: Experience-Conditioned Pre-Grasp Force Selection with Vision-Language Models

Exp-Force:基于视觉-语言模型的经验条件预抓取力选择

Siqi Shang, Minchao Huang, Bill Fan, Lillian Chin

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)

专题命中 其他VLM :vision-language model(title,abstract)

AI总结 Exp-Force通过经验条件框架利用视觉-语言模型实现基于单张图像的预抓取力选择,显著提升了抓取力的准确性和可靠性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18969 2026-03-02 cs.NE 78%

Increasing Computation Resolves Conflicts in Vision Language Models

提升计算能力缓解视觉语言模型中的冲突

Bingyang Wang, Yijiang Li, Yitong Qiao, Maijunxian Wang, Tianwei Zhao, Yucheng Sun, Binyue Deng, Hokin Deng, Nuno Vasconcelos, Dezhi Luo

专题命中 其他VLM :vision language model(title,abstract)

AI总结 本研究发现视觉语言模型通过增加计算资源能更有效解决冲突,展示了大规模神经网络中的优化动态如何产生类人认知控制。

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.00229 2026-02-09 cs.HC 78%

When and How to Integrate Multimodal Large Language Models in College Psychotherapy: Perspectives from Multi-stakeholders

何时以及如何将多模态大语言模型整合到大学生心理治疗中:来自多利益相关者的视角

Jiyao Wang, Youyu Sheng, Qihang He, Zian Zhang, Haolong Hu, Yumei Jing, Dengbo He

专题命中 其他VLM :multimodal large language model(title);MLLM(abstract)

AI总结 本研究探讨了多模态大语言模型在大学生心理治疗中的整合时机与方式,指出其作为辅助工具的作用,并揭示了用户接受度受社交身份和相对地位影响的关键因素。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.04355 2026-02-05 cs.CL 78%

Can Vision Replace Text in Working Memory? Evidence from Spatial n-Back in Vision-Language Models

视觉能否取代文本在工作记忆中的作用?来自视觉-语言模型中空间n-Back任务的证据

Sichu Liang, Hongyu Zhu, Wenwen Wang, Deyu Zhou

机构 * Southeast University(东南大学) Shanghai Jiao Tong University(上海交通大学) Carnegie Mellon University(卡内基梅隆大学)

专题命中 其他VLM :vision-language model(title,abstract)

AI总结 研究通过空间n-Back任务评估视觉-语言模型中视觉与文本对工作记忆的影响,发现文本条件下的表现优于视觉条件,揭示了视觉信息处理中的计算差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11807 2025-11-04 cs.CL 78%

Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?

Simeon Junker, Manar Ali, Larissa Koch, Sina Zarrieß, Hendrik Buschmeier

机构 * Bielefeld University(比勒菲尔德大学)

专题命中 其他VLM :multimodal large language model(title,abstract)

Comments To appear in ACL Findings 2025

Journal ref Findings of the Association for Computational Linguistics: ACL 2025, pp. 24101-24109

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.21214 2025-10-27 cs.CR 78%

Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses

Xingwei Zhong, Kar Wai Fok, Vrizlynn L. L. Thing

专题命中 其他VLM :MLLM(title);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21301 2025-09-26 cs.OS 78%

Nova: Real-Time Agentic Vision-Language Model Serving with Adaptive Cross-Stage Parallelization

Yuhang Xu, Shengzhong Liu, Dong Zhang, Bingheng Yan, Fan Wu, Guihai Chen

专题命中 其他VLM :vision-language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19662 2025-09-12 physics.ed-ph 78%

Multimodal large language models and physics visual tasks: comparative analysis of performance and costs

Giulia Polverini, Bor Gregorcic

专题命中 其他VLM :multimodal large language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12530 2025-08-20 cs.CL 78%

Basic Category Usage in Vision Language Models

Hunter Sawyer, Jesse Roberts, Kyle Moore

机构 * Computer Science, Tennessee Tech University(田纳西科技大学计算机科学系) Computer Science, Vanderbilt University(范德比大学计算机科学系)

专题命中 其他VLM :vision language model(title);vision-language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19196 2025-08-19 cs.RO cs.CL cs.HC 78%

Towards Multimodal Social Conversations with Robots: Using Vision-Language Models

Ruben Janssens, Tony Belpaeme

机构 * Ghent University–imec(根特大学–imec)

专题命中 其他VLM :vision-language model(title,abstract)

Comments Accepted at the workshop "Human - Foundation Models Interaction: A Focus On Multimodal Information" (FoMo-HRI) at IEEE RO-MAN 2025 (Camera-ready version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16781 2025-07-23 cs.DC 78%

Cooling Matters: Benchmarking Large Language Models and Vision-Language Models on Liquid-Cooled Versus Air-Cooled H100 GPU Systems

Imran Latif, Muhammad Ali Shafique, Hayat Ullah, Alex C. Newkirk, Xi Yu, Arslan Munir

专题命中 其他VLM :vision-language model(title,abstract)

Comments 11 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.00700 2025-07-02 cs.CL 78%

Contrasting Cognitive Styles in Vision-Language Models: Holistic Attention in Japanese Versus Analytical Focus in English

Ahmed Sabir, Azinovič Gasper, Mengsay Loem, Rajesh Sharma

机构 * University of Tartu(塔尔图大学) University of Ljubljana(卢布尔雅那大学) Sansan, Inc.(Sansan公司) Plaksha University(Plaksha大学)

专题命中 其他VLM :vision-language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.07818 2025-06-10 cs.CL 78%

WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code

Zhiyu Lin, Zhengda Zhou, Zhiyuan Zhao, Tianrui Wan, Yilun Ma, Junyu Gao, Xuelong Li

机构 * Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究院(TeleAI)、中国电信) Northwestern Polytechnical University(西北工业大学) Beijing Jiaotong University(北京交通大学) Nanjing University(南京大学)

专题命中 其他VLM :multimodal large language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09936 2025-05-16 cs.HC cs.GR cs.MA cs.MM 78%

CartoAgent: a multimodal large language model-powered multi-agent cartographic framework for map style transfer and evaluation

Chenglong Wang, Yuhao Kang, Zhaoya Gong, Pengjun Zhao, Yu Feng, Wenjia Zhang, Ge Li

专题命中 其他VLM :multimodal large language model(title,abstract)

Comments 57 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.09093 2025-04-23 cs.CR 78%

BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger

Yulin Chen, Haoran Li, Yirui Zhang, Zihao Zheng, Yangqiu Song, Bryan Hooi

专题命中 其他VLM :multimodal large language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.15213 2025-04-17 eess.SP 78%

Sig2text, a Vision-language model for Non-cooperative Radar Signal Parsing

Hancong Feng KaiLI Jiang Bin tang

专题命中 其他VLM :vision-language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.00457 2025-01-03 cs.LG cs.AI cs.CL cs.CV 78%

Differentiable Prompt Learning for Vision Language Models

Zhenhan Huang, Tejaswini Pedapati, Pin-Yu Chen, Jianxi Gao

专题命中 其他VLM :vision language model(title);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.10879 2024-12-11 cs.LG cs.AI cs.CL cs.CV 78%

Enhancing Vision-Language Model Pre-training with Image-text Pair Pruning Based on Word Frequency

Mingliang Liang, Martha Larson

专题命中 其他VLM :vision-language model(title);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.12058 2024-11-20 cs.SD eess.AS 78%

Vision Language Models Are Few-Shot Audio Spectrogram Classifiers

Satvik Dixit, Laurie M. Heller, Chris Donahue

专题命中 其他VLM :vision language model(title,abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏