arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1562 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1562 篇

2410.11201 2025-04-22 cs.CV cs.AI cs.LG 82%

Tree of Attributes Prompt Learning for Vision-Language Models

Tong Ding, Wanhua Li, Zhongqi Miao, Hanspeter Pfister

机构 * Harvard University(哈佛大学) Mass General Brigham(麻省总医院) Microsoft(微软公司)

专题命中 其他VLM :vision-language model(title);vision language model(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.02217 2025-04-04 cs.HC 82%

The Plot Thickens: Quantitative Part-by-Part Exploration of MLLM Visualization Literacy

Matheus Valentim, Vaishali Dhanoa, Gabriela Molina León, Niklas Elmqvist

专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract)

Comments 11 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.21785 2025-03-31 eess.AS cs.SD 82%

Lend a Hand: Semi Training-Free Cued Speech Recognition via MLLM-Driven Hand Modeling for Barrier-free Communication

Guanjie Huang, Danny Hin Kwok Tsang, Li Liu

专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10665 2025-03-17 cs.CV cs.AI cs.CL cs.LG 82%

Small Vision-Language Models: A Survey on Compact Architectures and Techniques

Nitesh Patnaik, Navdeep Nayak, Himani Bansal Agrawal, Moinak Chinmoy Khamaru, Gourav Bal, Saishree Smaranika Panda, Rishi Raj, Vishal Meena, Kartheek Vadlamani

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10079 2025-03-14 cs.CL 82%

Information Density Principle for MLLM Benchmarks

Chunyi Li, Xiaozhe Li, Zicheng Zhang, Yuan Tian, Ziheng Jia, Xiaohong Liu, Xiongkuo Min, Jia Wang, Haodong Duan, Kai Chen, Guangtao Zhai

专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07094 2025-03-11 cs.CL 82%

A Novel Ophthalmic Benchmark for Evaluating Multimodal Large Language Models with Fundus Photographs and OCT Images

Xiaoyi Liang, Mouxiao Bian, Moxin Chen, Lihao Liu, Junjun He, Jie Xu, Lin Li

专题命中 其他VLM :multimodal large language model(title,abstract);MLLM(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.18021 2024-12-10 cs.CV cs.AI cs.LG 82%

Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Shuo Chen, Zhen Han, Bailan He, Jianzhe Liu, Mark Buckley, Yao Qin, Philip Torr, Volker Tresp, Jindong Gu

专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments accepted by WACV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.10618 2024-11-05 cs.AI cs.CV cs.LG 82%

Private Attribute Inference from Images with Vision-Language Models

Batuhan Tömekçe, Mark Vero, Robin Staab, Martin Vechev

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.00309 2024-10-02 cs.CV cs.AI cs.LG 82%

Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models

Laura Bravo-Sánchez, Jaewoo Heo, Zhenzhen Weng, Kuan-Chieh Wang, Serena Yeung-Levy

专题命中 其他VLM :vision language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Project webpage: https://laubravo.github.io/apu_website/

详情

展开后加载摘要…

URL PDF HTML 收藏
2311.13951 2024-09-17 cs.CL 82%

MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria

Wentao Ge, Shunian Chen, Guiming Hardy Chen, Junying Chen, Zhihong Chen, Nuo Chen, Wenya Xie, Shuo Yan, Chenghao Zhu, Ziyue Lin, Song Dingjie, Xidong Wang, Anningzhe Gao, Zhang Zhiyi, Jianquan Li, Xiang Wan, Benyou Wang

专题命中 其他VLM :MLLM(title,abstract);multimodal large language model(abstract)

Comments 23 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.13979 2024-08-28 cs.CV cs.AI cs.CL cs.LG 82%

Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models

Shuai Fu, Xiequn Wang, Qiushi Huang, Yu Zhang

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at ICLR 2024 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.16414 2024-08-15 cs.CV cs.AI cs.LG 82%

AutoCLIP: Auto-tuning Zero-Shot Classifiers for Vision-Language Models

Jan Hendrik Metzen, Piyapat Saranrittichai, Chaithanya Kumar Mummadi

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments accepted at TMLR, Camera Ready Version

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.09143 2024-08-13 cs.AI cs.CE cs.CV cs.LG cs.NE 82%

Generative AI-based Prompt Evolution Engineering Design Optimization With Vision-Language Model

Melvin Wong, Thiago Rios, Stefan Menzel, Yew Soon Ong

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted and to be published in IEEE Congress on Evolutionary Computation (CEC) 2024. Copyright 2024 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses

Journal ref IEEE Congress on Evolutionary Computation (CEC), 2024, 1-8

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.00667 2024-06-05 eess.IV cs.AI cs.CL cs.CV cs.LG 82%

An Early Investigation into the Utility of Multimodal Large Language Models in Medical Imaging

Sulaiman Khan, Md. Rafiul Biswas, Alina Murad, Hazrat Ali, Zubair Shah

专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted in Fifth IEEE Workshop on Artificial Intelligence for HealthCare, IEEE 25th International Conference on Information Reuse and Integration for Data Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.07309 2024-05-14 cs.RO cs.AI cs.CV cs.LG 82%

DiffGen: Robot Demonstration Generation via Differentiable Physics Simulation, Differentiable Rendering, and Vision-Language Model

Yang Jin, Jun Lv, Shuqiang Jiang, Cewu Lu

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.01468 2024-05-03 cs.LG cs.AI cs.CV 82%

Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models

Yifei Ming, Yixuan Li

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments The paper is accepted at ICML 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08475 2024-04-19 cs.CL cs.AI cs.CV cs.LG cs.MM 82%

Can We Edit Multimodal Large Language Models?

Siyuan Cheng, Bozhong Tian, Qingbin Liu, Xi Chen, Yongheng Wang, Huajun Chen, Ningyu Zhang

专题命中 其他VLM :multimodal large language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments EMNLP 2023. Add the Exact Match/Accuracy results of Reliability and T-Generality

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.17589 2024-03-27 cs.CV cs.AI cs.LG cs.MM 82%

Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models

Yabin Zhang, Wenjie Zhu, Hui Tang, Zhiyuan Ma, Kaiyang Zhou, Lei Zhang

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments CVPR2024; Codes are available at \url{https://github.com/YBZh/DMN}

详情

展开后加载摘要…

URL PDF HTML 收藏
2303.16133 2024-02-23 cs.CV cs.AI cs.CL cs.LG 82%

Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models

Adyasha Maharana, Amita Kamath, Christopher Clark, Mohit Bansal, Aniruddha Kembhavi

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments TMLR 2024; Project Website: https://adymaharana.github.io/cococon/

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02998 2024-01-29 cs.CV cs.AI cs.CL cs.LG 82%

ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language Models

Yi-Lin Sung, Jaehong Yoon, Mohit Bansal

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments ICLR 2024 (project page: https://ecoflap.github.io/)

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.01390 2023-08-08 cs.CV cs.AI cs.LG 82%

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, Jenia Jitsev, Simon Kornblith, Pang Wei Koh, Gabriel Ilharco, Mitchell Wortsman, Ludwig Schmidt

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2109.01134 2022-10-07 cs.CV cs.AI cs.LG 82%

Learning to Prompt for Vision-Language Models

Kaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei Liu

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI、cs.LG

Comments International Journal of Computer Vision (IJCV), 2022. Update: Adds results on the DOSCO (DOmain Shift in COntext) benchmark

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.08064 2026-06-09 cs.MM cs.CV 版本更新 81%

PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning

PUMA: 基于层剪枝的语言模型,用于具有模态自适应学习的高效统一多模态检索

Yibo Lyu, Rui Shao, Gongwei Chen, Yijie Zhu, Weili Guan, Liqiang Nie

机构 * Harbin Institute of Technology(哈尔滨工业大学)

专题命中 其他VLM :MLLM(summary_cn,abstract_cn);multimodal large language model(abstract);分类 cs.CV

AI总结 提出PUMA,通过层剪枝自蒸馏减少MLLM参数,并设计模态自适应对比学习损失(MAC-Loss)提升检索效率,在降低资源消耗的同时保持性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.12333 2026-08-14 cs.CL cs.AI cs.CV 新提交 81%

Vision-Language Models are Fragile Multilingual Associators

视觉-语言模型是脆弱的多语言关联器

Ritabrata Chakraborty, Rajatsubhra Chakraborty, Shivakumara Palaiahnakote, Angelo Cangelosi, Umapada Pal

机构 * Manipal University Jaipur(斋浦尔马尼帕尔大学) University of North Carolina Charlotte(北卡罗来纳大学夏洛特分校) University of Salford(索尔福德大学) University of Manchester(曼彻斯特大学) Indian Statistical Institute Kolkata(印度统计研究所加尔各答分所)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 该研究针对视觉-语言模型的多语言关联稳定性问题,构建M²BIND基准,发现其跨语系/文字时会出现绑定崩溃,关系密切语言则表现较好,表明多语言部署的VLMs关联质量存疑。

Comments Preprint (under review). Project Page: https://ritabrata04.github.io/m2bind/

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.07550 2026-08-11 cs.CV cs.LG 新提交 81%

Auditing Medical Vision-Language Models on Chest Radiographs: Estimating Reference Agreement Across Institutions

胸部X光医学视觉语言模型审计:估算跨机构参考一致性

Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng, M. Tahar Kechadi

机构 * University College Dublin(都柏林大学学院) School of Computer Science, University College Dublin(都柏林大学学院计算机科学学院) The Third Affiliated Hospital of Southern Medical University(南方医科大学第三附属医院) Dublin City University(都柏林城市大学)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.LG

AI总结 该研究通过评估三个生成式视觉语言模型在胸部X光数据上的表现,估算跨机构参考一致性,发现无法推荐默认估算器,且参考一致性需按站点和接口重新评估。

Comments 10 pages, 3 figures, 3 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.05268 2026-07-17 cs.CV cs.LG 版本更新 81%

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models

几何在发挥作用吗?对双曲视觉-语言模型层次性的工作点审计

Jaeyoung Kim, Eunseok Kim, Dongsuk Jang

机构 * MADI

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.LG

AI总结 本研究针对双曲视觉-语言模型是否利用其几何特性的问题,提出多组诊断指标审计三类主流模型,发现其实际未激活双曲径向/锥机制,层次性表现与几何特性无关。

Comments 48 pages, 5 figures, Under review at TMLR

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.01876 2026-07-03 cs.CV cs.AI 新提交 81%

SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models

SAB-LVLM: 面向大型视觉-语言模型的重要性感知二值化

Qi Lyu, Jiahua Dong, Baichen Liu, Xudong Wang, Mingfei Han, Yulun Zhang, Fahad Shahbaz Khan, Salman Khan, Lianqing Liu, Zhi Han

机构 * State Key Laboratory of Robotics and Intelligent Systems(机器人学国家重点实验室) Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所) University of Chinese Academy of Sciences(中国科学院大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学) Shanghai Jiao Tong University(上海交通大学)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 提出SAB-LVLM方法,通过构建空间重要性图与模态引导整合策略,实现跨层跨模态权重重要性感知的二值化,在约1比特压缩下优于现有二值化方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2606.23897 2026-06-24 cs.CV cs.AI 新提交 81%

The Professor: Multi-Teacher Unsupervised Prompt Distillation for Vision-Language Models

The Professor: 面向视觉语言模型的多教师无监督提示蒸馏

Ahmad Algadhi, Ahmed Alzuhair, Omar Alkhulaif, Muzammil Behzad

机构 * King Fahd University of Petroleum and Minerals(法赫德国王石油矿产大学)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 提出多教师提示蒸馏方法TheProfessor,融合领域微调教师和零样本教师,在四个数据集上平均HM提升1.77点,尤其对域偏移数据集效果显著。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.22080 2026-05-29 cs.CV cs.AI 81%

JMed48k: A Multi-Profession Japanese Medical Licensing Benchmark for Vision-Language Model Evaluation

JMed48k:用于视觉语言模型评估的多专业日本医疗执照基准

Yue Xun, Junyu Liu, Qian Niu, Xinyi Wang, Zheng Yuan, Zirui Li, Zequn Zhang, Bowen Zhao, Shujun Wang, Irene Li, Kan Hatakeyama-Sato, Yusuke Iwasawa, Yutaka Matsuo

机构 * The Hong Kong Polytechnic University(香港理工大学) Kyoto University(京都大学) The University of Tokyo(东京大学) Hohai University(淮海大学) University of Science and Technology of China(中国科学技术大学) University of Toronto(多伦多大学)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出JMed48k,一个包含48,862道试题和20,142张图像的多专业日本医疗执照基准,通过评估21个模型并引入配对图像移除审计,发现专有和开源模型显著受益于图像,而医学专用模型对视觉证据利用有限。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.10576 2026-05-12 cs.CV cs.AI 81%

SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models

SenseBench: 一个用于遥感低级视觉感知与描述的大型视觉-语言模型基准

Chen Zhong, Xiao An, Jiaxing Sun, Zihan Gui, Guangyi Yang, Wei He

机构 * Wuhan University(武汉大学) Shanghai Artificial Intelligent Laboratory(上海人工智能实验室)

专题命中 其他VLM :vision-language model(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出SenseBench,首个专为遥感低级视觉感知与描述设计的基准,通过物理基础的分层分类体系,评估29种先进VLMs在遥感降质识别中的性能,揭示领域先验偏差、多退化崩溃、流畅性幻觉及感知-描述反转效应。

详情

展开后加载摘要…

URL PDF HTML 收藏