arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1565 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1565 篇

2505.12795 2025-09-30 cs.AI cs.LG 73%

FRABench and UFEval: Unified Fine-grained Evaluation with Task and Aspect Generalization

Shibo Hong, Jiahao Ying, Haiyuan Liang, Mengdi Zhang, Jun Kuang, Jiazheng Zhang, Yixin Cao

机构 * School of Computer Science, Fudan University(复旦大学计算机学院) Meituan Group(美团集团) School of Computer Science, Singapore Management University(新加坡管理学院)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08830 2025-08-13 cs.AI cs.CV cs.CY 73%

Silicon Minds versus Human Hearts: The Wisdom of Crowds Beats the Wisdom of AI in Emotion Recognition

Mustafa Akben, Vinayaka Gude, Haya Ajjan

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.07076 2025-07-31 cs.CV cs.AI 73%

StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification

Yichen He, Yuan Lin, Jianchao Wu, Hanchong Zhang, Yuchen Zhang, Ruicheng Le

专题命中 其他VLM :vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.02946 2025-07-08 cs.CV cs.AI 73%

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding

Chenglin Li, Qianglong Chen, fengtao, Yin Zhang

机构 * Zhejiang University(浙江大学)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.21184 2025-06-27 cs.CV cs.AI 73%

Task-Aware KV Compression For Cost-Effective Long Video Understanding

Minghao Qin, Yan Shu, Peitian Zhang, Kun Lun, Huaying Yuan, Juenjie Zhou, Shitao Xiao, Bo Zhao, Zheng Liu

机构 * Beijing Academy of Artificial Intelligence(北京人工智能研究院) Shanghai Jiao Tong University(上海交通大学) University of Trento(特伦特大学) Renmin University of China(中国人民大学) Beijing University of Posts and Telecommunications(北京邮电大学) Hong Kong Polytechnic University(香港理工大学) Institute of Automation CAS Beijing China(中国科学院自动化研究所)

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 14 pages, 3 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22613 2025-05-29 cs.CV cs.AI cs.CL 73%

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction

Yuchi Wang, Yishuo Cai, Shuhuai Ren, Sihan Yang, Linli Yao, Yuanxin Liu, Yuanxing Zhang, Pengfei Wan, Xu Sun

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments code: https://github.com/wangyuchi369/RICO

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01064 2025-05-05 cs.CV cs.LG 73%

Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs

Hari Chandana Kuchibhotla, Sai Srinivas Kancheti, Abbavaram Gowtham Reddy, Vineeth N Balasubramanian

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG

Comments preprint; earlier version accepted at NeurIPS 2024 Workshop on Adaptive Foundation Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12315 2025-04-18 cs.CL cs.AI cs.CV 73%

Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models

Xingguang Ji, Jiakang Wang, Hongzhi Zhang, Jingyuan Zhang, Haonan Zhou, Chenxi Sun, Yahui Liu, Qi Wang, Fuzheng Zhang

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.10529 2025-03-14 cs.CV cs.AI 73%

PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models

Zilu Guo, Hongbin Lin, Zhihao Yuan, Chaoda Zheng, Pengshuo Qiu, Dongzhi Jiang, Renrui Zhang, Chun-Mei Feng, Zhen Li

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Technical Report

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.14973 2025-03-07 cs.CL cs.AI cs.LG 73%

GenCeption: Evaluate Vision LLMs with Unlabeled Unimodal Data

Lele Cao, Valentin Buchner, Zineb Senane, Fangkai Yang

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.AI、cs.LG

Comments Published by Computer Speech & Language (https://doi.org/10.1016/j.csl.2025.101785). Source code and Leaderboard: https://github.com/llcresearch/GenCeption

Journal ref Computer Speech & Language 93 (2025) 101785

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.17540 2025-02-26 cs.CV cs.AI cs.CL 73%

PosterSum: A Multimodal Benchmark for Scientific Poster Summarization

Rohit Saxena, Pasquale Minervini, Frank Keller

专题命中 其他VLM :vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments This paper includes a dataset of research posters with abstracts. We provide two cited examples ( arXiv:2211.11880 and arXiv:2210.07571 ) to illustrate reference summaries

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.08737 2024-12-13 cs.CV cs.AI cs.CL 73%

Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions

Jiarui Zhang, Ollie Liu, Tianyu Yu, Jinyi Hu, Willie Neiswanger

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 33 pages, 22 figures, 5 tables, 7 algorithms

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.02611 2024-12-04 cs.CV cs.AI cs.CL cs.MM cs.SD eess.AS 73%

AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?

Kaixiong Gong, Kaituo Feng, Bohao Li, Yibing Wang, Mofan Cheng, Shijia Yang, Jiaming Han, Benyou Wang, Yutong Bai, Zhuoran Yang, Xiangyu Yue

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Project page: https://av-odyssey.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.08303 2024-11-26 cs.CV cs.AI 73%

DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Xiaotong Li, Fan Zhang, Haiwen Diao, Yueze Wang, Xinlong Wang, Ling-Yu Duan

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments Accepted by NeurIPS 2024. Project is available at https://github.com/baaivision/DenseFusion

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.12866 2024-11-13 cs.CL cs.AI cs.CV 73%

How Does the Textual Information Affect the Retrieval of Multimodal In-Context Learning?

Yang Luo, Zangwei Zheng, Zirui Zhu, Yang You

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments EMNLP 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13854 2024-10-18 cs.CL cs.AI cs.CV cs.CY 73%

Can MLLMs Understand the Deep Implication Behind Chinese Images?

Chenhao Zhang, Xi Feng, Yuelin Bai, Xinrun Du, Jinchang Hou, Kaixin Deng, Guangzeng Han, Qinrui Li, Bingli Wang, Jiaheng Liu, Xingwei Qu, Yifei Zhang, Qixuan Zhao, Yiming Liang, Ziqiang Liu, Feiteng Fang, Min Yang, Wenhao Huang, Chenghua Lin, Ge Zhang, Shiwen Ni

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments 32 pages,18 figures. Project Page: https://cii-bench.github.io/ Code: https://github.com/MING_X/CII-Bench Dataset: https://huggingface.co/datasets/m-a-p/CII-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2401.03253 2024-03-19 cs.CV cs.CL cs.LG 73%

VLLaVO: Mitigating Visual Gap through LLMs

Shuhao Chen, Yulong Zhang, Weisen Jiang, Jiangang Lu, Yu Zhang

专题命中 其他VLM :vision-language model(abstract);vision language model(abstract);分类 cs.CV、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2309.11499 2024-03-19 cs.CV cs.CL cs.LG 73%

DreamLLM: Synergistic Multimodal Comprehension and Creation

Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi, Zheng Ge, Jinrong Yang, Liang Zhao, Jianjian Sun, Hongyu Zhou, Haoran Wei, Xiangwen Kong, Xiangyu Zhang, Kaisheng Ma, Li Yi

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.LG

Comments ICLR 2024 (Spotlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.10805 2024-02-19 cs.MM cs.AI cs.CL cs.CV cs.IR 73%

Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond

Yongqi Li, Wenjie Wang, Leigang Qu, Liqiang Nie, Wenjie Li, Tat-Seng Chua

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2308.16741 2023-11-03 cs.AI cs.CV 73%

Socratis: Are large multimodal models emotionally aware?

Katherine Deng, Arijit Ray, Reuben Tan, Saadia Gabriel, Bryan A. Plummer, Kate Saenko

专题命中 其他VLM :vision-language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI

Comments ICCV 2023 WECIA

详情

展开后加载摘要…

URL PDF HTML 收藏
2305.03701 2023-09-29 cs.CV cs.AI 73%

LMEye: An Interactive Perception Network for Large Language Models

Yunxin Li, Baotian Hu, Xinyu Chen, Lin Ma, Yong Xu, Min Zhang

专题命中 其他VLM :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI

Comments working in progress

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.01957 2025-10-27 cs.CV cs.SD eess.AS 72%

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Chaoyou Fu, Haojia Lin, Xiong Wang, Yi-Fan Zhang, Yunhang Shen, Xiaoyu Liu, Haoyu Cao, Zuwei Long, Heting Gao, Ke Li, Long Ma, Xiawu Zheng, Rongrong Ji, Xing Sun, Caifeng Shan, Ran He

机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) Tencent Youtu Lab(腾讯优图实验室) XMU(厦门大学) CASIA(中国科学院自动化研究所)

专题命中 其他VLM :MLLM(abstract,comments);multimodal large language model(abstract);分类 cs.CV

Comments NeurIPS 2025 Spotlight, Code 2.4K Stars: https://github.com/VITA-MLLM/VITA

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.19037 2026-08-10 cs.SD cs.CL cs.MM eess.AS 71%

MLLM-based Speech Recognition: When and How is Multimodality Beneficial?

Yiwen Guan, Viet Anh Trinh, Vivek Voleti, Jacob Whitehill

机构 * Worcester Polytechnic Institute(沃斯特理工学院)

专题命中 其他VLM :MLLM(title)

Journal ref IEEE Transactions on Multimedia, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19267 2026-01-28 cs.CL 71%

DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models

DiaDem: 促进音频视频字幕中的对话描述以提升多模态大语言模型

Xinlong Chen, Weihong Lin, Jingyun Hua, Linli Yao, Yue Ding, Bozhou Li, Bohan Zeng, Yang Shi, Qiang Liu, Yuanxing Zhang, Pengfei Wan, Liang Wang, Tieniu Tan

机构 * New Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所模式识别新实验室) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) Kling Team, Kuaishou Technology(快手科技 Kling 团队) Peking University(北京大学) Nanjing University(南京大学)

专题命中 其他VLM :multimodal large language model(title)

AI总结 DiaDem通过合成高质量数据集和难度分区的两阶段GRPO策略,提升了音频视频字幕中的对话描述准确性,并在多种基准测试中表现出色。

Comments Project webpage: https://diadem-captioner.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23065 2025-12-16 cs.CL 71%

SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services

SNS-Bench-VL:在社交网络服务中评估多模态大语言模型的基准测试

Hongcheng Guo, Zheyong Xie, Shaosheng Cao, Boyang Wang, Weiting Liu, Anjie Le, Lei Li, Zhoujun Li

专题命中 其他VLM :multimodal large language model(title)

AI总结 SNS-Bench-VL是一个用于评估多模态大语言模型在社交网络服务中性能的基准测试,涵盖8个多模态任务,包含4001个问题-答案对,评估25种先进模型,揭示多模态社交理解的挑战。

Comments We found problems in the code while rechecking our implementation. These issues led to noticeable numerical discrepancies, making some of the reported results and conclusions potentially unreliable. Therefore, we request to withdraw this submission

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.13225 2025-11-18 cs.CL 71%

Seeing isn't Hearing: Benchmarking Vision Language Models at Interpreting Spectrograms

Tyler Loakman, Joseph James, Chenghua Lin

机构 * Department of Computer Science, University of Sheffield, UK(谢菲尔德大学计算机科学系) Department of Computer Science, University of Manchester, UK(曼彻斯特大学计算机科学系)

专题命中 其他VLM :vision language model(title)

Comments Accepted to IJCNLP-AACL 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.13911 2025-10-17 q-bio.QM 71%

OralGPT: A Two-Stage Vision-Language Model for Oral Mucosal Disease Diagnosis and Description

Jia Zhang, Bodong Du, Yitong Miao, Dongwei Sun, Xiangyong Cao

专题命中 其他VLM :vision-language model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02665 2025-10-06 cs.CL 71%

Self-Improvement in Multimodal Large Language Models: A Survey

Shijian Deng, Kai Wang, Tianyu Yang, Harsh Singh, Yapeng Tian

机构 * The University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Toronto(多伦多大学) University of Notre Dame(诺特丹大学) Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

专题命中 其他VLM :multimodal large language model(title)

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22447 2025-05-29 cs.CR 71%

Privacy-preserving Prompt Personalization in Federated Learning for Multimodal Large Language Models

Sizai Hou, Songze Li, Baturalp Buyukates

专题命中 其他VLM :multimodal large language model(title)

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.02777 2023-10-05 cs.CL 71%

The Role of Linguistic Priors in Measuring Compositional Generalization of Vision-Language Models

Chenwei Wu, Li Erran Li, Stefano Ermon, Patrick Haffner, Rong Ge, Zaiwei Zhang

专题命中 其他VLM :vision-language model(title)

详情

展开后加载摘要…

URL PDF HTML 收藏