arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 1566 信号源:cs.CV, cs.AI, cs.LG

1. 其他VLM 1566 篇

2506.08817 2025-06-13 cs.CV 57%

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought

Shuyi Zhang, Xiaoshuai Hao, Yingbo Tang, Lingfeng Zhang, Pengwei Wang, Zhongyuan Wang, Hongxuan Ma, Shanghang Zhang

机构 * Institute of Automation, CAS(中国科学院自动化研究所) School of Artifcial Intelligence, UCAS(中国科学技术大学人工智能学院) Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院) Shenzhen International GraduateSchool,Tsinghua University(深圳国际研究生院,清华大学) State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09175 2025-06-12 cs.CL cs.AI cs.SD eess.AS 57%

PHRASED: Phrase Dictionary Biasing for Speech Translation

Peidong Wang, Jian Xue, Rui Zhao, Junkun Chen, Aswin Shanmugam Subramanian, Jinyu Li

机构 * Microsoft USA(微软公司)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.12559 2025-06-10 cs.CV cs.CL cs.MM 57%

AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Xiao Wang, Qingyi Si, Jianlong Wu, Shiyu Zhu, Li Cao, Liqiang Nie

机构 * Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳)) Huawei Technologies Co., Ltd.(华为技术有限公司) Shandong University(山东大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.06173 2025-06-10 cs.CV 57%

VideoAuteur: Towards Long Narrative Video Generation

Junfei Xiao, Feng Cheng, Lu Qi, Liangke Gui, Jiepeng Cen, Zhibei Ma, Alan Yuille, Lu Jiang

机构 * Johns Hopkins University(约翰霍普金斯大学) ByteDance Project(字节跳动项目)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Preprint, https://videoauteur.github.io/; V2: Method is updated

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24870 2025-06-09 cs.CV 57%

GenSpace: Benchmarking Spatially-Aware Image Generation

Zehan Wang, Jiayang Xu, Ziang Zhang, Tianyu Pang, Chao Du, Hengshuang Zhao, Zhou Zhao

机构 * Zhejiang University(浙江大学) Sea AI Lab(海思人工智能实验室) The University of Hong Kong(香港大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.05080 2025-06-06 cs.CL cs.CV 57%

Parking, Perception, and Retail: Street-Level Determinants of Community Vitality in Harbin

HaoTian Lan

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 22 pages,5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03511 2025-06-05 astro-ph.EP astro-ph.IM cs.AI eess.IV 57%

POLARIS: A High-contrast Polarimetric Imaging Benchmark Dataset for Exoplanetary Disk Representation Learning

Fangyi Cao, Bin Ren, Zihao Wang, Shiwei Fu, Youbin Mo, Xiaoyang Liu, Yuzhou Chen, Weixin Yao

机构 * UC Riverside(加州大学河滨分校) OCA/MPIA(天文台/马克斯·普朗克研究所) UT Chattanooga(田纳西大学查塔努加分校) MGH(麻省总医院) Harvard(哈佛大学) UC San Diego(加州大学圣地亚哥分校) Adobe(Adobe公司)

专题命中 其他VLM :vision-language model(abstract);分类 cs.AI

Comments 9 pages main text with 5 figures, 9 pages appendix with 9 figures. Submitted to NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.15167 2025-06-05 cs.CV 57%

M3-AGIQA: Multimodal, Multi-Round, Multi-Aspect AI-Generated Image Quality Assessment

Chuan Cui, Kejiang Chen, Zhihua Wei, Wen Shen, Weiming Zhang, Nenghai Yu

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments 24 pages. This work has been submitted to the ACM for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03414 2025-06-04 cs.CV cs.CL 57%

Enhancing Target-unspecific Tasks through a Features Matrix

Fangming Cui, Yonggang Zhang, Xuan Wang, Xinmei Tian, Jun Yu

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted by ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01902 2025-06-03 cs.CV cs.CL 57%

Enhancing Biomedical Multi-modal Representation Learning with Multi-scale Pre-training and Perturbed Report Discrimination

Xinliu Zhong, Kayhan Batmanghelich, Li Sun

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments 6 pages, 1 figure, accepted by 2024 IEEE Conference on Artificial Intelligence (CAI)

Journal ref 2024 IEEE Conference on Artificial Intelligence (CAI), 2024, 480-485

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.01853 2025-06-03 cs.CV 57%

ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

Junliang Ye, Zhengyi Wang, Ruowen Zhao, Shenghao Xie, Jun Zhu

机构 * Tsinghua University(清华大学) Peking University(北京大学) ShengShu(盛舒)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Project page: https://github.com/JAMESYJL/ShapeLLM-Omni

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13882 2025-06-03 cs.CV 57%

Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model

Long Le, Jason Xie, William Liang, Hung-Ju Wang, Yue Yang, Yecheng Jason Ma, Kyle Vedder, Arjun Krishna, Dinesh Jayaraman, Eric Eaton

机构 * University of Pennsylvania(宾夕法尼亚大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments ICLR 2025. Project website and open-source code: https://articulate-anything.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.18672 2025-06-03 cs.CV 57%

FactCheXcker: Mitigating Measurement Hallucinations in Chest X-ray Report Generation Models

Alice Heiman, Xiaoman Zhang, Emma Chen, Sung Eun Kim, Pranav Rajpurkar

机构 * Stanford University, USA(斯坦福大学) Harvard University, USA(哈佛大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24158 2025-06-02 cs.CV 57%

Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Bo Fang, Wenhao Wu, Qiangqiang Wu, Yuxin Song, Antoni B. Chan

机构 * Department of Computer Science, City University of Hong Kong(香港城市大学计算机科学系) Baidu Inc.(百度公司) University of Sydney(悉尼大学)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22858 2025-05-30 cs.CV 57%

A Probabilistic Jump-Diffusion Framework for Open-World Egocentric Activity Recognition

Sanjoy Kundu, Shanmukha Vellamcheti, Sathyanarayanan N. Aakur

机构 * CSSE Department, Auburn University(安全科学与工程系,阿伯丁大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Extended abstract of arXiv:2504.03948 for CVPR 2025 EgoVis Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22150 2025-05-30 cs.CV cs.CL 57%

Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging

Runze Xia, Shuo Feng, Renzhi Wang, Congchi Yin, Xuyun Wen, Piji Li

机构 * College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(人工智能学院,南京航空航天大学) MIIT Key Laboratory of Pattern Analysis and Machine Intelligence(信息科技部模式分析与机器智能重点实验室) The Key Laboratory of Brain-Machine Intelligence Technology, Ministry of Education(教育部脑机智能技术重点实验室)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments CogSci2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.22481 2025-05-29 stat.ML cs.LG 57%

Hypothesis Testing in Imaging Inverse Problems

Yiming Xi, Konstantinos Zygalakis, Marcelo Pereyra

机构 * School of Mathematical and Computer Sciences, Heriot-Watt University(赫瑞斯泰学院数学与计算机科学系,赫瑞瓦特大学) School of Mathematics, University of Edinburgh(爱丁堡大学数学学院) Maxwell Institute for Mathematical Sciences(麦克斯韦数学科学研究所)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13953 2025-05-29 cs.CL cs.AI 57%

Redundancy Principles for MLLMs Benchmarks

Zicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, Guangtao Zhai

机构 * Shanghai AI Laboratory(上海人工智能实验室) Shanghai Jiaotong University(上海交通大学) Zhejiang University(浙江大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09441 2025-05-28 cs.CV eess.IV 57%

Structure-Accurate Medical Image Translation via Dynamic Frequency Balance and Knowledge Guidance

Jiahua Xu, Dawei Zhou, Lei Hu, Zaiyi Liu, Nannan Wang, Xinbo Gao

机构 * Xidian University(西安电子科技大学) Guangdong Provincial People’s Hospital(广东省人民医院) Chongqing University of Posts and Telecommunications(重庆邮电大学)

专题命中 其他VLM :visual language model(abstract);分类 cs.CV

Comments Medical image translation, Diffusion model, 16 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.08396 2025-05-27 cs.CV cs.CL 57%

Level Up Your Tutorials: VLMs for Game Tutorials Quality Assessment

Daniele Rege Cambrin, Gabriele Scaffidi Militone, Luca Colomba, Giovanni Malnati, Daniele Apiletti, Paolo Garza

机构 * Politecnico di Torino(托里尼理工大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted at ECCV 2024 CV2 Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.01263 2025-05-27 cs.CV cs.CL 57%

Generalizable Prompt Learning of CLIP: A Brief Overview

Fangming Cui, Yonggang Zhang, Xuan Wang, Xule Wang, Liang Xiao

机构 * Meituan(美团)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12746 2025-05-26 cs.AI 57%

Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs

Haruka Asanuma, Naoko Koide-Majima, Ken Nakamura, Takato Horii, Shinji Nishimoto, Masafumi Oizumi

机构 * The University of Tokyo, Graduate School of Arts and Sciences(东京大学艺术与科学研究生院) Center for Information and Neural Networks (CiNet), National Institute of Information and Communications Technology(信息与神经网络中心(CiNet),信息与通信技术国家研究所) The University of Osaka, Graduate School of Frontier Biosciences(大阪大学前沿生命科学研究生院) The University of Tokyo, Faculty of Engineering(东京大学工学部) The University of Osaka, Graduate School of Engineering Science(大阪大学工学研究院) The University of Osaka, Graduate School of Medicine(大阪大学医学研究院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.AI

Comments 25 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.15415 2025-05-23 cs.CV 57%

OCSU: Optical Chemical Structure Understanding for Molecule-centric Scientific Discovery

Siqi Fan, Yuguang Xie, Bowen Cai, Ailin Xie, Gaochao Liu, Mu Qiao, Jie Xing, Zaiqing Nie

机构 * Institute for AI Industry Research (AIR), Tsinghua University(人工智能产业研究所(AIR),清华大学) PharMolix Inc.(PharMolix公司)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15145 2025-05-22 cs.CV 57%

CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation

Xinran Wang, Songyu Xu, Xiangxuan Shan, Yuxuan Zhang, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, Zhanyu Ma

机构 * Beijing University of Posts and Telecommunications(北京邮电大学) China Mobile Research Institute(中国移动研究院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13669 2025-05-21 cs.CV cs.RO 57%

GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching

Barkin Dagda, Muhammad Awais, Saber Fallah

机构 * Connected and Autonomous Vehicles Lab (CAV-Lab)(连接与自动驾驶车辆实验室) University of Surrey(萨里大学) Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音和信号处理中心)

专题命中 其他VLM :vision language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.13082 2025-05-20 cs.SD cs.AI eess.AS 57%

MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers

Kyeongman Park, Seongho Joo, Kyomin Jung

机构 * Seoul National University(首尔国立大学)

专题命中 其他VLM :MLLM(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12660 2025-05-20 cs.CV 57%

Predicting Reaction Time to Comprehend Scenes with Foveated Scene Understanding Maps

Ziqi Wen, Jonathan Skaza, Shravan Murlidaran, William Y. Wang, Miguel P. Eckstein

机构 * Department of Computer Science University of California Santa Barbara(计算机科学系加州大学圣芭芭拉分校) Graduate Program in Dynamical Neuroscience University of California Santa Barbara(动态神经科学研究生项目加州大学圣芭芭拉分校) Department of Psychological and Brain Sciences University of California Santa Barbara(心理学与脑科学系加州大学圣芭芭拉分校)

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.13372 2025-05-20 cs.GR cs.CV 57%

MoVer: Motion Verification for Motion Graphics Animations

Jiaju Ma, Maneesh Agrawala

专题命中 其他VLM :vision-language model(abstract);分类 cs.CV

Comments Accepted to ACM Transactions on Graphics (SIGGRAPH 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.11865 2025-05-19 cs.CV 57%

From Image to Video, what do we need in multimodal LLMs?

Suyuan Huang, Haoxin Zhang, Linqing Zhong, Honggu Chen, Yan Gao, Yao Hu, Zengchang Qin

机构 * Intelligent Computing and Machine Learning Lab, School of ASEE, Beihang University(北京航空航天大学自动化学院智能计算与机器学习实验室) Xiaohongshu(小红书) School of Sino-French Engineer, Beihang University(北京航空航天大学中法工程师学院) College of Engineering and Computer Science, VinUniversity(Vin大学工程与计算机科学学院)

专题命中 其他VLM :multimodal large language model(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09659 2025-05-16 cs.LG cs.CL 57%

LAS: Loss-less ANN-SNN Conversion for Fully Spike-Driven Large Language Models

Long Chen, Xiaotian Song, Yanan Sun

机构 * Collage of Computer Science, Sichuan University(计算机科学学院,四川大学)

专题命中 其他VLM :vision-language model(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏