arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 5048 信号源:cs.CV, cs.AI, cs.LG

1. VLM训练与架构 5048 篇

2601.01322 2026-01-06 cs.CV cs.AI cs.LG cs.MM eess.IV 75%

LinMU: Multimodal Understanding Made Linear

LinMU: 使多模态理解线性化

Hongjie Wang, Niraj K. Jha

机构 * Princeton University(普林斯顿大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 LinMU通过线性复杂度设计实现多模态理解,无需二次注意力模块,提升视频处理效率。

Comments 23 pages, 7 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.03195 2025-12-01 cs.CV cs.AI cs.LG 75%

Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs

未标记数据提升多模态大语言模型在细粒度图像零样本分类中的性能

Yunqi Hong, Sohyun An, Andrew Bai, Neil Y. C. Lin, Cho-Jui Hsieh

机构 * Computer Science Department, University of California, Los Angeles(加州大学洛杉矶分校计算机科学系) Mechanical and Aerospace Engineering Department, University of California, Los Angeles(加州大学洛杉矶分校机械与航空航天工程系)

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 AutoSEP通过利用未标记数据提升多模态大语言模型在细粒度图像零样本分类中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.01082 2025-11-04 cs.CV cs.AI cs.LG 75%

GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction

Narges Ghasemi, Amir Ziashahabi, Salman Avestimehr, Cyrus Shahabi

机构 * of Computer Science, University of Southern California, Los Angeles, CA, USA Computer Engineering, University of Southern California, Los Angeles, CA, USA

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to IEEE International Conference on Data Mining (ICDM) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22370 2025-10-28 cs.RO cs.AI cs.CV cs.LG cs.SE 75%

BLIP-FusePPO: A Vision-Language Deep Reinforcement Learning Framework for Lane Keeping in Autonomous Vehicles

Seyed Ahmad Hosseini Miangoleh, Amin Jalal Aghdasian, Farzaneh Abdollahi

机构 * Department of Electrical Engineering, Amirkabir University of Technology (Tehran Polytechnic)(电气工程系,阿米尔卡比尔技术大学(德黑兰理工大学))

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments https://github.com/Amin-A96/BLIP-FusePPO-A-Vision-Language-Deep-Reinforcement-Learning-Framework-for-Lane-Keeping-in-Autonomous.git

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18502 2025-10-22 cs.CV cs.AI cs.CL cs.LG 75%

Zero-Shot Vehicle Model Recognition via Text-Based Retrieval-Augmented Generation

Wei-Chia Chang, Yan-Ann Chen

机构 * Yuan Ze University(元智大学)

专题命中 VLM训练与架构 :vision language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted by The 38th Conference of Open Innovations Association FRUCT, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.14968 2025-10-17 cs.RO cs.AI cs.CV cs.LG cs.SY eess.SY 75%

RDD: Retrieval-Based Demonstration Decomposer for Planner Alignment in Long-Horizon Tasks

Mingxuan Yan, Yuping Wang, Zechun Liu, Jiachen Li

机构 * University of California, Riverside(加州大学河滨分校) University of Michigan(密歇根大学) Meta AI

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 39th Conference on Neural Information Processing Systems (NeurIPS 2025); Project Website: rdd-neurips.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.11835 2025-10-15 cs.CV cs.AI cs.CL cs.LG cs.MM 75%

Data or Language Supervision: What Makes CLIP Better than DINO?

Yiming Liu, Yuhui Zhang, Dhruba Ghosh, Ludwig Schmidt, Serena Yeung-Levy

机构 * Stanford University(斯坦福大学) Tsinghua University(清华大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments EMNLP 2025 Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.26625 2025-10-01 cs.LG cs.AI cs.CV cs.MM 75%

Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training

Junlin Han, Shengbang Tong, David Fan, Yufan Ren, Koustuv Sinha, Philip Torr, Filippos Kokkinos

机构 * Meta Superintelligence Labs(Meta 超智能实验室) University of Oxford(牛津大学)

专题命中 VLM训练与架构 :visual reasoning(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Project page: https://junlinhan.github.io/projects/lsbs/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22697 2025-09-30 cs.CV cs.AI cs.LG 75%

Learning Hyperspectral Images with Curated Text Prompts for Efficient Multimodal Alignment

Abhiroop Chatterjee, Susmita Ghosh

机构 * Jadavpur University(贾瓦帕尔大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at the IEEE/CVF International Conference on Computer Vision (ICCV 2025), Workshop on Curated Data for Efficient Learning

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.22195 2025-09-29 cs.RO 75%

Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting

Asher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky, Anirudha Majumdar

机构 * Princeton University(普林斯顿大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);visual question answering(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.21107 2025-09-26 cs.RO cs.AI cs.CV cs.LG 75%

Cross-Modal Instructions for Robot Motion Generation

William Barron, Xiaoxiang Dong, Matthew Johnson-Roberson, Weiming Zhi

机构 * College of Connected Computing, Vanderbilt University(连接计算学院,范德比尔特大学) Robotics Institute, Carnegie Mellon University(机器人研究所,卡内基梅隆大学) School of Computer Science, The University of Sydney(计算机科学学院,悉尼大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18938 2025-09-24 cs.CV cs.AI cs.LG 75%

No Labels Needed: Zero-Shot Image Classification with Collaborative Self-Learning

Matheus Vinícius Todescato, Joel Luís Carbonera

机构 * Institute of Informatics(信息学院) UFRGS(乌拉圭国家研究学院)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments This paper was accepted at International Conference on Tools with Artificial Intelligence (ICTAI) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10344 2025-09-15 cs.CV cs.AI cs.LG 75%

GLAM: Geometry-Guided Local Alignment for Multi-View VLP in Mammography

Yuexi Du, Lihui Chen, Nicha C. Dvornek

机构 * Department of Biomedical Engineering(生物医学工程系) Department of Radiology & Biomedical Imaging(放射科与生物医学成像系) Yale University(耶鲁大学)

专题命中 VLM训练与架构 :VLM(abstract);visual language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted by MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.09955 2025-09-15 cs.LG cs.AI cs.CV eess.IV 75%

Adaptive Token Merging for Efficient Transformer Semantic Communication at the Edge

Omar Erak, Omar Alhussein, Hatem Abou-Zeid, Mehdi Bennis, Sami Muhaidat

机构 * Department of Electrical Engineering and Computer Science(电气工程与计算机科学系) Khalifa University(卡里马大学)

专题命中 VLM训练与架构 :LLaVA(abstract);visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Submitted to IEEE Journals

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08338 2025-09-11 cs.CV cs.AI cs.LG 75%

Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis

Jihyun Moon, Charmgil Hong

机构 * Handong Global University(-handong全球大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Medical Image Computing and Computer-Assisted Intervention (MICCAI) ISIC Skin Image Analysis Workshop (MICCAI ISIC) 2025; 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.01824 2025-08-28 cs.CV cs.AI cs.LG cs.MM 75%

X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

Zeyi Sun, Ziyang Chu, Pan Zhang, Tong Wu, Xiaoyi Dong, Yuhang Zang, Yuanjun Xiong, Dahua Lin, Jiaqi Wang

机构 * Shanghai Jiao Tong University(上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) CPII under InnoHK(创新香港科技促进会) MThreads AI

专题命中 VLM训练与架构 :vision-language model(abstract);vision language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments code: https://github.com/SunzeY/X-Prompt

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.04680 2025-08-20 cs.LG cs.AI cs.CV 75%

Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving Distillation

Wenhao Li, Xiu Su, Jingyi Wu, Feng Yang, Yang Liu, Yi Chen, Shan You, Chang Xu

专题命中 VLM训练与架构 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments In Figure 2, the correlation coefficient and the scatter plot do not match. I calculated this correlation using two sets of settings. I used the scatter plot from setting A, but accidentally wrote the correlation coefficient, r, from setting B

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.19378 2025-08-05 cs.CV cs.AI cs.CL cs.LG 75%

Libra: Leveraging Temporal Images for Biomedical Radiology Analysis

Xi Zhang, Zaiqiao Meng, Jake Lever, Edmond S. L. Ho

机构 * Information Retrieval Group(信息检索组) AI4BioMed Lab(AI4BioMed实验室) School of Computing Science(计算科学学院) University of Glasgow(格拉斯哥大学)

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments 30 pages, 5 figures, Adding Appendix

Journal ref Association for Computational Linguistics, 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20110 2025-07-29 cs.CV cs.AI cs.LG 75%

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding

Shiyu Liu, Lianlei Shan

机构 * School of Electrical and Electronic Engineering(电气与电子工程学院) Nanyang Technological University(南洋理工大学) School of Computer Science and Technology(计算机科学与技术学院) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 VLM训练与架构 :visual language model(abstract);multimodal large language model(abstract);分类 cs.CV、cs.AI、cs.LG

Comments **14 pages, 3 figures, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16572 2025-07-23 cs.CL 75%

Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models

Mohamad Ballout, Serwan Jassim, Elia Bruni

机构 * Institute of Cognitive Science, University of Osnabrück(认知科学研究所,奥斯纳布吕克大学)

专题命中 VLM训练与架构 :LLaVA(abstract);InternVL(abstract);multimodal large language model(abstract)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.07525 2025-07-23 cs.CV cs.AI cs.LG 75%

RadAlign: Advancing Radiology Report Generation with Vision-Language Concept Alignment

Difei Gu, Yunhe Gao, Yang Zhou, Mu Zhou, Dimitris Metaxas

机构 * Rutgers University(罗格斯大学) Stanford University(斯坦福大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19944 2025-05-27 cs.CV cs.AI cs.CL cs.LG 75%

Can Visual Encoder Learn to See Arrows?

Naoyuki Terashita, Yusuke Tozaki, Hideaki Omote, Congkha Nguyen, Ryosuke Nakamoto, Yuta Koreeda, Hiroaki Ozaki

机构 * Hitachi, Ltd.(日本日立株式会社) Kyoto Sangyo University(京都 Sangyo 大学) Gifu University(岐阜大学)

专题命中 VLM训练与架构 :vision language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments This work has been accepted for poster presentation at the Second Workshop on Visual Concepts in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.18578 2025-05-27 cs.LG cs.AI cs.CV 75%

Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding

Tianyu Chen, Xingcheng Fu, Yisen Gao, Haodong Qian, Yuecen Wei, Kun Yan, Haoyi Zhou, Jianxin Li

机构 * SKLCCSE, School of Computer Science and Engineering, Beihang University, China(信息与通信工程学院,北京航空航天大学) School of Software, Beihang University, China(软件学院,北京航空航天大学) Key Lab of Education Blockchain and Intelligent Technology, Guangxi Normal University, China(教育区块链与智能技术重点实验室,广西师范大学) Institute of Artificial Intelligence, Beihang University, Beijing, China(人工智能研究院,北京航空航天大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments CVPR(Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17316 2025-05-26 cs.CV cs.AI cs.CL cs.LG 75%

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Jiachen Jiang, Jinxin Zhou, Bo Peng, Xia Ning, Zhihui Zhu

机构 * Department of Computer Science and Engineering, The Ohio State University(计算机科学与工程系,俄亥俄州立大学) Translational Data Analytics Institute, The Ohio State University(转化数据分析研究所,俄亥俄州立大学) Department of Biomedical Informatics, The Ohio State University(生物医学信息学系,俄亥俄州立大学)

专题命中 VLM训练与架构 :grounding(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.19627 2025-05-20 cs.CL cs.AI cs.CV cs.LG 75%

VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning

Run Luo, Renke Shan, Longze Chen, Ziqiang Liu, Lu Wang, Min Yang, Xiaobo Xia

机构 * Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究所) University of Chinese Academy of Sciences(中国科学院大学) National University of Singapore(新加坡国立大学) MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(中国科学技术大学脑启发智能感知与认知教育部重点实验室)

专题命中 VLM训练与架构 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments VCM

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.02370 2025-05-06 cs.CV cs.AI cs.LG 75%

SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing

Ming Li, Xin Gu, Fan Chen, Xiaoying Xing, Longyin Wen, Chen Chen, Sijie Zhu

机构 * ByteDance Intelligent Creation (USA)(字节跳动智能创作(美国)) Center for Research in Computer Vision, University of Central Florida(计算机视觉研究中心,中央佛罗里达大学)

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Code, Data and Models are available at: https://github.com/bytedance/SuperEdit

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.02885 2025-04-09 cs.CV cs.AI cs.LG 75%

Expertized Caption Auto-Enhancement for Video-Text Retrieval

Baoyao Yang, Junxiang Chen, Wanyun Li, Wenbin Yao, Yang Zhou

专题命中 VLM训练与架构 :multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI、cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.07591 2025-04-08 cs.CV cs.AI cs.LG 75%

Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning

Bardia Safaei, Faizan Siddiqui, Jiacong Xu, Vishal M. Patel, Shao-Yuan Lo

专题命中 VLM训练与架构 :vision-language model(abstract);LLaVA(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted at CVPR 2025 (Highlight)

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11912 2025-04-01 cs.CV cs.AI cs.LG 75%

F$^3$OCUS -- Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-Heuristics

Pramit Saha, Felix Wagner, Divyanshu Mishra, Can Peng, Anshul Thakur, David Clifton, Konstantinos Kamnitsas, J. Alison Noble

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted in CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.11148 2025-03-25 cs.CV cs.AI cs.LG 75%

Few-Shot Recognition via Stage-Wise Retrieval-Augmented Finetuning

Tian Liu, Huixin Zhang, Shubham Parashar, Shu Kong

专题命中 VLM训练与架构 :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI、cs.LG

Comments Accepted to CVPR 2025. Website and code: https://tian1327.github.io/SWAT/

详情

展开后加载摘要…

URL PDF HTML 收藏