arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2510.03441 2025-10-07 cs.CV cs.AI cs.LG 62%

Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning

Chashi Mahiul Islam, Oteo Mamo, Samuel Jacob Chacko, Xiuwen Liu, Weikuan Yu

机构 * Florida State University(佛罗里达州立大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.01681 2025-10-03 cs.CV cs.AI 62%

Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning

Xuchen Li, Xuzhao Li, Jiahui Gao, Renjie Pi, Shiyu Hu, Wentao Zhang

机构 * CASIA(中国科学院自动化研究所) UCAS(中国科学技术大学) ZGCA(北京智感科技有限公司) HKU(香港大学) HKUST(香港科技大学) NTU(国立台湾大学) PKU(北京大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Preprint, Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.07675 2025-10-01 cs.LG cs.AI cs.CV 62%

Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization

Seongjae Kang, Dong Bok Lee, Hyungjoon Jang, Sung Ju Hwang

机构 * VUNO Inc.(VUNO公司) KAIST(韩国科学技术院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 38 pages, 17 figures, preprint

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.24192 2025-09-30 cs.CV cs.AI 62%

Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection

Sojung An, Kwanyong Park, Yong Jae Lee, Donghyun Kim

机构 * Korea University(韩国大学) University of Seoul(首尔大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 23 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15963 2025-09-30 cs.CV cs.CL 62%

OViP: Online Vision-Language Preference Learning for VLM Hallucination

Shujun Liu, Siyuan Wang, Zejun Li, Jianxiang Wang, Cheng Zeng, Zhongyu Wei

机构 * Fudan University(复旦大学) University of Southern California(南加州大学) ByteDance(字节跳动)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.23768 2025-09-29 cs.CL cs.CV 62%

Texture or Semantics? Vision-Language Models Get Lost in Font Recognition

Zhecheng Li, Guoxian Song, Yujun Cai, Zhen Xiong, Junsong Yuan, Yiwei Wang

机构 * University of California, San Diego(加州大学圣地亚哥分校) ByteDance(字节跳动) The University of Queensland(昆士兰大学) University of Southern California(南加州大学) University at Buffalo(布法罗大学) University of California, Merced(加州大学默塞德分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted to COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.23745 2025-09-25 cs.CV cs.AI cs.LG 62%

To Trust Or Not To Trust Your Vision-Language Model's Prediction

Hao Dong, Moru Liu, Jian Liang, Eleni Chatzi, Olga Fink

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.18189 2025-09-24 cs.CV cs.AI 62%

Qianfan-VL: Domain-Enhanced Universal Vision-Language Models

Daxiang Dong, Mingming Zheng, Dong Xu, Bairong Zhuang, Wenyu Zhang, Chunhua Luo, Haoran Wang, Zijian Zhao, Jie Li, Yuxuan Li, Hanjun Zhong, Mengyue Liu, Jieting Chen, Shupeng Li, Lun Tian, Yaping Feng, Xin Li, Donggang Jiang, Yong Chen, Yehua Xu, Duohao Qin, Chen Feng, Dan Wang, Henghua Zhang, Jingjing Ha, Jinhui He, Yanfeng Zhai, Chengxin Zheng, Jiayi Mao, Jiacheng Chen, Ruchang Yao, Ziye Yuan, Jianmin Wu, Guangjun Xie, Dou Shen

机构 * Qianfan Team, Baidu AI Cloud(百度AI云团队)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 12 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.16721 2025-09-23 cs.CV cs.AI cs.RO 62%

Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding

Haoyuan Li, Rui Liu, Hehe Fan, Yi Yang

机构 * College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 19 pages, 12 figures, 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.17974 2025-09-23 cs.CL cs.CV 62%

Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts

Xuyang Wu, Yuan Wang, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang

机构 * Santa Clara University(圣克拉拉大学) DOCOMO Innovations, Inc.(DOCOMO创新公司) Rochester Institute of Technology(罗切斯特理工学院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments EMNLP Findings

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.15490 2025-09-22 cs.CV cs.AI 62%

SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters

Abdarahmane Traore, Éric Hervet, Andy Couturier

机构 * Embia, Computer Science Department, Faculty of Science, Université de Moncton(Embia计算机科学系,科学学院,蒙特龙大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 3 figures, IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13282 2025-09-17 cs.CL cs.CV cs.LG 62%

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan, Mir Rayat Imtiaz Hossain, Leonid Sigal, Giuseppe Carenini

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments EMNLP 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10105 2025-09-17 cs.CV cs.CL 62%

VARCO-VISION-2.0 Technical Report

Young-rok Cha, Jeongho Ju, SunYoung Park, Jong-Hyeon Lee, Younghyun Yu, Youngjune Kim

机构 * NC AI

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 19 pages, 1 figure, 14 tables. Technical report for VARCO-VISION-2.0, a Korean-English bilingual VLM in 14B and 1.7B variants. Key features: multi-image understanding, OCR with text localization, improved Korean capabilities

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.13021 2025-09-17 cs.CL cs.CV 62%

Dynamic Relation Inference via Verb Embeddings

Omri Suissa, Muhiim Ali, Ariana Azarbal, Hui Shen, Shekhar Pradhan

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.15244 2025-09-17 cs.CV cs.AI 62%

Adversarial Prompt Distillation for Vision-Language Models

Lin Luo, Xin Wang, Bojia Zi, Shihao Zhao, Xingjun Ma, Yu-Gang Jiang

机构 * Shanghai Key Lab of Intell. Info. Processing, School of CS, Fudan University(上海智能信息处理实验室,计算机科学学院,复旦大学) The Chinese University of Hong Kong, Shatin, Hong Kong(香港中文大学,沙田,香港) The University of Hong Kong, Pokfulam, Hong Kong(香港大学,薄扶林,香港)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments This work has been submitted to the IEEE for possible publication

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.11895 2025-09-16 cs.CV cs.AI 62%

Integrating Prior Observations for Incremental 3D Scene Graph Prediction

Marian Renz, Felix Igelbrink, Martin Atzmueller

机构 * DFKI Niedersachsen(德克萨斯联合研究所(北莱茵威斯特法伦)) Cooperative and Autonomous Systems, DFKI Niedersachsen(合作与自主系统,DFKI北莱茵威斯特法伦) German Research Center for Artificial Intelligence(德国人工智能研究中心) Semantic Information Systems, Osnabrück University(语义信息系统,奥斯纳布吕克大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at 24th International Conference on Machine Learning and Applications (ICMLA'25)

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.08490 2025-09-11 cs.CV cs.AI 62%

A Structured Review of Underwater Object Detection Challenges and Solutions: From Traditional to Large Vision Language Models

Edwine Nabahirwa, Wei Song, Minghua Zhang, Yi Fang, Zhou Ni

机构 * Shanghai Ocean University(上海海洋大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments 72 Pages, 11 Figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.06130 2025-09-11 cs.CV cs.CL 62%

Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models

Ce Zhang, Zifu Wan, Zhehan Kan, Martin Q. Ma, Simon Stepputtis, Deva Ramanan, Russ Salakhutdinov, Louis-Philippe Morency, Katia Sycara, Yaqi Xie

机构 * School of Computer Science, Carnegie Mellon University(计算机科学系,卡内基梅隆大学)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted by ICLR 2025. Project page: https://zhangce01.github.io/DeGF/

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.06535 2025-09-09 cs.CV cs.AI cs.LG 62%

On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''

Hua Chang Bakker, Stan Fris, Angela Madelon Bernardy, Stan Deutekom

机构 * University of Amsterdam(阿姆斯特丹大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.16647 2025-09-03 cs.CV cs.AI 62%

Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models

Sushant Gautam, Michael A. Riegler, Pål Halvorsen

机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet), Norway(Simula数字工程中心(SimulaMet)) Oslo Metropolitan University (OsloMet), Norway(奥斯陆 Metropolitan 大学(OsloMet)) Simula Research Laboratory, Norway(Simula研究实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted as a full paper at the 38th IEEE International Symposium on Computer-Based Medical Systems (CBMS) 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06794 2025-09-03 cs.CV cs.CL 62%

Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study

Yizheng Sun, Hao Li, Chang Xu, Hongpeng Zhou, Chenghua Lin, Riza Batista-Navarro, Jingyuan Sun

机构 * University of Manchester(曼彻斯特大学) Microsoft Research(微软研究院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL

Comments Accepted to EMNLP 2025 Main Conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.21732 2025-09-01 cs.CV cs.AI 62%

CAD2DMD-SET: Synthetic Generation Tool of Digital Measurement Device CAD Model Datasets for fine-tuning Large Vision-Language Models

João Valente, Atabak Dehban, Rodrigo Ventura

机构 * Institute for Systems and Robotics(系统与机器人研究所) University of Lisbon(里斯本大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10583 2025-08-29 cs.CV cs.CL 62%

Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models

Diogo Freitas, Brigt Håvardstun, Cèsar Ferri, Darío Garigliotti, Jan Arne Telle, José Hernández-Orallo

机构 * Interactive Technologies Institute and NOVA LINCS Faculty of Exact Sciences and Engineering University of Madeira Portugal(互动技术研究所和NOVA LINCS精确科学与工程学院马德拉大学) Department of Informatics University of Bergen Norway(信息学院卑尔根大学挪威) Valencian Research Institute for Artificial Intelligence Universitat Politècnica de València Spain(瓦伦西亚人工智能研究机构瓦伦西亚理工大学西班牙) Leverhulme Centre for the Future of Intelligence and Valencian Research Institute for Artificial Intelligence Spain(未来智能中心和瓦伦西亚人工智能研究机构西班牙)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 54 pages (42 pages of appendix). Accepted for publication at the ECAI 2025 conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.19376 2025-08-28 cs.LG cs.AI cs.CV hep-ex 62%

Fine-Tuning Vision-Language Models for Neutrino Event Analysis in High-Energy Physics Experiments

Dikshant Sagar, Kaiwen Yu, Alejandro Yankelevich, Jianming Bian, Pierre Baldi

机构 * Department of Computer Science University of California, Irvine(计算机科学系加州大学伊文斯顿分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.16569 2025-08-25 eess.IV cs.AI cs.CV 62%

A Disease-Centric Vision-Language Foundation Model for Precision Oncology in Kidney Cancer

Yuhui Tao, Zhongwei Zhao, Zilong Wang, Xufang Luo, Feng Chen, Kang Wang, Chuanfu Wu, Xue Zhang, Shaoting Zhang, Jiaxi Yao, Xingwei Jin, Xinyang Jiang, Yifan Yang, Dongsheng Li, Lili Qiu, Zhiqiang Shao, Jianming Guo, Nengwang Yu, Shuo Wang, Ying Xiong

机构 * Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, 200032, China(复旦大学基础医学学院数字医学研究中心) Shanghai Key Laboratory of Medical Imaging Computing and Computer Assisted Intervention, Shanghai, 200032, China(上海市医疗影像计算与计算机辅助干预重点实验室) Department of Urology, Qilu Hospital of Shandong University, Jinan, Shandong, 250012, China(山东大学齐鲁医院泌尿科) Microsoft Research Asia, Shanghai, 200232, China(微软亚洲研究院) Department of Radiology, The First Affiliated Hospital, Zhejiang University School of Medicine, Hangzhou, 310006, China(浙江大学医学院附属第一医院放射科) Center of Health data science, Linyi People’s Hospital, Shandong, 276003, China(临沂人民医院健康数据科学中心) Shandong Open Laboratory of Data Innovation Application, Shandong, 276003, China(山东省数据创新应用开放实验室) Department of Radiology, the First People’s Hospital of Lianyungang, Lianyungang, 222002, China(连云港第一人民医院放射科) Department of Urology, Zhangye People’s Hospital affiliated to Hexi University, Zhangye, 734000, China(张掖人民医院(河西大学附属)泌尿科) Department of Urology, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, 200025, China(上海交通大学附属瑞金医院泌尿科) Department of Urology, Linyi People’s Hospital, Shandong, 276003, China(临沂人民医院泌尿科) Department of Urology, Zhongshan Hospital, Fudan University, Shanghai, 200032, China(复旦大学中山医院泌尿科)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.01225 2025-08-25 cs.CV cs.AI 62%

Multi-Cache Enhanced Prototype Learning for Test-Time Generalization of Vision-Language Models

Xinyu Chen, Haotian Zhai, Can Zhang, Xiupeng Shi, Ruirui Li

机构 * Shanghai University(上海大学) Beijing University of Chemical Technology(北京化工大学) University of Minnesota(明尼苏达大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted by ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.14045 2025-08-21 cs.CL cs.CV 62%

From Image Captioning to Visual Storytelling

Admitos Passadakis, Yingjin Song, Albert Gatt

机构 * TUDelft(代尔夫特理工大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments 16 pages (including references), 5 figures and 6 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12109 2025-08-19 cs.CV cs.AI 62%

Simple o3: Towards Interleaved Vision-Language Reasoning

Ye Wang, Qianglong Chen, Zejun Li, Siyuan Wang, Shijie Guo, Zhirui Zhang, Zhongyu Wei

机构 * Independent Researcher(独立研究者)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.11317 2025-08-18 cs.CV cs.MM 62%

Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models

Yuchen Zhou, Jiayu Tang, Shuo Yang, Xiaoyan Xiao, Yuqin Dai, Wenhao Yang, Chao Gou, Xiaobo Xia, Tat-Seng Chua

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.MM

详情

展开后加载摘要…

URL PDF HTML 收藏
2403.09346 2025-08-18 cs.CV cs.AI 62%

B-AVIBench: Towards Evaluating the Robustness of Large Vision-Language Model on Black-box Adversarial Visual-Instructions

Hao Zhang, Wenqi Shao, Hong Liu, Yongqiang Ma, Ping Luo, Yu Qiao, Nanning Zheng, Kaipeng Zhang

机构 * National Key Laboratory of Human-Machine Hybrid Augmented Intelligence(人机混合增强智能国家重点实验室) National Engineering Research Center for Visual Information and Applications(视觉信息与应用国家工程研究中心) Institute of Artificial Intelligence and Robotics(人工智能与机器人研究院) Xi’an Jiaotong University(西安交通大学) Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) Osaka University(大阪大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by IEEE Transactions on Information Forensics & Security

详情

展开后加载摘要…

URL PDF HTML 收藏