arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 46073 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4651 篇

2505.19031 2025-05-27 cs.CV cs.AI 62%

Medical Large Vision Language Models with Multi-Image Visual Ability

Xikai Yang, Juzheng Miao, Yuchen Yuan, Jiaze Wang, Qi Dou, Jinpeng Li, Pheng-Ann Heng

机构 * Dept. of Computer Science and Engineering, The Chinese University of Hong Kong, Hong Kong, China(计算机科学与工程系,香港中文大学,香港,中国) Institute of Medical Intelligence and XR, The Chinese University of Hong Kong, Hong Kong, China(医学智能与XR研究所,香港中文大学,香港,中国)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

Comments 10 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18434 2025-05-27 cs.CV cs.AI 62%

TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP

Yuliang Cai, Jesse Thomason, Mohammad Rostami

机构 * University of Southern California(南加州大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 15 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17625 2025-05-26 cs.CL cs.CV 62%

Enhancing Large Vision-Language Models with Layout Modality for Table Question Answering on Japanese Annual Securities Reports

Hayato Aida, Kosuke Takahashi, Takahiro Omi

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Accepted at IIAI AAI 2025, the 3rd International Conference on Computational and Data Sciences in Economics and Finance

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.17425 2025-05-26 cs.CV cs.CL 62%

Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads

Wei Jie Yeo, Rui Mao, Moloud Abdar, Erik Cambria, Ranjan Satapathy

机构 * Nanyang Technological University(南洋理工大学) The University of Queensland(昆士兰大学) IHPC(高性能计算中心)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.06442 2025-05-21 cs.CV cs.MM 62%

OT-DETECTOR: Delving into Optimal Transport for Zero-shot Out-of-Distribution Detection

Yu Liu, Hao Tang, Haiqi Zhang, Jing Qin, Zechao Li

机构 * School of Computer Science and Engineering(计算机科学与工程学院) Centre for Smart Health(智能健康中心)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.MM

Comments Accepted to the 34th International Joint Conference on Artificial Intelligence (IJCAI 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12835 2025-05-20 cs.CL cs.CV 62%

FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models

Hengxing Cai, Jinhan Dong, Jingjun Tan, Jingcheng Deng, Sihang Li, Zhifeng Gao, Haidong Wang, Zicheng Su, Agachai Sumalee, Renxin Zhong

机构 * School of Intelligent Systems Engineering, Sun Yat-Sen University(中山大学智能系统工程学院) DP Technology(DP技术公司) Beijing University Of Posts and Telecommunications(北京邮电大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所) Tongji University(同济大学) School of Integrated Innovation, Chulalongkorn University(朱拉隆功大学创新学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.07830 2025-05-20 cs.CV cs.AI cs.LG 62%

Captured by Captions: On Memorization and its Mitigation in CLIP Models

Wenhao Wang, Adam Dziedzic, Grace C. Kim, Michael Backes, Franziska Boenisch

机构 * CISPA Georgia Institute of Technology(佐治亚理工学院)

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

Comments Accepted at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.11060 2025-05-19 cs.CV cs.AI 62%

CUBIC: Concept Embeddings for Unsupervised Bias Identification using VLMs

David Méndez, Gianpaolo Bontempo, Elisa Ficarra, Roberto Confalonieri, Natalia Díaz-Rodríguez

机构 * Dept. of Computer Science and Artificial Intelligence, DaSCI Institute, University of Granada(计算机科学与人工智能系,DaSCI研究所,格拉纳达大学) Dept. of Engineering ”Enzo Ferrari”, University of Modena and Reggio Emilia(工程系,摩德纳和雷吉奥艾米利亚大学) Dept. of Mathematics ’Tullio Levi-Civita’, University of Padova(数学系,帕多瓦大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 3 figures, 5 tables. Accepted at IJCNN 2025; to appear in IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10664 2025-05-19 cs.CV cs.AI 62%

CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier

Ziyang Ou

机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) University of Rochester(罗切斯特大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 5 figures, not submitted to any conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.10453 2025-05-16 cs.CV cs.AI 62%

Vision language models have difficulty recognizing virtual objects

Tyler Tran, Sangeet Khemlani, J. G. Trafton

机构 * US Naval Research Laboratory(美国海军研究实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.08910 2025-05-16 cs.CV cs.CL 62%

Behind Maya: Building a Multilingual Vision Language Model

Nahid Alam, Karthik Reddy Kanjula, Surya Guthikonda, Timothy Chung, Bala Krishna S Vegesna, Abhipsha Das, Anthony Susevski, Ryan Sze-Yin Chan, S M Iftekhar Uddin, Shayekh Bin Islam, Roshan Santhosh, Snegha A, Drishti Sharma, Chen Liu, Isha Chaturvedi, Genta Indra Winata, Ashvanth. S, Snehanshu Mukherjee, Alham Fikri Aji

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments Accepted at VLMs4ALL CVPR 2025 Workshop; corrected workshop name spelling

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.09435 2025-05-15 cs.CV cs.AI 62%

Endo-CLIP: Progressive Self-Supervised Pre-training on Raw Colonoscopy Records

Yili He, Yan Zhu, Peiyao Fu, Ruijie Yang, Tianyi Chen, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang

机构 * Digital Medical Research Center, School of Basic Medical Sciences, Fudan University, Shanghai, China(上海复旦大学基础医学学院数字医学研究中心) University College London, London, UK(伦敦大学学院) Shanghai Key Laboratory of MICCAI, Shanghai, China(上海MICCAI重点实验室) Endoscopy Center and Endoscopy Research Institute, Zhongshan Hospital, Fudan University, Shanghai, China(复旦大学中山医院内窥镜中心和内窥镜研究所) Shanghai Collaborative Innovation Center of Endoscopy, Shanghai, China(上海内窥镜协同创新中心) Shanghai Institute for Advanced Study of Zhejiang University, Shanghai, China(浙江大学上海高级研究院) Alliance Manchester Business School, The University of Manchester, Manchester, UK(曼彻斯特大学曼彻斯特商业学校) Data Science Institute, Imperial College London, London, UK(伦敦帝国理工学院数据科学研究院)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments Early accepted to MICCAI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09425 2025-05-14 cs.CV cs.CL 62%

Vision-Language Models Do Not Understand Negation

Kumail Alhamoud, Shaden Alshammari, Yonglong Tian, Guohao Li, Philip Torr, Yoon Kim, Marzyeh Ghassemi

机构 * Institution1(机构1) Institution2(机构2)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

Comments CVPR 2025; project page: https://negbench.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05318 2025-05-09 cs.CV cs.AI cs.CY cs.HC cs.RO 62%

Mapping User Trust in Vision Language Models: Research Landscape, Challenges, and Prospects

Agnese Chiatti, Sara Bernardini, Lara Shibelski Godoy Piccolo, Viola Schiaffonati, Matteo Matteucci

机构 * Politecnico di Milano, Italy(米兰理工学院,意大利) University of Oxford, UK(牛津大学,英国) CODE University of Applied Sciences, Germany(德国应用科学大学)

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03981 2025-05-09 cs.AI cs.CL cs.LG 62%

X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains

Qianchu Liu, Sheng Zhang, Guanghui Qin, Timothy Ossowski, Yu Gu, Ying Jin, Sid Kiblawi, Sam Preston, Mu Wei, Paul Vozila, Tristan Naumann, Hoifung Poon

机构 * Microsoft Research(微软研究院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CL、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.03380 2025-05-07 cs.CV cs.AI eess.IV 62%

Reinforced Correlation Between Vision and Language for Precise Medical AI Assistant

Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang, Marawan Elbatel, Yi Qin, Huijun Hu, Baoxun Li, Wenhui Deng, Weifeng Qin, Hongrui Li, Jialin Liang, Jun Shen, Xiaomeng Li

机构 * Department of Electronic and Computer Engineering, HKUST(香港科技大学电子与计算机工程系) Department of Radiology, Guangdong Provincial Key Laboratory of Malignant Tumor Epigenetics and Gene Regulation, Sun Yat-Sen Memorial Hospital, Sun Yat-Sen University(中山大学放射科、广东省恶性肿瘤表观遗传与基因调控重点实验室、中山纪念医院) Department of Computer Science and Engineering, HKUST(香港科技大学计算机科学与工程系)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01958 2025-05-06 cs.CV cs.CL 62%

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models

Liqiang Jing, Guiming Hardy Chen, Ehsan Aghazadeh, Xin Eric Wang, Xinya Du

机构 * University of Texas at Dallas(德克萨斯大学达拉斯分校) University of Massachusetts at Amherst(马萨诸塞大学阿姆赫斯特分校) University of California, Santa Cruz(加州大学圣克鲁兹分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.16723 2025-04-24 cs.CV cs.AI 62%

Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering

Ali Anaissi, Junaid Akram, Kunal Chaturvedi, Ali Braytee

机构 * The University of Sydney, School of Computer Science(悉尼大学计算机科学学院) University of Technology Sydney, School of Computer Science(新南威尔士大学技术学院) University of Technology Sydney, TD School(新南威尔士大学TD学院) Australian Catholic University, Peter Faber Business School(澳大利亚天主教大学彼得·法伯商学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 13 pages, 2 figures, 2025 International Conference on Computational Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.15199 2025-04-22 cs.CV cs.AI cs.LG cs.PF 62%

Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning

Yassir Benhammou, Alessandro Tiberio, Gabriel Trautmann, Suman Kalyan

机构 * NstarX

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 9 pages, 2 tables, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.14848 2025-04-22 cs.CV cs.AI 62%

Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation

Yunpu Zhao, Rui Zhang, Junbin Xiao, Ruibo Hou, Jiaming Guo, Zihao Zhang, Yifan Hao, Yunji Chen

机构 * University of Science and Technology of China(中国科学技术大学) SKL of Processors, Institute of Computing Technology, CAS(中国科学院计算技术研究所处理器专项实验室) National University of Singapore(新加坡国立大学) University of Chinese Academy of Sciences(中国科学院大学) University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.12256 2025-04-17 cs.CV cs.AI cs.LG 62%

FLIP Reasoning Challenge

Andreas Plesner, Turlan Kuzhagaliyev, Roger Wattenhofer

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Published at First Workshop on Open Science for Foundation Models at ICLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.11038 2025-04-16 cs.CV cs.AI 62%

QAVA: Query-Agnostic Visual Attack to Large Vision-Language Models

Yudong Zhang, Ruobing Xie, Jiansheng Chen, Xingwu Sun, Zhanhui Kang, Yu Wang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Accepted by NAACL 2025 main

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.20826 2025-04-16 cs.CV cs.CL cs.LG eess.IV 62%

Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation

Zhiwei Yang, Yucong Meng, Kexue Fu, Feilong Tang, Shuo Wang, Zhijian Song

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL

Comments CVPR2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09480 2025-04-15 cs.CV cs.AI 62%

Vision-Language Model for Object Detection and Segmentation: A Review and Evaluation

Yongchao Feng, Yajie Liu, Shuai Yang, Wenrui Cai, Jinqing Zhang, Qiqi Zhan, Ziyue Huang, Hongxi Yan, Qiao Wan, Chenguang Liu, Junzhe Wang, Jiahui Lv, Ziqi Liu, Tengyuan Shi, Qingjie Liu, Yunhong Wang

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments A Review and Evaluation about Vision-Language Model for Object Detection and Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.09203 2025-04-15 cs.CV cs.AI 62%

AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images

Saikat Dutta, Akhil Vasim, Siddhant Gole, Hamid Rezatofighi, Biplab Banerjee

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

Comments Accepted at EarthVision workshop, CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.08974 2025-04-15 cs.AI cs.CV 62%

Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict

Pouya Pezeshkpour, Moin Aminnaseri, Estevam Hruschka

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.01589 2025-04-09 cs.CV cs.AI 62%

Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models

Zhaochen Wang, Bryan Hooi, Yiwei Wang, Ming-Hsuan Yang, Zi Huang, Yujun Cai

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Under review at COLM 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00114 2025-04-09 cs.CV cs.AI 62%

SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World Environments

Yue Cao, Yun Xing, Jie Zhang, Di Lin, Tianwei Zhang, Ivor Tsang, Yang Liu, Qing Guo

专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05305 2025-04-08 cs.CV cs.AI 62%

URECA: Unique Region Caption Anything

Sangbeom Lim, Junwan Kim, Heeji Yoon, Jaewoo Jung, Seungryong Kim

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments Project page: https://cvlab-kaist.github.io/URECA Code: https://github.com/cvlab-kaist/URECA

详情

展开后加载摘要…

URL PDF HTML 收藏
2404.09654 2025-04-08 cs.CV cs.MM 62%

Do LLMs Understand Visual Anomalies? Uncovering LLM's Capabilities in Zero-shot Anomaly Detection

Jiaqi Zhu, Shaofeng Cai, Fang Deng, Beng Chin Ooi, Junran Wu

专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.MM

Comments Accepted by MM'24 (Oral)

详情

展开后加载摘要…

URL PDF HTML 收藏