arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4644 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 4644 篇

2411.16044 2025-09-03 cs.CV 79%

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Haozhan Shen, Kangjia Zhao, Tiancheng Zhao, Ruochen Xu, Zilun Zhang, Mingwei Zhu, Jianwei Yin

机构 * Zhejiang University(浙江大学) Om AI Research(Om AI 研究所) Binjiang Institute of Zhejiang University(浙江大学滨江研究院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted by EMNLP-2025 Main. Project page: https://szhanz.github.io/zoomeye/

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.00329 2025-08-22 cs.CV cs.LG 79%

ABC: Achieving Better Control of Multimodal Embeddings using VLMs

Benjamin Schneider, Florian Kerschbaum, Wenhu Chen

机构 * University of Waterloo(滑铁卢大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments TMLR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12458 2025-08-19 cs.CL 79%

M3PO: Multimodal-Model-Guided Preference Optimization for Visual Instruction Following

Ruirui Gao, Emily Johnson, Bowen Tan, Yanfei Qian

机构 * University of Massachusetts, Amherst(马萨诸塞大学阿姆赫斯特分校)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12198 2025-08-19 physics.ao-ph cs.AI cs.LG 79%

Exploring Multimodal AI Reasoning for Meteorological Forecasting from Skew-T Diagrams

ChangJae Lee, Heecheol Yang, Jonghak Choi

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments 24 pages, 3 figures, 9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12137 2025-08-19 cs.CV 79%

Infusing fine-grained visual knowledge to Vision-Language Models

Nikolaos-Antonios Ypsilantis, Kaifeng Chen, André Araujo, Ondřej Chum

机构 * VRG, FEE, Czech Technical University in Prague(捷克布拉格技术大学)

专题命中 图文多模态 :multimodal(abstract,comments);cross-modal(abstract);image-text(abstract);分类 cs.CV

Comments ICCVW 2025 accepted paper. Workshop name: "What is Next in Multimodal Foundation Models?"

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.10444 2025-08-15 cs.CL 79%

DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales

Herun Wan, Jiaying Wu, Minnan Luo, Xiangzheng Kong, Zihan Ma, Zhi Zeng

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14976 2025-08-15 cs.CV 79%

Hierarchical Cross-modal Prompt Learning for Vision-Language Models

Hao Zheng, Shunzhi Yang, Zhuoxin He, Jinfeng Yang, Zhenhua Huang

机构 * South China Normal University(华南师范大学) Shenzhen Polytechnic University(深圳职业技术大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted by ICCV2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.06623 2025-08-12 cs.CV 79%

ContextGuard-LVLM: Enhancing News Veracity through Fine-grained Cross-modal Contextual Consistency Verification

Sihan Ma, Qiming Wu, Ruotong Jiang, Frank Burns

机构 * Inner Mongolia University of Science & Technology(内蒙古科技大学) Federal University of Rio de Janeiro(里约热内卢联邦大学)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.09688 2025-08-11 cs.CV 79%

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

Jiho Choi, Seonho Lee, Minhyun Lee, Seungho Lee, Hyunjung Shim

机构 * KAIST, Republic of Korea(韩国科学技术院) Samsung Electronics, Republic of Korea(三星电子)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.05383 2025-08-08 cs.AI 79%

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models

Xiangxiang Zhang, Jingxuan Wei, Donghong Zhong, Qi Chen, Caijun Jia, Cheng Tan, Jinming Gu, Xiaobo Qin, Zhiping Liu, Liang Hu, Tong Sun, Yuchen Wu, Zewei Sun, Chenwei Lou, Hua Zheng, Tianyang Zhan, Changbao Wang, Shuangzhi Wu, Zefa Lin, Chang Guo, Sihang Yuan, Riwei Chen, Shixiong Zhao, Yingping Zhang, Gaowei Wu, Bihui Yu, Jiahui Wu, Zhehui Zhao, Qianqian Liu, Ruofeng Tang, Xingyue Huang, Bing Zhao, Mengyang Zhang, Youqiang Zhou

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04868 2025-08-08 cs.CV 79%

Dual-Stream Attention with Multi-Modal Queries for Object Detection in Transportation Applications

Noreen Anwar, Guillaume-Alexandre Bilodeau, Wassim Bouachir

机构 * LITIV, Polytechnique Montréal(Polytechnique Montréal 的 LITIV) Data Science Laboratory, Université du Québec (TELUQ)(Université du Québec (TELUQ) 的 Data Science Laboratory)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.04028 2025-08-07 cs.CV cs.IR 79%

Dual Prompt Learning for Adapting Vision-Language Models to Downstream Image-Text Retrieval

Yifan Wang, Tao Wang, Chenwei Tang, Caiyang Yu, Zhengqing Zang, Mengmi Zhang, Shudong Huang, Jiancheng Lv

机构 * College of Computer Science, Sichuan University(四川大学计算机学院) Engineering Research Center of Machine Learning and Industry Intelligence, Ministry of Education, Chengdu, China(教育部机器学习与产业智能工程研究中心) Deep NeuroCognition Lab, I2R and CFAR, Agency for Science, Technology and Research, Singapore(深度神经认知实验室,I2R和CFAR,科技研究局,新加坡)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Comments 10 pages, 7figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.11435 2025-08-05 cs.CV 79%

GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts

Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu, Jun-Yan He, Chenyang Li, Hanyuan Chen, Jin-Peng Lan, Bin Luo, Yifeng Geng

机构 * Dalian University of Technology(大连理工大学) Alibaba Group(阿里巴巴集团)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments ACM Multimedia 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.04585 2025-07-31 cs.CL 79%

MuSciClaims: Multimodal Scientific Claim Verification

Yash Kumar Lal, Manikanta Bandham, Mohammad Saqib Hasan, Apoorva Kashi, Mahnaz Koupaee, Niranjan Balasubramanian

机构 * Stony Brook University(石溪大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.18300 2025-07-25 cs.CV 79%

LMM-Det: Make Large Multimodal Models Excel in Object Detection

Jincheng Li, Chunyu Xie, Ji Ao, Dawei Leng, Yuhui Yin

机构 * AI Research(360人工智能研究院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Accepted at ICCV 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16572 2025-07-23 cs.CL 79%

Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models

Mohamad Ballout, Serwan Jassim, Elia Bruni

机构 * Institute of Cognitive Science, University of Osnabrück(认知科学研究所,奥斯纳布吕克大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CL

详情

展开后加载摘要…

URL PDF HTML 收藏
2402.09816 2025-07-21 cs.CV 79%

Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment

Angelos Zavras, Dimitrios Michail, Begüm Demir, Ioannis Papoutsis

机构 * organization= Orion Lab, National Observatory of Athens \& National Technical University of Athens , country= Greece organization= Department of Informatics \& Telematics, Harokopio University of Athens , country= Greece organization= Faculty of Electrical Engineering organization= BIFOLD - Berlin Institute for the Foundations of Learning

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

Comments Accepted at the ISPRS Journal of Photogrammetry and Remote Sensing. Our code implementation and weights for all experiments are publicly available at https://github.com/Orion-AI-Lab/MindTheModalityGap

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.10053 2025-07-15 cs.CV 79%

CoSMo: A Multimodal Transformer for Page Stream Segmentation in Comic Books

Marc Serra Ortega, Emanuele Vivoli, Artemis Llabrés, Dimosthenis Karatzas

机构 * Computer Vision Center and Universitat Autònoma de Barcelona(计算机视觉中心和巴塞罗那自治大学) MICC, University of Florence, Italy(佛罗伦萨大学MICC)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2408.12246 2025-07-11 cs.CV 79%

RT-OVAD: Real-Time Open-Vocabulary Aerial Object Detection via Image-Text Collaboration

Guoting Wei, Xia Yuan, Yu Liu, Zhenhao Shang, Xizhe Xue, Peng Wang, Kelu Yao, Chunxia Zhao, Haokui Zhang, Rong Xiao

机构 * Nanjing University of Science and Technology(南京理工大学) Northwestern Polytechnical University(西北工业大学) Zhejiang Lab(浙江实验室) Intellifusion Inc.(Intellifusion公司)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.24164 2025-07-08 cs.MM 79%

SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation

Ngoc Dung Huynh, Mohamed Reda Bouadjenek, Imran Razzak, Hakim Hacid, Sunil Aryal

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.MM

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.19812 2025-06-30 cs.CV cs.IR 79%

Image-text matching for large-scale book collections

Artemis Llabrés, Arka Ujjal Dey, Dimosthenis Karatzas, Ernest Valveny

机构 * Computer Vision Center, UAB(计算机视觉中心,巴塞罗那大学)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV

Journal ref Document Analysis Systems, Lecture Notes in Computer Science, vol. 14994, pp. 89-102, Springer, 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.20944 2025-06-27 cs.MM cs.CR 79%

E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs

Van-Hoang Phan, Long-Khanh Pham, Dang Vu, Anh-Duy Tran, Minh-Son Dao

专题命中 图文多模态 :cross-modal(title);multimodal(abstract);分类 cs.MM

Comments Accepted to AsiaCCS 2025 @ SCID

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.17901 2025-06-24 cs.CV 79%

PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs

Yixuan Wu, Yang Zhang, Jian Wu, Philip Torr, Jindong Gu

机构 * University of Oxford(牛津大学) Zhejiang University(浙江大学) National University of Singapore(新加坡国立大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.16673 2025-06-23 cs.CV 79%

Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge

Ruiming Chen, Junming Yang, Shiyu Xia, Xu Yang, Jing Wang, Xin Geng

机构 * School of Computer Science and Engineering, Southeast University, Nanjing 210096, China(东南大学计算机科学与工程学院) Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其跨学科应用关键实验室)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.00473 2025-06-19 cs.CV 79%

Jailbreak Large Vision-Language Models Through Multi-Modal Linkage

Yu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang, Tianxing He

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院) Shanghai Qi Zhi Institute(上海启智研究所) University of Chicago(芝加哥大学)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.14212 2025-06-18 cs.AI 79%

What's in the Box? Reasoning about Unseen Objects from Multimodal Cues

Lance Ying, Daniel Xu, Alicia Zhang, Katherine M. Collins, Max H. Siegel, Joshua B. Tenenbaum

机构 * Massachusetts Institute of Technology(麻省理工学院) Harvard University(哈佛大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

Comments Paper published at CogSci 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11155 2025-06-16 cs.CV 79%

Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search

Linhao Yu, Xinguang Ji, Yahui Liu, Fanheng Kong, Chenxi Sun, Jingyuan Zhang, Hongzhi Zhang, V. W., Fuzheng Zhang, Deyi Xiong

机构 * TJUNLP Lab, College of Intelligence and Computing, Tianjin University(天津大学智能计算学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments 28 pages; ACL 2025(main)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.09473 2025-06-12 cs.CV 79%

Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning

Cheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao, Bolin Ding, Jia Li

机构 * State Key Laboratory of Virtual Reality Technology and Systems(虚拟现实技术与系统国家重点实验室) Tongyi Lab, Alibaba Group(阿里云实验室)

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 10 pages, 6 figures, CVPR 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2407.18941 2025-06-06 cs.CV cs.LG 79%

LEMoN: Label Error Detection using Multimodal Neighbors

Haoran Zhang, Aparna Balagopalan, Nassim Oufattole, Hyewon Jeong, Yan Wu, Jiacheng Zhu, Marzyeh Ghassemi

机构 * Massachusetts Institute of Technology(麻省理工学院)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV

Comments Published in ICML 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2405.15228 2025-06-04 cs.LG cs.CV 79%

Learning from True-False Labels via Multi-modal Prompt Retrieving

Zhongnian Li, Jinghao Xu, Peng Ying, Meng Wei, Xinzheng Xu

专题命中 图文多模态 :multi-modal(title,abstract);分类 cs.CV

Comments 15 pages, 5 figures

详情

展开后加载摘要…

URL PDF HTML 收藏